Cairo AI logo
Platform
Map
EvaluationsIndicatorsTeamsIntegrations
Execute
Learning PathsCoursesScenariosAnalytics
Predict
Ask CairoSimulationsRecommendationsFit Prediction
Person typing on a laptop
Get the full platform
See pricing
Science
Research
Cairo Rosetta
Data
Benchmarks
Lab
Cairo Labs
Tool
ROI Calculator
Resources
BlogResearchFAQ
Person wearing a blue sweater typing on a laptop at a wooden table with sunlight casting shadows.
Featured
The latest from the Cairo blog
Read the blog
Company
About CairoCareers
Get in touch
Contact
Trust
SecurityPrivacy
Book a Demo
Blog
Insight
October 10, 2026
Fernando Flores
Written by
Fernando Flores
Share this article
Link copied

What Is a Large Behavior Model (LBM)? Aaru, Simile and Cairo Rosetta Explained

A large behavior model (LBM) is an AI model trained to predict what a specific person, group or population will do in a given situation. Where a large language model (LLM) predicts the next word, an LBM predicts the next decision: which option someone picks, how they react to a message, whether they adopt a product, or whether they stay in a job after a change.

Behavior simulation went from research curiosity to one of the most heavily funded categories in AI in less than two years. This guide explains what an LBM is, how it differs from an LLM, how it works, and how three companies are applying the idea in very different ways: Aaru, Simile and Cairo Rosetta. We're Cairo, so we're one of the three. We've tried to be precise about what each one is built for.

Key takeaways
  • An LBM predicts behavior (choices, reactions, actions) from a description of the person and the situation. An LLM predicts text.
  • General-purpose LLMs simulate the average person reasonably well, but struggle to stay consistent with one specific individual. That gap is why specialized behavior models exist.
  • Aaru simulates whole populations with thousands of AI agents for market research and forecasting. Simile is building a foundation model of human behavior for enterprise decisions. Cairo Rosetta predicts how individual people inside organizations will react, grounded in their psychometric profile.
  • Judge any LBM by one thing: how well it predicts real decisions it has never seen, with honest confidence intervals.

What is a large behavior model?

A large behavior model is a model whose output is a behavior rather than a piece of content. You give it two things:

  • The person: who they are, described through a profile, a history of past decisions, survey answers, transactions or psychometric traits.
  • The environment: the situation they face and the options available, such as a price change, a new policy, a negotiation offer or a message from their manager.

It returns the most likely behavior, ideally with a probability attached. This framing goes back to psychologist Kurt Lewin, who proposed that behavior is a function of the person and their environment: B = f(P, E). Recent LBM research makes that formula explicit. A July 2026 paper from Amity AI, for example, models retail customers exactly this way, grounding each prediction in a shopper's transaction history plus the current decision environment.[1]

Person (P)profile, traits, historyEnvironment (E)situation and optionsBehavior (B)predicted choice + probabilityLBMB = f(P, E)
An LBM maps a person and a situation to a predicted behavior.

A note on the term

“Large behavior model” is also used in robotics, where it describes models that learn physical actions (like grasping or folding) from demonstrations. That's a different field. In this guide, LBM means a model of human decision-making and social behavior, the sense used by behavior simulation companies and by recent papers such as OMGene AI Lab's February 2026 LBM for predicting individual strategic choices.[2]

LBM vs. LLM: what's the difference?

Large language model (LLM)Large behavior model (LBM)
PredictsThe next token of textThe next decision or action of a person or group
Trained onText from the web, books and codeRecords of real behavior: choices, survey answers, transactions, interviews, outcomes
Unit of analysisA promptA person (or population) in a situation
Measured byQuality, reasoning and task benchmarksAccuracy against real, held-out human decisions
Typical failurePlausible but generic answers; drifts toward the average personOverconfidence outside the situations it was trained on
Main useWriting, coding, search, assistantsSimulation, forecasting and testing decisions before making them

Most LBMs today are built on top of language models: they start from an LLM and are fine-tuned or continually trained on behavioral data. The difference isn't the architecture. It's what the model is optimized to get right.

Why aren't LLMs enough to simulate behavior?

You can prompt any LLM with “You are a 42-year-old operations manager who scores high on conscientiousness” and ask what that person would do. It will give you a fluent answer. The problem is that fluent isn't the same as accurate.

Research on behavior simulation keeps finding the same pattern: LLMs capture population-level tendencies but have trouble producing consistent, individual-specific behavior, especially when the right answer depends on how a person's traits interact with a particular situation.[2] Personas written in natural language get diluted by the model's general priors, and the simulated person drifts toward the average.

That's why a second line of work trains models directly on large datasets of real human decisions. Examples include Centaur, a foundation model of human cognition, and the LBMs described above.[2] The bet is simple: if you want to predict behavior, train on behavior.

How does a large behavior model work?

Implementations vary, but credible LBMs share four ingredients:

  1. Ground truth about real people. Interviews, surveys, psychometric assessments, transaction logs or recorded decisions. Without it, a simulation is just an LLM's imagination.
  2. A structured representation of the person. A profile the model can condition on consistently, such as trait scores, purchase history or a vector of behavioral indicators.
  3. A model that maps person + situation to behavior. Usually a language model fine-tuned on behavioral data, sometimes combined with many interacting agents.
  4. Validation on held-out decisions. The model is tested on real choices it never saw in training, and ideally reports how confident it is in each prediction.

Aaru, Simile and Cairo Rosetta: three approaches to behavior models

All three companies work on predicting human behavior, but they answer different questions at different scales.

Aaru: simulating entire populations

Aaru, founded in March 2024 by Cameron Fink, Ned Koh and John Kessler, builds a prediction engine that generates thousands of AI agents to simulate how populations behave, using public and proprietary data.[3] Its flagship model is called Lumen, and Accenture Song committed to integrating it into its AI offering for product development, marketing and customer service.[4]

Aaru's use cases sit at the population level: pricing decisions, political and election forecasting, disaster-response planning and adoption-rate prediction.[5] Its polling methodology drew attention after it correctly predicted the outcome of the New York Democratic primary, according to Semafor. Partners include Accenture, EY and Interpublic Group, and in December 2025 it raised a Series A led by Redpoint Ventures at a reported $1 billion headline valuation.[3]

Best described as: multi-agent population simulation for market research and forecasting.

Simile: a foundation model of human behavior

Simile comes out of Stanford. Its CEO, Joon Sung Park, was the lead author of the 2023 Generative Agents paper, which placed 25 AI agents in a simulated town where they formed relationships and organized a party on their own.[6] His co-founders include Stanford professors Michael Bernstein and Percy Liang, plus Lainie Yallen on go-to-market.

Simile describes itself as a human simulation company building a foundation model of human behavior, grounded in rich information about real people rather than only what's available on the web.[7] A 2024 study by Park and colleagues found that agents built from two-hour interviews replicated people's survey answers about 85% as accurately as those people replicated their own answers two weeks later.[8] Simile also says it trained a separate confidence model that estimates how reliable each simulation is.[6]

Simile closed a $100 million Series A led by Index Ventures in February 2026, then a $200 million Series B at a $2 billion valuation led by Greenoaks five months later, with CVS Health among its customers.[9][10] Its stated mission is to accurately simulate all eight billion people on Earth.

Best described as: a general foundation model for simulating individuals, segments and markets, sold to large enterprises for product, marketing and strategy decisions.

Cairo Rosetta: behavior models for people inside organizations

Cairo Rosetta is the large behavior model developed by Cairo Labs. Where Aaru and Simile mostly simulate consumers, voters and markets, Rosetta focuses on a narrower and harder question: how will this specific person react to this specific situation at work? A new manager, a reorganization, a policy change, a difficult piece of feedback, a negotiation.

Rosetta takes a structured profile of the person and a description of the situation, and returns the option that best fits that profile. In Cairo's platform, that profile comes from Cairo Index, which describes each employee across 80+ behavioral indicators.

We test Rosetta in public through Rosetta Arena. In the first experiment, on 30 held-out real decisions from economic, negotiation and social dilemmas, Rosetta reached 60.0% accuracy versus 56.7% for the best frontier LLM, with about 41% lower latency and no external API cost. With only 30 cases, the accuracy gap isn't statistically significant, and we say so. The full results and limitations are in our Rosetta Arena write-up.[11]

Best described as: an individual-level behavior model for organizations, grounded in psychometric profiles and designed for people development, not for hiring or termination decisions.

Side-by-side comparison

AaruSimileCairo Rosetta
Core questionHow will a population respond?How will people, segments and markets respond?How will this person at work respond?
Main unitSynthetic populationsAgentic twins and panelsIndividual employee profiles and teams
Grounding dataPublic and proprietary dataInterviews, surveys, transactions, behavioral sciencePsychometric profiles (80+ indicators) and real decision data
Typical usesMarket research, pricing, polling, forecastingProduct concepts, customer insight, strategyChange management, communication, team design, development
BuyerConsultancies, agencies, campaigns, enterprisesFortune 100 enterprisesHR and leadership teams in growing companies

What are large behavior models used for?

  • Market and product research: testing a concept, price or message against a simulated audience in hours instead of weeks.
  • Forecasting: elections, adoption curves and reactions to public events.
  • Policy and public sector: estimating how citizens might respond to a new rule before it's rolled out.
  • Organizations: anticipating how teams will react to a reorganization, which version of an announcement lands best, or where a change is most likely to create friction, so leaders can prepare people instead of reacting after the fact.
  • Training: conversational simulations where people practice a hard conversation with a realistic counterpart before having it for real.

How to evaluate a large behavior model

Behavior simulation is easy to demo and hard to validate. Before trusting any LBM, ask:

6 questions to ask any behavior model vendor
  1. What real behavior was it trained on, and how was it collected?
  2. Is accuracy measured on decisions the model has never seen?
  3. Is it compared against a simple baseline, like “predict the most common answer”?
  4. Does it report confidence for each prediction, and is that confidence calibrated?
  5. Does it predict individuals, or only aggregates?
  6. What is it explicitly not allowed to be used for?

Two caveats apply to the whole category. First, human behavior is noisy: people don't even answer the same question the same way twice, which puts a ceiling on how accurate any model can be. Second, a model that predicts individual behavior is powerful, so how it's used matters as much as how good it is. At Cairo, results are aggregated by team for leaders, individual data stays confidential, and predictions are used to develop people, never to decide who gets hired or fired.

Frequently asked questions

What does LBM stand for in AI?

LBM stands for large behavior model (sometimes written large behavioral model). It's an AI model trained to predict human decisions and actions rather than to generate text. In robotics, the same acronym refers to models that learn physical robot actions.

What is the difference between an LBM and an LLM?

An LLM predicts the next word in a piece of text. An LBM predicts the next decision a person or group will make, given who they are and the situation they face. Many LBMs are built by fine-tuning LLMs on records of real behavior.

Is a large behavior model the same as a digital twin?

They're related. A digital twin of a person is a simulated version of that individual. An LBM is the model that powers it: the engine that predicts what the twin, or a whole population of twins, will do.

Which companies are building large behavior models?

Notable companies include Aaru (population-scale simulation and forecasting), Simile (a foundation model of human behavior for enterprises) and Cairo with Cairo Rosetta (individual behavior prediction inside organizations). Others in the space include CulturePulse and Artificial Societies, plus academic labs publishing LBM research.

How accurate are large behavior models?

It depends on the task and the data. Simile's founding research reported agents that replicated people's survey answers about 85% as well as people replicate themselves. Cairo Rosetta reached 60% on a small set of real multi-option decisions. Always look for accuracy on held-out data, with confidence intervals and a baseline.

Can a large behavior model predict how employees will react to change?

That's what Cairo Rosetta is designed for. Given an employee's psychometric profile and a description of a change, it predicts the most likely reaction, so leaders can prepare teams and communication in advance. Predictions are probabilistic and should support human judgment, not replace it.

The bottom line

LLMs taught machines to talk like people. Large behavior models are trying to predict what people actually do. Aaru does it at the scale of populations, Simile at the scale of a general foundation model, and Cairo Rosetta at the scale where most real decisions about people happen: one person, one team, one change at a time.

Want to see what behavior intelligence looks like for your organization? Create your Cairo workspace and get your team's first profiles within weeks.

Sources

  1. Modecrua, W., Pachtrachai, K., & Kraisingkorn, T., Large Behavior Model: A Promptable Digital Twin of the Retail Customer, arXiv:2607.06993, July 2026.
  2. Yellin, B., Ezra, E., Foreman, M., & Grinapol, S., Decoding the Human Factor: High Fidelity Behavioral Prediction for Strategic Foresight, arXiv:2602.17222, February 2026.
  3. TechCrunch, AI synthetic research startup Aaru raised a Series A at a $1B ‘headline’ valuation, December 5, 2025.
  4. Research Live / DRNO, Further Funding for Consumer Simulation Platform Aaru, December 8, 2025.
  5. NeuronFeed, Aaru company profile.
  6. Unite.AI, Simile Raises More Than $200 Million at a $2 Billion Valuation to Scale Human Behavior Simulations, 2026.
  7. Toolradar, Simile company profile.
  8. Park, J. S., et al., Generative Agent Simulations of 1,000 People, arXiv:2411.10109, 2024.
  9. Latham & Watkins, Latham & Watkins Advises Simile AI, Inc. on US$100 Million Series A, February 19, 2026.
  10. RuntimeWire, Simile raises $200M at a $2B valuation to simulate human behavior, July 2026.
  11. Cairo Labs, Cairo Rosetta vs. Frontier LLMs: Predicting What a Specific Person Will Actually Choose (Rosetta Arena, Experiment 1), October 2026.
Blog
Insight
October 10, 2026
Fernando Flores
Written by
Fernando Flores
Share this article
Link copied
Keep in the loop
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Monthly updates • No spam ever
Abstract decorative illustration
Cairo AI logo
Product
OverviewFor SchoolsHCM SoftwarePerformance Management
Product
Request demo
Company
BlogContactPrivacy
© Copyright Cairo Labs Inc.