Predicting what people do on average is a solved problem for most businesses. Predicting what one specific person will choose, given who they are and the situation in front of them, is much harder. It's also the question behind almost every people decision: how will this employee react to a change, a message or a new policy?
Large language models (LLMs) are increasingly used to simulate human behavior, and they work reasonably well in aggregate. But at the individual level, there's far less evidence. So at Cairo Labs we built Rosetta Arena, an evaluation environment that tests different AI systems on the same question: given a person's profile and a closed set of options, which one did they actually pick?
This post summarizes the first experiment. It's a preliminary technical report, and we want to be clear about what it shows and what it doesn't.
What we tested
The question was simple: can a model built specifically for human behavior compete with general-purpose frontier LLMs at predicting individual choices, while being faster and cheaper?
We compared five competitors:
- Cairo Rosetta: a specialized behavior model (a Large Behavior Model, or LBM) developed by Cairo AI. It takes a structured description of the person and the situation and returns the option that best fits that profile.
- Claude Sonnet 5.5, GPT-5.6 Terra and Gemini 3.8 Flash: three general-purpose LLMs, accessed through OpenRouter, each given the same structured prompt and asked to return the option they predict the person chose.
- A frequency baseline: for each event, it simply predicts the option most people historically chose. If a system can't beat this, it isn't extracting anything useful from the individual's profile.
The data and the protocol
The test set is 30 real decisions made by real people in economic, negotiation and social dilemmas. They come from two public datasets: Twin-2K-500, which collects responses from more than 2,000 participants to more than 500 questions, and CaSiNo, a corpus of negotiation dialogues about dividing camping resources.
To keep the evaluation clean:
- Every case came from data held out of Rosetta's training and out of the baseline's reference data.
- Each case had between 2 and 12 discrete options, shuffled with a fixed seed so no model benefits from favoring the first or last option.
- Every competitor received the same information: the person's standardized cognitive and personality profile, a description of the situation and the list of options. The real choice stayed hidden until scoring.
The results
Accuracy: a directional edge, not a proven win
Rosetta scored highest, followed closely by Claude Sonnet 5.5. But the confidence intervals are wide, roughly ±16.5 percentage points, and they overlap heavily. Rosetta's advantage over Claude Sonnet 5.5 amounts to a single case out of 30.
The honest reading is competitive parity with a directional advantage for Rosetta. The larger gaps (20 points over the baseline, 23 points over Gemini) don't reach statistical significance in an unpaired test either, though a case-by-case paired analysis could change that. We'll run it as the sample grows.
Evidence that Rosetta reads the individual
The frequency baseline got 12 of 30 right, which means in the other 18 cases the person chose something other than the most common answer. Those are the interesting decisions, because you can only get them right by using information about the individual. By construction, the baseline gets none of them. Rosetta got at least 6 of those 18 right (a third or more), and Claude Sonnet 5.5 and GPT-5.6 Terra got at least 5 and 3 respectively.
It's also worth noting that Gemini 3.8 Flash landed below the baseline. Being a frontier model doesn't guarantee that you extract useful signal from a person's profile.
Efficiency: where the gap is clearest
Rosetta answered in 1.73 seconds on average, against about 2.93 seconds for GPT-5.6 Terra and Claude Sonnet 5.5, and 4.73 seconds for Gemini 3.8 Flash. It also generated no external API charges, while the commercial LLMs cost between about USD 1.90 and USD 3.03 per 1,000 predictions.
At small scale that's pocket change. At the scale Rosetta is designed for, it isn't. Simulating 10,000 profiles across 100 scenarios means a million predictions, which would cost roughly USD 1,900 to USD 3,030 in commercial API fees.
*This figure excludes the compute, hosting and amortized training costs of Rosetta's dedicated infrastructure, which this telemetry doesn't capture. A full total-cost-of-ownership comparison is still pending.
On the combined view, Rosetta sits ahead of all three LLMs on accuracy, latency and direct cost at the point estimates. The accuracy part of that claim carries the uncertainty described above. The latency and cost parts are large differences.
What this result does not tell us
We'd rather list the limits ourselves than have you find them:
- Small sample. 30 decisions gives about ±16.5 points of uncertainty per accuracy figure. Detecting a 10-point difference with 80% power would take roughly 390 cases per competitor in an unpaired design.
- Home-field advantage. Rosetta and the baseline had training or reference data from the same sources as the test cases (in separate partitions), while the LLMs ran zero-shot. This is a specialist in its domain against generalists, and Rosetta's generalization to unseen data sources is still untested.
- Possible LLM contamination. Both datasets are public, so they may be in the LLMs' pretraining data. If so, that would favor the LLMs, not Rosetta.
- One prompt. The LLMs got a single structured prompt. Few-shot examples, extended reasoning or tuned sampling could change their accuracy, latency and cost.
- Infrastructure differences. LLM latency includes OpenRouter routing and each provider's network, while Rosetta runs on dedicated infrastructure. Part of the speed gap comes from infrastructure, not only from the model. We also report averages without the p50/p95 distribution.
- One run, one seed. We haven't measured variability across seeds or LLM non-determinism, and the number of options per case (2 to 12) means chance-level accuracy differs from case to case.
What comes next
The next iteration of Rosetta Arena is designed to answer exactly these questions:
- Scale to several hundred cases, stratified by source and type of dilemma.
- Publish the case-by-case results and run paired McNemar tests, with accuracy split between common and uncommon decisions.
- Add probabilistic metrics (log-loss, Brier score, calibration) and a chance-adjusted accuracy.
- Test the LLMs with stronger prompting protocols and include other specialized models as competitors.
- Measure Rosetta on data sources it hasn't seen, including domains closer to real use cases, like reactions to messages and communications.
- Report latency distributions and full infrastructure cost.
- Run an ablation where every competitor gets the cases without the individual profile, to quantify how much accuracy truly comes from the psychometric information.
The bottom line
This first run shows competitive accuracy with a directional edge for Rosetta, and a clear, large advantage in speed and cost. It does not show that a specialized behavior model is more accurate than frontier LLMs. That claim needs a bigger, paired evaluation, and that's the work we're doing next.
What the experiment does validate is the infrastructure: a reproducible, leak-free arena where we can keep testing Rosetta in public as it evolves. Want to follow along or see what behavior simulation could do for your organization? Create your Cairo workspace and talk to us.
Sources
- Cairo Labs, Cairo Rosetta vs. frontier language models in predicting individual human choices: Rosetta Arena, Experiment 1 (preliminary technical report), October 2026.
- Toubia, O., Gui, G. Z., Peng, T., Merlau, D. J., Li, A., & Chen, H., Twin-2K-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions, Marketing Science, 2025.
- Chawla, K., Ramirez, J., Clever, R., Lucas, G., May, J., & Gratch, J., CaSiNo: A corpus of campsite negotiation dialogues for automatic negotiation systems, NAACL-HLT 2021.




