Cairo AI logo
Platform
Map
EvaluationsIndicatorsTeamsIntegrations
Execute
Learning PathsCoursesScenariosAnalytics
Predict
Ask CairoSimulationsRecommendationsFit Prediction
Person typing on a laptop
Get the full platform
See pricing
Science
Research
Cairo Rosetta
Data
Benchmarks
Lab
Cairo Labs
Tool
ROI Calculator
Resources
BlogResearchFAQ
Person wearing a blue sweater typing on a laptop at a wooden table with sunlight casting shadows.
Featured
The latest from the Cairo blog
Read the blog
Company
About CairoCareers
Get in touch
Contact
Trust
SecurityPrivacy
Book a Demo
Blog
Insight
October 7, 2026
Fernando Flores
Written by
Fernando Flores

Cairo Rosetta vs. Frontier LLMs: Predicting What a Specific Person Will Actually Choose

Download PDF

Predicting what people do on average is a solved problem for most businesses. Predicting what one specific person will choose, given who they are and the situation in front of them, is much harder. It's also the question behind almost every people decision: how will this employee react to a change, a message or a new policy?

Large language models (LLMs) are increasingly used to simulate human behavior, and they work reasonably well in aggregate. But at the individual level, there's far less evidence. So at Cairo Labs we built Rosetta Arena, an evaluation environment that tests different AI systems on the same question: given a person's profile and a closed set of options, which one did they actually pick?

This post summarizes the first experiment. It's a preliminary technical report, and we want to be clear about what it shows and what it doesn't.

Key takeaways
  • Cairo Rosetta, our specialized behavior model, got the highest raw accuracy: 60.0% (18 of 30 decisions).
  • With only 30 cases, that lead over the best LLM (56.7%) is not statistically distinguishable. Read it as competitive parity with a directional edge, not proven superiority.
  • The clearest difference is efficiency: about 41% lower latency than the fastest LLMs and no external API charges.
  • Rosetta beat a naive "always pick the most common answer" baseline by 20 points, early evidence that it uses information about the individual.

What we tested

The question was simple: can a model built specifically for human behavior compete with general-purpose frontier LLMs at predicting individual choices, while being faster and cheaper?

We compared five competitors:

  • Cairo Rosetta: a specialized behavior model (a Large Behavior Model, or LBM) developed by Cairo AI. It takes a structured description of the person and the situation and returns the option that best fits that profile.
  • Claude Sonnet 5.5, GPT-5.6 Terra and Gemini 3.8 Flash: three general-purpose LLMs, accessed through OpenRouter, each given the same structured prompt and asked to return the option they predict the person chose.
  • A frequency baseline: for each event, it simply predicts the option most people historically chose. If a system can't beat this, it isn't extracting anything useful from the individual's profile.

The data and the protocol

The test set is 30 real decisions made by real people in economic, negotiation and social dilemmas. They come from two public datasets: Twin-2K-500, which collects responses from more than 2,000 participants to more than 500 questions, and CaSiNo, a corpus of negotiation dialogues about dividing camping resources.

To keep the evaluation clean:

  • Every case came from data held out of Rosetta's training and out of the baseline's reference data.
  • Each case had between 2 and 12 discrete options, shuffled with a fixed seed so no model benefits from favoring the first or last option.
  • Every competitor received the same information: the person's standardized cognitive and personality profile, a description of the situation and the list of options. The real choice stayed hidden until scoring.

The results

CompetitorCorrectAccuracy95% CIAvg. latency
Cairo Rosetta18 / 3060.0%42.3–75.4%1.73 s
Claude Sonnet 5.517 / 3056.7%39.2–72.6%2.94 s
GPT-5.6 Terra15 / 3050.0%33.2–66.8%2.93 s
Frequency baseline12 / 3040.0%24.6–57.7%~0 s (precomputed)
Gemini 3.8 Flash11 / 3036.7%21.9–54.5%4.73 s

Accuracy: a directional edge, not a proven win

Rosetta scored highest, followed closely by Claude Sonnet 5.5. But the confidence intervals are wide, roughly ±16.5 percentage points, and they overlap heavily. Rosetta's advantage over Claude Sonnet 5.5 amounts to a single case out of 30.

The honest reading is competitive parity with a directional advantage for Rosetta. The larger gaps (20 points over the baseline, 23 points over Gemini) don't reach statistical significance in an unpaired test either, though a case-by-case paired analysis could change that. We'll run it as the sample grows.

Evidence that Rosetta reads the individual

The frequency baseline got 12 of 30 right, which means in the other 18 cases the person chose something other than the most common answer. Those are the interesting decisions, because you can only get them right by using information about the individual. By construction, the baseline gets none of them. Rosetta got at least 6 of those 18 right (a third or more), and Claude Sonnet 5.5 and GPT-5.6 Terra got at least 5 and 3 respectively.

It's also worth noting that Gemini 3.8 Flash landed below the baseline. Being a frontier model doesn't guarantee that you extract useful signal from a person's profile.

Efficiency: where the gap is clearest

1.73 s
Rosetta's average latency per prediction
-41%
latency vs. the two fastest LLMs (1.69x faster)
-63%
latency vs. Gemini 3.8 Flash (2.73x faster)
$0
external API charges for Rosetta*

Rosetta answered in 1.73 seconds on average, against about 2.93 seconds for GPT-5.6 Terra and Claude Sonnet 5.5, and 4.73 seconds for Gemini 3.8 Flash. It also generated no external API charges, while the commercial LLMs cost between about USD 1.90 and USD 3.03 per 1,000 predictions.

At small scale that's pocket change. At the scale Rosetta is designed for, it isn't. Simulating 10,000 profiles across 100 scenarios means a million predictions, which would cost roughly USD 1,900 to USD 3,030 in commercial API fees.

*This figure excludes the compute, hosting and amortized training costs of Rosetta's dedicated infrastructure, which this telemetry doesn't capture. A full total-cost-of-ownership comparison is still pending.

On the combined view, Rosetta sits ahead of all three LLMs on accuracy, latency and direct cost at the point estimates. The accuracy part of that claim carries the uncertainty described above. The latency and cost parts are large differences.

What this result does not tell us

We'd rather list the limits ourselves than have you find them:

  • Small sample. 30 decisions gives about ±16.5 points of uncertainty per accuracy figure. Detecting a 10-point difference with 80% power would take roughly 390 cases per competitor in an unpaired design.
  • Home-field advantage. Rosetta and the baseline had training or reference data from the same sources as the test cases (in separate partitions), while the LLMs ran zero-shot. This is a specialist in its domain against generalists, and Rosetta's generalization to unseen data sources is still untested.
  • Possible LLM contamination. Both datasets are public, so they may be in the LLMs' pretraining data. If so, that would favor the LLMs, not Rosetta.
  • One prompt. The LLMs got a single structured prompt. Few-shot examples, extended reasoning or tuned sampling could change their accuracy, latency and cost.
  • Infrastructure differences. LLM latency includes OpenRouter routing and each provider's network, while Rosetta runs on dedicated infrastructure. Part of the speed gap comes from infrastructure, not only from the model. We also report averages without the p50/p95 distribution.
  • One run, one seed. We haven't measured variability across seeds or LLM non-determinism, and the number of options per case (2 to 12) means chance-level accuracy differs from case to case.

What comes next

The next iteration of Rosetta Arena is designed to answer exactly these questions:

  • Scale to several hundred cases, stratified by source and type of dilemma.
  • Publish the case-by-case results and run paired McNemar tests, with accuracy split between common and uncommon decisions.
  • Add probabilistic metrics (log-loss, Brier score, calibration) and a chance-adjusted accuracy.
  • Test the LLMs with stronger prompting protocols and include other specialized models as competitors.
  • Measure Rosetta on data sources it hasn't seen, including domains closer to real use cases, like reactions to messages and communications.
  • Report latency distributions and full infrastructure cost.
  • Run an ablation where every competitor gets the cases without the individual profile, to quantify how much accuracy truly comes from the psychometric information.

The bottom line

This first run shows competitive accuracy with a directional edge for Rosetta, and a clear, large advantage in speed and cost. It does not show that a specialized behavior model is more accurate than frontier LLMs. That claim needs a bigger, paired evaluation, and that's the work we're doing next.

What the experiment does validate is the infrastructure: a reproducible, leak-free arena where we can keep testing Rosetta in public as it evolves. Want to follow along or see what behavior simulation could do for your organization? Create your Cairo workspace and talk to us.

Sources

  1. Cairo Labs, Cairo Rosetta vs. frontier language models in predicting individual human choices: Rosetta Arena, Experiment 1 (preliminary technical report), October 2026.
  2. Toubia, O., Gui, G. Z., Peng, T., Merlau, D. J., Li, A., & Chen, H., Twin-2K-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions, Marketing Science, 2025.
  3. Chawla, K., Ramirez, J., Clever, R., Lucas, G., May, J., & Gratch, J., CaSiNo: A corpus of campsite negotiation dialogues for automatic negotiation systems, NAACL-HLT 2021.
Free download
Rosetta Arena technical report (PDF)
joincairo.com
Download
Blog
Insight
October 7, 2026
Fernando Flores
Written by
Fernando Flores
Keep in the loop
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Monthly updates • No spam ever
Abstract decorative illustration
Cairo AI logo
Product
OverviewFor SchoolsHCM SoftwarePerformance Management
Product
Request demo
Company
BlogContactPrivacy
© Copyright Cairo Labs Inc.