benchmarks evals
Human Partners Read the Hint That Wasn’t There
Hanabi logs exposed a 26-point gap between literal information and human play—almost absent in AI-only pairs.
Summary
Hanabi logs exposed a 26-point gap between literal information and human play—almost absent in AI-only pairs.
The proposed convention gap compares failure predicted from a message’s literal content with what players actually do. Across about 101,000 Hanabi actions, the gap measured 26.2 points for human pairs, minus 0.7 for AI pairs and 16.4 for human–AI teams. Unhinted human plays drove a 46-point gap, suggesting shared conventions carried information beyond explicit clues. Within human–AI games, partners with similar literal information produced sharply different human outcomes. The authors argue that AI–AI score alone can miss compatibility with people.
Why it matters
Hanabi logs exposed a 26-point gap between literal information and human play—almost absent in AI-only pairs.
Limits and context
No additional limitation was separately recorded.
Key claims
Hanabi logs exposed a 26-point gap between literal information and human play—almost absent in AI-only pairs.
Evidence: source-2026-09-12-007
Sources
- arXiv preprint 2609.11489arXiv · primary research
Corrections
No corrections have been recorded for this story.