TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Agent Groups Agreed—Mostly on the Wrong Answer

Matched language-model groups exceeded human consensus by 34.0 to 44.4 percentage points in sensitivity analyses.

Published Updated Story ID: mp-2026-09-19-004
Read the complete editionStory JSON

Summary

Matched language-model groups exceeded human consensus by 34.0 to 44.4 percentage points in sensitivity analyses.

Researchers replayed 100 held-out human Wason-task groups with language-model agents anchored to each participant's initial answer. Human consensus estimates moved from 24% to 57% depending on who counted as participating, yet two sensitivity analyses still found model groups 34.0 to 44.4 percentage points more consensual. Reasoning-mode agents agreed nearly unanimously after a change designed to remove a memorizable answer, but usually agreed on the wrong answer. In this setting, synthetic agreement did not estimate human deliberation or collective accuracy.

Why it matters

Matched language-model groups exceeded human consensus by 34.0 to 44.4 percentage points in sensitivity analyses.

Limits and context

  • In this setting, synthetic agreement did not estimate human deliberation or collective accuracy.

Key claims

  1. Matched language-model groups exceeded human consensus by 34.0 to 44.4 percentage points in sensitivity analyses.

    Qualification: In this setting, synthetic agreement did not estimate human deliberation or collective accuracy.

    Evidence: source-2026-09-19-004

Sources

  1. arXiv preprint 2609.20543arXiv · primary research

Corrections

No corrections have been recorded for this story.