TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Correct Answer Was Present and Still Lost the Vote

Fixed candidate-pool replays isolated how frequency and judge signals determine the answer a multi-agent system reports.

Published Updated Story ID: mp-2026-08-27-012
Read the complete editionStory JSON

Summary

Fixed candidate-pool replays isolated how frequency and judge signals determine the answer a multi-agent system reports.

Across 81,390 fixed pools from 16,278 questions, combining answer frequency with judge evaluation changed only terminal selection and raised accuracy from 63.82 percent to 70.82–70.95 percent. The gains mainly rescued correct answers outnumbered by popular errors; judge reliability varied with task, generator and answer rarity.

Why it matters

Fixed candidate-pool replays isolated how frequency and judge signals determine the answer a multi-agent system reports.

Limits and context

  • Across 81,390 fixed pools from 16,278 questions, combining answer frequency with judge evaluation changed only terminal selection and raised accuracy from 63.82 percent to 70.82–70.95 percent.

Key claims

  1. Fixed candidate-pool replays isolated how frequency and judge signals determine the answer a multi-agent system reports.

    Qualification: Across 81,390 fixed pools from 16,278 questions, combining answer frequency with judge evaluation changed only terminal selection and raised accuracy from 63.82 percent to 70.82–70.95 percent.

    Evidence: source-2026-08-27-012

Sources

  1. arXiv preprint 2608.25937arXiv · primary research

Corrections

No corrections have been recorded for this story.