TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Live-State Wrappers Cut Grounding Failures About 78%

A 200-session game-coaching benchmark separated correct answers from correct direction.

Published Updated Story ID: mp-2026-09-24-016
Read the complete editionStory JSON

Summary

A 200-session game-coaching benchmark separated correct answers from correct direction.

State-grounded wrappers lifted session-level grounded accuracy to 83.5% from 20.0% for prompting and 26.5% for a tool-use baseline. The authors do not establish independent effects for each wrapper.

Why it matters

A 200-session game-coaching benchmark separated correct answers from correct direction.

Limits and context

  • The authors do not establish independent effects for each wrapper.

Key claims

  1. A 200-session game-coaching benchmark separated correct answers from correct direction.

    Qualification: The authors do not establish independent effects for each wrapper.

    Evidence: source-2026-09-24-018

Sources

  1. arXiv preprint 2609.27606arXiv · primary research

Corrections

No corrections have been recorded for this story.