benchmarks evals
The Security Agent Failed Before the Tested Capability Appeared
Checkpoint instrumentation separates long-horizon failures that happen before and after an agent reaches the relevant state.
Summary
Checkpoint instrumentation separates long-horizon failures that happen before and after an agent reaches the relevant state.
In one 92-seed study, protocol-disambiguation guidance raised state observation for Gemini 2.5 Flash from 65.5 to 95.4 percent, but repeating the design with Gemini 3.7 Flash produced the opposite effect. The shifting bottleneck shows why final success alone cannot reveal which capability failed or whether it was exercised at all.
Why it matters
Checkpoint instrumentation separates long-horizon failures that happen before and after an agent reaches the relevant state.
Limits and context
- The shifting bottleneck shows why final success alone cannot reveal which capability failed or whether it was exercised at all.
Key claims
Checkpoint instrumentation separates long-horizon failures that happen before and after an agent reaches the relevant state.
Qualification: The shifting bottleneck shows why final success alone cannot reveal which capability failed or whether it was exercised at all.
Evidence: source-2026-08-24-007
Sources
- arXiv preprint 2608.20563arXiv · primary research
Corrections
No corrections have been recorded for this story.