benchmarks evals
Harder Workflows Flattened Every Judge
AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.
Summary
AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.
Across six workflow graph shapes and three difficulty levels, alignment with the programmatic reference declined as tasks grew harder and fell faster when ground truth was hidden. On hard no-ground-truth cases, six judges converged in a narrow 77-to-82-percent band despite scale differences. Structured rubrics improved alignment by as much as 6.5 points, while reasoning traces and temperature changes had little effect. Ground truth sometimes reduced alignment through apparent over-anchoring.
Why it matters
AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.
Limits and context
No additional limitation was separately recorded.
Key claims
AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.
Evidence: source-2026-08-30-010
Sources
- arXiv preprint 2608.26623arXiv · primary research
Corrections
No corrections have been recorded for this story.