benchmarks evals
The Graph Solver Chose Its Own Representation
GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.
Summary
GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.
Graph Theory Bench tests 24 classical graph problems in 44 task-structure settings and more than 100,000 examples represented as natural language, structured language, adjacency lists or adjacency matrices. Across eight language models, accuracy changed with graph density, size, topology and encoding; no single representation stayed best. The accompanying Graph Theory Agent chooses a representation, plans and decomposes around a frozen executor. On the reported benchmark it raised Phi-4 from 53.5% to 69.1% on the easy split and from 33.0% to 41.5% on the hard split, then transferred without retraining to two other graph benchmarks.
Why it matters
GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.
Limits and context
No additional limitation was separately recorded.
Key claims
GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.
Evidence: source-2026-09-14-007
Sources
- arXiv preprint 2609.12265arXiv · primary research
Corrections
No corrections have been recorded for this story.