TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Graph Solver Chose Its Own Representation

GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.

Published Updated Story ID: mp-2026-09-14-007
Read the complete editionStory JSON

Summary

GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.

Graph Theory Bench tests 24 classical graph problems in 44 task-structure settings and more than 100,000 examples represented as natural language, structured language, adjacency lists or adjacency matrices. Across eight language models, accuracy changed with graph density, size, topology and encoding; no single representation stayed best. The accompanying Graph Theory Agent chooses a representation, plans and decomposes around a frozen executor. On the reported benchmark it raised Phi-4 from 53.5% to 69.1% on the easy split and from 33.0% to 41.5% on the hard split, then transferred without retraining to two other graph benchmarks.

Why it matters

GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. GT Bench spans 100,000 examples and four graph encodings; a selector-and-planner lifted Phi-4 from 33.0% to 41.5% on the hard split.

    Evidence: source-2026-09-14-007

Sources

  1. arXiv preprint 2609.12265arXiv · primary research

Corrections

No corrections have been recorded for this story.