benchmarks evals
Harder Programs Looked Equivalent When They Were Not
PolyHuman tests functional equivalence across human-written C++, Java and Python.
Summary
PolyHuman tests functional equivalence across human-written C++, Java and Python.
Models increasingly mislabeled non-equivalent programs as equivalent as difficulty rose, and the strongest tested model showed Python sensitivity plus run-to-run instability under identical settings. Manual analysis of 81 systematic disagreements found partial reliance on similarity cues rather than reliable semantic judgment.
Why it matters
PolyHuman tests functional equivalence across human-written C++, Java and Python.
Limits and context
No additional limitation was separately recorded.
Key claims
PolyHuman tests functional equivalence across human-written C++, Java and Python.
Evidence: source-2026-08-26-020
Sources
- arXiv preprint 2608.23961arXiv · primary research
Corrections
No corrections have been recorded for this story.