TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Harder Programs Looked Equivalent When They Were Not

PolyHuman tests functional equivalence across human-written C++, Java and Python.

Published Updated Story ID: mp-2026-08-26-018
Read the complete editionStory JSON

Summary

PolyHuman tests functional equivalence across human-written C++, Java and Python.

Models increasingly mislabeled non-equivalent programs as equivalent as difficulty rose, and the strongest tested model showed Python sensitivity plus run-to-run instability under identical settings. Manual analysis of 81 systematic disagreements found partial reliance on similarity cues rather than reliable semantic judgment.

Why it matters

PolyHuman tests functional equivalence across human-written C++, Java and Python.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. PolyHuman tests functional equivalence across human-written C++, Java and Python.

    Evidence: source-2026-08-26-020

Sources

  1. arXiv preprint 2608.23961arXiv · primary research

Corrections

No corrections have been recorded for this story.