benchmarks evals
Lean Checked the First Broken Step
FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.
Summary
FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.
The system decomposes natural-language proofs into local units, extracts typed obligations and uses a semantic alignment gate before accepting Lean validation. With a GPT-5.4 backbone, it reports 81.43 percent exact first-error accuracy on 350 Olympiad problems versus 72.29 percent for direct judging, and 84.5 percent on 200 university-level problems versus 75 percent. The benchmarks are expert-verified but finite and model-specific.
Why it matters
FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.
Limits and context
No additional limitation was separately recorded.
Key claims
FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.
Evidence: source-2026-08-30-015
Sources
- arXiv preprint 2608.26310arXiv · primary research
Corrections
No corrections have been recorded for this story.