TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Lean Checked the First Broken Step

FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

Published Updated Story ID: mp-2026-08-30-026
Read the complete editionStory JSON

Summary

FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

The system decomposes natural-language proofs into local units, extracts typed obligations and uses a semantic alignment gate before accepting Lean validation. With a GPT-5.4 backbone, it reports 81.43 percent exact first-error accuracy on 350 Olympiad problems versus 72.29 percent for direct judging, and 84.5 percent on 200 university-level problems versus 75 percent. The benchmarks are expert-verified but finite and model-specific.

Why it matters

FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

    Evidence: source-2026-08-30-015

Sources

  1. arXiv preprint 2608.26310arXiv · primary research

Corrections

No corrections have been recorded for this story.