safety security
Thirty-One of Thirty-Five Pipelines Missed the Bar
Only four open document-extraction configurations cleared 0.5 F1 on a high-risk student-application task.
Summary
Only four open document-extraction configurations cleared 0.5 F1 on a high-risk student-application task.
Vision-language models generally led OCR-plus-LLM systems, yet about three quarters of all configurations scored below 0.25. Preserving document structure mattered independently of model size, warning against zero-shot deployment in consequential workflows.
Why it matters
Only four open document-extraction configurations cleared 0.5 F1 on a high-risk student-application task.
Limits and context
No additional limitation was separately recorded.
Key claims
Only four open document-extraction configurations cleared 0.5 F1 on a high-risk student-application task.
Evidence: source-2026-08-20-020
Sources
- arXiv preprint 2608.18289arXiv · primary research
Corrections
No corrections have been recorded for this story.