benchmarks evals
The Patent Judge Improved Drafts—and Disagreed by Metric
Iterative judge feedback helped a low-reasoning drafting agent approach a costlier agent, but attorney agreement varied sharply with the metric.
Summary
Iterative judge feedback helped a low-reasoning drafting agent approach a costlier agent, but attorney agreement varied sharply with the metric.
Vibe Patenting tests an agent that drafts patents and a separately invoked language-model judge that critiques each revision. Judge-guided iteration consistently raised judge-assessed quality, while unguided revision saturated; a lower-reasoning agent approached the score of a more expensive high-reasoning configuration. A professional patent attorney’s review found meaningful agreement with the automated judge, but calibration and agreement changed substantially by metric. The result supports structured critique as an optimization signal while warning that self-scoring is not a substitute for qualified legal judgment.
Why it matters
Iterative judge feedback helped a low-reasoning drafting agent approach a costlier agent, but attorney agreement varied sharply with the metric.
Limits and context
- The result supports structured critique as an optimization signal while warning that self-scoring is not a substitute for qualified legal judgment.
Key claims
Iterative judge feedback helped a low-reasoning drafting agent approach a costlier agent, but attorney agreement varied sharply with the metric.
Qualification: The result supports structured critique as an optimization signal while warning that self-scoring is not a substitute for qualified legal judgment.
Evidence: source-2026-09-15-005
Sources
- arXiv preprint 2609.13422arXiv · primary research
Corrections
No corrections have been recorded for this story.