developer tools
Access Rules Worked Better When the Agent Had Tests
A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.
Summary
A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.
Researchers translated natural-language access rules into Rego policies for Open Policy Agent using a pipeline that detects policy components, validates a schema, lints, compiles and generates positive and negative tests. On 372 annotated access-control statements, 50.3% of outputs passed the full correctness gate, versus 15.3% from a direct single-prompt baseline. The improvement is large, but roughly half still failed even under the structured approach. These are benchmark results, not a claim of production-ready access control.
Why it matters
A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.
Limits and context
- These are benchmark results, not a claim of production-ready access control.
Key claims
A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.
Qualification: These are benchmark results, not a claim of production-ready access control.
Evidence: source-2026-09-22-015
Sources
- arXiv preprint 2609.24036arXiv · primary research
Corrections
No corrections have been recorded for this story.