TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

Access Rules Worked Better When the Agent Had Tests

A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.

Published Updated Story ID: mp-2026-09-22-026
Read the complete editionStory JSON

Summary

A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.

Researchers translated natural-language access rules into Rego policies for Open Policy Agent using a pipeline that detects policy components, validates a schema, lints, compiles and generates positive and negative tests. On 372 annotated access-control statements, 50.3% of outputs passed the full correctness gate, versus 15.3% from a direct single-prompt baseline. The improvement is large, but roughly half still failed even under the structured approach. These are benchmark results, not a claim of production-ready access control.

Why it matters

A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.

Limits and context

  • These are benchmark results, not a claim of production-ready access control.

Key claims

  1. A structured Rego pipeline reached 50.3% strict correctness versus 15.3% for a direct prompt.

    Qualification: These are benchmark results, not a claim of production-ready access control.

    Evidence: source-2026-09-22-015

Sources

  1. arXiv preprint 2609.24036arXiv · primary research

Corrections

No corrections have been recorded for this story.