safety security
The Security Agent Handed the Network Graph to a Smaller Policy
Sentinel-RL reserves graph topology and constrained actions for specialized models while the LLM writes an analyst-facing account.

Summary
Sentinel-RL reserves graph topology and constrained actions for specialized models while the LLM writes an analyst-facing account.
A graph-attention encoder compresses a live authentication subgraph, a PPO policy chooses from constrained investigation steps, and an LLM may explain recommendations only after critic review and before human approval. In tests using the LANL security dataset and Indiana University's Quartz cluster, the authors report 0.91 precision, 0.87 recall and a median 6.3-second detect-to-approval loop; a 24-million-edge Neo4j load completed in 14.2 minutes. These are system-specific evaluation results, not proof of safe autonomous containment.
Why it matters
Sentinel-RL reserves graph topology and constrained actions for specialized models while the LLM writes an analyst-facing account.
Limits and context
- A graph-attention encoder compresses a live authentication subgraph, a PPO policy chooses from constrained investigation steps, and an LLM may explain recommendations only after critic review and before human approval.
- These are system-specific evaluation results, not proof of safe autonomous containment.
Key claims
Sentinel-RL reserves graph topology and constrained actions for specialized models while the LLM writes an analyst-facing account.
Qualification: A graph-attention encoder compresses a live authentication subgraph, a PPO policy chooses from constrained investigation steps, and an LLM may explain recommendations only after critic review and before human approval.
Evidence: source-2026-09-04-008
Sources
- arXiv preprint 2609.04159arXiv · primary research
Corrections
No corrections have been recorded for this story.