safety
The Safety Gate Rejected Every Chance to Learn
A range-based confidence gate admitted none of a common update stream at 2,000 episodes, while paired checks admitted 31.6%.
Summary
A range-based confidence gate admitted none of a common update stream at 2,000 episodes, while paired checks admitted 31.6%.
Independent validation can stop harmful policy updates, but an overconservative test can also freeze a continual agent. In a constructed one-step pushing diagnostic with 32 seeds, fresh paired-binomial checks admitted 31.6% of a common update stream at 2,000 episodes per stage; a range-based confidence gate admitted none. Unconditional replay still learned better in closed-loop runs. The paper proposes auditing both error control and missed learning opportunities, plus controlled historical-reference promotion. Physical-robot and vision-language-action validation remain open.
Why it matters
A range-based confidence gate admitted none of a common update stream at 2,000 episodes, while paired checks admitted 31.6%.
Limits and context
- Physical-robot and vision-language-action validation remain open.
Key claims
A range-based confidence gate admitted none of a common update stream at 2,000 episodes, while paired checks admitted 31.6%.
Qualification: Physical-robot and vision-language-action validation remain open.
Evidence: source-2026-09-13-004
Sources
- arXiv preprint 2609.10873arXiv · primary research
Corrections
No corrections have been recorded for this story.