safety
Explaining Whether to Speak Changed the Decision to Speak
Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.
Summary
Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.
The strongest direct policy performed better but exposed no reasoning; a reasoning policy offered an inspectable trace at lower quality, especially on true intervention opportunities. Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.
Why it matters
Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.
Limits and context
- Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.
Key claims
Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.
Qualification: Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.
Evidence: source-2026-08-24-012
Sources
- arXiv preprint 2608.20670arXiv · primary research
Corrections
No corrections have been recorded for this story.