TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

Explaining Whether to Speak Changed the Decision to Speak

Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.

Published Updated Story ID: mp-2026-08-24-012
Read the complete editionStory JSON

Summary

Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.

The strongest direct policy performed better but exposed no reasoning; a reasoning policy offered an inspectable trace at lower quality, especially on true intervention opportunities. Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.

Why it matters

Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.

Limits and context

  • Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.

Key claims

  1. Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.

    Qualification: Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.

    Evidence: source-2026-08-24-012

Sources

  1. arXiv preprint 2608.20670arXiv · primary research

Corrections

No corrections have been recorded for this story.