TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety security

Red Teaming Learned From Each Failure

CART adaptively chose new tests instead of replaying a fixed prompt set.

Published Updated Story ID: mp-2026-09-24-019
Read the complete editionStory JSON

Summary

CART adaptively chose new tests instead of replaying a fixed prompt set.

Across three evaluation families, the closed-loop challenger found more failures and higher average risk than static replay for every target with a baseline. The study measures discovered weaknesses, not deployment frequency.

Why it matters

CART adaptively chose new tests instead of replaying a fixed prompt set.

Limits and context

  • The study measures discovered weaknesses, not deployment frequency.

Key claims

  1. CART adaptively chose new tests instead of replaying a fixed prompt set.

    Qualification: The study measures discovered weaknesses, not deployment frequency.

    Evidence: source-2026-09-24-021

Sources

  1. arXiv preprint 2609.27336arXiv · primary research

Corrections

No corrections have been recorded for this story.