TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

policy regulation

Worst-Case Scoring Moved AI Risk Ratings by Up to 37 Points

An open dashboard lets users trace four EU Code of Practice risk categories back to 19 benchmarks.

Published Updated Story ID: mp-2026-09-24-003
Read the complete editionStory JSON

Summary

An open dashboard lets users trace four EU Code of Practice risk categories back to 19 benchmarks.

The Systemic Risk Index organizes 19 public benchmarks into CBRN, cyber offense, harmful manipulation and loss-of-control categories from the EU GPAI Code of Practice. Across 18 models, changing from average to worst-case aggregation lowered scores by 14 to 37 points, showing how a summary choice can hide weak areas. A blind audit found 83% of sampled benchmark transformations preserved the original harm. The tool is an open evidence interface, not an official EU compliance determination.

Why it matters

An open dashboard lets users trace four EU Code of Practice risk categories back to 19 benchmarks.

Limits and context

  • The tool is an open evidence interface, not an official EU compliance determination.

Key claims

  1. An open dashboard lets users trace four EU Code of Practice risk categories back to 19 benchmarks.

    Qualification: The tool is an open evidence interface, not an official EU compliance determination.

    Evidence: source-2026-09-24-003

Sources

  1. arXiv preprint 2609.28335arXiv · primary research

Corrections

No corrections have been recorded for this story.