TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Eighty-Two Hard Tasks Cut Through Fifty-Four Agent Benchmarks

Harbor Adapters ports more than 80 agent benchmarks into one infrastructure, while Harbor-Index distills 29 of them into a smaller audited set.

Published Updated Story ID: mp-2026-09-07-004
Read the complete editionStory JSON

Summary

Harbor Adapters ports more than 80 agent benchmarks into one infrastructure, while Harbor-Index distills 29 of them into a smaller audited set.

The project validates benchmark adapters through code review and parity experiments, then evaluates eight models across 54 benchmarks using a shared agent plus native harnesses. Harbor-Index selects 82 difficult, diverse tasks spanning 29 benchmarks after difficulty filtering and human-and-AI audit. No evaluated model-harness configuration exceeded 30 percent pass rate; the strongest reported result was 28.0 percent. The index lowers evaluation cost, but its conclusions still depend on the selected tasks, adapters and harnesses.

Why it matters

Harbor Adapters ports more than 80 agent benchmarks into one infrastructure, while Harbor-Index distills 29 of them into a smaller audited set.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Harbor Adapters ports more than 80 agent benchmarks into one infrastructure, while Harbor-Index distills 29 of them into a smaller audited set.

    Evidence: source-2026-09-07-004

Sources

  1. arXiv preprint 2609.04298arXiv · primary research

Corrections

No corrections have been recorded for this story.