developer tools
The Harness Stopped Turning Recovery Into a Model Failure
Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.

Summary
Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.
The open-source harness pairs a persistent IPython environment with standardized execution, recovery, verification and resource accounting. Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.
Why it matters
Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.
Limits and context
- Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.
Key claims
Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.
Qualification: Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.
Evidence: source-2026-08-25-008
Sources
- arXiv preprint 2608.23552arXiv · primary research
Corrections
No corrections have been recorded for this story.