research
Five Thousand GPU-Hours Searched the Folding Model
AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.

Summary
AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.
Starting from ESMFold, the system proposed and debugged code changes, evaluated variants and stored both successful and failed interventions. The authors report roughly 80 variants, 170 million language-model tokens and about 5,000 GPU-hours. Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.
Why it matters
AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.
Limits and context
- Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.
Key claims
AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.
Qualification: Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.
Evidence: source-2026-08-30-003
Sources
- arXiv preprint 2608.26747arXiv · primary research
Corrections
No corrections have been recorded for this story.