research
The Genome Model Knew More Than Its Last Layer Could Say
Frozen probes recovered nearly all promoter performance but exposed much larger access gaps for splice sites and some regulatory tasks.
Summary
Frozen probes recovered nearly all promoter performance but exposed much larger access gaps for splice sites and some regulatory tasks.
A unified comparison of five genomic language models found that frozen features recovered 95 to 100 percent of fine-tuned performance on promoter tasks, while average splice-site recovery fell to 60 to 88 percent. Layer probes and sequence perturbations suggested that some local biological signal exists inside the models but is not reliably accessible through final pooled embeddings. The result is a representation-accessibility audit, not evidence that any one model understands genomic function generally.
Why it matters
Frozen probes recovered nearly all promoter performance but exposed much larger access gaps for splice sites and some regulatory tasks.
Limits and context
- Layer probes and sequence perturbations suggested that some local biological signal exists inside the models but is not reliably accessible through final pooled embeddings.
- The result is a representation-accessibility audit, not evidence that any one model understands genomic function generally.
Key claims
Frozen probes recovered nearly all promoter performance but exposed much larger access gaps for splice sites and some regulatory tasks.
Qualification: Layer probes and sequence perturbations suggested that some local biological signal exists inside the models but is not reliably accessible through final pooled embeddings.
Evidence: source-2026-08-09-004
Sources
- arXiv preprint 2608.05329arXiv · primary research
Corrections
No corrections have been recorded for this story.