infrastructure
Attention Clustered Once, Then Ran Sparse
ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.
Summary
ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.
The authors report two-to-six-times speedups on a tabular foundation model while retaining at least 99 percent of dense accuracy, plus a 1.8-times video-generation speedup with outputs closer to dense attention than a tested specialist method.
Why it matters
ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.
Limits and context
No additional limitation was separately recorded.
Key claims
ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.
Evidence: source-2026-08-29-017
Sources
- arXiv preprint 2608.26965arXiv · primary research
Corrections
No corrections have been recorded for this story.