developer tools
A Four-Billion-Parameter Model Recovered Nearly All the UI Teacher's Score
Catalog-conditioned fine-tuning reached roughly 98 percent of teacher semantic quality and 97 percent of visual quality at far lower reported cost.

Summary
Catalog-conditioned fine-tuning reached roughly 98 percent of teacher semantic quality and 97 percent of visual quality at far lower reported cost.
The study evaluates declarative interface generation, where a model selects approved components and binds data rather than writing arbitrary frontend code. Across two React and TypeScript domains, the 4B student retained nearly all measured teacher quality at more than an order of magnitude lower cost. Perturbed-catalog and constrained-ground-truth training each improved the quality-cost frontier in different ways. The result is specific to the tested component catalogs, domains, checkpoints and scoring system.
Why it matters
Catalog-conditioned fine-tuning reached roughly 98 percent of teacher semantic quality and 97 percent of visual quality at far lower reported cost.
Limits and context
No additional limitation was separately recorded.
Key claims
Catalog-conditioned fine-tuning reached roughly 98 percent of teacher semantic quality and 97 percent of visual quality at far lower reported cost.
Evidence: source-2026-09-06-008
Sources
- arXiv preprint 2609.04184arXiv · primary research
Corrections
No corrections have been recorded for this story.