benchmarks evals
Synthetic Curves Taught the Model How Trends Combine
TimeThink trained only on generated time-series primitives and outperformed strong baselines on synthetic and real-world compositional questions.

Summary
TimeThink trained only on generated time-series primitives and outperformed strong baselines on synthetic and real-world compositional questions.
Time-series language models can answer familiar questions while failing when trends, seasonality and other temporal primitives must be composed in a new way. TimeThink generates atomic and composite question-answer pairs with deterministic ground truth, then uses reinforcement learning with verifiable rewards to train explicit reasoning. The model was trained only on synthetic data yet outperformed strong baselines on both synthetic and real-world benchmarks in the authors’ experiments. The finding isolates a useful training mechanism; it does not establish clinical readiness for the high-stakes applications that motivate the work.
Why it matters
TimeThink trained only on generated time-series primitives and outperformed strong baselines on synthetic and real-world compositional questions.
Limits and context
- The model was trained only on synthetic data yet outperformed strong baselines on both synthetic and real-world benchmarks in the authors’ experiments.
- The finding isolates a useful training mechanism; it does not establish clinical readiness for the high-stakes applications that motivate the work.
Key claims
TimeThink trained only on generated time-series primitives and outperformed strong baselines on synthetic and real-world compositional questions.
Qualification: The model was trained only on synthetic data yet outperformed strong baselines on both synthetic and real-world benchmarks in the authors’ experiments.
Evidence: source-2026-09-15-008
Sources
- arXiv preprint 2609.13457arXiv · primary research
Corrections
No corrections have been recorded for this story.