infrastructure
The Offline Test Predicted Who Would Get the Impressions
A counterfactual evaluation task estimates how a candidate ranker would redistribute traffic before it reaches an online A/B test.

Summary
A counterfactual evaluation task estimates how a candidate ranker would redistribute traffic before it reaches an online A/B test.
Accuracy metrics can improve while a ranking model shifts impressions among click, video-view or other objective buckets in ways that hurt downstream utility. The proposed task models those shares from observational data using candidate confidence and delivery capacity. A random forest cut L1 error by 49 percent for model families seen in training, but failed against the baseline during the hardest first hour for held-out models; a two-hour rollout architecture recovered a 22 percent gain there. The result exposes both the promise and the cold-start limit of offline traffic-allocation forecasts.
Why it matters
A counterfactual evaluation task estimates how a candidate ranker would redistribute traffic before it reaches an online A/B test.
Limits and context
No additional limitation was separately recorded.
Key claims
A counterfactual evaluation task estimates how a candidate ranker would redistribute traffic before it reaches an online A/B test.
Evidence: source-2026-08-18-008
Sources
- arXiv preprint 2608.16872arXiv · primary research
Corrections
No corrections have been recorded for this story.