research
Many A/B Tests Shared the Same Reward
A tree-coupled design preserved each policy's standalone trajectory law while reusing matched feedback across comparisons.
Summary
A tree-coupled design preserved each policy's standalone trajectory law while reusing matched feedback across comparisons.
Directly comparing J adaptive policies for T rounds consumes JT reward-bearing interactions. The proposed exact coupling connects policy histories with a predictable tree, shares one reward within matched components and retains each policy's finite-horizon law. Its query count becomes T plus cumulative edge mismatches and can approach T rather than JT when policies converge. Experiments on reward models, language-model evaluation and adaptive search reported a better cost-precision frontier; practical gains depend on how closely the policies' actions can be coupled.
Why it matters
A tree-coupled design preserved each policy's standalone trajectory law while reusing matched feedback across comparisons.
Limits and context
No additional limitation was separately recorded.
Key claims
A tree-coupled design preserved each policy's standalone trajectory law while reusing matched feedback across comparisons.
Evidence: source-2026-08-16-010
Sources
- arXiv preprint 2608.12831arXiv · primary research
Corrections
No corrections have been recorded for this story.