benchmarks evals
More Agent Searches Did Not Guarantee Better Answers
Four conversational platforms differed in when they searched, what domains surfaced and which results they cited.
Summary
Four conversational platforms differed in when they searched, what domains surfaced and which results they cited.
A combined study of real interactions and controlled API experiments found platform-specific query strategies and search-domain preferences. Responses were largely grounded, but some claims depended on uncited results; search frequency alone did not predict response quality.
Why it matters
Four conversational platforms differed in when they searched, what domains surfaced and which results they cited.
Limits and context
- Responses were largely grounded, but some claims depended on uncited results; search frequency alone did not predict response quality.
Key claims
Four conversational platforms differed in when they searched, what domains surfaced and which results they cited.
Qualification: Responses were largely grounded, but some claims depended on uncited results; search frequency alone did not predict response quality.
Evidence: source-2026-09-20-021
Sources
- arXiv preprint 2609.19244arXiv · primary research
Corrections
No corrections have been recorded for this story.