benchmarks evals
The Serving Stack Changed the Tool Score
Harness and serving behavior masqueraded as model failure in local evaluations.
Summary
Harness and serving behavior masqueraded as model failure in local evaluations.
Identical tool-use requests behaved differently across Ollama, llama.cpp, vLLM and SGLang, and missing failure metadata could turn harness rejection into an apparent model non-call. Turn-pooled and per-instance estimates differed by as much as about 55 points.
Why it matters
Harness and serving behavior masqueraded as model failure in local evaluations.
Limits and context
No additional limitation was separately recorded.
Key claims
Harness and serving behavior masqueraded as model failure in local evaluations.
Evidence: source-2026-09-23-018
Sources
- arXiv preprint 2609.26693arXiv · primary research
Corrections
No corrections have been recorded for this story.