Writing ·
How to separate model failures from serving-stack failures in tool-use evaluations
Record request, serving, model, parsing, and retry outcomes separately before comparing local agents on tool-call fidelity.
By Youssef Hemimy · agent evaluation · tool use · AgentOps
A failed tool-use turn is not automatically a model failure. Record whether the request reached the model, whether the server accepted and rendered the tool schema, whether a generated call was parsed, and whether the harness recovered or exhausted retries. Only then compute a model tool-call rate, with transport and non-response rates beside it. This is the measurement lesson from Tang and Zheng's local tool-use study, not a claim that one serving stack or model is best.
01
Request
Accepted or rejected before inference?
Record: HTTP status + server error
02
Serving
Were tools rendered in the intended channel?
Record: version + template + parser
03
Model
Did a response contain a tool call?
Record: raw response + call
04
Harness
Was the call parsed, retried, or exhausted?
Record: terminal reason + episode
Rejected before inference → model call fidelity is undefined, not zero.
In the authors' specific Ollama 0.30.8 setup, tools= requests for selected Phi-3 and Gemma-3 models were rejected with HTTP 400 before inference. Their particular ReAct harness stored exhausted requests as generic assistant-message text, so analysis could count those turns as model non-calls unless it recognized the error string. The paper does not claim all harnesses lose structured error metadata.
Put an attribution boundary in the trajectory
A tool-use evaluation passes through independently fallible layers: the harness sends a request; the server accepts or rejects it; a template renders tool specifications; the model may respond; a parser may extract a call; and the harness may retry. vLLM's guide makes serving configuration explicit, including the auto-tool-choice switch, parser, and chat template. Ollama's guide shows the tools request and tool_calls response shape; it does not promise identical behavior for every model.
| Record per attempted turn | Distinction preserved |
|---|---|
| Model and exact server version/configuration | The tested model–server pair, parser, and template |
| Request acceptance and HTTP/error type | Pre-inference rejection versus a model response |
| Tool-schema channel | Native tools field versus instructions in prompt text |
| Raw response and parsed call | No-call prose, valid call, unknown tool, or unparseable text |
| Retry count and terminal reason | Recovered failure versus exhausted request or no response |
| Episode ID, turn, completion and timeout | One looping episode versus many independent tasks |
This is an instrumentation pattern derived from the study's per-turn taxonomy and error-recording failure, not a prescribed API schema. Keep transport errors as structured data even if the trajectory includes assistant-readable error text. If the server rejected the request before inference, model tool-call fidelity is undefined for that turn, not zero.
Hold the task fixed while changing one interface variable
Replay the same task and tool schema through named conditions: native tool channel, native channel plus explicit text guidance, and a text-only protocol if your system can parse it. Record server version, checkpoint, prompt, template, parser, decoding settings, and retry policy. Compare outcome composition before a single aggregate score. This follows the paper's experimental design; it is a diagnostic test, not a production recommendation.
Adding a text hint while retaining the native channel substantially changed fidelity for several accepted Qwen models. But a uniform text-only protocol performed worse for Llama-3.2 than native-plus-hint in the authors' setup: 44% versus 82% per seed. These selected conditions are not a portable ranking. In limited cross-stack probes, vLLM's default launch rejected auto-tool requests without relevant flags; SGLang accepted requests but returned calls as text without a parser. That check covered only two models on those servers.
Report denominators that reveal loops
A turn-pooled success rate can reward one long episode for emitting the same valid call repeatedly. In the study's text-tools condition, Qwen-0.5B showed 85% turn-pooled protocol fidelity but 34% mean per episode: one 41-turn looping episode dominated the pooled turns. The primary comparison used only eight task-instance seeds and sampled decoding, so these rates are too thin and noisy for precise model ranking.
- 85%
- turn-pooled protocol fidelity in that condition
- 34%
- mean per-episode fidelity in the same condition
Report request-rejection share, non-response and retry-exhaustion share, valid-call share over produced model turns, per-episode mean and turn-pooled rate, and episode completion, timeout, and loop share. A call that parses is only a protocol success: it does not establish appropriate tool choice, correct arguments, or task completion. The AI agent rollout checklist is the next decision point once measurement is trustworthy.
Do not mistake a format fix for an agent fix
The paper tested JSON-schema-constrained output as a baseline. All nine local models emitted an in-schema call in 8/8 one-step trials, but forcing a call every turn made weaker models loop for hundreds of turns. Separate schema validity, semantic tool choice, and workflow completion. If a run fails, use the earliest typed failure—not the final assistant message—as the starting point for agent failure root-cause analysis. A transport rejection, parser miss, bad tool choice, and failed execution need different remedies.
What this evidence does and does not establish
This v1 preprint was accepted at the REALM workshop at EMNLP 2026. Its strongest contribution is a measurement caveat demonstrated in one instrumented ReAct harness, selected local models, Ollama 0.30.8, and limited cross-stack probes. The authors warn that version-level tool gating can change and reject general claims about model family or scale. Their replication repository is public, but this article does not claim independent reproduction or a validated production fix.
Sources
FAQ
Is a rejected tool request a model failure?
No. If the server rejected the request before inference, the model did not get a chance to call a tool. Record the error separately and leave model tool-call fidelity undefined for that turn.
Does a valid JSON tool call prove the agent succeeded?
No. Schema validity does not establish correct tool choice or arguments, appropriate stopping, or task completion.
Should tool-use evaluations report turn-pooled or per-episode success?
Report both, plus completion, timeout, and loop rates. A long looping episode can dominate a pooled turn rate.
Can this study rank local models or serving stacks?
No. It probes selected models and configurations in one instrumented ReAct harness, with limited cross-stack checks. Use it to design attribution records, not a universal leaderboard.
Building something that has to hold up?
We do this work for teams — agent reliability hardening, custom MCP servers, and full-stack AI systems built to survive production.