Writing ·

How to separate model failures from serving-stack failures in tool-use evaluations

Record request, serving, model, parsing, and retry outcomes separately before comparing local agents on tool-call fidelity.

By Youssef Hemimy · agent evaluation · tool use · AgentOps

A failed tool-use turn is not automatically a model failure. Record whether the request reached the model, whether the server accepted and rendered the tool schema, whether a generated call was parsed, and whether the harness recovered or exhausted retries. Only then compute a model tool-call rate, with transport and non-response rates beside it. This is the measurement lesson from Tang and Zheng's local tool-use study, not a claim that one serving stack or model is best.

A tool-use score crosses four independently fallible boundaries. Preserve the earliest typed failure; a request rejected before inference is not evidence that the model chose not to call a tool.

In the authors' specific Ollama 0.30.8 setup, tools= requests for selected Phi-3 and Gemma-3 models were rejected with HTTP 400 before inference. Their particular ReAct harness stored exhausted requests as generic assistant-message text, so analysis could count those turns as model non-calls unless it recognized the error string. The paper does not claim all harnesses lose structured error metadata.

Put an attribution boundary in the trajectory

A tool-use evaluation passes through independently fallible layers: the harness sends a request; the server accepts or rejects it; a template renders tool specifications; the model may respond; a parser may extract a call; and the harness may retry. vLLM's guide makes serving configuration explicit, including the auto-tool-choice switch, parser, and chat template. Ollama's guide shows the tools request and tool_calls response shape; it does not promise identical behavior for every model.

Record per attempted turnDistinction preserved
Model and exact server version/configurationThe tested model–server pair, parser, and template
Request acceptance and HTTP/error typePre-inference rejection versus a model response
Tool-schema channelNative tools field versus instructions in prompt text
Raw response and parsed callNo-call prose, valid call, unknown tool, or unparseable text
Retry count and terminal reasonRecovered failure versus exhausted request or no response
Episode ID, turn, completion and timeoutOne looping episode versus many independent tasks

This is an instrumentation pattern derived from the study's per-turn taxonomy and error-recording failure, not a prescribed API schema. Keep transport errors as structured data even if the trajectory includes assistant-readable error text. If the server rejected the request before inference, model tool-call fidelity is undefined for that turn, not zero.

Hold the task fixed while changing one interface variable

Replay the same task and tool schema through named conditions: native tool channel, native channel plus explicit text guidance, and a text-only protocol if your system can parse it. Record server version, checkpoint, prompt, template, parser, decoding settings, and retry policy. Compare outcome composition before a single aggregate score. This follows the paper's experimental design; it is a diagnostic test, not a production recommendation.

Adding a text hint while retaining the native channel substantially changed fidelity for several accepted Qwen models. But a uniform text-only protocol performed worse for Llama-3.2 than native-plus-hint in the authors' setup: 44% versus 82% per seed. These selected conditions are not a portable ranking. In limited cross-stack probes, vLLM's default launch rejected auto-tool requests without relevant flags; SGLang accepted requests but returned calls as text without a parser. That check covered only two models on those servers.

Report denominators that reveal loops

A turn-pooled success rate can reward one long episode for emitting the same valid call repeatedly. In the study's text-tools condition, Qwen-0.5B showed 85% turn-pooled protocol fidelity but 34% mean per episode: one 41-turn looping episode dominated the pooled turns. The primary comparison used only eight task-instance seeds and sampled decoding, so these rates are too thin and noisy for precise model ranking.

85%
turn-pooled protocol fidelity in that condition
34%
mean per-episode fidelity in the same condition

Report request-rejection share, non-response and retry-exhaustion share, valid-call share over produced model turns, per-episode mean and turn-pooled rate, and episode completion, timeout, and loop share. A call that parses is only a protocol success: it does not establish appropriate tool choice, correct arguments, or task completion. The AI agent rollout checklist is the next decision point once measurement is trustworthy.

Do not mistake a format fix for an agent fix

The paper tested JSON-schema-constrained output as a baseline. All nine local models emitted an in-schema call in 8/8 one-step trials, but forcing a call every turn made weaker models loop for hundreds of turns. Separate schema validity, semantic tool choice, and workflow completion. If a run fails, use the earliest typed failure—not the final assistant message—as the starting point for agent failure root-cause analysis. A transport rejection, parser miss, bad tool choice, and failed execution need different remedies.

What this evidence does and does not establish

This v1 preprint was accepted at the REALM workshop at EMNLP 2026. Its strongest contribution is a measurement caveat demonstrated in one instrumented ReAct harness, selected local models, Ollama 0.30.8, and limited cross-stack probes. The authors warn that version-level tool gating can change and reject general claims about model family or scale. Their replication repository is public, but this article does not claim independent reproduction or a validated production fix.

Sources

FAQ

Is a rejected tool request a model failure?

No. If the server rejected the request before inference, the model did not get a chance to call a tool. Record the error separately and leave model tool-call fidelity undefined for that turn.

Does a valid JSON tool call prove the agent succeeded?

No. Schema validity does not establish correct tool choice or arguments, appropriate stopping, or task completion.

Should tool-use evaluations report turn-pooled or per-episode success?

Report both, plus completion, timeout, and loop rates. A long looping episode can dominate a pooled turn rate.

Can this study rank local models or serving stacks?

No. It probes selected models and configurations in one instrumented ReAct harness, with limited cross-stack checks. Use it to design attribution records, not a universal leaderboard.

Building something that has to hold up?

We do this work for teams — agent reliability hardening, custom MCP servers, and full-stack AI systems built to survive production.