· 7 min read

When small models break tool‑call JSON

A coding agent lives and dies by tool calls. With a frontier model you barely think about them. With a 7B model running on your laptop, the JSON is where things fall apart. This is what we ran into in oxi, what is safe to fix on the client, and why everything else should go back to the model instead of being hidden.

How a tool call gets to your code

In an OpenAI-compatible API, the model doesn't run anything. It returns a message with a tool_calls array: a tool name and an arguments field. That field is a string, and the string is supposed to contain a JSON object that matches the tool's schema:

{
  "type": "function",
  "function": {
    "name": "read",
    "arguments": "{\"path\": \"src/main.rs\"}"
  }
}

With local models there is one more step. The model writes plain text in whatever format it was trained on, and the server (llama-server, Ollama, LM Studio) uses the model's chat template to recognise a tool call in that text and turn it into the structure above. oxi relies on that native tool calling rather than inventing its own text protocol, because the model was trained on that format and follows it best.

The weak spot is the arguments string. Nothing guarantees it is valid JSON. The server extracts what the model wrote, and a small model doesn't always write what it meant.

What broken arguments look like

The failures fall into a few recognisable shapes:

The first three are formatting mistakes: the model knows what it wants to call, and the intent is all there. The last two are real errors, and no amount of cleverness on the client can recover the missing half of a file.

The worst fix: pretend it was empty

The tempting code, and the code oxi had until this week, is one line:

let args: Value = serde_json::from_str(&tc.arguments).unwrap_or(json!({}));

It never crashes, which is why it looks fine. But think about what the model sees next. It asked to read src/main.rs, the arguments didn't parse, the tool ran with {}, and the result that came back was:

missing path

That is a lie about what happened. The model did send a path. A strong model shrugs and tries again. A small one reads it literally: maybe the file doesn't exist, maybe it should list the directory first, maybe it should apologise and stop. It burns rounds solving a problem it doesn't have, and from the outside it looks like the model is dumb, when the client threw away the evidence.

Repair what is unambiguous, report the rest

The rule we settled on: only repair a call when there is exactly one reasonable reading of it. Fences, surrounding prose and double encoding all qualify, because the intended object is sitting right there. Everything else is sent back to the model as a tool error that says what went wrong and quotes what it sent, and the tool doesn't run at all.

This is the whole function in oxi (Rust, serde_json):

pub(crate) fn parse_tool_args(name: &str, raw: &str) -> Result<Value, String> {
    let trimmed = raw.trim();
    if trimmed.is_empty() {
        return Ok(Value::Object(Default::default()));
    }
    let first_err = match serde_json::from_str::<Value>(trimmed) {
        Ok(v @ Value::Object(_)) => return Ok(v),
        // Double-encoded: a JSON string that itself holds the object.
        Ok(Value::String(inner)) => match serde_json::from_str::<Value>(inner.trim()) {
            Ok(v @ Value::Object(_)) => return Ok(v),
            _ => "expected a JSON object, got a string".to_string(),
        },
        Ok(_) => "expected a JSON object".to_string(),
        Err(e) => e.to_string(),
    };
    // Fences or prose: try the span from the first `{` to the last `}`.
    if let (Some(start), Some(end)) = (trimmed.find('{'), trimmed.rfind('}'))
        && start < end
        && let Ok(v @ Value::Object(_)) = serde_json::from_str::<Value>(&trimmed[start..=end])
    {
        return Ok(v);
    }
    const MAX_ECHO: usize = 500;
    let echo: String = trimmed.chars().take(MAX_ECHO).collect();
    let ellipsis = if trimmed.chars().count() > MAX_ECHO { "…" } else { "" };
    Err(format!(
        "The arguments for `{name}` were not valid JSON ({first_err}). Received: {echo}{ellipsis}\n\
         Call `{name}` again with a single JSON object that matches its parameters."
    ))
}

A few details matter more than they look. The repair only accepts a result that is a JSON object, since that's the only thing a tool's arguments can be. The error quotes the raw input, capped at 500 characters, so a truncated 20 KB write doesn't flood the context. And it names the tool, because the model may have made several calls in the same turn.

Here is what the model now gets back for a truncated call:

The arguments for `read` were not valid JSON (EOF while parsing a string at line 1 column 21).
Received: {"path": "src/main.rs
Call `read` again with a single JSON object that matches its parameters.

That message is true, specific and actionable. The model can see its own broken output and exactly where parsing stopped, which is everything it needs to fix the call on the next try. It's the same reason compilers quote the offending line.

Where the check runs matters

Parsing happens before anything else touches the call. A call with bad arguments:

oxi talks to three wire formats (OpenAI-style chat completions, Anthropic messages and the Codex responses API), and all three loops go through the same function, even though it's the small local models that need it most.

Why not constrain the output instead?

The stronger fix is constrained decoding: llama.cpp can apply a grammar or a JSON schema at the sampler, so the model physically cannot emit invalid JSON. It's a good tool, and worth using when you control the server. We didn't start there for two reasons. oxi also runs against Ollama, LM Studio, remote boxes over SSH and hosted APIs, so the client needs a sensible answer for malformed output regardless of the backend. And constraints fix the syntax, not the intent: a model forced into valid JSON can still send the wrong field names, and that also has to come back as a clear error.

The two approaches stack well: constrain where you can, and report clearly everywhere.

If you're building on small models

The change shipped in oxi v1.11.2. It came from a question on X about small models drifting off the tool-call JSON, which is the kind of feedback we're always happy to get.

This post was drafted with help from Claude and reviewed by the oxi maintainer.