Skip to content

Fix GPT OSS tool calls on Databricks - #124

Merged
wch merged 4 commits into
mainfrom
fix/databricks-gpt-oss-tool-calls
Sep 29, 2026
Merged

wch merged 4 commits into
mainfrom
fix/databricks-gpt-oss-tool-calls

Conversation

@blairj09

Copy link
Copy Markdown
Collaborator

Fix GPT OSS tool calls on Databricks.

Problem

GPT OSS models on Databricks lose their tool call arguments. The tool then runs with {} and fails schema checks. There are two causes:

  1. On chat completions, Databricks sends delta.content as an array of reasoning parts. The OpenAI spec says it must be a string or null. The AI SDK rejects the whole chunk, and the tool arguments are in that same chunk.
  2. On the AI Gateway, we route GPT OSS to the unified Responses API. That route streams GPT OSS tool calls with arguments: "" in every event. Qwen and Llama stream their arguments correctly.

Changes

  • openai-compat-fetch.ts: new transform 7. It turns array content into its text parts, or null, and drops reasoning parts. The SDK already rejected array content for all providers, so this cannot break a provider that works today.
  • databricks-helpers.ts: gpt-oss* models do not get the mlflow-responses protocol. They use openai-chat, which they also advertise.
  • memory-bank/aiConfig.md: records the behaviour and the date we observed it.

Known limits

  • Reasoning text is still dropped on chat completions.
  • Chat completions can also drop GPT OSS output completely on large requests with many tools. That is a Databricks issue, and we are reporting it to them.

Testing

  • New tests: captured Databricks chunks through the middleware and streamText, and the GPT OSS routing rule. Both fail without the fix.
  • Checked by hand in Positron with databricks-gpt-oss-120b through the AI Gateway.

blairj09 and others added 4 commits September 29, 2026 10:41
Databricks streams reasoning models (e.g. GPT OSS) with delta.content as
an array of reasoning/text parts instead of a string. The AI SDK rejects
the whole chunk, and Databricks sends tool call arguments in the same
chunk as reasoning, so tool calls completed with {} input.

Collapse array content to its text parts (or null) in the shared
OpenAI-compatible SSE transform, dropping reasoning parts.
The Unity AI Gateway's unified Responses API streams GPT OSS tool calls
with empty arguments in every event, response.completed included, so
every tool call reached the client as {} (or not at all). The same
request non-streamed, or streamed over gateway chat completions, carries
the arguments; Qwen and Llama stream them fine on Responses.

Withhold the mlflow-responses stamp from gpt-oss identities so they fall
back to openai-chat, which they advertise.
Keep output_text parts as well as text parts when collapsing array
content, and test parts whose text is not a string. Note that gateway
chat completions can also drop GPT OSS output on large requests, so it
is the better route, not a clean one. Add Databricks to the fetch
wrapper's provider list.
@wch
wch merged commit 25bd7ca into main Sep 29, 2026
4 checks passed
@wch
wch deleted the fix/databricks-gpt-oss-tool-calls branch September 29, 2026 19:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants