Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 32 additions & 12 deletions memory-bank/aiConfig.md
Original file line number Diff line number Diff line change
Expand Up @@ -581,14 +581,15 @@ Classification rules, in order:
not read — splits change independently of the configuration observed here);
the native route requires every entity to resolve to the same protocol
(mixed/empty/ambiguous prefers unified MLflow Responses when every entity
advertises it, then falls back to `openai-chat` when chat is unanimously
advertised); capabilities aggregate conservatively across
advertises it and is eligible, then falls back to `openai-chat` when chat is
unanimously advertised); capabilities aggregate conservatively across
same-protocol entities (minimum numeric limits, intersection of
media-type/effort-level sets, boolean AND, `vendor`/`family` only when
unanimous); and the whole computation is entity-order invariant.
- **Fallback stamps are always explicit**, never `undefined`: a gateway endpoint
gets `mlflow-responses` when unanimously advertised, otherwise `openai-chat`
wherever chat exists. `undefined` stays reserved for providers that made no
gets `mlflow-responses` when every entity advertises it and can stream tool
arguments on that route (unlike GPT OSS), otherwise `openai-chat` wherever
chat exists. `undefined` stays reserved for providers that made no
routing decision at all.

**Two documented limitations, not fixed here:**
Expand Down Expand Up @@ -633,14 +634,33 @@ table's levels.
completions.** The gateway exposes `/ai-gateway/mlflow/v1/responses` alongside
`/ai-gateway/mlflow/v1/chat/completions`. It is a different API from the
`openai/v1/responses` **native passthrough**, which Databricks refuses for
models it does not proxy natively. Endpoints advertising the
`mlflow/v1/responses` api_type are therefore stamped `mlflow-responses`, which
routes to `{host}/ai-gateway/mlflow/v1` in the SDK's Responses mode. The chat
surface loses on all three counts that matter: it rejects `store`, rejects
`max_completion_tokens`, and streams reasoning as a `delta.content` block array
that the OpenAI chunk schema cannot represent (so reasoning is discarded),
while Responses accepts the same thinking controls and carries reasoning as
first-class items.
models it does not proxy natively. When every served entity advertises
`mlflow/v1/responses` and can stream tool arguments on it, the endpoint is
stamped `mlflow-responses`. This routes to `{host}/ai-gateway/mlflow/v1` in the
SDK's Responses mode. The chat surface loses on all three counts that matter:
it rejects `store` and `max_completion_tokens`, and streams reasoning as a
`delta.content` block array that the OpenAI chunk schema cannot represent,
while Responses accepts those thinking controls and carries reasoning as
first-class items. On the chat
surface the OpenAI-compatible fetch (transform 7) collapses that array to its
text parts: the AI SDK otherwise rejects the whole chunk, and Databricks can
send tool call arguments in the same chunk as reasoning (GPT OSS on classic
serving completed every tool call with `{}`). With that transform in place the
block array no longer breaks tool calls, so the remaining reason to prefer
Responses here is reasoning fidelity: chat still discards the reasoning.

**GPT OSS is the exception: it stays on chat completions on the gateway.** Its
streamed `mlflow/v1/responses` output carries `function_call` items with
`arguments: ""` in every event (including `response.completed`) and no
argument deltas, so every tool call reaches the client as `{}`; the same
request non-streamed carries the arguments, and other hosted families (Qwen,
Llama) stream them fine. The classifier therefore withholds the
`mlflow-responses` stamp from `gpt-oss*` identities, and they fall back to
`openai-chat` (which they advertise). Chat completions is the better of two
imperfect routes rather than a clean one: it streams the arguments, but it has
also been observed to drop GPT OSS output entirely on large, tool-heavy
requests (observed 2026-09, on both surfaces; re-test before relaxing the
rule).

The stamp is **gateway-only**: classic serving has no unified Responses route
(`/serving-endpoints/responses` is native passthrough and refuses
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -379,6 +379,11 @@ describe("inferDatabricksModelProfile — unified MLflow Responses route", () =>
]);
expect(responseProtocol("gateway", [claude])).toBe("anthropic-messages");
});

it("keeps GPT OSS off unified Responses (empty streamed tool arguments)", () => {
const gptOss = foundationEntity("system.ai.gpt-oss-120b", [GATEWAY_CHAT, GATEWAY_RESPONSES]);
expect(responseProtocol("gateway", [gptOss])).toBe("openai-chat");
});
});

describe("inferDatabricksModelProfile — gateway chat availability", () => {
Expand Down
38 changes: 32 additions & 6 deletions packages/ai-config/src/model-capabilities/databricks-helpers.ts
Original file line number Diff line number Diff line change
Expand Up @@ -19,9 +19,10 @@
* The rules are deliberately a **positive identification**: a native protocol is
* stamped only when the endpoint's structure, the model's identity, and (on the
* gateway surface) the advertised `api_types` all agree. A non-native gateway
* endpoint uses unified MLflow Responses when every entity advertises it, then
* falls back to an explicit `openai-chat` stamp when every entity advertises
* chat. `undefined` is never returned as a protocol.
* endpoint uses unified MLflow Responses when every entity advertises it and
* streams tool arguments on it (GPT OSS does not), then falls back to an
* explicit `openai-chat` stamp when every entity advertises chat. `undefined`
* is never returned as a protocol.
*
* Two outcomes exist because unavailability is real: on the gateway surface an
* endpoint whose entities cannot all serve any supported route must not be
Expand Down Expand Up @@ -221,6 +222,21 @@ function isResponsesCompatibleOpenAIId(modelId: string): boolean {
return /^gpt-5/.test(bare) || /^gpt-4o/.test(bare);
}

/**
* Whether the gateway's unified Responses API streams this model's tool calls
* intact. GPT OSS does not: its streamed `function_call` items carry
* `arguments: ""` in every event, `response.completed` included, with no
* `function_call_arguments.delta` events — so every tool call arrives with `{}`
* input. The same request non-streamed carries the arguments, and other hosted
* families (Qwen, Llama) stream them fine on Responses. Chat completions is the
* better of two imperfect routes for GPT OSS: it streams the arguments, though
* it has also been observed to drop output entirely on large, tool-heavy
* requests (observed 2026-09; re-test before relaxing this rule).
*/
function streamsUnifiedResponsesToolArguments(modelId: string): boolean {
return !/^gpt-oss/.test(modelId.toLowerCase());
}

// ---------------------------------------------------------------------------
// Per-entity resolution
// ---------------------------------------------------------------------------
Expand All @@ -240,6 +256,8 @@ interface EntityResolution {
readonly vendor: string | undefined;
readonly apiTypes: readonly string[];
readonly gatewayV2Supported: boolean;
/** Whether the unified MLflow Responses route may serve this entity. */
readonly unifiedResponsesEligible: boolean;
}

/** OpenAI table capabilities with the shared-context input reservation applied. */
Expand Down Expand Up @@ -353,6 +371,7 @@ function resolveEntity(
vendor: recognizedVendor(normalizedIdentity),
apiTypes: foundationModel?.api_types ?? [],
gatewayV2Supported: foundationModel?.ai_gateway_v2_supported === true,
unifiedResponsesEligible: streamsUnifiedResponsesToolArguments(normalizedIdentity),
} as const;

if (!nativeAllowed) {
Expand Down Expand Up @@ -574,8 +593,8 @@ export function inferDatabricksModelProfile(

// --- Unified MLflow Responses, preferred over chat completions ---
// The gateway exposes a unified Responses API alongside chat completions.
// For everything that did not qualify for a native vendor route it is the
// better target: chat completions rejects `store`, rejects
// For eligible models that did not qualify for a native vendor route it is
// the better target: chat completions rejects `store`, rejects
// `max_completion_tokens`, and streams reasoning as a block array that the
// OpenAI chunk schema cannot represent (so reasoning is dropped), while
// Responses carries reasoning as first-class items and accepts the same
Expand All @@ -586,7 +605,14 @@ export function inferDatabricksModelProfile(
// Gateway-only: classic serving has no unified Responses route
// (`/serving-endpoints/responses` is native passthrough and refuses
// non-passthrough models), so serving keeps chat completions.
if (input.surface === "gateway" && everyGatewayEntityAdvertises(GATEWAY_RESPONSES_API_TYPE)) {
//
// Advertising the route is not enough for every family: GPT OSS streams
// empty tool arguments on it (see `streamsUnifiedResponsesToolArguments`).
if (
input.surface === "gateway" &&
everyGatewayEntityAdvertises(GATEWAY_RESPONSES_API_TYPE) &&
resolutions.every((entity) => entity.unifiedResponsesEligible)
) {
return {
excluded: false,
protocol: "mlflow-responses",
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
/*---------------------------------------------------------------------------------------------
* Copyright (C) 2026 Posit Software, PBC. All rights reserved.
*--------------------------------------------------------------------------------------------*/

import { createOpenAI } from "@ai-sdk/openai";
import { jsonSchema, streamText, tool } from "ai";
import { describe, expect, it } from "vitest";

import { createOpenAICompatibleFetchMiddleware } from "../openai-compat-fetch";

/**
* Databricks streams reasoning models (e.g. `databricks-gpt-oss-120b`) with
* `delta.content` as an array of `{type: "reasoning", summary: [...]}` blocks
* instead of a string. The AI SDK rejects such a chunk whole — and Databricks
* sends a tool call's arguments in the same chunk as its reasoning, so the tool
* call completed with `{}` input. The tool-call and reasoning chunks below are
* captured from the wire (ids shortened); the `{type: "text"}` part is modeled
* on Databricks' non-streaming content shape.
*/
function chunk(delta: Record<string, unknown>, finishReason: string | null = null): string {
return JSON.stringify({
id: "chatcmpl_1",
object: "chat.completion.chunk",
created: 1790444153,
model: "gpt-oss-120b-080525",
choices: [{ index: 0, delta, finish_reason: finishReason, logprobs: null }],
});
}

function reasoning(text: string) {
return [{ type: "reasoning", summary: [{ type: "summary_text", text }] }];
}

const toolCallStream = [
chunk({
role: "assistant",
content: "",
tool_calls: [
{
id: "call_1",
index: 0,
type: "function",
function: { name: "get_weather", arguments: "" },
},
],
}),
chunk({
content: reasoning('We need to call the get_weather function with city "Paris".'),
tool_calls: [{ index: 0, function: { arguments: '{\n "city": "Paris"\n}' } }],
}),
chunk({ content: "", tool_calls: [{ index: 0, function: { arguments: "" } }] }, "tool_calls"),
];

const textStream = [
chunk({ role: "assistant", content: "" }),
chunk({ content: reasoning("We need") }),
chunk({ content: reasoning(" to say hello.") }),
chunk({ content: [{ type: "text", text: "Hello" }] }),
chunk({ content: "!" }),
chunk({ content: "" }, "stop"),
];

function run(sseChunks: string[]) {
const wire = async () =>
new Response(sseChunks.map((c) => `data: ${c}\n\n`).join("") + "data: [DONE]\n\n", {
headers: { "content-type": "text/event-stream" },
});
const fetch = createOpenAICompatibleFetchMiddleware("Databricks", "key")(wire);
const provider = createOpenAI({ apiKey: "key", baseURL: "https://example.invalid", fetch });
return streamText({
model: provider.chat("databricks-gpt-oss-120b"),
prompt: "What's the weather in Paris?",
tools: {
get_weather: tool({
inputSchema: jsonSchema<{ city: string }>({
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
}),
}),
},
});
}

async function parts(sseChunks: string[]) {
const result = run(sseChunks);
const collected = [];
for await (const part of result.fullStream) {
collected.push(part);
}
return collected;
}

describe("createOpenAICompatibleFetch — array delta.content", () => {
it("keeps tool arguments that share a chunk with reasoning content", async () => {
const collected = await parts(toolCallStream);

expect(collected.filter((p) => p.type === "error")).toEqual([]);
const toolCalls = collected.flatMap((p) => (p.type === "tool-call" ? [p.input] : []));
expect(toolCalls).toEqual([{ city: "Paris" }]);
});

it("keeps co-chunked tool arguments when content contains a non-object part", async () => {
const malformed = [...toolCallStream];
malformed[1] = chunk({
content: [null, ...reasoning("Calling get_weather")],
tool_calls: [{ index: 0, function: { arguments: '{"city":"Paris"}' } }],
});
const collected = await parts(malformed);

expect(collected.filter((p) => p.type === "error")).toEqual([]);
const toolCalls = collected.flatMap((p) => (p.type === "tool-call" ? [p.input] : []));
expect(toolCalls).toEqual([{ city: "Paris" }]);
});

it("keeps text parts and drops reasoning parts without stream errors", async () => {
const collected = await parts(textStream);

expect(collected.filter((p) => p.type === "error")).toEqual([]);
const text = collected.flatMap((p) => (p.type === "text-delta" ? [p.text] : [])).join("");
expect(text).toBe("Hello!");
});

it("keeps output_text parts and ignores parts without string text", async () => {
const collected = await parts([
chunk({ role: "assistant", content: "" }),
chunk({ content: [{ type: "output_text", text: "Hi" }] }),
chunk({ content: [{ type: "text", text: 42 }] }),
chunk({ content: "" }, "stop"),
]);

expect(collected.filter((p) => p.type === "error")).toEqual([]);
const text = collected.flatMap((p) => (p.type === "text-delta" ? [p.text] : [])).join("");
expect(text).toBe("Hi");
});
});
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@
/**
* Shared OpenAI-compatible fetch wrapper.
*
* Many OpenAI-compatible providers (Snowflake Cortex, MS Foundry, generic
* endpoints) return responses that deviate from the OpenAI Chat Completions
* Many OpenAI-compatible providers (Snowflake Cortex, MS Foundry, Databricks,
* generic endpoints) return responses that deviate from the OpenAI Chat Completions
* spec in small but breaking ways. The AI SDK's Zod schema validation
* rejects these malformed chunks, crashing the stream.
*
Expand Down Expand Up @@ -49,9 +49,19 @@
* 6. Empty tool `type` `""` → `"function"` in tool call chunks
* Spec requires `type` to be `"function"`. Some providers send `""`.
*
* 7. Array `content` → string (or `null`) in delta chunks
* Spec requires `content` to be a string or null. Databricks streams
* reasoning models (e.g. GPT OSS) with an array of content parts —
* `{type: "reasoning", summary: [...]}` and `{type: "text", text}`. The AI
* SDK rejects the whole chunk, and Databricks can send tool call arguments
* in the same chunk as reasoning, so the tool call completes with `{}`.
* Text parts (`text` or `output_text`) are kept; every other part type,
* reasoning included, is dropped (the chat completions path has no
* reasoning channel to carry them).
*
* ## Auth
*
* 7. Strip `Authorization` header when `apiKey === ""`
* 8. Strip `Authorization` header when `apiKey === ""`
* For unauthenticated endpoints (e.g., local servers with no auth).
* Only matches empty string — `undefined` means the caller manages
* auth separately (e.g. Foundry injects its own token).
Expand Down Expand Up @@ -90,11 +100,12 @@ interface MalformedToolCall {

/**
* A delta where `role` may be empty string instead of `"assistant"`,
* and `tool_calls` may contain malformed entries.
* `content` may be an array of parts instead of a string, and `tool_calls`
* may contain malformed entries.
*/
interface MalformedDelta {
role?: "assistant" | ""; // may be "" instead of "assistant"
content?: string | null;
content?: string | null | unknown[]; // may be an array of untrusted content parts
tool_calls?: MalformedToolCall[];
}

Expand Down Expand Up @@ -375,6 +386,16 @@ function fixMalformedChunk(chunk: MalformedChatCompletionChunk, noArgTools: stri
delta.role = "assistant";
}

// Transform 7: Array content → string (or null).
// Spec: content is a string or null.
// Broken: Databricks streams reasoning models with an array of
// `{type: "reasoning"}` / `{type: "text"}` parts.
// Impact: AI SDK Zod validation rejects the chunk, dropping any tool
// call arguments it can carry alongside the reasoning.
if (Array.isArray(delta.content)) {
delta.content = flattenContentParts(delta.content);
}

if (!Array.isArray(delta.tool_calls)) continue;

for (const tc of delta.tool_calls) {
Expand Down Expand Up @@ -403,3 +424,25 @@ function fixMalformedChunk(chunk: MalformedChatCompletionChunk, noArgTools: stri
}
}
}

/**
* Collapse array-valued delta `content` to the string the spec requires:
* the concatenated text of its text parts, or `null` when it has none.
* Every other part type (reasoning included) is dropped deliberately.
*/
function flattenContentParts(parts: unknown[]): string | null {
let text = "";
for (const part of parts) {
if (
typeof part === "object" &&
part !== null &&
"type" in part &&
(part.type === "text" || part.type === "output_text") &&
"text" in part &&
typeof part.text === "string"
) {
text += part.text;
}
}
return text === "" ? null : text;
}
Loading
Loading