ollama/middleware
Bruce MacDonald 8edecb5c69
openai: match openai's streaming wire format for chat completions (#17485)
Reworked our /v1/chat/completions streaming to match what api.openai.com actually sends,
chunk-for-chunk, based on captures I took of real OpenAI traffic.

What changed:
 - finish_reason now goes on its own chunk with an empty delta {}, instead of riding on the last content
   chunk. Precedence is length > tool_calls > the response's done reason > stop.
 - role is only sent on the first chunk of a stream, not on every chunk.
 - With stream_options.include_usage, usage goes out on its own chunk with choices: [] after the finish
   chunk.
 - A truncated response keeps finish_reason: "length" even when tool calls were streamed — it used to get
   overwritten with "tool_calls". Fixed in both streaming and non-streaming paths.
 - The metrics-only trailer response (empty message at end of stream) no longer produces a stray
   delta:{"content":""} chunk before the finish chunk. A wholly empty completion still opens with a role
   chunk.
 - Every chunk in a stream shares one timestamp, from the response's CreatedAt.
2026-08-03 15:36:57 -07:00
..
anthropic.go anthropic: fix KV cache reuse degraded by tool call argument reordering 2026-03-27 14:30:16 -07:00
anthropic_test.go anthropic: fix KV cache reuse degraded by tool call argument reordering 2026-03-27 14:30:16 -07:00
openai.go openai: match openai's streaming wire format for chat completions (#17485) 2026-08-03 15:36:57 -07:00
openai_encoding_format_test.go
openai_test.go openai: match openai's streaming wire format for chat completions (#17485) 2026-08-03 15:36:57 -07:00
test_home_test.go anthropic: enable websearch (#14246) 2026-02-13 19:20:46 -08:00