ollama

mirror of https://github.com/ollama/ollama.git synced 2026-05-13 14:27:00 +00:00

Author	SHA1	Message	Date
Eva H	bad32c7244	launch/docs: fix title for pool (#15883 )	2026-04-29 17:18:44 -04:00
Eva H	ab2e005bf7	app: align the app launch page with ollama launch (#15753 )	2026-04-29 14:45:19 -04:00
Parth Sareen	321cc8a2ba	server/launch: add model recommendations cache endpoint (#15868 )	2026-04-28 17:09:04 -07:00
Daniel Hiltgen	87288ced4f	New models (#15861 ) * mlx: add laguna model support * convert: support fp8 safetensors import Decode HF F8_E4M3 safetensors with block scale companions into GGUF-supported tensor types, and record which output tensors came from FP8 source weights. Use that source-precision metadata during create quantization: default FP8-sourced GGUFs to Q8_0, keep non-FP8 tensors at their original precision for Q8_0, and promote non-FP8 quantizable tensors to Q8_0 for Q4_K requests. * ggml: add laguna model support * server: preserve generate logprobs with builtin parsers Generate requests were dropping logprob-only chunks whenever a builtin parser buffered visible content. Chat already handled this case, but generate only forwarded chunks with visible response, thinking, or tool-call output. Keep generate chunks that carry logprobs even when the builtin parser has not flushed visible content yet, and add a regression test that exercises the behavior with a generic thinking parser. * review comments - perf improvements * ggml: implement nemotron 3 nano omni * add poolside integration * update poolside doc * adapt to new cache setup * fix test * fix test --------- Co-authored-by: Eva Ho <hoyyeva@gmail.com>	2026-04-28 11:50:12 -07:00
Jesse Gross	2bbe2405fe	mlxrunner: decouple models from attention cache storage layout Models build their own attention masks and read K/V directly from the cache's buffers, which ties them to the cache's storage layout. That blocks multi-sequence batching — right-padded rows need a query-padding mask composed onto every model — and rules out variants like paged attention where K/V isn't one contiguous tensor. Caches now hand back a per-layer KVHistory holding post-update K, V, and a MaskApplier that merges the cache's storage restrictions into the model's logical mask. Models describe their mask in logical terms; SDPA composes model, padding, and applier contributions and dispatches to the kernel's causal or no-mask fast path when it can. KVHistory still exposes K, V, and the composed mask for manual attention paths (e.g. CUDA prefill at head_dim > 128). Performance for single-sequence inference is unchanged.	2026-04-27 20:04:46 -07:00
Jesse Gross	bd21678b16	mlxrunner: apply RoPE at per-row positions Switch RoPE from the scalar-offset kernel (mlx_fast_rope) to the array-offset one (mlx_fast_rope_dynamic) so each batch row can start at its own position. The pipeline tracks the current position locally and passes it to the model through Batch.SeqOffsets; each model materializes that slice into an int32 array for the RoPE call. Single-sequence behavior is unchanged; this is the wiring needed before the runner can batch independent sequences.	2026-04-27 20:04:46 -07:00
Jesse Gross	088dfd89a8	mlxrunner: wrap model forward inputs in a Batch struct Gives a single extension point for per-call context (positions, sequence IDs, masks) as multi-sequence batching grows, without having to churn every model's Forward signature again.	2026-04-27 20:04:46 -07:00
Eva H	3cab8a7b02	app/server: fix desktop app startup killing active `ollama launch` sessions (#15657 )	2026-04-27 22:52:53 -04:00
Daniel Hiltgen	03aee88186	mlx: Support NVIDIA TensorRT Model Optimizer import (#15566 ) * mlx: Support NVIDIA TensorRT Model Optimizer import * x/create: support FP8 safetensors import Decode HF F8_E4M3 safetensors with block scale companions into MLX-importable tensor blobs, including compressed-tensors weight_scale metadata, packed NVFP4 layouts, and mixed-precision tensor headers. Use that source-precision metadata during create quantization: default FP8-sourced imports to mxfp8, allow source FP8 to target MLX low-bit formats, preserve source-quantized NVFP4 layouts, selectively keep or promote tensors based on their source precision, and detect quantized dtype from mixed-precision safetensors manifests. * review comments	2026-04-27 18:28:10 -07:00
Daniel Hiltgen	ec9b4e9e47	tokenizer: fix multi-regex BPE offset handling (#15844 ) Use the current fragment offset when emitting unmatched spans during multi-regex BPE splitting. This avoids duplicating earlier prompt text and inflating token counts for multi-stage BPE tokenizers.	2026-04-27 14:14:27 -07:00
Jesse Gross	4656a07e56	mlxrunner: batch the sampler across multiple sequences Register sequences with Add/Remove; each Sample call takes any subset of registered slots and samples one token per row, appending to each slot's ring-buffer history. When all slots share Options and penalty rings are full, one fused transform pass runs over the whole batch via a persistent pooled history tensor; otherwise calls fall back to per-slot serial processing indexed against the same pool. Performance is unchanged for a single sequence, which is all that is exposed for now.	2026-04-25 09:53:53 -07:00
Jesse Gross	30f86cb9dd	mlxrunner: track sampler history in a fixed-size ring buffer AppendToken used to concatenate the new token onto the history tensor and slice it back to RepeatLastN every decode step, churning the graph shape and reallocating a fresh tensor each call. The stateful penalties don't care about order within the window, so a fixed-capacity ring with one SliceUpdate per append keeps the tensor shape constant across steps.	2026-04-25 09:53:53 -07:00
Parth Sareen	ea01af6f76	openai: map responses reasoning effort to think (#15789 )	2026-04-24 02:49:36 -07:00
Parth Sareen	c2ebb4d57c	api: accept "max" as a think value (#15787 )	2026-04-24 01:49:39 -07:00
Parth Sareen	590109c835	launch: harden OpenClaw onboarding flow (#15777 )	2026-04-23 16:47:20 -07:00
Eva H	b4442c6d17	launch: resave managed integration config when live config drifts (#15776 )	2026-04-23 19:32:36 -04:00
Eva H	85ff8e4a21	launch: keep launch recommended models in a fixed canonical order (#15750 )	2026-04-23 16:33:00 -04:00
Parth Sareen	160660e572	launch: use bundled OpenClaw ollama web search (#15757 )	2026-04-22 16:34:19 -07:00
madflow	3b43b9bc4b	docs: update structured outputs doc for cloud (#15733 ) --------- Co-authored-by: Parth Sareen <parth.sareen@ollama.com>	2026-04-22 00:42:39 -07:00
Parth Sareen	21883571b7	launch: replace kimi-k2.5 with k2.6 as top recommended model (#15737 )	2026-04-21 15:13:20 -07:00
Jesse Gross	ce99f24731	mlxrunner: tokenize prompts in request handler goroutines Move tokenization out of the single GPU processing goroutine and into each request's HTTP handler goroutine. This allows the next request's prompt to be tokenized on the CPU while the current request is executing on the GPU.	2026-04-21 14:38:49 -07:00
Jesse Gross	04f5f0cdb4	mlx: improve thread safety of array management Use atomic.Int32 for Array.pinned and a sync.Mutex for the global arrays slice so MLX arrays can be created and pinned from multiple goroutines without racing on those structures. Convert Array value receivers to pointer receivers and struct fields from Array to *Array to avoid copying the atomic. This does not fully achieve thread safety even when building completely independent graphs. The tracing flag and traceScratch slice in compile.go are unprotected, so concurrent Compile calls will race. MLX itself is not fully thread-safe either although it is working to improve.	2026-04-21 14:38:49 -07:00
Matteo Celani	fb36a01ffe	app/ui: fix model picker showing stale model after switching chats (#15280 ) * app/ui: fix model picker showing stale model after switching chats Optimistic messages created during streaming were storing the full Model object instead of the model name string. When switching back to a chat with cached streaming data, the restore effect read an object where it expected a string, causing the model picker to fail matching and remain stuck on the previous chat's model. * app/ui: fix two more instances of Model object passed as model name Fix the same bug at lines 523 and 536 in the assistant_with_tools event handler, where selectedModel (object) was used instead of selectedModel.model (string).	2026-04-21 15:08:06 -04:00
Michael Verrilli	0c65ed33bc	cmd: populate model capabilities in launchInteractiveModel (#15712 ) launchInteractiveModel was introduced in PR #14609 without the client.Show() capability-detection block that RunHandler uses. This left opts.MultiModal always false in the TUI path, causing image/audio file paths to always be treated as unknown commands instead of being loaded as multimodal attachments. Mirror the Show() call, pull-on-404 fallback, cloud auth handling, and MultiModal/Think population from RunHandler into launchInteractiveModel. Fixes #15711	2026-04-21 14:37:36 -04:00
Jesse Gross	22d6c817f8	mlxrunner: fuse top-P and top-K into a single sort pass When both filters are active, avoid paying for a full sort in top-P and a partial sort in top-K. Single-filter paths are unchanged. Improves generation throughput on gemma4:e4b by 1.5%.	2026-04-20 17:43:00 -07:00
Jesse Gross	ca01373b28	mlxrunner: use MaxAxis in the min-P sampler One reduction op instead of Argmax + TakeAlongAxis.	2026-04-20 17:43:00 -07:00
Jesse Gross	24e038d56a	mlxrunner: add logprobs support Match the ollamarunner and OpenAI semantics: raw, full-vocab log-softmax with the top-K ranked by probability. Skipped on the GPU when the request doesn't ask for logprobs so decode doesn't pay for it otherwise.	2026-04-20 17:43:00 -07:00
Parth Sareen	5d1021603a	server: apply format when think=false for gemma4 (#15678 )	2026-04-20 17:42:29 -07:00
Parth Sareen	8e05d734b9	launch: add kimi cli integration with installer flow (#15723 )	2026-04-20 15:33:32 -07:00
Jesse Gross	05e0f21bec	mlx: fuse sigmoid router head in glm4_moe_lite DeepSeek-V2-style aux-loss-free routing computes sigmoid(gates) once but needs it twice: the raw sigmoid output is gathered after top-k, while the post-bias negation is the argpartition key. Fuse into a single multi-output Compiled kernel returning both, saving two launches on the routing path per token. Exposed as a general SigmoidRouter since the same pattern is shared across DeepSeek-V2 descendants. Improves glm4.7 generation performance by approximately 1%.	2026-04-20 15:02:14 -07:00
Daniel Hiltgen	ff23dd343f	mlx: apply repeat penalties in sampler (#15631 )	2026-04-18 07:49:38 -07:00
Parth Sareen	123b300af6	docs: update hermes (#15655 )	2026-04-17 14:20:59 -07:00
Parth Sareen	57653b8e42	cmd/launch: show WSL guidance on Windows instead of handing off (#15637 )	2026-04-16 17:18:04 -07:00
Parth Sareen	a50ce61c54	launch: skip unchanged managed-single rewrite (#15633 )	2026-04-16 16:20:42 -07:00
Daniel Hiltgen	2bb7ea00d2	create: avoid gc race with create (#15628 ) If you have a long running create, and start another ollama server with the same model dir, the GC algorithm deletes the pending blobs and breaks the create. This adds a 1h grace period to avoid deleting in-flight creation operations.	2026-04-16 13:29:16 -07:00
Daniel Hiltgen	55fa80d07a	mlx: additional gemma4 cache fixes (#15607 ) Harden additional corner cases	2026-04-16 13:07:19 -07:00
Daniel Hiltgen	b9cb535407	mlx: fix gemma4 cache to use logical view (#15617 )	2026-04-16 11:54:30 -07:00
Daniel Hiltgen	031baef094	mlx: fix imagegen lookup (#15588 ) * mlx: fix imagegen lookup Fixes #15533 - imagegen had fallen out of sync with the new layout for multiple mlx libraries on Metal. * review comments	2026-04-16 10:39:00 -07:00
Mike Wallio	7d271e6dc9	cmd/launch: add Copilot CLI integration (#15583 ) --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2026-04-15 17:22:53 -07:00
Devon Rifkin	c88dae2d6b	Merge pull request #15612 from ollama/drifkin/gemma4-split-templates gemma4: render differently based on model size	2026-04-15 17:15:35 -07:00
Devon Rifkin	9e3618d663	make empty block conditional	2026-04-15 15:35:25 -07:00
Daniel Hiltgen	5d920cc6bc	Keep Gemma4 router projection in source precision (#15613 )	2026-04-15 15:04:23 -07:00
Devon Rifkin	e585ecd11f	gemma4: render differently based on model size Following up on #15560, this change now has e2b/e4b render differently from 26b/31b. For backwards compatibility, we take the existing renderer name `gemma4` and make it do dynamic resolution based on the model name/size, but the intended use is for the models to be republished with the renderer variant specified explicitly: `gemma4-small` or `gemma4-large`.	2026-04-15 14:37:16 -07:00
Eva H	cdddea0592	launch: always list cloud recommendations first (#15593 )	2026-04-15 13:17:35 -07:00
Parth Sareen	43f90def04	launch: add hermes (#15569 )	2026-04-15 12:00:23 -07:00
Daniel Hiltgen	06ae6367bd	mlx: fix RotatingKVCache.concat() dropping context on mid-rotation (#15591 ) After the rotating buffer has wrapped (c.offset > c.maxSize) a subsequent L>1 Update() went through a slice-to-[0, c.idx) path that discarded all slots in [c.idx, Dim), losing the older-but-still-in-window tokens the first Q of the new batch needs for its sliding-window attention. Linearize the circular buffer to logical order in that wrapped case so the existing trim + concat preserves the last (maxSize - 1) old tokens. When the buffer has not yet wrapped (c.offset <= c.maxSize), slots [c.idx, Dim) are grow padding or stale post-rewind data, so keep dropping them.	2026-04-14 18:29:06 -07:00
Daniel Hiltgen	48ad7085c4	mlx: Improve gemma4 performance with fused operations (#15587 ) * mlx: Improve gemma4 performance with fused operations * review comments	2026-04-14 18:04:04 -07:00
Jesse Gross	e1e3cec8d0	models: fuse MLP activation functions via mlx_compile Converts SiLU/GELUApprox to compiled kernels and adds SwiGLU, matching upstream mlx/mlx_lm's activations pattern. Routes llama, qwen3, qwen3_5 (dense + MoE), and glm4_moe_lite MLP paths through mlx.SwiGLU so each MLP invocation runs as one fused Metal/CUDA kernel rather than a chain of per-op launches.	2026-04-14 16:38:32 -07:00
Jesse Gross	d3e67e305c	mlx: add compiled closure support Wraps MLX's mlx_compile API so Go functions can be traced into fused kernels. Contiguous elementwise chains collapse into a single Metal/CUDA kernel instead of launching one per op. Exposes Compile plus arity helpers (Compile1/2/3) that mirror Python's @mx.compile decorator shape, lazily building the closure on first call so package-level declarations work before the MLX dylib loads.	2026-04-14 16:38:32 -07:00
Eva H	698e04a14b	launch: OpenCode inline config (#15586 )	2026-04-14 15:08:42 -07:00

1 2 3 4 5 ...

5356 commits