More

jmorgan · 2026-02-24T22:33:52 1771972432

I've been using Pi day to day recently for simple, smaller tasks. It's a great harness for use with smaller parameter size models given the system prompt is quite a bit shorter vs Claude or Codex (and it uses a nice small set of tools by default).

hrmtst93837 · 2026-02-25T14:22:32 1772029352

That's interesting; I've found Pi really shines for rapid prototyping. Balancing minimalism and functionality is tricky, but it sounds like they're nailing it with these constraints.

rpastuszak · 2026-02-24T23:58:21 1771977501

Which models do you use and what for? I'm looking for ideas to play with.

jmorgan · 2026-02-25T06:51:39 1772002299

For local models I've been trying it with GLM-4.7-Flash and the new LFM2 24B model. I'm excited to try it with the new Qwen3.5 models that came out today as well.

jmorgan · 2026-01-25T21:05:04 1769375104

That's not good, sorry. I work on Ollama - shoot me an email (jeff@ollama.com) and we can help debug

jmorgan · 2026-01-19T22:28:31 1768861711

It's available (with tool parsing, etc.): https://ollama.com/library/glm-4.7-flash but requires 0.14.3 which is in pre-release (and available on Ollama's GitHub repo)

jmorgan · 2025-12-22T03:59:33 1766375973

The source is available here: https://github.com/ollama/ollama/tree/main/app

ekianjo · 2025-12-22T16:06:26 1766419586

Thanks, I stand corrected.

jmorgan · 2025-11-08T04:13:47 1762575227

The gpt-oss weights on Ollama are native mxfp4 (the same weights provided by OpenAI). No additional quantization is applied, so let me know if you're seeing any strange results with Ollama.

Most gpt-oss GGUF files online have parts of their weights quantized to q8_0, and we've seen folks get some strange results from these models. If you're importing these to Ollama to run, the output quality may decrease.

jmorgan · 2025-09-25T20:18:04 1758831484

We did consider building functionality into Ollama that would go fetch search results and website contents using a headless browser or similar. However we had a lot of worries about result quality and also IP blocking from Ollama creating crawler-like behavior. Having a hosted API felt like a fast path to get results into users' context window, but we are still exploring the local option. Ideally you'd be able to stay fully local if you want to (even when using capabilities like search)

jmorgan · 2025-08-14T18:32:09 1755196329

Amazing work. This model feels really good at one-off tasks like summarization and autocomplete. I really love that you released a quantized aware training version on launch day as well, making it even smaller!

canyon289 · 2025-08-14T18:36:46 1755196606

Thank you Jeffrey, and we're thrilled that you folks at Ollama partner with us and the open model ecosystem.

I personally was so excited to run ollama pull gemma3:270b on my personal laptop just a couple of hours ago to get this model on my devices as well!

blitzar · 2025-08-14T19:09:19 1755198559

> gemma3:270b

I think you mean gemma3:270m - Its Dos Comas not Tres Comas

freedomben · 2025-08-14T19:26:54 1755199614

Maybe it's 270m after Hooli's SOTA compression algorithm gets ahold of it

canyon289 · 2025-08-14T21:37:44 1755207464

Ah yes thank you. Even I still instinctively type B

jmorgan · 2025-08-06T05:42:39 1754458959

It should open ollama.com/connect – sorry about that. Feel free to message me jeff@ollama.com if you keep seeing issues

jmorgan · 2025-08-05T19:25:10 1754421910

Sorry about this. Re-downloading Ollama should fix the error

nodesocket · 2025-08-06T10:17:01 1754475421

Thanks for the reply and speedy patch Jeffery. Seems to be working now, except my 4060ti can’t hang lacking enough vram.

jmorgan · 2025-06-11T03:50:08 1749613808

Working on adding tool calling support to Magistral in Ollama. It requires a tokenizer change and also uses a new tool calling format. Excited to see the results of combining thinking + tool calling!