Local AI is not only getting better, it is also getting smaller
Everything needed to answer whether local language models are still catching up has been on the table since last Monday. llama.cpp shipped its first major release in two years, vLLM version 0.30.0, Xiaomi followed up with MiMo v2.6, an open model at 1.02 trillion total parameters, and PrismML pushed the lower bound of model size a further step down with ternary weights. The four items address different readers. Taken together they still form one picture.
llama.cpp v0.5.0: the operations release
Version 0.5.0 of 23 September is an operations release, not a model release, which makes it more useful to us than any column of numbers. --host now accepts comma-separated addresses and UNIX sockets in a single call, CUDA gets conv2d as implicit GEMM, and the ROCm path has its fp32 accumulation on MFMA GPUs plus AllReduce switched back on. Then the two lines that carry an operational signal: the RPC protocol stands at major v7, and the build palette has moved entirely to CUDA 13.4.
Major v7 yields one rule for everyone distributing llama.cpp over RPC across several machines: no mixed versions. From v0.5.0 an old worker beside a new client is simply a connection error. And 13.4 in the CUDA builds means the host driver has to move before the container update does. Neither of those sentences stands in the changelog in that form; both follow from it.
vLLM v0.30.0: the weight cache and host RAM
762 commits from 315 contributors, dated 22 September. Two features change the architecture question in a self-built setup:
The weight-cache daemon keeps post-quantized, TP-sharded weights permanently in GPU memory. An engine restart remaps them through CUDA IPC (--load-format ipc_cache) instead of loading them from disk, now for FP4 checkpoints too and across several nodes. Anyone who has waited two minutes for the first request to go through knows what that means in an environment with several models: the restart disappears as a cost item. That version is already running here, the feature is not: our production lane on an RTX 5090 has been showing the image vllm/vllm-openai:v0.30.0 since the weekend, measured via docker ps on 24 September 2026. A cache-backed start is therefore, for the first time, no reason to hold back the update.
HiSparse (via HiSparseConnector, #53781) spills KV pages into pinned host RAM under pressure. With that, the server memory becomes an official part of the KV store, Prometheus counters included. Anyone running 32 GB cards knows the problem: the KV cache fills up first, the weights last. Whether the spill path stays clean under load is a measurement question for us rather than a matter for commentary.
The breaking points that every configuration has to walk through before the update: scale-out endpoints are opt-in only via --enable-scale-out, GPTQ without g_idx, the 0.29 deprecations are out, gRPC is now called vllm serve --grpc, and YaRN no longer auto-scales max_model_len.
MiMo v2.6: 1.02 trillion total, 42 active, and where that stops for us
Xiaomi put the MiMo v2.6 line on Hugging Face on 21 September, MIT license, sparse MoE with 1.02 trillion total parameters and 42 active per token. The model card lists 70 layers, 384 routed experts with 8 active per token, 1,048,576 positions of context and hybrid attention with a 128-token window. On Sunday 27 September at 03:58 UTC two more repos landed: MiMo-V2.6-Pro-MOPD and Flash-MOPD, post-trained on the RL checkpoints and aimed, per the card, at tool calls repeating themselves in agent runs.
The number that moved this release from a changelog to a headline comes from outside. Artificial Analysis scores MiMo-V2.6-Pro at 46 on its Intelligence Index, first among the open-weight models in its class, ahead of GLM-5.3 at 45 and Kimi K3 at 44. The best proprietary model on the same list sits at 58. For comparison, that same measurement gave MiMo-V2.5-Pro, released in April 2026, a 26. The jump from 26 to 46 is independently measured. The benchmark figures on the model card are not: DeepSWE 71.9 against Claude Opus 5 at 74.0 is Xiaomi's own number.
At our hardware the enthusiasm stops at a file size. The weights of MiMo-V2.6-Pro-RL occupy 534 gigabytes on Hugging Face, spread over 130 safetensors files. Our inference runs on four RTX 5090 at 32,607 MiB each, 130,428 MiB of video memory in total, three of the four cards filled to about 31.9 GiB by the lane already running there, nvidia-smi on 192.168.180.3 on 27 September 2026. The weights are 4.2 times what fits in our cards, and quantising smaller does not help, since they are already stored quantised. MiMo-V2.6-Flash-RL brings 166 gigabytes, the official ggml-org GGUF builds start at 117 gigabytes for Q2_K, and that does not fit next to the lane already running on the same machine.
What remains is the small member of the family, the 9-billion checkpoint: 9 billion parameters, 5.4 gigabytes as Q4_K_M, a fine-tune on Qwen3.5-9B with 262,144 positions of context. That size is the one on our list, as a comparison candidate against our Qwen lane. The rest of the family we watch through an API for as long as our cards are too small for it. A second candidate from the same week sits even lower: Ternary Bonsai 2 as GGUF, 16 September, Apache-2.0, 3.34 million downloads as of 27 September. Its card works with 1.72 bits per weight and about 5.9 gigabytes for a 27-billion model, but links its own KNOWN_ISSUES.md, and the llama.cpp builds it needs come from a PrismML fork rather than the main branch. The quality figures on that card, 98.2 percent of FP16 intelligence and 47 tokens per second on an M5 Max, are self-reported, and no independent measurement was available on Sunday. Our Bonsai post from this week carries the assessment.
What we take from this
The gap between an API subscription and a self-built stack has rarely been smaller, and these days it sits in operations: RPC version discipline, KV spilling under load, weights in a cache instead of on the disk. The week belongs to whoever maintains llama.cpp and vLLM the way other people maintain their backup software. Our own stack fits the picture: LibreChat, which we run ourselves, added 949 stars in a single week, more than in most weeks. If you host your own models, you end up maintaining the toolchain around them as well. The next test for the new engines is an ordinary weekday with full chat logs, repeated over several weeks.
Further reading
- llama.cpp v0.5.0, 23 September 2026
- vLLM release page, v0.30.0 of 22 September 2026
- MiMo-V2.6-Pro-RL, Flash-RL, the 9-billion checkpoint and the MOPD repos from 27 September, model cards and file sizes, retrieved 27 September 2026
- Artificial Analysis on MiMo-V2.6-Pro for the Intelligence Index and the comparison against the open field, retrieved 27 September 2026
- ggml-org GGUF builds of MiMo-V2.6-Flash-RL for the file sizes
- Ternary-Bonsai-2-27B-gguf, model card including download count, retrieved 27 September 2026
- Video: vLLM against llama.cpp (AllesTested)
- Bonsai 2 and the ternary weights, our post from this week
- Qwen3.8-27B self-hosted
- llama.cpp and GGUF explained
Do we have to update to llama.cpp v0.5.0 or vLLM v0.30.0 right away?+
No. llama.cpp v0.5.0 brings a new RPC protocol (major v7), and mixed versions across the inference hosts break the communication with it. vLLM v0.30.0 has four breaking points, the scale-out endpoints among them. Plan both updates first, then apply them.
Does this week's NVFP4 performance story apply to our RTX 5090?+
No. The new FlashInfer backend replaces Marlin as the default only on SM100 and SM103. The RTX 5090 is SM120, and there the previous path stays in place. Anyone who applies that headline to our cards gets a number for somebody else's hardware.
Are ternary models ready for production yet?+
The architecture is there, the independent measurement is not. Ternary Bonsai 2 has been on Hugging Face since 16 September with 3.3 million downloads, but with its own known-issues file and a llama.cpp fork as a requirement. We test before we judge.
senn-tech