senn-techsenn-tech
KI-Hardware
KI-Hardware2026-09-29· By Franz Senn

CUDA 13 becomes the floor: vLLM 0.30.0 and llama.cpp 0.5.0 moved their build base

Two engine releases in one week moved their build environment, and both turned the same screw: CUDA 13 becomes the floor you plan against. vLLM v0.30.0 arrived on 22 September, 762 commits from 315 contributors, standard build now against CUDA 13.0. llama.cpp v0.5.0 followed on 23 September, the first major release since v0.4.0, with release builds including the Docker image on CUDA 13.4.1. If you run inference hosts, check the NVIDIA driver before you pull images, not after a container sits in a crash loop.

NVIDIA logo
Both engines now build against CUDA 13. Per host the only open question is the driver level, and that is one line to measure. (Quelle: Simple Icons, file nvidia.svg (CC0))

What vLLM 0.30.0 carries

Two features earn attention because they address operational problems every lane with slow disk starts knows. First, the weight-cache daemon: a persistent daemon per GPU keeps post-quantized, TP-sharded weights in video memory, restarts map them via CUDA IPC with --load-format ipc_cache (#54921) instead of reading from disk, now also for FP4 checkpoints (#55465) and multi-node tensor parallelism (#55468). Second, HiSparse: a host-RAM tier for sparse-MLA decode that pages KV pages into pinned memory under pressure through HiSparseConnector (#53781), with Prometheus counters (#56061). On 32-gigabyte cards that is the difference between aborting and paging out.

The breaking points, all from the release notes: scale-out endpoints only exist opt-in via --enable-scale-out (#54579/#55176), the environment variable is gone. GPTQ loses the activation order g_idx (#54809). The deprecations announced for 0.29 are removed, including VLLM_PREFIX_CACHE_RETENTION_INTERVAL and VLLM_MM_HASHER_ALGORITHM (#55353). gRPC runs through vllm serve --grpc (#56746). And YaRN no longer scales max_model_len (#56446), context lengths you bent together with YaRN bend back on update.

What llama.cpp 0.5.0 carries

CUDA conv2d as implicit GEMM (#29135). Multi-address binding: --host accepts comma-separated addresses and UNIX sockets (#28690). Vulkan MMQ kernels for IQ4_XS and IQ3_S, ROCm AllReduce re-enabled, a tensor-parallel fix against capacity under-provisioning with fused QKV (#29160). For shared setups the most important bugfix: child processes no longer receive the --api-key-file (#29279), the API key no longer sits half-exposed in every child process command line. RPC protocol major v7 and ggml 0.25 sit underneath, that is the line that ends version mixing across hosts.

Measured on ours

Measured on 29 September shortly before eleven: our fallback lane answers on port 8080 with {"version":"0.30.0"}, the host sits on driver 610.57.04, read with nvidia-smi --query-gpu=driver_version. It runs, and the image tag behind it is built on a CUDA-13 base. The second inference host reports a nightly state 0.1.dev20073+g8e685d198, that is not a release and comparable to no version line, so it stays out of this account. Recommendation for your own fleet: measure the driver level per host in one line (nvidia-smi), then pull. With llama.cpp across several hosts run one update round instead of stepwise, since RPC v7 mixing is a functional defect, not a gray zone. And the scale-out endpoints now opt-in: where nobody needs them, the update is a small gain in attack surface.

Further sources

Questions?
Does vLLM 0.30.0 still run under CUDA 12.9?+

As an image variant, yes, the cu129 build ships as an explicit extra tag while the default is CUDA 13.0. Anyone unwilling to touch the driver can use that tag. It is a bridge across one update window, not a perspective, those extra tags age out.

Does NVFP4 help on an RTX 5090?+

Not this week. The new FlashInfer path that replaces Marlin as default applies only to SM100 and SM103 per the release notes. The RTX 5090 is SM120 and stays on the Marlin path. This week's NVFP4 success stories explicitly do not apply to our cards.

Why does the llama.cpp RPC protocol version matter?+

RPC major moved to v7. Mixed hosts then means new and old builds no longer speak on that path. Anyone running llama.cpp servers across several machines has to update them in one round.