senn-techsenn-tech
Infrastructure
Infrastructure2026-10-09· By Franz Senn

GPU inference on Proxmox: what LXC instead of a VM actually buys you

In early October Bart Dworzanczyk published the piece "Why Should You Run Your GPU Inference Box on Proxmox LXC Instead of a VM?" on LinkedIn. His machine: one RTX 5090, an EPYC 7352, Proxmox 9.1, LVM-thin underneath, eight containers on top, seven of them allowed to open the card, all of them privileged, with Docker and vLLM inside (corrected 9 Oct, the first version said six). His argument in one line: passthrough into a VM wastes a card, LXC shares it, and snapshots cost seconds rather than hours. His numbers come from a single machine, so they are self-reported. He put the raw outputs and the scripts in a public GitHub repository, which keeps them checkable (corrected 9 Oct, the first version said most of his evidence sat behind the login). We ran the check anyway, against the Proxmox documentation, against NVIDIA documents, and against our own three GPU hosts.

The short version: the technology is real and the documentation supports it. We still are not rebuilding, for measured reasons. Every number below is our own measurement from 9 October 2026, and every guest and host was read-only while we took it.

What the Proxmox documentation does and does not say

His strongest exhibit is a reference to the documentation, and it holds. In the chapter on virtual machines: "if you pass through a device to a virtual machine, you cannot use that device anymore on the host or in any other VM." Elsewhere, in the wiki article on the same topic: "Note that VMs with passed-through devices cannot be migrated." That second sentence decides our architecture. Our GPU guests are tied to one node each, live migration is not available to them, and moving to LXC would not lower that price, only redistribute it.

A second exhibit gets thin. He writes that the Proxmox documentation states that memory ballooning does not work with PCIe passthrough. That statement is not in the documentation. What the QEMU chapter says about ballooning is this: "When the host is running low on RAM, the VM will then release some memory back to the host, swapping running processes if needed and starting the oom killer in last resort." And on fixed allocation: "When setting memory and minimum memory to the same amount Proxmox VE will simply allocate what you specify to your VM." So the cost logic behind his claim is sound, but it is not quoted correctly. That matters to us because our AI guest carries 24,576 MiB and never sets the balloon field at all. Per the documentation the balloon device gets added anyway, and that device can take memory back from the guest under host pressure.

Two more items from his list we pinned down in the documentation as well. Bind mounts are indeed missing from one kind of backup: "The contents of bind mount points are not backed up when using vzdump." The same goes for device mount points. For replication the rule is the opposite, there additional mount points are replicated by default whenever the root disk is replicated. Anyone who generalises that sentence loses the wrong half in an emergency. And on privileged containers the documentation is harsher than the post: "The LXC team considers this kind of container as unsafe, and they will not consider new container escape exploits to be security issues worthy of a CVE and quick fix. That's why privileged containers should only be used in trusted environments."

The sentence that actually counts

His most honest paragraph appears far down the page: shared access is not shared memory. He demonstrates it on his own container, where vLLM holds 27 GiB of 32 GiB and a start requesting 24 GiB dies. We measured the same condition on our own fallback host on the morning of 9 October, without anyone experimenting.

HostRealityVRAM measured 9 OctoberSnapshotBackup
.180.211, guest 111VM with a passed-through RTX 4060 Ti12,024 of 16,380 MiB, four services, 73.4 %yesyes, PBS
.180.202bare metal, 1x RTX 509030,440 of 32,607 MiB, 2,167 MiB freenono
.180.3bare metal, 4x RTX 509031,442 to 31,524 MiB per card, tensor-parallel jobnono

The row with 2,167 MiB free shows his argument from its worst side. That host carries our standby model at 27,534 MiB and our embedding fallback at 2,906 MiB. A third model that wants 24 GiB cannot come up there. That is a budget question, and no hypervisor decision answers it.

What the post does show us

The two bare-metal hosts have neither snapshot nor backup. No vzdump job knows them, because they are not guests. For the cards themselves that is survivable, the model weights live on disk anyway. For everything around them his sentence about bind mounts applies in full: configuration, Compose files and environment variables exist exactly once. A total loss means rebuilding the lane from scratch, pulling the driver from a DKMS tree, re-pulling images, and reassembling the serving profiles we had tuned.

Block snapshots are not retrofittable there either without becoming reckless. On .180.3, vgs reports VFree = 0 for the volume group. The single linear logical volume consumes the entire 3.49 TB group. The 1.7 TB that df reports as free is free inside the filesystem. A thin pool needs free extents, and we would only get them by shrinking a 3.49 TB filesystem while it is mounted. On .180.202 there is no LVM at all, there XFS sits directly on a partition.

Then there is one finding that appears in neither camp of this debate. The NVIDIA driver licence text ships on our hosts inside the libnvidia-compute package, and section 2.8 reads: "You agree that GeForce or Titan SOFTWARE: (i) is licensed for use only on GeForce or Titan hardware products you own, and (ii) is not licensed for datacenter deployment." His setup is a homelab, that clause does not concern him. Our cards sit in our own rack and carry an internal company workload. What exactly counts as datacentre deployment for a rack in Kufstein is a legal question, not a technical one. We have not answered it, and we are not going to answer it in a blog post. Fittingly, NVIDIA CUDA Forward Compatibility restricts that mechanism as well: the document limits it to systems with data centre GPUs, certain NGC server variants and Jetson. On GeForce cards no forward compatibility package helps, which is why our driver maintenance is as rigid as it is.

What we take from it

Nothing gets rebuilt. Our only GPU virtual machine is the box that needs his LXC advantage least: three backup jobs cover the cluster, two of them run with all = 1 across every guest, and its guest name appears in no exclusion list. We can freeze it today as well. On nesting, the Proxmox documentation states that for use cases demanding maximum isolation and the ability to live-migrate, packaging containers inside a Proxmox QEMU virtual machine remains a recommended practice. We are running exactly that variant. On .180.3, with four cards running at 96 % of capacity, splitting one card helps nothing, there is nothing left to split.

We take two things away, and both are small. First: a restic backup of the configuration on both bare-metal hosts, with Compose files and environment variables also going into a git repository. Second: a VRAM budget per lane with a start gate that stays closed when the free space does not cover the request. His own start script built exactly that spot as a warning and launches vLLM anyway after three minutes, which makes it a sign rather than a gate. Our case has to be decided the other way round. A start requesting 24 GiB on a lane with 2,167 MiB free must be rejected, and the reason written to the log so we can see who asked.

Further reading

Questions?
Is PCI passthrough really exclusive?+

Yes, and the Proxmox documentation says so in so many words: once you pass a device through to a virtual machine, you cannot use that device on the host or in any other VM. The post leaves out a second sentence that matters more to us: VMs with passed-through devices cannot be migrated. Our inference guest is therefore pinned to one node.

Why are you not rebuilding on LXC?+

Because the one box that would benefit already has what LXC is supposed to deliver. Our GPU virtual machine is covered by two of our three backup jobs, the ones that run across every guest, and it can be frozen at any time. Our two other GPU hosts are not guests at all, so LXC cannot help them. And the Proxmox documentation still recommends nesting containers inside a QEMU virtual machine when you need isolation and migration.

Can I run several containers on one GeForce card?+

Yes, that part is not a datacentre privilege. But shared access is shared memory: on our single 5090 host we measured 30,440 MiB used out of 32,607 MiB on the morning of 9 October, which leaves 2,167 MiB. A second model that asks for 24 GiB will not start there, no matter which virtualisation model sits underneath it.