Migrating from VMware to Proxmox
Reference client: Logistics and trading group in Tyrol, Austria · ~100 employees · 4 sites
Customer data in copy and images has been neutralised.
Starting point
Virtualization had run on VMware for years, until the new licensing policy drove costs to a level no mid-sized company could justify. Add the lock-in: proprietary storage, proprietary management, no easy way out.
The goal was an exit: without a big bang, without weekend outages, and with high availability that lives up to the name.
The hard part
Converting a virtual machine takes minutes. What costs planning is the order: a service whose database has already moved but whose application still runs on the old side is not half a migration; it is an outage. So the first thing built was a dependency list, and the migration followed that list rather than convenience.
But the real decision was never the hypervisor. That part is replaceable. A project like this is decided at the storage layer, and there were five serious candidates, which we worked through with arithmetic rather than argument.
A central storage server running ZFS, attached over iSCSI, is the most obvious route and was the first to drop out: the attachment has only one head. One box owns every volume, which makes it the single point of failure for both sites. That would delete the very goal the whole design exists for. Nor can two ordinary nodes be declared a highly available pair: that needs dual-controller hardware with a shared backplane and not two servers with internal drives.
Two mirrored storage servers would have worked, and on closer inspection are the same thing that already runs, merely relocated onto extra appliances with an extra network hop in between. That is where the comparison became surprising: two mirrored storage servers and switching the local disk layout from mirroring to parity deliver the same gain of roughly one third. Because it is the same geometry: parity locally, two copies across the sites. One route costs a storage appliance and a new failure domain; the other is a rebuild on disks we already own.
Ceph is the name somebody always raises in these conversations. Run across two sites it needs four copies plus an arbiter, and that makes its storage efficiency identical to today, so the space gain is zero. And latency decides against it: computed against our own database workload it lands 40 to 90 percent slower. Six nodes is also small for Ceph; one node is a sixth of the cluster, and a recovery hits proportionally hard.
Vitastor was the most honest case, because almost everything spoke for it. Measured, it is equally fast. The widespread claim that it is twice as slow does not survive measurement. Space usage is a wash too. And it even has the better arrangement: four independent copies on four hosts survive a site loss and a subsequent node failure, whereas two copies plus an arbiter leave exactly one copy after a site loss, with no margin at all. In the end it dropped out on a single axis: trust in the write path. That is not a verdict on the project but a statement about where we are willing to place a bet and where we are not.
The second site is why the arithmetic comes out this way. The nodes are deliberately interleaved rather than split blockwise, which is what lets the cluster survive the loss of an entire fire compartment. The price is a rule you have to know: two nodes may only reboot together if they stand at the same site. Go by number instead of by location and you remove the ground from under both halves of a pair at once.
Solution
The migration went to a six-node Proxmox cluster spread across two fire compartments. Storage replicates synchronously via DRBD/LINSTOR over dedicated 100-gigabit links. If a node or an entire site fails, the machines carry on from the other side.
The roughly 60 production VMs moved during live operation, service by service, with a rollback path. Today more than 300 replicated storage resources run in the cluster, checked nightly for silent data errors via checksum verification.
The cluster is watched by the in-house SIEM, from the NVMe layer down to the guest file system.
How it is built
Six nodes, two fire compartments, interleaved. Storage replicates synchronously between node pairs over a dedicated 100-gigabit network fully separated from production traffic. Replication and application load never compete for the same link. More than 300 replicated storage resources carry the production machines.
On that foundation sits the actual high availability: around 57 virtual machines are managed by the cluster manager. Not a handful flagged as important but effectively the entire production estate. Each carries its own limits for restart and relocation attempts; if a machine repeatedly fails to come up healthy, the cluster stops trying rather than running an endless loop.
Placement follows rules rather than hands. Node affinity rules bind each group of machines strictly to exactly the node pair that physically holds its data. Without that rule the scheduler would be free to place a machine on a node where its disk is only reachable over the network. That works, but it costs every read a network hop. The rule turns a placement that functions into one that is correct.
The second kind of rule is what separates real redundancy from nominal redundancy: negative affinity. The two name servers, the two directory services and the two reverse proxies are each declared as a pair that may never run on the same node. Two DNS servers on one machine are not a second DNS server; they are two processes with one shared outage.
For maintenance, a node that is shutting down is configured to relocate its machines rather than stop them. A node empties itself before it goes. Relocation runs over the same separated fabric as replication and is deliberately unencrypted there, because the wire is private, and the encryption saved is exactly the difference between a relocation in minutes and one in a quarter of an hour.
Backups go to a dedicated backup server with deduplication and verified restore points. The cluster feeds the in-house SIEM, from disk wear all the way up to the file system inside the guest.
Day-to-day operation
Nodes are updated one at a time, and because the maintenance policy relocates machines rather than stopping them, the procedure is unspectacular: let the node empty itself, update, bring it back, wait for replication, only then the next one. It takes longer than a joint reboot, and it is the reason updates here are no longer an event.
The cluster continuously assesses its own distribution and reports imbalance as a number. In normal operation it reads zero. That is the figure which tells you whether a relocation after a failure was tidied up again, or whether a machine simply stayed where the incident threw it.
Every night a verify run checks stored data against checksums to catch silent corruption: the failure type no service notices, because nothing crashes and the file still reads.
Restores are practised rather than assumed. A backup nobody has ever pulled anything out of is a hypothesis.
Outcome
Recurring virtualization licence costs are down to zero. The hardware belongs to the company and the stack is open, no more vendor lock-in.
Effectively the entire production estate runs under high availability, and placement is recorded as a rule rather than as knowledge in one person's head. The loss of a node is therefore a procedure the cluster works through itself, and the loss of an entire site one the other side carries.
Node updates have gone from a weekend appointment to routine: the node empties itself, gets updated, and takes its machines back. Availability target 99.9 percent, without anyone having to work at night for it.
What we learned
The number that turned the whole decision was not storage latency but its share. A database operation here takes around two milliseconds; roughly 0.15 of that is storage and the rest is the database engine, the transaction and the network round-trip. Storage that is twice as fast therefore improves nothing measurable, and storage that is ten times slower shows up immediately. Compare layers in isolation and you compare numbers that never reach the user. Measure at application level, or the wrong option wins regularly.
The most dangerous setting in the whole build sits in the defaults and reads: what does a node do when it loses quorum? By default the virtual disk reports an error. The guest then switches its file system to read-only and stays broken even when the network returns twenty seconds later. Suspending is the better answer: the machine waits and carries on as soon as quorum is back. But that is not a free win either: on a permanent loss of quorum the I/O then hangs indefinitely. It is only correct together with an alert on exactly that.
There is a command that brings a diverged replication back in step by declaring one of the two sides invalid. It is on our forbidden list. It does work, which is the problem: you reach for it in exactly the moment you are under pressure, and then discard the good side with equal probability.
Verifying and repairing are two operations, and their order is not a matter of taste. An automatic repair run can overwrite the very finding the verification produced the night before, leaving a cluster that looks clean because the evidence is gone. Repair may only run after somebody has read the finding.
And the sentence that stood over the whole review at the end: every problem found was a setting in the existing solution and not a property of it. That is the most common expensive mistake in infrastructure: replacing a product because it is misconfigured. The replacement carries the old faults along and adds new, unfamiliar ones.
Other projects we run

SIEM & security monitoring with real-time correlation
Company-wide logs in one place, and only the alerts that matter.

NIS2 compliance platform
55 controls, clear owners, automatic reminders: compliance as a living process.
Similar problem?
Tell us what you're planning, a short call clarifies whether it pays off.
senn-tech