Vitastor 3.2, Second Look: Upstream Now Tests Power Loss
We have so far ruled out Vitastor for our production, because we did not trust the write path enough. Upstream has since worked on exactly that. Vitastor 3.2.0 of 31 August ships a simulated power-loss test, and 3.2.1 (13 September) and 3.2.2 (27 September) fixed more bugs. The figures from the release notes come from the maintainer, and we did not measure them ourselves. Where we measured or experienced something ourselves, we say so.
What the new test does
A process killed with kill -9 loses no writes that already sit in the operating system cache. A power failure loses them. The integration tests Vitastor shipped up to 3.1 ran against files, so they could not reach this class of bug. That is how we read the test suite on 22 August. The 3.2.0 test simulates the disk instead: at a simulated outage the volatile write cache loses a random part of its contents, and writes that were in flight land completely, partly (torn at a sector boundary) or not at all. On top of that, io_uring delivers completions in random order and with a random delay. The store itself is the real one: journal, metadata, flusher, compaction and checksums.
The CI runs 200 runs, called seeds, on every build across 24 configurations: old and new store, server and desktop drives (with and without capacitors), two-phase (EC) and instant (replicated) writes, checksums off or on with 4 KB or 16 KB blocks.
83 of 200 runs failed against 3.1.0. To us this is the number that matters: as far as we read the test suite, nothing in that version could find this class of bug before.
What the test found
According to the 3.2.0 release notes, the new store had these problems:
- After a power outage an object became unreadable with a checksum error, and in one case it came back holding the data of another object.
SYNCcould return success without flushing anything. If two clients synced at the same time, the second consumed the first one's counter, and data acknowledged as synced could be lost on a power failure.- A write could be lost while a newer write to the same object survived, leaving the object with a mix of two versions.
- An object could vanish completely after a restart if power failed in the middle of compacting it. The maintainer writes that this is rare on server SSDs but possible in theory.
In the old store the OSD aborted with an internal error on every seed of all six EC configurations. Journal replay could resurrect stale writes of a deleted object, two writes to the same journal sector could be in flight at once, and the maintainer also lists fixes for drives without capacitors.
What came after
3.2.1 made the test more asynchronous and found about six more bugs in the old store and about nine in the new one. It also fixed a snapshot bug: rm with the "inverse" optimisation deleted a layer's data before renaming it, so child images could read wrong data. 3.2.2 brought five fixes. They include replicated objects wrongly turning INCOMPLETE after several failed partial writes, a rejected compaction, and a crash of vitastor-cli rm that 3.2.1 had introduced itself. The 3.2.2 notes say nothing about testing.
Since 29 September the master branch carries another new-store fix that is in no release yet: "fill the checksum slot of zero-length small writes" (commit f3e047fb05). 3.2.2 is therefore not the latest state.
What stays open
Issue #75, an AI-generated data-integrity audit from 9 June, is still open. It has five comments, all from the maintainer, the last on 19 June. Together with #72 to #74 these are the four open issues on the main tracker, all from external AI audits. According to the reporter of #75, vitastor-cli modify --resize to a smaller size deletes all image data, there is a heap overflow in the OSD reachable without authentication, and on the ublk path a guest fsync is acknowledged without reaching the OSD. These are the reporter's claims, and some carry a note there that the maintainer still has to confirm them. We found no confirmation from the maintainer. Individual points have been fixed since, and we do not find the others in the release notes up to 3.2.2. We compared keywords only and did not compare code. Proxmox uses the QEMU driver and does not use ublk, so for us the fsync point is indirect.
The test also runs in a simulation. It shows what the code does under the faults the simulation knows. Firmware that acknowledges writes it did not make is not among them.
More points from our notes and the trackers:
- The scrub that finds bad copies is off by default (
auto_scrubisfalse,scrub_intervalis 30 days). - On GitHub, issues #138 (is TRIM passed through the QEMU driver?) and #143 (RDMA-CM much slower than RDMA) are still open.
- 3.2.1 added the
rxbounceoption for Windows guests with TCP and checksums. The maintainer names no cause. We suspect Windows modifies buffers while they are being sent, and that is a guess. - Vitastor has no authentication. The documentation requires keeping guests off the OSD and etcd network.
- If copies are split evenly across two sites (2:2), a majority decides nothing.
scrub_find_bestpicks the version with the most matching copies. Whether checksums tip the balance is not in the documentation, and we have not tested it.
How mature the project is
The maintainer has 2915 commits, the runner-up 9 (as of 22 August 2026, GitHub). The 3.0 line saw sixteen patch releases between December 2025 and July 2026. In the notes for 3.0.14 the maintainer writes that more than 50 bugs were fixed, most found with LLM analysis, and that most fixes now come with regression tests. 3.0.16 of 19 July fixed, among other things, CAS writes skipping their built-in fsync, the OSD carrying on after a data fsync error in batched writes, and reads from degraded objects possibly returning zeros. Six days later 3.1.0 put encryption into the write path.
What went wrong in our own test
We built a black-box test that hits Vitastor with faults and then checks the result against a log of acknowledged writes. When we reviewed it on 22 August, it had never run to the end:
- The device paths were off by one: the data disk pointed at the metadata disk, and the journal device did not exist.
- A build from source installs no systemd units and no udev rule. Only the Debian package ships them, so neither an OSD nor a monitor started.
- The monitor was given a flag that does not exist (
--etcd_urlinstead of--etcd_address), so no PGs were assigned. vitastor-clihas noscrubcommand. The second check the test promised never ran.- The verifier did not compare the block index. A block with the right sequence number in the wrong place counted as clean. We reproduced that: the verifier said
clean. The self-test had no case for it and has one now (7 of 7 instead of 6 of 6). - The configured "queue depth" was none, because the writer wrote with blocking calls and never had more than one request open.
The lesson: bash -n checks syntax and says nothing about whether a /dev/ path exists. The test has produced no result so far. The next run against 3.2.2 needs hardware of its own.
How fast it was
In our test of 31 July with 100,000 mixed write and read operations, Vitastor was level with DRBD at about 200 seconds. Of roughly 2 ms per operation, about 0.15 ms is storage replication. Speed was never the reason for our caution.
How we rate it
As a production sign-off we had noted four conditions, and all four have to be met: issue #75 is fully closed, online resharding is back on with regression tests, a second independent audit exists without findings, and two releases in a row bring no new serious data-loss report. In addition we had noted an upstream power-loss test or an independent audit as the trigger for a second look. The trigger is met since 3.2.0. Of the four conditions, none is met as far as we read the release notes.
The test shows the gap was real, and the maintainer has closed it. Whether the bug class is finished, the next releases will tell: 3.2.1 found about 15 more bugs once the test got more asynchrony, 3.2.2 brought further fixes, and the next one is already on master. Our production sign-off stays withheld for that reason.
More copies do not protect against this: a bug in the write path would corrupt all copies the same way. That is our conclusion, not a measurement.
What we do next
Anyone evaluating Vitastor now should use the newest version, switch on auto_scrub, and run a power-loss test of their own on real hardware before customer data goes on it. We do the same: once we have hardware, we run our test against 3.2.2 and also check whether the right copy wins at 2:2 when one is corrupted.
Further reading
- Vitastor 3.2.0 release notes (31 Aug 2026)
- Vitastor 3.2.1 release notes (13 Sep 2026)
- Vitastor 3.2.2 release notes (27 Sep 2026)
- Vitastor 3.0.14 release notes (21 Jun 2026)
- Vitastor 3.0.16 release notes (19 Jul 2026)
- Vitastor 3.1.0 release notes (24 Jul 2026)
- Vitastor issue #75, data-integrity audit
- Vitastor commit f3e047fb05 of 29 Sep 2026
- Vitastor OSD parameters, scrub (documentation)
- Vitastor contributors (GitHub)
- Storage Journey, Part 2: Vitastor, fast, but young
- Storage Journey, Part 3: Arrived at DRBD & LINSTOR
What is new in Vitastor 3.2?+
Since 3.2.0 the CI simulates power loss. The simulated disk's write cache drops a random part of its contents, and writes in flight land completely, partly or not at all. It runs 200 seeds per build across 24 configurations. Against 3.1.0, 83 of those 200 runs failed, according to the maintainer.
Does that make Vitastor production-ready?+
We cannot confirm that. Our trigger for a second look, an upstream power-loss test, is met. Of our four sign-off conditions, none is met as far as we read the release notes: issue #75 is open, we found no independent audit, online resharding is still switched off according to the notes, and 3.2.1 and 3.2.2 fixed further bugs in the write path.
Why is killing a process with kill -9 not enough as a test?+
Writes in the operating system cache survive the death of the process, and a power failure destroys them. A test against files therefore cannot trigger fsync-path bugs, because the process never loses its disk.
senn-tech