Two vLLM cache advisories and what operators of their own AI stack take from them
In late September 2026 the vLLM team published two security advisories, both updated again on 5 October 2026: GHSA-ph3r-5jfg-f84f carrying the CVE number CVE-2026-105753 and GHSA-935w-9g4m-p28p with CVE-2026-105752. The time depends on which layer of the reporting platform you query, and we name both: the advisory in the repository carries 28 September 2026 as its creation date and 5 October 2026 as its last change, the same entry in the global security database of the same platform stands at 6 October 2026, 00:02 UTC. The repository version is the one that counts for this text, because it carries the content of the report. Both are fixed in current versions, neither one executes code. For operators of their own inference fleets they still describe two places worth knowing, because both sit inside mechanisms that every shared instance switches on by default: the multimodal cache and the shared prefix cache. According to the advisory entry, both findings come from the Patch the Planet programme, a collaboration between Trail of Bits and OpenAI, discovered with the model GPT-5.5-Cyber.
CVE-2026-105753: One rejected request is enough for an outage
vLLM runs the multimodal cache mirrored across two processes by default: the advisory calls the frontend P0, where the metadata lives, and the engine core P1, where the payload data lives. The design assumes that both sides keep the cache in lockstep at every point in time, so that P0 may claim without asking back that a medium is already present in P1.
That assumption breaks inside one specific window. A chat request carrying an image is rendered first and gets registered in P0 during that step, the length check against max_model_len comes afterwards. If the check comes out negative, the request is rejected and the P0 entry stays. P0 now treats the image as cached, P1 never saw it. The next request that sends the same image hits in P0 and therefore receives no payload along, only the reference to the cache. P1, which has nothing, runs straight into the line assert mm_item is not None. The advisory traces this to all versions before 0.28.0 running the standard configuration mm_processor_cache_type="lru"; the alternative paths processor_only, a switched-off cache and shm are not affected according to the advisory.
The severity sits at CVSS 6.5 (AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H). The letter A:H is the point: it is availability alone. In the build the advisory tested, vLLM catches the assertion as an error of that one request, public reports show the same assertion in other builds as a cascade of further engine errors. The reported effect is a service outage that anyone allowed to use the API can repeat. The remedy proposed in the advisory (atomic adoption into both caches after all admission checks, withdrawal of the P0 entry on rejection, and the assertion handled as an error per request) stands in the advisory text as a proposal; the linked pull request #51897 was closed without being merged. Upstream the problem has still become smaller, the files themselves show that, more on that below.
CVE-2026-105752: The salt applies only to the first round
The shared prefix cache was already a reported weakness once: CVE-2025-46570 from May 2025 described the timing oracle through which a tenant can guess whether a prompt is already sitting in the cache. The documented countermeasure is called cache_salt: every tenant sets its own value, and from that follows its own cache namespace.
Exactly this separation dissolves in parts on the responses path POST /v1/responses. This interface uses the Harmony model format (the GPT-OSS family) and drives requests with attached tools as a multi-turn loop: after every tool call vLLM rebuilds the prompt of the next round and submits it into the engine again. The first pass hands over the cache_salt of the caller correctly, the continuation calls the engine entry without a salt. The continuation then sits in the unsalted, global namespace even though the caller had switched salting on.
A second tenant who can reconstruct the post-tool history well enough submits the same continuation unsalted and reads the exact number of cached tokens per turn out of the response usage. That is the membership oracle again, the one cache_salt is supposed to protect against, this time from a different source: the response usage itself delivers the exact numbers, a timing measurement is not needed for that. The preconditions are heavy, which is why the report sits at low (CVSS 3.1, AC:H): a Harmony model, the tool server activated, prefix caching on (standard), the victim sets a salt and triggers at least one tool continuation, and the attacker has to be able to guess the history. It is fixed since 0.30.0: there the path is rewritten so that it reads the salt back out of the engine input and carries it along in the further steps. The patch proposed in the advisory (#51818) was never merged, the implementation went through the restructuring.
The self-test: our lanes measured on 11 October 2026
The first question at any report is the one about your own estate. On 11 October 2026 we queried the four inference endpoints through their own API:
- Primary lane on
192.168.180.3:8000: reports0.1.dev20073+g8e685d198 - Fallback lane on
192.168.180.202:8080: reports0.30.0 - Embedding and rerank lanes on
192.168.180.211:8010and:8011: report0.30.0, the pulled image isvllm/vllm-openai:v0.30.0
Read the way the advisory metadata states it (affected < 0.28.0 and < 0.30.0), the three 0.30.0 endpoints are out for both reports. On the primary lane the version string says nothing: a nightly tree with a dev number and a commit from our own build branch cannot be compared against any release line. So we looked into the running file instead of guessing, inside the container of the production image:
- In
multimodal/cache.pythe fix is in. Our running file is identical to the release tagv0.28.0, both checksums are the same. That means the file contains what 0.28.0 introduced: the classMultiModalCacheMissErrorwith five occurrences, twoinvalidatemethods and one_cache.pop, so the handled error together with the withdrawal of the shadow entry. The assertionExpected a cached itemstill stands three times in the file, and three times at the tagv0.28.0as well. Counting only the assertions misses the fix in both directions. - In
responses/serving.pythe fix is not. Our file has 1563 lines, the tagv0.27.1has 1561, the checksums differ, the structure is the same. The continuation stands on line 722 astokens_input(token_ids)without a salt. The buildv0.30.0is rebuilt at this point and some 180 lines shorter: it reads the salt back out of the engine input withcache_salt = engine_input.get("cache_salt")and passes it again when re-rendering the Harmony messages. Both are missing in our build.
So the in-house build carries one of the two fixes and not the other. The version string 0.1.dev20073+g8e685d198 of the image would have informed us wrongly in both directions, two checksums against the upstream tags say it in seconds. For us that means, for report one: the path is healed, we are out. For report two the finding stays standing without hitting us, because the Harmony path is not reachable on our side, we operate no GPT-OSS, and the gateway uses the chat-completions interface. The production image has not been rebuilt since the measurement and still carries the build date 3 October 2026 unchanged. Catching up would be the clear answer, and it is open.
The third entry of the same wave of reports also fits the picture, and it touches our document processing: GHSA-q43m-vhcp-mhvm (CVE-2026-105750) reports that Docling does not enforce the barrier enable_local_fetch in the HTML browser rendering mode, affected are 2.82.0 up to before 2.118.1 and the same span for the package docling-slim from 2.92.0. Our Docling container reports 2.43.0 and therefore sits outside the affected span, measured on 11 October. A second reason, independent of the version: the advisory explicitly excludes the standard configuration without page rendering, the command line and docling-serve, and that is exactly how our instance runs. The reporting platform lists the report at 5.9, the NVD metadata at 7.5, the difference lies solely in the attack effort.
What operators take from this
The version boundaries for your own instance, sorted by exposure:
- Multimodal requests from users: at least
0.28.0, preferably0.31.0today (released 5 October 2026). The assertion still stands there, the difference is that the drift is handled as an error in that revision and the shadow entry gets withdrawn. - Tenant separation through
cache_salt: at least0.30.0(released 22 September 2026). - Text only, internal only, no tool path: on a pure text stack both reports pass by, a current build still remains the better starting position.
Two points beyond version numbers. The first: cache_salt is a promise per call, not a switch on the server. A gateway that does not pass the value through runs unsalted, and internal resubmissions can lose that same value, as this report shows. Anyone who genuinely wants to separate tenants will not get around their own instance per tenant, or at least their own cache namespace, in the long run. The second: with self-built images the code you read says more than the version string, and in both directions. Our in-house build was built on 3 October 2026 and reports 0.1.dev20073+g8e685d198, so no number you could compute against an advisory span. We checked with two checksums: multimodal/cache.py is identical to the tag v0.28.0 and therefore patched, responses/serving.py matches the structure of v0.27.1 and therefore is not. In substance the tree of our file then stands at two different places, and the version string would have revealed neither of them.
One observation belongs on record at the margin as well: both findings come from a programme that couples machine search for vulnerabilities with a human-written report. Projects that patch their own inference engine will have to treat the advisory list of that engine the way they treat the DSA list of an operating system.
Further reading
- GHSA-ph3r-5jfg-f84f / CVE-2026-105753, published 28 September 2026, updated 5 October 2026
- GHSA-935w-9g4m-p28p / CVE-2026-105752, published 28 September 2026, updated 5 October 2026
- GHSA-4qjh-9fv9-r85r / CVE-2025-46570, the original membership oracle in the prefix cache from 28 May 2025
- vLLM releases: 0.28.0 of 26 August 2026, 0.30.0 of 22 September 2026, 0.31.0 of 5 October 2026
- GHSA-q43m-vhcp-mhvm / CVE-2026-105750 for Docling, affected 2.82.0 up to before 2.118.1
- NVD entry CVE-2026-105752 and NVD entry CVE-2026-105753
Do I have to lift vLLM to 0.31.0 right away?+
Staged by urgency. Anyone accepting multimodal input from users should go to at least 0.28.0, better 0.30.0 or 0.31.0, because there the crash has become a handled error that withdraws the cache entry again. Anyone separating several tenants through cache_salt needs at least 0.30.0, otherwise the separation does not hold inside tool continuations. A text-only instance that is reachable internally and has no tool endpoint is in less of a hurry. Both reports are rated low and moderate in severity, neither one executes code.
Does that make cache_salt useless as protection?+
No, but you have to treat it as a contract per call. The advisory for CVE-2026-105752 shows an internal resubmission that loses the salt. Gateways and upstream proxies belong in a check for whether they pass the value on at all, and the clean separation stays an instance per tenant or at least a separate cache namespace.
Why is the version number not enough for my Docker image?+
Because self-built images often carry a nightly identifier such as 0.1.devNNNN, which cannot be compared against any release line. For our image we read the two affected code places directly in the running file. That is ten minutes of work and the only statement that counts for in-house builds.
senn-tech