senn-techsenn-tech
Security
Security2026-10-11· By Franz Senn

Two vLLM cache advisories and what operators of their own AI stack take from them

In late September 2026 the vLLM team published two security advisories, both updated again on 5 October 2026: GHSA-ph3r-5jfg-f84f carrying the CVE number CVE-2026-105753 and GHSA-935w-9g4m-p28p with CVE-2026-105752. The time depends on which layer of the reporting platform you query, and we name both: the advisory in the repository carries 28 September 2026 as its creation date and 5 October 2026 as its last change, the same entry in the global security database of the same platform stands at 6 October 2026, 00:02 UTC. The repository version is the one that counts for this text, because it carries the content of the report. Both are fixed in current versions, neither one executes code. For operators of their own inference fleets they still describe two places worth knowing, because both sit inside mechanisms that every shared instance switches on by default: the multimodal cache and the shared prefix cache. According to the advisory entry, both findings come from the Patch the Planet programme, a collaboration between Trail of Bits and OpenAI, discovered with the model GPT-5.5-Cyber.

How one rejected request poisons the multimodal cacheRequest with an imageThe prompt gets renderedAdmission rejectsLength above max_model_lenFrontend stayspoisonedMetadata claims the image iscachedLater the same imagefileHit on the frontend, nopayload to the coreAssertion in theengine coreMessage: Expected a cacheditem
Sequence per advisory GHSA-ph3r-5jfg-f84f. The frontend side holds only metadata, the engine core holds the payload data. When a request is rejected after rendering, the frontend is left holding an entry that never existed at the back. (Quelle: Advisory GHSA-ph3r-5jfg-f84f, vLLM, 28 September 2026)

CVE-2026-105753: One rejected request is enough for an outage

vLLM runs the multimodal cache mirrored across two processes by default: the advisory calls the frontend P0, where the metadata lives, and the engine core P1, where the payload data lives. The design assumes that both sides keep the cache in lockstep at every point in time, so that P0 may claim without asking back that a medium is already present in P1.

That assumption breaks inside one specific window. A chat request carrying an image is rendered first and gets registered in P0 during that step, the length check against max_model_len comes afterwards. If the check comes out negative, the request is rejected and the P0 entry stays. P0 now treats the image as cached, P1 never saw it. The next request that sends the same image hits in P0 and therefore receives no payload along, only the reference to the cache. P1, which has nothing, runs straight into the line assert mm_item is not None. The advisory traces this to all versions before 0.28.0 running the standard configuration mm_processor_cache_type="lru"; the alternative paths processor_only, a switched-off cache and shm are not affected according to the advisory.

The severity sits at CVSS 6.5 (AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H). The letter A:H is the point: it is availability alone. In the build the advisory tested, vLLM catches the assertion as an error of that one request, public reports show the same assertion in other builds as a cascade of further engine errors. The reported effect is a service outage that anyone allowed to use the API can repeat. The remedy proposed in the advisory (atomic adoption into both caches after all admission checks, withdrawal of the P0 entry on rejection, and the assertion handled as an error per request) stands in the advisory text as a proposal; the linked pull request #51897 was closed without being merged. Upstream the problem has still become smaller, the files themselves show that, more on that below.

CVE-2026-105752: The salt applies only to the first round

The shared prefix cache was already a reported weakness once: CVE-2025-46570 from May 2025 described the timing oracle through which a tenant can guess whether a prompt is already sitting in the cache. The documented countermeasure is called cache_salt: every tenant sets its own value, and from that follows its own cache namespace.

Exactly this separation dissolves in parts on the responses path POST /v1/responses. This interface uses the Harmony model format (the GPT-OSS family) and drives requests with attached tools as a multi-turn loop: after every tool call vLLM rebuilds the prompt of the next round and submits it into the engine again. The first pass hands over the cache_salt of the caller correctly, the continuation calls the engine entry without a salt. The continuation then sits in the unsalted, global namespace even though the caller had switched salting on.

A second tenant who can reconstruct the post-tool history well enough submits the same continuation unsalted and reads the exact number of cached tokens per turn out of the response usage. That is the membership oracle again, the one cache_salt is supposed to protect against, this time from a different source: the response usage itself delivers the exact numbers, a timing measurement is not needed for that. The preconditions are heavy, which is why the report sits at low (CVSS 3.1, AC:H): a Harmony model, the tool server activated, prefix caching on (standard), the victim sets a salt and triggers at least one tool continuation, and the attacker has to be able to guess the history. It is fixed since 0.30.0: there the path is rewritten so that it reads the salt back out of the engine input and carries it along in the further steps. The patch proposed in the advisory (#51818) was never merged, the implementation went through the restructuring.

The self-test: our lanes measured on 11 October 2026

The first question at any report is the one about your own estate. On 11 October 2026 we queried the four inference endpoints through their own API:

  • Primary lane on 192.168.180.3:8000: reports 0.1.dev20073+g8e685d198
  • Fallback lane on 192.168.180.202:8080: reports 0.30.0
  • Embedding and rerank lanes on 192.168.180.211:8010 and :8011: report 0.30.0, the pulled image is vllm/vllm-openai:v0.30.0

Read the way the advisory metadata states it (affected < 0.28.0 and < 0.30.0), the three 0.30.0 endpoints are out for both reports. On the primary lane the version string says nothing: a nightly tree with a dev number and a commit from our own build branch cannot be compared against any release line. So we looked into the running file instead of guessing, inside the container of the production image:

  • In multimodal/cache.py the fix is in. Our running file is identical to the release tag v0.28.0, both checksums are the same. That means the file contains what 0.28.0 introduced: the class MultiModalCacheMissError with five occurrences, two invalidate methods and one _cache.pop, so the handled error together with the withdrawal of the shadow entry. The assertion Expected a cached item still stands three times in the file, and three times at the tag v0.28.0 as well. Counting only the assertions misses the fix in both directions.
  • In responses/serving.py the fix is not. Our file has 1563 lines, the tag v0.27.1 has 1561, the checksums differ, the structure is the same. The continuation stands on line 722 as tokens_input(token_ids) without a salt. The build v0.30.0 is rebuilt at this point and some 180 lines shorter: it reads the salt back out of the engine input with cache_salt = engine_input.get("cache_salt") and passes it again when re-rendering the Harmony messages. Both are missing in our build.
Occurrences of MultiModalCacheMissError in the multimodal cachev0.27.10 · affected per the advisoryv0.28.05 · fix linev0.31.09 · code ships as a subpackageour production image5 · identical to v0.28.009
Our own count across the sources of the release tags and inside the running container, as of 11 October 2026. The file vllm/multimodal/cache.py of our production image carries the same checksum as the tag v0.28.0. From v0.31.0 the same code ships as the subpackage cache/ across five files, counted over all five. (Quelle: Our own measurement, sources of the tags v0.27.1, v0.28.0, v0.31.0 and the running container)

So the in-house build carries one of the two fixes and not the other. The version string 0.1.dev20073+g8e685d198 of the image would have informed us wrongly in both directions, two checksums against the upstream tags say it in seconds. For us that means, for report one: the path is healed, we are out. For report two the finding stays standing without hitting us, because the Harmony path is not reachable on our side, we operate no GPT-OSS, and the gateway uses the chat-completions interface. The production image has not been rebuilt since the measurement and still carries the build date 3 October 2026 unchanged. Catching up would be the clear answer, and it is open.

The third entry of the same wave of reports also fits the picture, and it touches our document processing: GHSA-q43m-vhcp-mhvm (CVE-2026-105750) reports that Docling does not enforce the barrier enable_local_fetch in the HTML browser rendering mode, affected are 2.82.0 up to before 2.118.1 and the same span for the package docling-slim from 2.92.0. Our Docling container reports 2.43.0 and therefore sits outside the affected span, measured on 11 October. A second reason, independent of the version: the advisory explicitly excludes the standard configuration without page rendering, the command line and docling-serve, and that is exactly how our instance runs. The reporting platform lists the report at 5.9, the NVD metadata at 7.5, the difference lies solely in the attack effort.

What operators take from this

The version boundaries for your own instance, sorted by exposure:

  • Multimodal requests from users: at least 0.28.0, preferably 0.31.0 today (released 5 October 2026). The assertion still stands there, the difference is that the drift is handled as an error in that revision and the shadow entry gets withdrawn.
  • Tenant separation through cache_salt: at least 0.30.0 (released 22 September 2026).
  • Text only, internal only, no tool path: on a pure text stack both reports pass by, a current build still remains the better starting position.

Two points beyond version numbers. The first: cache_salt is a promise per call, not a switch on the server. A gateway that does not pass the value through runs unsalted, and internal resubmissions can lose that same value, as this report shows. Anyone who genuinely wants to separate tenants will not get around their own instance per tenant, or at least their own cache namespace, in the long run. The second: with self-built images the code you read says more than the version string, and in both directions. Our in-house build was built on 3 October 2026 and reports 0.1.dev20073+g8e685d198, so no number you could compute against an advisory span. We checked with two checksums: multimodal/cache.py is identical to the tag v0.28.0 and therefore patched, responses/serving.py matches the structure of v0.27.1 and therefore is not. In substance the tree of our file then stands at two different places, and the version string would have revealed neither of them.

One observation belongs on record at the margin as well: both findings come from a programme that couples machine search for vulnerabilities with a human-written report. Projects that patch their own inference engine will have to treat the advisory list of that engine the way they treat the DSA list of an operating system.

Further reading

Questions?
Do I have to lift vLLM to 0.31.0 right away?+

Staged by urgency. Anyone accepting multimodal input from users should go to at least 0.28.0, better 0.30.0 or 0.31.0, because there the crash has become a handled error that withdraws the cache entry again. Anyone separating several tenants through cache_salt needs at least 0.30.0, otherwise the separation does not hold inside tool continuations. A text-only instance that is reachable internally and has no tool endpoint is in less of a hurry. Both reports are rated low and moderate in severity, neither one executes code.

Does that make cache_salt useless as protection?+

No, but you have to treat it as a contract per call. The advisory for CVE-2026-105752 shows an internal resubmission that loses the salt. Gateways and upstream proxies belong in a check for whether they pass the value on at all, and the clean separation stays an instance per tenant or at least a separate cache namespace.

Why is the version number not enough for my Docker image?+

Because self-built images often carry a nightly identifier such as 0.1.devNNNN, which cannot be compared against any release line. For our image we read the two affected code places directly in the running file. That is ten minutes of work and the only statement that counts for in-house builds.