senn-techsenn-tech
Security
Security2026-10-08· By Franz Senn

Agent Sandboxes Are GA, Management Is Not: What Microsoft's MXC Actually Enforces

On 7 October 2026 Microsoft released Windows Execution Containers (MXC), and the same day the project received version 1.0.0 in its repository, tagged at 05:54 UTC. The event also brought devices with Nvidia's RTX Spark, a demo of local coding models, and the claim that Windows should become the safest platform for agents.

The sentence that matters for us came from the repository. MXC is open under the MIT licence, ships SDKs for Rust, .NET and Node, and its default Linux backend is Bubblewrap. We read the repository, the backend documentation, the developer blog and the GitHub docs, and counted what holds.

How MXC enforces a policy

MXC is an execution layer. A program hands over a container type, rules and a command, MXC validates, picks the backend and starts the workload inside it. The policy sits outside the agent, so a model cannot widen it mid-task. Files, network, processes and desktop get contained.

PlatformDefault backendOther backends
Windows 11 x64 and ARM64processcontainerisolation_session, wslc, plus windows_sandbox, microvm, hyperlight
Linux x64 and ARM64bubblewraplxc, microvm, hyperlight
macOS ARM64 and x64seatbeltnone

Taken from the overview in the MXC repository, state of 8 October 2026. Microsoft marks windows_sandbox, microvm and hyperlight as experimental, which are the containment types that place an extra layer between agent and machine.

Three modes come with it. Learning mode blocks a denied action and writes a JSON activity file, enforcement mode blocks without that report, permissive mode allows while logging.

The Linux backend is further along than its name

The Bubblewrap documentation carries the line "Status: Stable, the default Linux backend". Needed are a kernel from 3.8 with user namespaces and Bubblewrap from 0.5.0, because the template uses --ro-bind-try and the environment scrub --clearenv. Network containment through a proxy or direction rules needs slirp4netns, nsenter, an iptables tooling with nftables backend and the module nf_conntrack, which an unprivileged container cannot load afterwards. On a host with iptables-legacy the validate check refuses to run, since the unprivileged helper cannot take the /run/xtables.lock lock. IPv6 stays unreachable in those modes even where a rule would allow it.

What is missing on our side we looked up rather than assumed, on our own agent host.

Check on dsh-pve6, 7 October 2026Value
Kernel7.0.0-38-generic, Ubuntu
unprivileged_userns_clone1
max_user_namespaces61366
bwrapnot installed
slirp4netnsnot installed

The kernel is ready, the packages are not. A trial run needs two packages and one policy, and the operating system has nothing missing.

What the release covers and what it does not

Engine, policy schema and SDKs are generally available. Intune distribution, the separation of agent activity from user activity through Entra and the Agent 365 controls for local agents are described by Microsoft as coming. An independent agent does run today as its own Windows account with its own desktop and storage path, and with no link to a directory identity. With Windows 365 for MXC the container runs in a cloud session.

Four points come out of Microsoft's own texts and belong in front of any pilot.

Enforcement mode ships no activity reports. What an agent needed during a productive task then appears in no listing from the container. The policy has to come from a preceding phase, and that phase is what you archive.

Capabilities a backend does not offer are dropped. Microsoft writes that the container may then run with reduced containment and recommends reviewing the effective policy that comes back. A written policy is therefore not yet a valid policy, and that review belongs at the start of a pilot description.

The checking switch --audit turns off all sandbox security for the workload under analysis, according to the warning in the repository. A switch that releases containment before start protects nothing in a running sequence.

The GitHub documentation for Copilot names two limits. The built-in file tools run inside Copilot itself, where an in-process check point tests their requests against the policy, while OS isolation of the child process is absent. Attached MCP servers sit outside the container. The stated default is outbound internet access on and local network access off.

The open backlog is a fair measure of maturity: on the Linux branch a race in the network setup of short-lived processes (1378) and a missing default-deny seccomp profile (1443). On Windows, MSYS2 and git-bash runtimes fail to initialise in the process container (1061), and in the experimental micro VM backend blockedHosts overrides a policy of defaultPolicy=block (786). Closed on 7 October: a Seatbelt leak that read sizes and timestamps on blocked paths through stat().

The contradiction since February

On 19 February 2026 Microsoft's Defender Security Research team described self-hosted agent runtimes using OpenClaw as the example. The quote there: OpenClaw is "not appropriate to run on a standard personal or enterprise workstation". Evaluating it, they wrote, needs a fully isolated environment, a dedicated virtual machine, unprivileged credentials of its own, monitoring and a rebuild plan. Earlier in the same text sits the sentence that installing a skill is basically installing privileged code.

On 7 October 2026 OpenClaw appears on the list of agents that already use MXC, next to GitHub Copilot, OpenAI Codex, Replit, LM Studio and Unsloth. Both texts are from Microsoft.

The backend list resolves the contradiction. A process container is a shielded process, no separate machine. It shrinks file space, network space and desktop, while the credentials the agent works with and the supply road for skills stay untouched. MXC limits the damage and does not replace the isolation demanded in February.

Local models are not an open bill

The hardware half of the event is called Surface Laptop Ultra with RTX Spark, orderable since 7 October, shipping from 16 October. For the local coding tier Microsoft names a model of its own, MAI Code 1.1 Flash, a mixture-of-experts model with 137 billion total parameters and 6.8 billion active. Its device build is described as mixed quantisation at roughly 3.3 bits per weight, which is 53 GB and 80 percent smaller, plus speculative decoding and a llama.cpp execution path.

BuildSWE-Bench Verified (500 tasks)Terminal-Bench 2.1 (89 tasks)
MAI Code 1.1 Flash72.6%62.9%
GPT OSS 120B as comparison32.0%23.6%
MAI Code 1.1 Flash, quantised70.80%66.29%

These figures come from Microsoft's own post of 7 October 2026, measured by the manufacturer. On Terminal-Bench the quantised build scores above the unquantised one with no explanation in the text, and the comparison runs against a community copy of GPT OSS at 80 percent quantisation. The footnote on throughput names 923.5 and 769.8 tokens per second at 64K and 128K context for prompt processing, while describing the same results as decode throughput. That number only comes with a caveat.

Our own arithmetic on it, using Nvidia's published bandwidth of 273 GB/s for DGX Spark with GB10, while the RTX Spark N1X figure is unpublished. Reading all 53 GB per output token would land near five tokens per second. At 6.8 billion active parameters and 3.3 bits you are near 2.5 GB per token, so the ceiling sits around a hundred tokens per second, before any losses. "Frontier locally" means here: a mixture-of-experts model with few active parameters plus one shared memory pool. The fast laptop alone does not explain the number. At 256K context Microsoft puts peak memory at 75.5 GB, of which up to 22 GB is key-value cache.

We looked for the open weights. DeepSeek V4 Flash sits on Hugging Face as deepseek-ai/DeepSeek-V4-Flash, licence MIT, around 996,000 downloads. Its configuration lists 43 layers, 256 expert networks with 6 active per token and FP8 weights, which makes every 2-bit step a further quantisation by third parties. The 1.6-bit build shown on stage comes from the write-up of the demo, it is not in the blog post. Microsoft's own MAI model is missing there, the microsoft organisation lists only maira-2 and MAI-DS-R1, and those weights ship with the product. Nvidia's model above 70 billion parameters at 2 bits stays unnamed, and the comparison figures against the Mac come, per the notes of that post, from commissioned testing in September 2026 on preproduction hardware with a single text model.

The build decides, the product name does not

MXC does not run on every Windows 11. The repository documentation fixes minimum levels per release, and the table shows that two cumulative updates may be needed, since session isolation asks for a higher build than process isolation.

Windows 11Process isolationSession isolation
24H226100.9278 (KB5120998)26100.9550 (KB5124010)
25H226200.9278 (KB5120998)26200.9550 (KB5124010)
26H226300.9550 (KB5124010)26300.9550 (KB5124010)
26H128000.2804 (KB5120996)28000.3086 (KB5124006)

Taken from the support page in the MXC repository, state of 8 October 2026.

Our inventory held six guests carrying the Windows 11 tag and fifteen running Windows server guests on 7 October 2026, among them two domain controllers, four terminal servers, two ERP systems and the MailStore instance. We counted through the Proxmox API, and the homelab runs no Windows guest. The tags are our own classification, the cumulative updates on the six devices are readable and still unread.

What we take from this

Read the effective policy first. Microsoft says itself that unsupported capabilities fall away, so a trial that never evaluates the policy coming back proves only that a container started.

The learning report comes before enforcement. Learning mode writes a JSON listing of every allowed and denied operation, which is the most honest road to a policy that fits the work and the only one you can archive while enforcement writes nothing down.

On our own Linux hosts the same mechanism is testable without Windows: two packages, a policy that denies outbound traffic by default, and a look at nf_conntrack. We did not build it on 7 October, the state is still visible: kernel ready, packages absent.

Identity stays our own work. Microsoft does not deliver the separation of agent and person through Entra yet, and we run agents on service accounts at the LiteLLM gateway. An agent holding our credentials stays our risk, however tight its container is.

In Windows environments measure first, approve second: read the build levels, pilot in learning mode on one device, hold fleet approval until Intune controls and a report exist in enforcing mode. For the six devices in our inventory that is one hour of work.

The open question is the order of work. Registering agents under a directory identity needs a directory service, and we run that for people. An agent with its own account and its own rights is an identity decision, no container question.

Further reading

Questions?
Is MXC only useful on Windows machines?+

No. MXC is a Rust project under the MIT licence with SDKs for Rust, .NET and Node. On Linux the default backend is Bubblewrap, which the repo marks Stable, and on macOS it is Seatbelt. What still has to arrive is the Windows side of management through Intune, Entra and Agent 365. The container engine itself runs on all three platforms.

Why are the three modes a problem then?+

Because Microsoft provides no activity reporting for the enforcing mode, Enforcement, in its own developer blog table of modes. The activity report is listed for Learning mode and for the permissive level only. Whoever enforces hard therefore never learns from the container which file or which destination an agent needed. The policy has to come from an earlier learning phase instead.

Can we use this on our Linux hosts for our own agents?+

The mechanism is there, the tooling is not, at least on our side. On our agent host dsh-pve6 there was neither bwrap nor slirp4netns on 7 October 2026, looked up with command -v. The kernel allows unprivileged user namespaces: unprivileged_userns_clone reads 1 and max_user_namespaces reads 61366. A trial needs two packages and one policy. Nothing is missing in the operating system.