Local AI in a Field Test: Seven Days Without Cloud Models on 64 GB
On 1 October 2026, Adrian Twarog published a seven-day experiment on YouTube: a full week of working, coding and running agents on local language models instead of cloud subscriptions. Sponsor ASUS supplied the machine, a NUC 16 Pro with 64 GB of memory. Sponsored, therefore, and still worth reading, because Twarog does the arithmetic on the two numbers that such videos usually wave away: the machine's memory bandwidth and the context a coding harness burns before your question is even asked. His verdict after seven days: local models are fine as agent underlayers, and not ready for codebases with tens of thousands of lines. This post lays his field test next to our production stack, because the failure points are the same ones.
What he chose to run
Model selection from the leaderboard is a reasonable first step. In the text arena it is topped by Anthropic's Claude Fable 5, at 1506 in ModelCap's snapshot of 10 August 2026. The open models right behind it would have taken the mini PC far past its memory: Twarog names Kimi K3, GLM 5.2 and Qwen3.8-Max, then drops them for triple-digit gigabyte requirements. What survived the cut is three well-documented models:
- Qwen3.8-27B: a dense 27-billion-parameter model, Apache 2.0, weights on Hugging Face since 13 August 2026, native context of 262,144 tokens per its config.json
- Muse Glimmer: Meta's first open model from its Superintelligence Labs, 30 billion parameters, Apache 2.0, released 10 August 2026, built for agents on a Mac or a single consumer GPU
- GLM-4.7-Flash: Z.ai's mixture-of-experts model, 30 billion parameters with roughly 3 billion active per token, released 19 January 2026
Ollama and LM Studio serve as runtimes. Two of his practical notes are correct and still routinely ignored: pick a GGUF that allows GPU offloading or the speed is useless, and the quantization class (Q4, Q6, Q8) is a trade between memory and accuracy, nothing more mystical.
On speed: measured against the ceiling
His measurements come from the LM Studio debug panel: Qwen3.8-27B around 12 tokens/s, Muse Glimmer around 6, GLM-4.7-Flash around 22. Two GLM-4.7-Flash instances run side by side, and small specialised models reach over 100 tokens/s once tuned. One more observation on Muse Glimmer: a large share of its reasoning tokens goes to re-checking its own guidelines while it thinks. He noted this from the reasoning traces without measuring it, and anyone building agents on that model should check it themselves.
What carries the comparison is memory bandwidth. The Core Ultra X7 358H supports LPDDR5X up to 9600 MT/s, ASUS ships the machines populated with LPDDR5X-8533 depending on model, and TechPowerUp measured above 100 GB/s in its review. A dense 27B model at 4-bit pulls close to 14 GB of weights per token generated. Counted without overhead, that leaves 8 to 11 tokens/s depending on the memory configuration. His measured 12 are in the same order, the residual difference coming from how the figure was read and from cache effects. The ceiling is the point: this hardware keeps no headroom for a dense model of this class, whatever the runtime.
More interesting than the raw values is the distance to what should be possible. With 3 billion active parameters, GLM-4.7-Flash moves only about 1.5 GB of weights per token, so counted from the measured 100 GB/s the ceiling sits near 60 tokens/s. He measured 22. On the DGX Spark, whose 273 GB/s are more than double this bandwidth, the same model class is documented at 89.3 tokens/s under llama.cpp with CUDA. The runtime and graphics stack on the Intel machine therefore harvests a much smaller share of the available bandwidth. Buying local hardware means buying the maturity of that stack along with the spec sheet.
Where the coding harnesses broke
The coding leg is where the experiment earned its keep. Twarog wires the local models in without any cloud account: VS Code ships its own model manager ("Manage language models"), where you enter the endpoint LM Studio exposes, URL and model name, and the models appear in the chat pane. The requests fail anyway at first. His own diagnosis: VS Code attaches its tool calls to every request, fills the small context window completely, and leaves nothing for the actual chat. Two changes fix it, disabling the tools and raising the context.
The agent CLIs behave differently. Claude Code takes more than two minutes on a hello world, and the answer finally arrives after about three: "How can I help you?" Twarog pulls up the prompt the harness hands to the model and gives up on small local models for this use: the harness loads a model that carries an 8,000-token window with a load it was never meant to bear. OpenClaw takes two to three minutes on the same hello world in his test. His interim verdict, worth more after five days in the field than any benchmark table: small local models are not harness-ready yet.
The same failure point, in our numbers
Our main lane runs Qwen3.8-Flash-Next on four RTX 5090, with a maximum of 220,000 tokens per session. That figure is not a flex: an agent harness re-sends its system prompt, tool schemas, file listings and sub-agent state on every pass, and tens of thousands of tokens are gone before a developer's first real question lands in the context. A field test configured at the 8,000-token default is therefore not competing against the model's intelligence. It is competing against the harness's bookkeeping.
The price of long context is the KV cache, and at our place that is now the tightest resource on the lane. Since 3 October 2026 we run the KV cache in NVFP4, which holds 8.2 full 220,000-token sessions in the pool, against 5.2 with FP8. The field test shows the same arithmetic from the other side: Qwen3.8-27B offers 262,144 tokens natively, LM Studio provisioned 8,000. The whole harness chapter lives between those two numbers.
For work with no interactive claim, local models have long been permanent fixtures here. A 0.6-billion-parameter embedding model with our own classification head decides spam on our corpus, measured AUC 0.97, thousands of calls a day at zero per-token cost. A Qwen3.8-27B quantized with AWQ idles as a fallback lane for days when the main lane is down. Services of that kind get a shadow phase before they ever speak: the candidate runs along, says nothing, and only switches on after the verdict round.
Where he himself lands after a week: hybrid
By the end of the week Twarog runs a cloud model as orchestrator and the local models as sub-agents for cron jobs, memory maintenance and repetitive checks, on a second machine that belongs to the agent alone. That is precisely our division of labour, with a gateway in front instead of cables behind the desk. The market has noticed the arrangement too: the NUC 16 Pro product page lists OpenClaw and Hermes Agent as use cases, ASUS is selling the mini PC as an always-on agent box.
Half of his closing hardware thesis, a future with "huge RAM allocations in most machines," is right. Capacity decides which model fits; bandwidth decides how fast it answers. The platform offers bigger machines anyway, soldered LPDDR5X up to 96 GB and socketed DDR5 variants up to 128 GB, so his wish was buyable. What more gigabytes would not change is the roughly 100 GB/s ceiling that sets the dense 27B model's speed. That a memory upgrade is expensive anyway is visible in the price sheets: NVIDIA raised the DGX Spark from 3,999 to 4,699 US dollars on 23 February 2026, citing memory shortages, and Framework has lifted its desktop price in several steps since February 2025. The full accounting sits in our comparison of the desktop boxes.
What stays local and what does not
The field test's real answer is a line of separation we can confirm from production: classification, embeddings, cron sub-agents and the fallback lane run on small hardware and should stay there. Interactive agent coding in large codebases currently fails on the multiplication of harness context and the bandwidth ceiling, even though the open models keep closing on Claude Fable 5 in the leaderboards. Twarog's closing line, that for codebases with thousands of lines "we might not be there yet," is the most honest sentence in the video by the evidence we have measured. For small, self-contained jobs that run overnight and bother nobody while waiting, the math has already worked out today.
Further reading
- Adrian Twarog: I Tried Coding with Local AI Models for 7 Days, YouTube, 1 October 2026
- ASUS NUC 16 Pro, product page with tech specs
- TechPowerUp review of the ASUS NUC 16 Pro, Core Ultra X7 358H and Arc B390
- Intel Core Ultra X7 358H, Intel ARK product brief, product number 245527 (the page sits behind a bot wall for scripts, reachable in a browser)
- ModelCap: Arena text leaderboard, snapshot of 10 August 2026
- NVIDIA developer forum: DGX Spark price change, 23 February 2026
Why do local models fail inside Claude Code or VS Code when the chat works?+
A coding harness sends its full system prompt and the schema of every tool with each request. With LM Studio set to the default 8,000-token context, the window was full before the actual question arrived. Twarog disabled the tools and raised the context limit, and the VS Code chat worked after that.
Which models ran on the 64 GB machine, and how fast were they?+
Qwen3.8-27B at roughly 12 tokens/s, Meta's Muse Glimmer at roughly 6, and GLM-4.7-Flash at roughly 22, all measured by the author inside LM Studio. The 12 tokens/s for the dense 27B model sit in the same order as the bandwidth ceiling of this machine class: 8 to 11 tokens/s at 4-bit by the math, depending on the memory configuration.
When does running agents on local models make sense?+
For repetitive background work: cron jobs, memory maintenance, classification, embeddings. Interactive agent coding in large codebases is another matter, because harness context and memory bandwidth decide the wait. In our stack, small local models handle classification and large open models handle interactive traffic, and new services start in shadow mode.
senn-tech