senn-techsenn-tech
KI-News
KI-News2026-09-29· By Franz Senn

Holo4: open computer-use models, and the best of them is non-commercial

On 28 September 2026 H Company released Holo4, a family of models that drive software through every open door: the screen, code, MCP tools, APIs. There is a dense model with 27 billion parameters, a mixture-of-experts model with 35 billion, three of them active, and as a third candidate Holotron4-30B-A3B, called Holotron 4 Nano in the blog, built on Nvidia Nemotron 3 Nano Omni. All three take images and text together, every config file lists 262,144 tokens of context, and the Holo4 weights sit on Hugging Face in BF16, FP8, NVFP4 and GGUF. The license of the model with the best numbers rules out commercial use.

Where the numbers come from

On OSWorld 2.0, the benchmark for long computer workflows, H Company reports 61.7 partial points for Holo4 27B, a 41.5 percent success rate and 1.22 US dollars per task. The company measured this itself, in a harness it rebuilt, in a single run.

OSWorld 2.0 partial points per the table in the blog post of 28 SeptemberOpus 5.581.8 · 81.8 at 8.48 USD per taskGPT-6 Astra73.5 · 73.5 at 9.07 USDMuse Spark 1.366.9 · 66.9, cost not statedHolo4 27B61.7 · 61.7 at 1.22 USD, own measurementQwen3.8 27B (base)48 · 48.0 at 3.49 USD090
Comparison values come from different harnesses, effort levels and task subsets, as the footnote of the blog post states. Source: H Company blog post. (Quelle: hcompany.ai/newsroom/holo4)

One day after the release, no independent re-measurement exists for any Holo4 number. What H Company does publish is every agent trajectory behind the scores, at trajectories.hcompany.ai. Anyone who doubts a number can replay, step by step, how the agent reached it.

The comparison missing from the table

The LinkedIn announcement, marked as edited, presents the 61.7 percent as a lead: just ahead of “GPT-6 Sol-Xhigh” at 60.5 percent, at roughly a quarter of the estimated cost. That model with that number appears nowhere in the linked blog post, in no table and in no dot of the cost chart. Those tables put Opus 5.5 at 81.8 and GPT-6 Astra at 73.5 percent in front, and the intro of the blog itself says Holo4 trails the strongest closed models. The closest dot in the chart is GPT-5.6 Sol at 66.2 percent, measured on another subset taken from the OpenAI launch chart. The habit for announcement posts stays the same: read the linked tables too.

The official OSWorld 2.0 leaderboard deserves a second look. Its best model solves 20.6 percent of the tasks under the binary metric (Opus 4.8, 54.8 partial). The 41.5 percent success rate of Holo4 would sit well above that. Direct comparison does not work, because the company rebuilt the harness, among other changes with a shell directly on the desktop machine and its own memory across hundreds of steps. What is provable: an official re-run is outstanding, and the 61.7 percent is the partial score of one own measurement.

What the fine-tuning bought

Holo4 27B builds on Qwen3.8-27B. In the same harness the base model reaches 84.3 percent on the older OSWorld, Holo4 85.2. On the long OSWorld 2.0 the values are 48.0 against 61.7 partial points, and the in-house MCP suite jumps from 74.2 to 89.4 percent. The pattern: on the short, already saturated benchmark the fine-tuning adds almost nothing, on long workflows and on tool calling it adds a lot. Training used 127 trillion tokens of supervised fine-tuning, about three quarters of them successful agent trajectories from the company task factory, which per the provider builds ten thousand verified tasks out of product documentation and screenshots. Two LoRA experts then went through online reinforcement learning, one for desktop and web, one for terminal, MCP and API, and both were merged back into the model at equal weight.

The AutomationBench number with its footnote

For business automation over MCP tools H Company names 45.4 percent on AutomationBench. The footnote of the post: 480 of the 600 public tasks sat inside the company training pool, and the blog itself rates the 120 held-out tasks at 49.3 percent for Holo4 27B. The private set of the benchmark has not been run for Holo4. Credit where due: the company writes this down itself. Quote the number only together with that note.

Which license covers which weights

ModelBaseLicense of the weights
Holo4-27BQwen3.8-27BCC BY-NC 4.0
Holo4-35B-A3BQwen3.6-35B-A3BApache 2.0
Holotron4-30B-A3BNemotron 3 Nano OmniNVIDIA Open Model License

The licenses stand in the header of the model repos on Hugging Face; neither the announcement post nor the blog mentions them with a single word. The base of Holo4 27B is licensed Apache 2.0, the fine-tune carries CC BY-NC 4.0 and therefore forbids commercial use. A pilot planned on the 27B table values is a plan for something you may not operate that way. The commercially usable member is Holo4-35B-A3B, and with 30.9 partial points and 12.3 percent success on OSWorld 2.0 it sits clearly behind the flagship. Holotron4 Nano, and H Company is a member of the Nvidia Nemotron Coalition, lifts its Nemotron base on GUI tracks strongly per the provider, from 21.0 to 76.3 percent on the older OSWorld, but stays at 7.9 percent on OSWorld 2.0.

The loop all three models are meant to driveSend a screenshot plus tool resultsModel answers with a click, a keystroke, code or a tool callThe harness runs the step on the machineThe new observation goes back to the modelup to 500 steps per task
One model call across every interface. The difference to a chat model sits in the loop around it. (Quelle: senn-tech, after the blog post and the model repos)

What running it costs

The BF16 weights of the 27B model fill about 55 gigabytes across 18 shard files; the provider ships FP8, NVFP4 and GGUF for both sizes. The Qwen3.8 27B class runs in our setup as a reserve lane on a single RTX 5090, quantized, on vLLM. Whether the same service handles the image-input variant with the same start parameters is untested on our side. The price sits in the loop around the model: per the provider, the benchmark harness works through OSWorld 2.0 tasks in up to 500 steps and several hours per task. At a 41.5 percent success rate the majority of tasks fails on its own. Productive use therefore needs the same human in the loop and the same cost cap per task as any other agent workflow. We have measured nothing ourselves, and for licensing reasons a commercial test would only come together with the Apache model. Checking whether the published trajectories fit your own applications works with the 27B model, as long as nothing commercial gets built on it.

Further reading

Questions?
Can we use Holo4 commercially?+

Not the 27B model. It carries CC BY-NC 4.0. Holo4-35B-A3B is Apache 2.0 and Holotron4-30B-A3B is under the NVIDIA Open Model License, so those two are commercially usable, though clearly weaker on long workflows. The licenses sit in the model repos on Hugging Face; the announcement text does not mention them.

How solid is the 61.7 percent OSWorld 2.0 score?+

It is a provider number: one run, in a harness the company modified, scored as partial points, with a 41.5 percent success rate. The comparison values come from other harnesses and task subsets. No third party has re-measured it yet; H Company does publish every agent trajectory behind the run.

Why is the release still worth a look?+

The open trajectories, four serving formats including FP8 and NVFP4, and the training documentation. The pattern in the tables, a big gain on long workflows and none on a saturated short benchmark, is something you can check against your own applications with the published material.