Model Releases Hugging Face Blog

Holo4: powering generalist computer-use agents

Holo4computer-use agentsH CompanyOSWorld benchmark

Holo4 is a new series of agentic models released in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts — both available on the H Models API, alongside an updated Holotron 3 release called Holotron4 Nano. The models build on the previous generation and interact with software through any available interface: GUIs, code, MCP and APIs. They were trained through supervised and reinforcement learning on a large set of environments and tasks, including ones generated by the company's Agentic Task Factory. Full weights are offered in FP16, FP8 and GGUF formats, and the trajectory viewer and dataset for the public benchmarks are open-sourced on Hugging Face.

The central design claim is interface agnosticism: Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, choosing whichever fits the task. Most agentic models are trained for a single interface, the post argues — GUI-focused models are blind without a screen, while tool-calling models stall in applications with no API — whereas real business tasks often require combining approaches. The same model and the same calling convention run on desktops, on the web, on Android, in a code sandbox and against business APIs, so users do not need to pick a different model per platform.

On benchmarks the models improve significantly over their Qwen base and trail only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% versus 81.8% for Opus 5.5, while Holo4 35B-A3B reaches 30.9% — achieved with orders of magnitude fewer parameters and at much lower cost. The company also published methodology notes for its cost-performance charts: OSWorld 2.0 costs are estimated from input and output tokens per agentic run, with Holo4 priced at H Models API rates for a single run, Qwen models priced at Alibaba Cloud list rates (with cache hits at 20% of input price for Qwen3.6 35B-A3B), GPT/Opus effort sweeps taken from OpenAI launch data, and other points from the official leaderboard, with the caveat that releases, harnesses and task subsets differ. AutomationBench measurements use v1.0.6 with scores and costs from an internal harness.

The upshot: Holo4 positions itself as a cheaper, open-weight alternative to frontier closed agents on desktop control (OSWorld 2.0) and API use (AutomationBench), with full trajectory transparency so users can inspect every step behind the reported scores.

Read original →

← Back to home