Open Source Hacker News (LLM)

Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp

MetalVirtualization.frameworkllama.cppApple Silicon

The work builds on Lume, a macOS virtualization stack originally launched via Show HN, and is the first result from connecting that Virtualization.framework foundation to Cua Driver, Cua Cloud, and Fleets. It addresses a long-standing limitation: Apple's Virtualization.framework provides macOS guests with a paravirtualized GPU, not direct passthrough, so LLM workloads run far slower than on bare metal. Tart, another Virtualization.framework CLI, has an open issue asking whether usable graphics and decent LLM performance are possible in a macOS guest. The new compatibility layer keeps Apple's virtual GPU but exposes newer Metal capability paths on that device, closing part of that gap.

The mechanism is a small, process-scoped compatibility layer that alters the Metal capability responses a macOS guest sees. In a stock Tahoe VM, the paravirtualized device reported roughly an Apple 5-era GPU family, only 32 KB of maximum threadgroup memory, and no SIMD-group matrix support. Those reports make llama.cpp select slower kernels even though the physical GPU could execute newer ones. By adjusting these capability queries, the layer lets Metal software choose faster kernel paths.

On an M1 Ultra with TinyLlama 1.1B, prompt processing ran 11.08× faster and token generation 16.36× faster than in the same stock VM, with prompt processing reaching 98% of bare-metal speed. Repeating with Google's Gemma 4 12B QAT Q4_0 (a 6.98 GB model) improved prompt processing by 7.20× and token generation by 14.54×, reaching 99.59% of bare-metal prompt speed and 94.82% of bare-metal generation speed.

The project is released today as a research release under the same permissive license as Lume and Cua, including source, build scripts, a capability probe, and raw benchmark logs so others can reproduce the results and map which Apple Silicon chips, macOS releases, and Metal workloads benefit.

Read original →

← Back to home