Open Source Hacker News (LLM)

From the creator of Redis; run LLM locally with ds4

antirezds4local inferencequantizationDeepSeek

ds4 (DwarfStar 4) is a narrow inference engine written in C by antirez, the creator of Redis, aimed at local frontier inference on high-memory Macs plus CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, covers both text and vision models, and ships local APIs, a CLI and a native agent in a single stack under the MIT license, targeting C/Metal/CUDA/ROCm backends.

The project's core premise is that DeepSeek V4 Flash — a 284-billion-parameter mixture-of-experts model normally served remotely — can instead be run locally. ds4 does this via asymmetric quantization that compresses the routed experts while keeping critical shared paths precise, so the model stays compressed rather than "lobotomized" and becomes practical on high-memory machines. The engine is deliberately not a generic GGUF runner: it follows a small, opportunistic set of model families and validates each supported layout end to end against official model outputs.

Three design choices define the stack. First, asymmetric 2-bit quantization (with imatrix) compresses routed experts while preserving critical shared paths, which is how the supported routed-MoE builds fit their target machines. Second, the KV cache is treated as a "disk citizen": long prefixes are saved to SSD and resumed by prompt hash, so restarts don't require a full re-prefill. Third, one engine exposes three interfaces — ./ds4 for interactive chat, ./ds4-server for OpenAI- and Anthropic-compatible local APIs, and ./ds4-agent for persistent coding sessions.

The architecture also includes SSD streaming, tensor parallelism, session batching, DSPARK + MTP and vision input, with memory classes spanning personal to distributed setups. Running it takes three steps: clone github.com/antirez/ds4, fetch weights with ./download_model.sh ds4f-q2, build for your backend (make, or make cuda-spark), then start ./ds4 or ./ds4-server --ctx 100000. Generic GGUF files are explicitly not the target — only the project's own GGUFs.

A hardware fit check positions V4 Flash Q2 as the baseline. At 128 GB of memory, GLM 5.3 Q2 and Qwen Q4 also fit, while V4.1 Q2 streams rather than sitting fully resident.

Read original →

← Back to home