Open Source Hacker News (LLM)

Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

P2Pmixture-of-expertscolibridistributed inference

Lumabri is a new open-source tool for running huge mixture-of-experts (MoE) models over a peer-to-peer swarm, built on the Colibri engine. It is written in pure C with no dependencies, and unlike networks that pool GPUs from the few, it is designed to recruit from everyone: any machine, GPU or not, can join. The engine is built for CPU and SSD first; a GPU only makes it faster, never different, and output is byte-for-byte identical.

The system splits the model into byte ranges served by a 'maintainer' process, with a lightweight 'tracker' index of who holds which files. A chat client mounts the model through an LD_PRELOAD shim (liblumabri.so) that intercepts standard libc calls like open, fopen, opendir, and pread. Files appear as sparse local mirrors of their true size, so fstat, readdir, and the page cache work natively. A missing block is fetched from a peer, written to the mirror, and then the engine's own pread proceeds; a warm read is just a table lookup plus a normal local read, with no FUSE and no daemon on the read path. Each verified MiB is stored by SHA-256 in a local content-addressed store (~/.lumabri/cas) shared across checkpoints, so equal chunks are downloaded once and can rebuild a different sparse mirror without a byte server. The model is strictly read-only: writing returns EROFS, and a block no peer can serve yields a loud EIO, never silent zeros, with byte identity verified cold, warm, and with every peer dead.

Quick start is minimal: run `make`, then `./lumabri serve --model /path/to/model` on the machine that has a model, and `./lumabri chat --tracker <server-ip>:7300 --engines-dir /path/to/colibri/c` on a client that wants to chat. A `make fixture` builds a tiny synthetic model so the whole flow can be tested. The first answer is slower while the working set crosses the network, but afterwards the mirror in ~/.lumabri keeps serving even if the server goes offline. A no-argument terminal UI asks for the swarm address and operator public key, finds engines itself, and remembers them in ~/.lumabri/config; inside the chat, `/swarm` shows the live anonymous network (peers numbered, never named) and `/model` lists the models on the swarm for switching on the fly.

By shifting from a few GPU servers to any machine with storage, Lumabri makes large MoE models more accessible and resilient. Because the engine binary is never modified and byte identity is guaranteed, the swarm provides a distributed cache that behaves like a local disk, with the network deciding only where bytes come from, never which bytes.

Read original →

← Back to home