llama.cpp
The post highlights pairing llama.cpp with a local coding agent. Users can run `llama serve` to host a model, install the `pi-llama` plugin via `pi install git:github.com/huggingface/pi-llama`, and then launch Pi, which automatically discovers the local model. This setup requires no configuration or API keys, and files stay on the machine with requests never leaving it.
llama.cpp is positioned as optimized for any hardware, running the same binary, models, and hand-tuned kernels across a broad range of GPUs and CPUs. The listed hardware includes Apple Silicon (M Ultra, M Pro, M Max), NVIDIA GPUs (RTX 5090, RTX 4090, RTX 3090, H100, A100, T4, B200), plus CPU, Jetson, MI300, Radeon RX, Intel Arc, and DGX Spark.