Model Releases Hacker News (LLM)

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

needle2on-device LLMfunction callingedge AI

Needle2 is an open 45M-parameter model for tool calling, device use, and structured extraction. The entire model compresses to a single 14MB binary using CQ2-bit compression with Cactus Quants, and runs in its own engine with a peak session RAM of about 28MB. On tool-call and mobile-device benchmarks, it trades wins with FunctionGemma 270M, LFM2.5 230M, and Apple FM, despite being 5–70x smaller and using 2-bit weights rather than f16. Decode speed reaches 500 tokens/sec on a Raspberry Pi 5, 400–1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and 300–700 tokens/sec on sub-$200 phones such as Samsung A-Series. It can even run on newer microcontrollers like the ESP32-S3.

The design bet is that edge AI should target devices under $200, not just PCs and Macs: there are over 21 billion connected IoT devices versus roughly 1.5 billion PCs, and most phones in emerging markets ship below $200. Around four in five edge devices cost under $200, and Needle is built for those—no GPU or NPU, just a few hundred MB of RAM. The core task is function calling: devices expose abilities as functions with typed parameters, and the hard part is mapping a messy sentence to the right function and values. That framing needs no world knowledge or open-ended prose, so 45M parameters suffice where chat models need billions.

For extraction, the schema is the interface: a schema plus a paragraph returns typed fields, an enum field acts as a classifier, and an array field collects a list in one call. This is enforced as a contract rather than a convention—every turn returns a call envelope, the empty call is the refusal, and a byte-level grammar compiled from the declared schemas constrains every token. The grammar carries the syntax, so all 45M parameters are spent on choosing functions and grounding arguments in the user's words.

Since no small model covers everything, Needle returns a learned confidence score for every response, and off-topic requests yield the empty call (refusal). Above the threshold it acts; below it, it re-asks or escalates to the cloud. Most device requests are routine control, so escalation stays rare. The model is released under Apache 2.0 with weights on Hugging Face, and the playground lets users test it for wearables, robots, smart homes, phones, and automotive.

Read original →

← Back to home