Cactus Compute has released Needle 2, an open tool-calling model with 45 million parameters that fits into a single 14MB binary and occupies roughly 28MB of RAM during a full session.
In this article
The team trains and deploys the weights at CQ2-bit using their Cactus Quants system. The model runs inside a proprietary C++ engine, meaning there is no separate runtime to install and no downloads required at inference time. Reported decode speeds hit 500 tokens per second on a Raspberry Pi 5, between 400 and 1,500 tokens per second on Meta Quest 3S and Apple Vision Pro, and 300 to 700 tokens per second on phones priced under $200.
The design premise is narrow and stated plainly by the team: mapping a messy sentence onto a typed function signature needs no world knowledge and no open-ended prose. That framing is why 45M parameters are enough here, and why the model targets hardware with no GPU and no NPU.
Is it deployable?
Yes, Needle 2 ships as prebuilt binaries and a static library for macOS, Linux (x86-64, ARM64, ARMv7, RISC-V, MIPS32el), Windows, Android, iOS/watchOS/tvOS, and WebAssembly. Cactus says Pebble already runs Needle locally in the Index 01 app for offline voice actions.
- Which companies: Any team shipping firmware or apps on constrained hardware. Seed-stage wearable and IoT startups, mid-market consumer-electronics OEMs, robotics teams, and large device makers needing an offline fallback. Cloud-first SaaS teams gain less.
- Industries: smart home, wearables, low-end mobile, automotive in-cabin control, service robotics, retail kiosks and POS, routers and IP cameras, and regulated settings where audio cannot leave the device.
- Applications: voice-to-action on screenless devices, offline appliance control, receipt and invoice field extraction, enum tagging, and local routing that escalates to the cloud only on low confidence.
Architecture: Simple Attention Network
Needle 2 uses what Cactus team calls a Simple Attention Network. The recipe replaces the FFN with a Hadamard MLP, keeps GQA attention, adds engram key-value memory from hashed n-gram tables, and uses multi-lane hyper-connections. The network is 27 layers and 512 wide. The underlying study is on arXiv as A Controlled Study of Attention-Only Transformers.
Pretraining used a proprietary 115B-token corpus, with 38B tokens of post-training. The research team notes LFM2.5-230M was pretrained on 19 trillion tokens.
Needle 2 spends 70 MFLOPs per token, with 35M of 45M parameters matmul-active. LFM2.5 230M spends 460, FunctionGemma 270M spends 540, and Apple FM sits near 6,000.
Engine, grammar, retrieval, and confidence
Weights never decompress into RAM. The 2-bit codes expand inside vector registers and fuse into integer dot products, so the arithmetic path stays int8. One binary probes the CPU at startup and selects a kernel tier: SDOT, NEON, AVX2, RISC-V vectors, wasm SIMD, or scalar.
A byte-level grammar compiled from your JSON schemas constrains every emitted token. Because the matcher knows which tokens are legal before logits exist, the engine skips up to 98% of the vocabulary projection on structural tokens.
Attention uses a 256-token sliding window, and the system turn plus tool declarations are pinned as KV sinks. Memory stays near 28MB regardless of conversation length.
Declare five or fewer tools and they render directly. Above five, a contrastive retrieval head embeds each schema once, scores the query per turn, and admits only the top five. Unselected tools are unreachable, not merely unlikely.
Every response carries a confidence value, the minimum of a calibrated post-hoc head and the decoding probability of the call tokens. Off-topic requests return the empty call []. The contract is a threshold: act above it, re-ask or escalate below it.
Evaluation
Cactus team evaluates on five public function-calling benchmarks using ordered strict exact match, where names, call order, and every argument must match. Needle 2 runs end-to-end through the shipped engine at CQ2-bit with retrieval on; baselines run f16 under vLLM.
| Benchmark | Needle 2 (CQ2) | LFM2.5 230M | FunctionGemma 270M | Apple FM |
|---|---|---|---|---|
| Mobile Actions (961) | 63.7 | 69.1 | 64.0 | 57.6 |
| DroidCall (200) | 17.0 | 11.0 | 17.5 | — |
| Seal-Tools in-domain (700) | 32.6 | 26.9 | 16.3 | — |
| Seal-Tools OOD (654) | 28.7 | 17.0 | 15.6 | — |
| BFCL v4 single-turn (3,641), overall | 42.6 | 60.8 | 46.1 | 61.7 |
Needle 2 leads both Seal-Tools splits and posts 98.3 function-name accuracy on Mobile Actions. It trails on BFCL v4, which Cactus attributes to distribution: its corpus is consumer device actions, not general or enterprise APIs. Well-formed output rate across the 3,641 BFCL rows is 93.4. The team states two asymmetries upfront: f16 baselines favor them, and task specialization favors Needle.
What it means
For the people building things on limited hardware, this removes the need for a separate inference server. You can ship a single binary that handles the logic locally, keeping data private and latency low. The strict grammar approach ensures the output is usable code or commands without needing a post-processing filter to fix syntax errors.
Developers on embedded systems can now integrate tool-calling capabilities directly into firmware, relying on the model to handle the complexity of mapping natural language to specific function signatures without requiring a high-end processor.




