Cactus-Compute/needle3
Summary
Needle 3 is a compact AI foundation model optimized for edge devices like mobiles and wearables, offering tool calling, structured extraction, and text embedding in a single 8-29 MB file.
View Cached Full Text
Cached at: 09/21/26, 08:56 PM
Cactus-Compute/needle3 · Hugging Face
Source: https://huggingface.co/Cactus-Compute/needle3
A foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29 MB file, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.
Needle does three jobs, all of them on the device:
- Tool calls: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.
- Structured extraction: declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses, and extraction generalises to classification.
- Text embedding: the same model returns a vector for a sentence, so an app can search, match and route locally.
https://huggingface.co/Cactus-Compute/needle3#modelModel
Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on therelease page. The repo holds the 20-layerneedle3\.cact, theneedle3\.safetensorscheckpoint to fine-tune, and an engine per platform.
https://huggingface.co/Cactus-Compute/needle3#benchmarksBenchmarks
Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.
The interactive frontier plot, the architecture and the fine-tuning results are atcactuscompute.com/needle.
https://huggingface.co/Cactus-Compute/needle3#get-startedGet started
pip install cactus-needle
Try it in the browser atcactuscompute.com/needle; the Python package and the source are onGitHub.
import needle
@needle.tool
def get_weather(city: str):
"Get the current weather for a city."
return {"city": city, "temp_c": 27, "sky": "clear"}
agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]
Every turn returns one JSON object withfunction\_calls, the model’sreasoningand a calibratedconfidence; an off-topic request returns an empty list rather than a guess. The engine and the weights are fetched from this repo once and cached.
https://huggingface.co/Cactus-Compute/needle3#guidesGuides
- How to design tools for Needle 3: one tool per action, names users would say, formats in descriptions, constraints in the grammar, triggers.
- Leveraging Needle’s confidence: what the score measures, what the engine withholds, and routing on act, confirm or refuse.
- Structured JSON extraction with Needle: the record as the only tool, typed results, classification with enums.
- Fine-tuning Needle: the data format, the commands, reading the loss, sizing the dataset.
- Needle Python docs: the API, the response shape, the behaviour contract, system facts, tool retrieval, offline devices, environments, the CLI.
- What devices are supported on Needle: every platform folder, the CLI runner, the C API, the browser, WASI, air-gapped setup.
- The .cact format: the file the engine maps and reads in place, Cactus Quants at 2.125 bits per weight, and how to parse it yourself.
- Porting Needle 3: notes for writing your own runtime, the oracle to test against, the tensor order the container promises, the prompt on the wire, the ladder rule, retrieval with
needle\_embed.
https://huggingface.co/Cactus-Compute/needle3#customisationCustomisation
Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product’s tools, runs optimally on devices far smaller than the full model needs. Fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters.
The Python package fine-tunes with LoRA on the frozen base at the full 20 layers, thenneedle build \[\-\-layers N\]merges the adapter, slices any subnetwork from 2 to 20 layers and exports a 4-bit\.cactthat runs on the same engine. The 2-bit post-training and quantisation behind the shipped model, enriched with Cactus proprietary datasets, run on theCactus Platform.
https://huggingface.co/Cactus-Compute/needle3#deployDeploy
Every platform folder in this repo holds an engine under 1 MB that loadsneedle3\.cactat start.needle build \-\-platform <folder\> \[\-\-layers N\]fetches the engine and header and puts the weights beside them at any depth, or download the folder here:
./needle --model needle3.cact --tools tools.json --prompt "dim the living room to 30"
./needle --model needle3.cact --tools tools.json --serve
Tool schemas share the model context with the system prompt and conversation.needle\_initreturns the tokenized static-prefix length on success and fails when that prefix does not fit the model context. In the C API, callneedle\_last\_error\(\)after a negative return to get the measured prefix tokens and context limit. Reduce tool descriptions/schema text, split large catalogues, or declare more than five tools so Needle can keep only retrieved tools in the per-turn prefix. At completion time,max\_new\_tokensalso reserves context room, so larger generation caps leave less space for the current turn and history.
Thedevices guidecovers the runner flags, the C API, the WASI component and air-gapped setup.
https://huggingface.co/Cactus-Compute/needle3#citationCitation
Needle 3 is built by the Cactus Compute team. If you use it in your work, please cite:
@misc{needle3_2026,
title = {Needle: Foundation Tool-Calling Model for Tiny Devices},
author = {Ndubuaku, Henry and Mosoyan, Karen and Mroz, Jakub and Cylich, Noah and
Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
year = {2026},
organization = {Cactus Compute, Inc.},
howpublished = {\url{https://github.com/cactus-compute/needle}}
}
Reach out on[email protected]for partnerships, collaborations, synergies and deploying Needle in your product.
Similar Articles
Cactus-Compute/needle
Cactus-Compute releases Needle, a 26M parameter distilled model from Gemini 3.1, using a pure attention architecture optimized for on-device inference and local fine-tuning.
Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash
Cactus Needle 3 is a small, sliceable foundation model for automation tasks that runs on-device, achieving performance comparable to larger models on function calling and structured extraction.
@cactuscompute: Needle’s best-kept secret isn’t function calling. It’s structured extraction. Long-form text in → valid JSON out. Small…
Needle 2 is an open, 45M-parameter AI model for tool calling and structured extraction, optimized to run in browsers at 14MB with guaranteed JSON output via constrained sampling.
Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.
Cactus releases Needle 2, a 14MB agentic LLM for phones, wearables, smart home devices, and robots, achieving fast inference on low-end hardware and supporting structured extraction and fine-tuning.
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
Cactus Compute releases Needle 2, a 45M-parameter agentic LLM compressed to a 14MB binary for phones, wearables, smart home and robots, achieving 500+ tokens/sec on a Raspberry Pi 5 and running in 28MB RAM.