Pulling the tape…
Pulling the tape…
What it looks like
No readme screenshots yet
Videos of what it's doing
Cactus Needle - The 26M Function Calling Model
I Can't Believe This AI Model Fits in 14 Megabytes (Needle 2)
On-Device Coding Agents With Cactus
Why it's moving
Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
Hey HN, Henry here from Cactus. We open-sourced Needle, a 26M parameter function-calling (tool use) model. It runs at 6000 tok/s prefill and 1200 tok/s decode on consumer devices. We were always frustrated by the little effort made towards building agentic models that r
HenryNdubuaku · 776
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the su
HenryNdubuaku · 537
Show HN: Cactus v2 – On-device AI with cloud fallback
Hi HN, Roman and Henry here from Cactus ( https://github.com/cactus-compute/cactus ). We just shipped the biggest upgrade to our on-device inference platform: - Built-in model confidence-based routing to hand off inference runs to the cloud - Converter for any
rshemet · 1
Repo
cactus-compute
Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
Stars
12k
Velocity
—
Change
—
Heat
70