---
title: "Cactus Compute ships 45M-parameter tool-calling model in 14MB"
url: https://www.parallelquant.com/posts/cactus-compute-ships-45m-parameter-tool-calling-model-in-14mb-340e3d
source_name: "MarkTechPost"
source_url: https://www.marktechpost.com/2026/08/13/cactus-compute-needle-2-45m-parameter-tool-calling-model/
published: 2026-08-14T05:45:47.000Z
topics: ["open source", "products"]
publisher: "Parallel Quant"
---

# Cactus Compute ships 45M-parameter tool-calling model in 14MB

*2026-08-14 · Source: [MarkTechPost](https://www.marktechpost.com/2026/08/13/cactus-compute-needle-2-45m-parameter-tool-calling-model/)*

Cactus Compute released Needle 2, an open 45-million-parameter model specialized for tool calling, on-device use, and structured data extraction. The full model ships as a single 14MB binary and runs a complete session in about 28MB of RAM, requiring no GPU or NPU. It leads both splits of the Seal-Tools benchmark among models targeting this class of hardware.

**Why it matters:** Tool-calling and agentic workflows have mostly assumed access to large cloud models, but Needle 2 shows a sub-50M-parameter model can handle structured tool use on ordinary CPUs. That matters for edge and embedded agents — IoT devices, browser extensions, offline apps — where latency, cost, or connectivity rule out calling a frontier API. It's part of a broader trend of shrinking specialized capabilities out of general-purpose LLMs into tiny, purpose-built models.

**Topics:** open source, products

---
Read the original: https://www.marktechpost.com/2026/08/13/cactus-compute-needle-2-45m-parameter-tool-calling-model/
Canonical: https://www.parallelquant.com/posts/cactus-compute-ships-45m-parameter-tool-calling-model-in-14mb-340e3d
Published by Parallel Quant — https://www.parallelquant.com
