---
title: "Perplexity open-sources Lily, a fast local inference engine for Apple Silicon"
url: https://www.parallelquant.com/posts/perplexity-open-sources-lily-a-fast-local-inference-engine-for-apple-sil-4d2ec4
source_name: "MarkTechPost"
source_url: https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon/
published: 2026-09-03T06:57:07.000Z
topics: ["open source", "llms"]
publisher: "Parallel Quant"
---

# Perplexity open-sources Lily, a fast local inference engine for Apple Silicon

*2026-09-03 · Source: [MarkTechPost](https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon/)*

Perplexity open-sourced Lily, a Rust-based inference engine with custom Metal kernels built specifically to run the Qwen3.6-35B-A3B model on Apple Silicon. In testing on a 40-core, 128GB Apple M5 Max chip, it reached roughly 1.23x the prefill throughput and 1.35x the decode throughput of MLX-LM.

**Why it matters:** This adds to a growing ecosystem of specialized, hardware-tuned local inference engines that let large open models run efficiently on consumer devices instead of the cloud, reducing reliance on API providers for some use cases. Perplexity releasing it as open source also signals continued investment in on-device AI as a complement to its cloud search products.

**Topics:** open source, llms

---
Read the original: https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon/
Canonical: https://www.parallelquant.com/posts/perplexity-open-sources-lily-a-fast-local-inference-engine-for-apple-sil-4d2ec4
Published by Parallel Quant — https://www.parallelquant.com
