Skip to content
Where new products land first
M

Maple-Preview

An open ternary 20B-A1B reasoning model built to run on your own device.

Other17d ago
Category
Other
Language
English
Listed
17d ago

Ternary weights mean every parameter is stored as one of three values, and the shrinkage is the whole point: Maple-Preview is an open-source 20B-A1B reasoning model small enough to run inference on hardware you already own rather than a rented GPU. Mixture of experts routing keeps roughly 1B parameters active per token, so decoding stays cheap next to the total parameter count.

It comes from DeepGrove, an independent lab whose stated aim is frontier intelligence that runs on any device. The lab's own figure is 218 tokens per second on a 16 GB M4 Mac mini, using ternary kernels together with its FlashHead work; the Hacker News submission carried a headline claiming 120 tokens per second on an iPhone. Published 2026-08-04.

Like this product?