Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

Your phone might be getting a brain boost — but the comments are already calling cap

TLDR: DeepGrove says its new open-source AI can run unusually fast on an iPhone and even solve elite math problems, a big step toward powerful offline AI on consumer gadgets. Commenters were impressed by the speed but immediately argued over whether it’s genuinely smart or just confidently wrong at turbo speed.

A tiny internet stampede broke out after DeepGrove showed off Maple-Preview, an open-source AI model it says can run shockingly fast on everyday devices — including an iPhone — while still tackling seriously hard math. On paper, it sounds like sci-fi: a model small enough to fit in about 5.3 GB, fast enough to spit out text at 127 tokens per second on an iPhone and more than 200 on a Mac mini, and flashy enough to solve an International Mathematical Olympiad problem. Unsurprisingly, some readers were instantly sold. One early reaction was the pure, wholesome “This is cool!”, while another cheered, “Edge is edging closer!” — basically the local-AI fan club declaring the future has entered the chat.

But this being the internet, the applause lasted about three seconds before the side-eye began. The biggest complaint? Speed is nice, but what if it’s fast and wrong? One commenter said the model “hallucinates knowledge quite aggressively,” accusing the online demo of hiding the problem behind web search tools. Another was even harsher, begging so-called “small” AI tools to stop being “confidently very incorrect.” And then came the bonus drama: one user noticed several low-karma accounts praising the launch and flatly called it “Suspicious.” So now the real story isn’t just that your phone may soon run powerful AI offline — it’s that the crowd is split between “holy wow” and “nice demo, but is this thing bluffing at light speed?”

Key Points

  • DeepGrove introduced Maple-Preview as an open-source 20B-A1B ternary-weight reasoning LLM.
  • The article reports Maple-Preview running at 218 tokens/s on a Mac mini M4, with a 5.31 GB checkpoint and 131,072-token context window.
  • DeepGrove claims Maple-Preview is 5–16× faster than models such as Gemma 4, Qwen3.5, and gpt-oss on the referenced hardware.
  • The article includes a full proof of IMO 2024 Problem 1 and says Maple-Preview achieved 7/7 on that problem at 281.5 tokens/s on a MacBook Pro with M5 Pro.
  • For mobile inference, the article says Maple-Preview runs at 127 tokens/s on iPhone versus 9.6 tokens/s for Bonsai 27B on the same prompt.

Hottest takes

"hallucinates knowledge quite aggressively" — Havoc
"Edge is edging closer!" — zooloo99
"Suspicious." — jsphweid
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.