Show HN: Local text, image, video, music and 3D from one CLI, no Python

One app claims it can make pics, music, video and even 3D — commenters are split

TLDR: mere.run says one local tool can handle text, images, music, video and 3D on your own computer, no extra Python setup required. Commenters zeroed in on two things: one user’s impressive speed test numbers, and another user dunking on the confusing title.

A new project called mere.run is pitching a very big dream in very simple packaging: one command-line tool that does almost everything creative on your own machine. Images, chat, speech, music, video, even turning a picture into a 3D object — all without needing Python, which is the programming setup many AI tools usually drag in. On paper, that sounds like catnip for people tired of messy installs and cloud subscriptions. And in the comments, one early tester came in swinging with the receipts, posting a full speed list from an Apple M4 Max machine: image making in under a minute, sound effects in seconds, music in 15 seconds, and short video with audio in under three minutes. That kind of benchmark post is basically comment-section currency.

But the thread also got a tiny splash of title drama almost immediately. One commenter bluntly complained that the headline "makes no sense in isolation," which is classic Hacker News energy: before the crowd debates the future of local AI, somebody has to roast the wording. That tension really sums up the mood here. On one side: "finally, one tool to rule them all" excitement, especially for people who want everything to stay local on their own computer. On the other: a skeptical eyebrow at whether the pitch is too broad, too buzzword-heavy, or just confusingly presented. The funniest part is that the biggest flex in the thread wasn’t marketing copy — it was a nerdy speedrun of real-world timings, which landed like a mic drop for believers and fresh ammo for skeptics.

Key Points

  • mere.run is presented as a local-first inference runtime for Apple Silicon and headless Linux with one public CLI covering multimodal generation, analysis, training, and API serving.
  • The article provides setup and usage examples, including machine inspection, model capability checks, model downloads, offline guides, and an initial image-generation workflow.
  • Supported image, text, and vision functions include image generation, LoRA training, chat, code generation, embeddings, anonymization, OCR, grounding, segmentation, tracking, pose, and optical flow.
  • The runtime also claims support for depth, geometry, 3D reconstruction, video generation, world serving, music generation, sound effects, speech synthesis, transcription, and diarization using a range of named model technologies.
  • Automation and deployment features described include OpenAI-compatible local API serving, resident model pooling, memory guards, typed preflight actions, machine-readable progress, replayable plans, immutable typed graphs, and local or remote workflow execution via SSH and relay executors.

Hottest takes

"some real numbers on my m4 max" — sawfwair
"video gen w/ audio ... 4s,768x512: 2m 48s" — sawfwair
"the title makes no sense in isolation" — ks2048
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.