August 12, 2026

Metal, mayhem, and a comeback?

Automatic1111 for Apple metal, 40% speed up sd1.5

Apple users got a big speed boost, but the comments are roasting the app's comeback

TLDR: Automatic1111 got a real speed boost on Apple machines, cutting many image generations down from roughly 8 to 10 seconds to about 3 to 7. But commenters stole the spotlight by arguing over whether this once-popular app is enjoying a comeback — or just getting polished after the crowd already moved on.

A longtime image-generation app just got a surprisingly juicy speed glow-up on Apple computers, with the creator saying some jobs that used to take around 8 to 10 seconds now finish in roughly 3 to 7. In plain English: the same old tool, same workflow, same favorite add-ons and styles — just much snappier. That alone would be nice news for Mac users, but the real show was the comment section, where people immediately turned this into a debate about whether anyone should even care anymore.

One camp was impressed by the effort to make an aging favorite feel fast without forcing users to switch apps. Another camp basically said, "Cute, but why are we optimizing a relic?" One commenter called Stable Diffusion 1.5 “insanely old,” while another declared Automatic1111 a total throwback and mourned the era “back before the massive amount of slop” when AI images felt fun instead of exhausting. Ouch. Then came the migration discourse: one user flatly claimed people had already moved on to Forge Neo, ComfyUI, and Maestro, turning the whole post into less of a victory lap and more of a nostalgia-soaked reunion tour.

And yes, the jokes were flying. The opening “Trigger warning: AI content” set the tone perfectly: half meme, half eye-roll, fully internet. Even the practical questions had attitude, with users asking whether the speedup still holds when using LoRAs — the custom style add-ons many people actually care about. So while the technical upgrade is real, the comments made it clear the bigger story is a familiar one: faster software is nice, but the community is still arguing over whether this old king should have stayed retired.

Key Points

  • The article reports observed generation-time reductions for Automatic1111 on Apple hardware, from roughly 8–10 seconds to 3–7 seconds on M3 Pro and from 13–20 seconds to 8–10 seconds on M1 Mac Mini for the author's workloads.
  • The optimization effort was designed to preserve the existing Automatic1111 workflow, including the WebUI, checkpoints, LoRAs, samplers, extensions, API, and prompt syntax.
  • The target workload is specifically Stable Diffusion 1.x using DPM++ SDE, Karras, 5 steps, CFG around 1.15, common image sizes, FP16 UNet on MPS, and FP32 VAE by default.
  • A custom Metal Flash Attention path was added selectively for SD 1.x attention shapes that tested faster, while other cases fall back to PyTorch SDPA for broader compatibility.
  • The article identifies repeated MPS command buffer commits after each attention call as a major source of overhead and addresses this by integrating the Metal extension into PyTorch's current MPS stream.

Hottest takes

"Trigger warning: AI content." — jurgenburgen
"SD 1.5 is so insanely old" — vanillax
"A1111 is such a throwback" — yogorenapan
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.