Qwen/Qwen3.8-2.4T-A95B

AI fans are freaking out over a giant new model that’s brilliant, absurd, and hilariously huge

TLDR: Qwen just released one of the biggest open AI models ever, promising top-tier performance and huge memory for long tasks. The community reaction is a mix of awe and side-eye: some call it a historic leap, while others joke that it’s so massive only rich enthusiasts can really run it.

Qwen has dropped Qwen3.8-2.4T-A95B, an open AI model the company is pitching as its biggest, smartest public release yet — built for coding, research, long tasks, and giant text windows that can swallow whole novels. On paper, it looks like a monster: stronger performance, more tool use, and a version on Qwen’s cloud with even more bells and whistles. But in the comments, the real headline was much simpler: this thing is enormous.

The community instantly turned the launch into a size contest mixed with a reality check. One of the loudest reactions came from people staring at the storage numbers in disbelief. “A ~5TB model,” one user gasped, and that pretty much set the tone. Another said the benchmark chart looked “almost too good to be true,” which kicked off the classic AI thread drama: is this a breakthrough, or are people being dazzled by shiny scoreboards again? Others were stunned that a heavily compressed version could still fit into hardware a wealthy hobbyist might actually buy, prompting excited talk that this could bring premium-level AI closer to regular users.

Still, not everyone was ready to throw confetti. Skeptics called it a “chonker” and warned that, at launch, it may be harder to run than rival models unless someone with very deep pockets steps in to slim it down. So the vibe is deliciously split: historic open release or impractical flex? Either way, the comments have already crowned it the internet’s newest beautiful beast.

Key Points

  • Qwen released Qwen3.8-2.4T-A95B as an open-weight post-trained causal language model in Hugging Face Transformers format.
  • The release is compatible with vLLM, SGLang, and TokenSpeed, while managed inference is offered through Qwen Cloud.
  • Qwen describes Qwen3.8 as its most capable open-model generation so far and the first open release of a Qwen-Max-class model.
  • The model has 2.4 trillion total parameters, 95 billion activated parameters, 92 layers, a Mixture-of-Experts design, and native 262,144-token context extensible to 1,010,000 tokens.
  • The article includes benchmark results across coding, general agent, and general capability evaluations such as SWE-bench Pro, PaperBench, Automation-Bench, GPQA Diamond, and LongBench v2.

Hottest takes

"A ~5TB model" — volf_
"The card looks almost too good to be true" — PunchTornado
"Bit of a chonker" — NitpickLawyer
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.