Kimi K3-256k

Same brain, smaller memory bill? Fans cheer while skeptics squint

TLDR: Kimi added K3-256k, a cheaper version of its coding model for people who don’t need massive chat history. The community is split between celebrating the savings, questioning whether it’s really new at all, and doing the usual excited internet yelling.

Kimi just dropped K3-256k, a new option that promises the same results within a smaller memory limit while using about half the quota of the big 1M version. In plain English: if you don’t need the giant, extra-long chat history, this version is supposed to be the cheaper daily driver for coding, question-answering, and ordinary work. The catch? It doesn’t support video, and switching can get messy if your current conversation is already too huge, so users may need to “compact” things first like they’re stuffing an overpacked suitcase.

But let’s be honest: the real show is in the comments. One camp instantly declared this a quiet win, with one user basically shrugging, saying they stay under 200k anyway. That vibe? Very “works for me, save me the credits.” Another group immediately went full detective mode, asking the spicy question: is this actually a new model, or just the same one wearing a cheaper price tag? That sparked the mini-drama of the thread, with people quoting the docs back like courtroom evidence and asking if this is just a smaller context window, not some magical upgrade.

And then there was the pure chaos energy: “omg! new model!!” A tiny comment, but honestly the perfect internet reaction. So the mood is split between practical savers, skeptical nitpickers, and hype goblins thrilled by anything shiny and new. Classic tech community behavior, no notes.

Key Points

  • Kimi Code introduced `k3-256k`, a 256k-context version of Kimi K3 that the documentation says provides the same results within 256k context while using less quota than `k3`.
  • The `k3` model is described as Kimi’s flagship coding model with 2.8T parameters and up to a 1M context window, while `k3-256k` is limited to 256k context and does not support video input.
  • When switching from `k3` to `k3-256k`, users may need to compact session context first, especially if the session exceeds 256k or contains video files.
  • The page says switching from `k3-256k` to `k3` can be done directly to avoid information loss near the 256k limit, and that this switch currently does not affect cache.
  • Model access depends on membership tier: K3 requires Moderato or above, 1M context and HighSpeed require Allegretto or higher, and capability mismatches can return a 401 error.

Hottest takes

"stay below 200k context anyway" — madihaa
"just an API level change" — hawtads
"omg! new model!!" — ibuildproducts
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.