Kimi-K3 Releases on HuggingFace 7/27

(huggingface.co)

163 points | by nateb2022 2 hours ago

18 comments

  • NitpickLawyer 1 hour ago
    This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

    Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.

    Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)

    • walrus01 12 minutes ago
      It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.

      Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.

      Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.

      I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.

      • sandworm101 8 minutes ago
        Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.
        • walrus01 4 minutes ago
          The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).

          But the memory bus speed is fully committed when generating tokens or thinking.

    • woctordho 45 minutes ago
      Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.

      I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.

      Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.

    • vb-8448 13 minutes ago
      > Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

      No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.

      Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

      • amelius 9 minutes ago
        > Without training cost you can infer only the marginal cost of serving this kind of models.

        Which is by far the most interesting number of the two.

        > Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

        If you get close in output quality, then does that matter?

    • dist-epoch 48 minutes ago
      > if "labs are subsidising tokens on API pricing"

      > SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%

      Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...

      https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...

      https://finance.biggo.com/news/02d45650-b569-4d12-b44d-8d6d8...

      • NitpickLawyer 45 minutes ago
        Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
      • knollimar 23 minutes ago
        That numner blends in training or no?
  • KronisLV 16 minutes ago
    I feel like most hardware to run LLMs on is shaped wrong for individuals.

    It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming like kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful.

    Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.

    How unfortunate.

  • gorgmah 34 minutes ago
    We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers

    I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong

    • markasoftware 23 minutes ago
      Press x to doubt on the 45% number. The cheaper providers on open router are fp4 vs fp8 for official zai. There are some cheap fp8 ones (like novita) but the ui makes it seem like it's a temporary promotion, with their normal prices being almost equal to official zai (idk much about open router so not really sure what's going on with these discounts)
    • wanick 16 minutes ago
      [dead]
  • maelito 54 minutes ago
    Did someone run censorship and political bias tests on this ? Must be interesting.
    • walrus01 6 minutes ago
      As a completely one person, single sample anecdote, the 'heretic' uncensored Q8 GGUF variants several people have published or Qwen 3.5-122, 3.6-27B and 3.6-35B-A3B will very happily discuss just about any controversial topic that the CCP hates. Including lots of things that would get you thrown into prison if you published them in Mandarin on the domestic Chinese internet.

      https://github.com/p-e-w/heretic

  • davidkunz 57 minutes ago
    This is historic. For the first time, an open-weights LLM is right at the top.

    We won't be able to run this ourselves, but many providers can.

    • embedding-shape 19 minutes ago
      > For the first time, an open-weights LLM is right at the top.

      Hmm, not quite true, I think that honor, for better or worse, goes to OpenAI. When they released GPT2 (or GPT1 for that matter) is was quite literally the SOTA in the ecosystem when it was released.

  • wwwhizz 1 hour ago
    That would be 7/27.
    • broodbucket 55 minutes ago
      20270727 if we're improving dates :)
      • c7b 45 minutes ago
        270727 if we want to save tokens :)
        • darkwater 34 minutes ago
          Y2K100 bug for you, then ;)
      • fragmede 42 minutes ago
        ISO 8601 ftw
      • m00dy 47 minutes ago
        Looks like another LeetCode problem about checking for anagrams.
    • ra 58 minutes ago
      or 27/7 for the rest of the world
      • LeonM 39 minutes ago
        No, 27-7 for the rest of the world.

        The separator is often the only way to distinguish American notation from ISO, so please use a dash for dd-mm-yy and a forward slash for mm/dd/yy

        • casparvitch 36 minutes ago
          Have never seen 27-7, as someone in rest-of-world
        • pavo-etc 4 minutes ago
          This is so confidently wrong it's funny. In Australia dd/mm/yy is the default.
        • distances 25 minutes ago
          I've never seen dd-mm-yy. It's usually dd.mm.yy, dd.mm.yyyy, or yyyy-mm-dd, with some dd/mm/yy sprinkled in for general confusion.
        • ashwoods 12 minutes ago
          Not really. Spain's traditional format is dd/mm/yyyy with slashes. This applies for a good chunk of Europe. Germany/Austria uses dot, I think nordic countries embraced the dash. While you might see more adoption in offical/digital contexts, I just double checked a few popular spanish websites, all slashes.
        • letier 37 minutes ago
          Or 27.7. in some other places.
        • kuboble 30 minutes ago
          Nope,

          There are quite some countries around the world using d/m/y

          https://en.wikipedia.org/wiki/List_of_date_formats_by_countr...

          Algeria, Belgium, Brazil, Chile...

    • linzhangrun 38 minutes ago
      China uses YYYY/MM/DD, which is logical.
      • _zoltan_ 34 minutes ago
        the only logical format.

        signed: a hungarian :)

        • egeozcan 10 minutes ago
          For me, the only format that doesn't make sense is the MM/DD/YYYY, together with its rarely seen worse sibling, MM/DD/YY (07/27/26).
        • JSR_FDED 15 minutes ago
          lpszReleaseDate ;-)
  • colortiles 23 minutes ago
    This looks really promising. Excited to see where this goes. Looking forward to trying it out!
  • pmg1991 34 minutes ago
    Hoping no issues on Huggingface due to download rush.
    • someguyornotidk 8 minutes ago
      For huge models like these, the only reasonable way to host them is via torrents. I don't understand why hf doesn't offer this as an option.

      Linux distributions got this right: Offer both HTTP and Torrents. Let the user decide.

  • rs38 23 minutes ago
    is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
    • embedding-shape 21 minutes ago
      Yeah, why not. Toughest part is running the hardware so you can create the traces for downstream training, but once over that hump, nothing would stop you from doing that no.
  • CodeCompost 58 minutes ago
    Why is there a countdown?
    • broodbucket 55 minutes ago
      You're not having a party?
      • InsideOutSanta 26 minutes ago
        I think it's shameful that Moonshot isn't providing us with party kits like Microsoft did with the Windows 7 Launch Party kit. How am I supposed to properly celebrate this without fun Kimi-themed quizzes for my guests?
    • mythz 47 minutes ago
      It's a release party
  • sreekanth850 40 minutes ago
    how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?
  • marvinLuck 52 minutes ago
    The weightings should be released on July 27.
  • throwaw12 39 minutes ago
    Strange communists, giving away such an expensive model to the public.

    On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"

  • minimaxir 21 minutes ago
    ...does Hugging Face have enough bandwidth to let people download en masse however much file size a 2 trillion parameters model is?
  • m00dy 1 hour ago
    There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.
    • Iolaum 54 minutes ago
      As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
    • torginus 47 minutes ago
      I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.

      I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)

  • extra-AI 19 minutes ago
    [flagged]
  • youre-wrong3 39 minutes ago
    [dead]
  • threerouter 2 hours ago
    [dead]