[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model
By SemiAnalysis
Summary
Topics Covered
- Google should feel embarrassed right now
- Closed-source AI margins are mindboggling
- Western open-source is shockingly behind
- The coding harness is now part of the product
- We're still incredibly early
Full Transcript
All right, Max, we're going to do a podcast. We're going to talk everything
podcast. We're going to talk everything about Kimmy K3 and maybe some other models that just came out. How you
doing?
Doing great. Uh, looking forward to it and thanks for having me, Jordan.
I'm not having you. Thanks for having me. Okay.
me. Okay.
On the docket, uh, let's say Kimmy K3 hot takes. Is this the third best model
hot takes. Is this the third best model in the world? impact on OpenAI, an entropic architecture changes, personal usage that we've had so far, what we think about their open source strategy,
and maybe more. Um, all right, Max, quick hot take. Is this the third best model in the world right now?
Um, I think the answer is a clear yes.
uh people people love shooting on benchmarks and I think benchmarks definitely have their problems but I think sort of if you take a composite of all the main benchmarks and just look at
the model rankings uh they have been directionally correct uh over time uh and I think if you look at that composite today there's like a pretty
clear top three with Fable uh Soul 5.6 six and now commun 3 and there's for just like always above everyone else uh which includes of course other open
source guys like you know DeepCm and whoever but also like very notably it includes Google and Meta and SpaceX. Um
which I think is honestly it's an extremely impressive and a very remarkable feat from the moonshot guys.
Uh Google in particular I think should feel incredibly embarrassed right now.
uh that at one point guys remember as as recent as like you know November December 2025 everyone thought that like the clear AI big three was Google
Anthropic and open AI and even when I talked to like you know boomers today they still seem to think that the clear top three is Google Anthropic and Open
AAI and it's just like clearly not the case anymore. Um so yeah I'd say
case anymore. Um so yeah I'd say definitely third best mall in the world.
I do think it's overall still worse than Fable and Soul 5.6. Uh kind of funny that they explicitly said that in their like model release blog post. Um maybe
it's some like good oldfashioned Chinese humility. Maybe it's like they don't
humility. Maybe it's like they don't want to, you know, incur scrutiny from the US government or anything because obviously there were some uh like delays
the Fable 56 release. Um but overall very impressed with them all.
Yeah, they in the limitation sections of the blog post they said despite being a highly competitive model overall K3 nonetheless exhibits a noticeable gap and user experience
compared with Fable 5 and GPT 5.6.
So my experience using this personally is that it is good. It's really slow which is really annoying. It's motivated
me to try um open source harnesses for the first time. And so I feel like I'm learning more about the harnesses than I am about the models because frankly all these models are like good enough to do
the basic work that I've been doing so far. I I can't really find a lot of
far. I I can't really find a lot of complicated stuff that it can't do which in is in and of itself is a bit of a feat. Um here here's here's my hot take
feat. Um here here's here's my hot take for me. This might be the second best
for me. This might be the second best model in the world right now because every time I try and do something meaningful with Fable, I get rejected and I get sent down to Opus. And
even though I don't know if this is better than Opus, um it is less annoying to not get rejected whenever I'm trying
to do something. Um however I am getting uh I'm not getting rejected when I use my API key and pay for per token but I
am hitting limits whenever I try and uh you know just use the web console or deep research or like the coding plan.
Um, I haven't used a coding play, but some other guys that some analysis have.
And so it leads me to be like, what what is the strategy here? Because these guys clearly just do not have enough GPUs to serve the demand that they're got for this model. And previously that was
this model. And previously that was solved by an open source strategy where they just drop the weights and then other people serve it and serve that demand. But they haven't dropped the
demand. But they haven't dropped the weights yet. So I think they said waits
weights yet. So I think they said waits in 10 days or something.
What do you think the strategy is for the delay between announcement of the model, the API being available and no weights yet?
Yeah. I mean, I think to be clear, this is all just pure speculation on my part.
Uh, but I think one big reason is like they need to give the VLM and SU Lane guys enough time to make sure they can serve like this model performantly. Um
because if they just like dropped it today, you have all this hype. Uh but
then like everyone else serving the model is only giving you like 20 tokens a second or something. That's probably
really bad uh for their brand is really sort of like capture uh so they have this like incredible opportunity to get a bunch of huge PR and a bunch of adoption and that was sort of like uh
kneecap it a little bit. Um I think another possibility is that they are actively talking to the together fireworkses nebuses like cories of the
world to figure out how uh they can sign some sort of licensing deal and have them serve like uh you know the incremental capacity on you know GB300s
or whatever. Um in my mind those are
or whatever. Um in my mind those are sort of like the two main reasons why you would wait 10 days to actually drop the model weights.
Yeah, makes sense. And I think um functionally this is really interesting because just to talk about the model architecture for a second, it's 2.8
trillion parameters. Um this does not
trillion parameters. Um this does not fit in a B200. So you need to have B300, GB300 or I guess MI355X
in order to be able to serve this model on a single system like a single 8way HGX server. Of course, you can do um
HGX server. Of course, you can do um unique things where you have pipeline parallelism across multiple nodes and stuff, but that's going to really impact performance. So, um I think there's a
performance. So, um I think there's a lot of recipes being cooked up and only people who have the latest and greatest chips are going to be able to serve this model. Um just a let's let's go back to
model. Um just a let's let's go back to your comment on on Google for a second.
The idea that this model is truly competitive in at the frontier at 2.8 8 trillion parameters um kind of gives us some insight into how big the closed source frontier
models are, right? Like it would be even more embarrassing if they're hitting these levels of performance and they're being compared to 10 trillion parameter
models with a lot more active, right? It
we have to be we have to assume that this is in the same range as what soul and fable are, right? Yeah, I I think that's a great point and uh you just
like have to be correct. I think I I still believe in sort of uh the competence and the crackness of all the openthropic researchers and if there are
some people on Twitter who like to claim that like oh you know current uh close source is like 10 trillion total parameters or something if that's actually true like guys it's time to pack up the bags like you know spies
probably should crash like 50% like tomorrow like you know it's over um I'm pretty confident that like Kimmy K3 cannot be much bigger uh or sorry much
smaller. If anything, it might even be
smaller. If anything, it might even be slightly bigger uh than the leading open source models today. Um sorry, the leading closed source models today. Uh
and if that's true, this actually further highlights a point we've been harping on for a while on some analysis, which is that the margins for these like closed source labs have to be absolutely
mindboggling. Because if you're telling
mindboggling. Because if you're telling me, you know, Kim is probably not at negative margins when they're serving K3 at 315 uh $3 per million input tokens,
$15 per million app tokens. Um that's
sort of the same price as Sonnet. And so
if you're telling me that like Fable is probably similarly sized and then can charge $10 per million Epo tokens and $50 per million APO tokens uh then this
should just like it may dispel any of the remaining concerns people have about the AI labs being these uh unprofitable businesses like um selling tokens at API
prices is just it might be even better than like SAS honestly just it's an incredible business today.
Yeah, makes sense.
No uh no cost of employees, just the GPUs. So, can you compare this pricing
GPUs. So, can you compare this pricing strategy to the previous stuff? Cuz you
said it's at 315. The previous version from Moonshot directly was at 95 and $4.
So, we're talking about Yeah.
more than 3x pricing increase from 2.7 code to Kim K3. Um,
do they have more even more room to increase pricing? Like what's the curve
increase pricing? Like what's the curve to get the Frontier open-source intelligence or Frontier soon to be open
weight yet yet to be determined what the license will be intelligence?
Honestly, I don't think they have that much more room to push pricing up. Uh like I would
guess that even at like this 315 there will be a lot of people who are like this is a little too expensive for me.
Uh my task is easy enough for like a G1 5.2 or a Miniax M3 and I might just like use one of those models instead. Um, and
I actually think sort of uh so on on one end sort of you have like the semi analysis of the world, right? Where we don't really care how
right? Where we don't really care how much money we're costing Dylan when we like burn tokens all day. Like we're
very happy, you know, using Fable for even like a relatively easy clask that we're like pretty confident uh one of these open source models can do pretty well. Um, and then on the other end, you
well. Um, and then on the other end, you have people who are like extremely costconcious, like you know, you maybe
only get like $200 worth of tokens per week, like as we heard some large companies like, you know, Tesla and Uber are implementing. Um, and I think like
are implementing. Um, and I think like pretty much all the people in that second bucket are going to be wanting to use like the GLM kind of pricing tier models because they are already good
enough for uh like most everyday tasks and then everyone in the semi analysis bucket is still using like 56 soul and fable. Um, so I think there actually is
fable. Um, so I think there actually is a pretty interesting question of who is the user that is actually going to be switching towards Kimmy K3. Um, I think it might just be a lot of people who
like philosophically love open source and are excited to like try this new hype model and support it. Um, but it wouldn't surprise me at all if there's
like not actual serious adoption amongst say like large enterprises um of this model.
Okay. What do you think about where we go from here? Like this is obviously a new base model. It's a fully new architecture for these guys. 2.8
trillion parameters. Um they've got Kimmy delta attention potential residuals the stable latente that they keep using the um it's a it's like a
scaled up bigger version of the previous models clearly about about two times bigger um but previously with K 2.5 we
saw cursor train composer based on just continued pre-training as well as MRL and then we saw Kimmy give us 2.5
2.6 2.6 7 checkpoints as they just continued the RL. Um, this is a new base model. It seems pretty complete. Like in
model. It seems pretty complete. Like in
my usage, it's working pretty well. It's
not screwing up anything basic when it comes to writing a PR description or like totally going off the rails. The
way that we've actually seen some other models that are kind of raw without a bunch of RL have some rough edges at the beginning. I'm not seeing those yet. So,
beginning. I'm not seeing those yet. So,
where do we go from here? When does 3.1 come out? How does pricing change over
come out? How does pricing change over time? Like does composer do we get a
time? Like does composer do we get a composer based on Kimmy K3?
Uh I mean composer based on Kimmy and K3 definitely not because I think the cursor guys are pretty set on trading the Roma from scratch now. Um as for
when like you know K3.1 K 3.2 whatever come out um I imagine you probably see like two or three updates uh within like the next each like a month or two apart
or something as they just continue post training this thing. Uh I would guess sort of like pricing stays about the same. Um just because
same. Um just because I don't like they're not going to be able to run it on new hardware in the next like two to three months. So
they're not going to get like a huge you know uh throughput increase there to reduce pricing. Uh maybe it's possible
reduce pricing. Uh maybe it's possible some like really craft I don't know kernel engineers figure out how to uh reduce the cost to server this thing so it's
closer to like uh like DFC v4 pricing or something. Uh that would be
something. Uh that would be really impressive but given that it's a 3 trillion parameter model I think I'm a little skeptical. I would guess that
little skeptical. I would guess that like the current pricing we see for like the miniaxes and the GLM is kind of already pushing the limits of like what
you can serve like a 1T to 1.5T model at and not have like just embarrassingly bad margins. Um so I feel like this
bad margins. Um so I feel like this pricing is probably here to stay at least for the next you know few months.
Um, I think the most interesting question is like whether or not the open versus closed gap is like going to continue shrinking and if open will ever
like fully match closed source like true frontier level parody. Uh, curious what your thoughts are there Jordan. I think
it has a serious implications for our whole industry if it actually happened obviously. Yeah, I mean my view is that
obviously. Yeah, I mean my view is that I believe the reason that this gap has closed right now is squarely put on the US government imposing restrictions on
anthropic and resulting in us not getting the actual best models that these guys have. And so they've artificially caught up basically. Um
interesting.
Clearly we see this with with mythos versus fable. Uh I can't use mythos. I
versus fable. Uh I can't use mythos. I
can only use fable sometimes if I ask it nicely. 5.6 six soul. I think our house
nicely. 5.6 six soul. I think our house view is that it's not the biggest model Open has ever trained. It's not the size of 4.5. To me, they have a bigger model
of 4.5. To me, they have a bigger model somewhere.
And I think the result is that we're going to be able to only access frontier intelligence if government entities allow us to.
Um, and that's a very interesting change to the the setup going forward. um
because I think it represents an opportunity for many of the players that are in fourth, fifth, 6th, 7th place to catch up to a limit at which point it's
okay to release everything and start to battle for user share without really being able to find the frontiers um and
have the frontier dominate. I think it's possible that you know we see the frontier take another big big step towards the end of the summer. It's
possible the politics change a little bit.
Yeah.
Um it's possible that we start to find other you know modalities beyond coding uh at which these guys can really
improve and they start exploring those like areas. Um we didn't you know intend
like areas. Um we didn't you know intend to talk about this right away but I just loved the uh release of Inkling by Thinking
Machines. um thought that the native
Machines. um thought that the native audio input would be super interesting, super useful in the future and kind of a sign of what's to come. But anyway,
yeah I there's on the topic of inkling, there is definitely like huge huge huge demand for a western open source model that
does not suck. Like it is like I'm I'm shocked how markets are still so inefficient and like we haven't had a single American company that's at least on par with like the fifth best Chinese
company. Uh, but if you just like I mean
company. Uh, but if you just like I mean one it's probably just a matter of time until the US government like bans Chinese open- source models like entirely. Uh, and maybe that's a can of
entirely. Uh, and maybe that's a can of warnings you don't have to go down on this conversation. Uh, but two, even if
this conversation. Uh, but two, even if that like doesn't happen, uh, I feel like the average large American enterprise is simply unwilling to like
put all of their proprietary data through a Chinese open- source model.
even though you can make sort of like all the logical arguments of like dude you know you're like you're just like loading their weights in like your air gap data center or whatever there's like no way the CCP is actually going to like
see any of your data but like I don't think the executives will actually buy that and I don't think they really care and I think there are a lot of people who like a care about token budgeting and b uh are only interested in running
uh a western model or like a non-Chinese model a non-Chinese model um and so it is like shocking to me that I guess Inkling is the best one that we have now. Um, but it's shock to me that we're
now. Um, but it's shock to me that we're not like actually closer to the open source frontier in America.
Yeah, I mean it was Neotron and then it was an Inkling and I'm I'm Yeah, it's really inspiring to see Inkling go for it. I think they have two business
it. I think they have two business opportunities there. They've got to be
opportunities there. They've got to be better than the bulk of Chinese open source. Like they have to be in the game
source. Like they have to be in the game there to be considered.
Yeah.
But then they also need to be better than Sonnet or better than Terra Luna.
the tier two, tier three models from the Frontier Labs because you know you can build a bunch of cheap applications um using close to frontier intelligence
using bedrock or foundry or whatever and just get access to the anthropic or open AI models and and save your money there by going with their second best model. So I I I never really understand
model. So I I I never really understand the western open- source angle of saving people money. I think it it is real and
people money. I think it it is real and getting those models into the ecosystem of companies like fireworks and together and base 10 and anybody who's serving open source is a good thing because it
it is a market but to me the bulk of the market is is government like one of the views on the Chinese models actually that's interesting is that um Xi has
been encouraging the Chinese companies to keep the models open source that is the view from their uh party and uh I think the the big
reason for that is that uh a bunch of the Chinese government wants to download the weights and run it on servers that they own and they want the support of
the local ecosystem. And I I think that the American government should work the exact same way. I I think that's a pretty pragmatic strategy. It's like you
need to give the people in your country um access and support to run this stuff.
Um maybe the the the other thing worth commenting on is that in the K3 blog they mentioned the post- training uh
sorry the the quantization during the SFT stage, right? And in that they they were commenting on natively
um using MX MXFP4 and MXFP8 weights and activations respectively for quote broad hardware compatibility.
Um what what other hardware do you think uh Moonshot cares about? Jordan,
I got a list of 11 Chinese accelerators.
You should subscribe to the semi analysis accelerator model and learn more. Yeah, Huawei Ascend, you've got
more. Yeah, Huawei Ascend, you've got BU, you've got Klugson, uh you've got the more threads guys, there's all sorts of different chips that are being
uh they're showing up in papers. We're
seeing code like it is a national priority for China to get these frontier models. These are frontier models now
models. These are frontier models now um running on their domestic accelerators.
Yeah. Yeah, I mean if uh I guess if we were calling Google a frontier lab, you know, end of 2025, we got to call they call Moonshot a frontier lab now.
Kind of crazy, dude.
It's like vanity.
You know, you're going to I'm still I'm still a 34 waist. Yeah. Yeah. And so are the seven other Chinese labs.
Yeah. No, no. Funny enough, my my dad is actually visiting China right now and he's telling me how like the hotel he's currently staying at is totally booked because like Xi Jinping is going to be
in the area soon and he's going to give a speech about how AI is a top priority um for China. So, I think a lot what you said is right. Um, circling back to what
you said earlier about the US government and how if they keep like uh kneecapping uh the the frontier models opens, you know, forcing them to delay them,
forcing them to only have like their second best model actually publicly label and therefore giving, you know, all the other players, the Googles, the SpaceXes, Metas, whoever time to catch
up. Uh, do you think that just
up. Uh, do you think that just completely destroys like the Frontier Lab business model? Like if you're open Enthropic, you just lose all pricing
power at that at that point, right? Like
I don't see how Enthropic can still uh accelerate net new ARR if their model is on par or comparable with like the Meta model, the SpaceX model, the Google
model, the Moonshot model, the Deep Seek model. like what happens to our industry
model. like what happens to our industry at that point, Jordan?
Yeah. I mean, uh first of all, no, I don't think that's going to happen and I think I can explain why, but uh second of all, I I don't I don't know for sure.
So, we'll have to see it play out. Um
interesting to think about. So I think the the biggest thing that I've realized in my personal usage of this stuff is one how hard it's getting to differentiate between using the absolute
frontier model and the max thinking mode versus high versus medium effort.
Yeah.
On those models. It's it's really really hard for me to find day-to-day tasks that these models can't figure out. Um,
and I'm just my behavior is default to the biggest and hardest thinking because I don't care about Dylan's budget. But
budget. But yeah, when it comes to actually um using this, there is an aspect of the harness being part of the product. So testing Kimmy K3
requires me to take a serious look at Open Code and Hermes and Pi. And so the harness is totally part of the product
still. Um, simple simple things can
still. Um, simple simple things can cause me to want to use one model over the other, like can I install it on my remote SSH server? How easy are the
keystrokes to get stuff in? Can I edit previous commands? I mean, you know,
previous commands? I mean, you know, it's it's silly, but these these little tiny features in the harness actually impact where I'm going to send my tokens, which results in where I'm
going to send my budget, right?
Um, yeah, that's an interesting point because I think a lot of people they talked about token buding, right? But I
think, and correct me if I'm wrong, but what I'm hearing from your description of your own workflow is that like even for tasks where I'm pretty confident
that like a GLM could successfully do it. like I'm happy routing it to Fable,
it. like I'm happy routing it to Fable, doing it on max intelligence because like the ROI of that task is still worth like the Fable price to me and so and like there's obviously there's like
always going to be some risk in the back of your mind where it's like oh if I you know use GM instead even on like on medium thinking mode and it's way cheaper like maybe it's not actually like as high quality as fable actually
would have been right and so uh even if the benchmarks claim that like oh a lot of your tasks can move to GLM. Uh you're
still fine by keeping them on enthropic models or open eye models for for the foreseeable future.
Um mostly yes, but I use a lot of Slack bots right now and I actually don't know what model's running behind the scenes on that Slackbot. So specifically at
computer with perplexity and I think if it starts routing it to Kim K3, if it starts routing it to GLM, starts routing it to sonnet and I know it's doing it today because I looked at my usage a few
weeks ago and found how much of the OpenAI models I was using because it was making that decision. Um, first cut at a PR before I go in and actually fix some
stuff up. I don't really care which
stuff up. I don't really care which model they're using, right? And
um, that is about the that is about the quality of the harness there for what I'm what I'm using. uh the the justification.
This might actually be a this might be a pitch for outcome based pricing if anything like one of these labs could potentially just like get 95% plus margins if they do outcome based pricing
for you because uh yeah as you said all these tasks you're like happy to pay even stable pricing for when they could probably like get it done at a fraction of the price even today.
Yeah. Yeah. 100%. Certainly with the dialing in the thinking mode which is where a ton of the expense ends up going. Um, I can totally totally imagine
going. Um, I can totally totally imagine them building a router.
The second thing though, just on the competition thing you said earlier, is like I I don't think that um we're out of use cases or ideas for these guys to work on. I don't think I
think they can continue to train incredible models to try and hit RSI on the coding side without ever releasing it to us in the uh in the proletariat and and keep their permanent under class.
Yeah. keep their bourgeois models training each other um and keep distilling them and and you know whatever giving us the the little tastes of it uh while still pursuing a research
objective that includes all sorts of other uses of AI like we're we're really exploring coding right now but we do some video generation stuff we do a lot
of audio to audio stuff we do lots of like deep research that's really doesn't look like coding in some ways um I I think there's lots of lots of use cases that they can continue to explore
without really encountering like the cyber security issues. Um, robotics and world models is a simple one, right?
What if anthropics sets their sights on automating away a whole bunch of manual labor jobs instead of knowledge work jobs? I mean, the idea that there's no
jobs? I mean, the idea that there's no way for them to um build a sustainable business with great ROI for their, you know, greatest technology the world's ever seen, I I don't believe that at
all.
Um yeah, it doesn't pass the smell tests.
No, I I not at all. Um
but uh even even beyond that the I use these models so much every day and I see first of all how much my friends who work in technology who are software
engineers spend 10 times less than me and use it 10 times less right now.
Right. Uh one person using Fable is like you know a person using Sonnet and a person using Fable they use it both the same amount the same day. The person
using Fable spends 10 times more at 90% margins. They make up the bulk bulk of
margins. They make up the bulk bulk of the business. Right? As soon as those
the business. Right? As soon as those people of which I would say there's maybe maybe I'm in the top 10% maybe even the 1% of the industry start using
the bigger models, they use them more.
That's just more demand for all of this business. And and the models don't even
business. And and the models don't even need to get any better for them to discover. They can use them for the
discover. They can use them for the really important tasks or the the bigger ideas that they have. And then I need to go and talk to my neighbors who don't work in technology. And there I'm
certainly in the 1% probably in the 0.1% maybe 01% and maybe we've got a thousand times more to go from here. And so
I I I get back to uh you know Masaasan golden goose exponential chart to the right you know sort of this point
of just like it's an exponential hop on why are you okay just to take this totally off off the rails
well actually before we go there uh I want to say that I think what she said about being early is totally right and this is like exactly why I don't think
Kimmy K3 is going to cause like net new ARR at Enthropic and Open AAI to accelerate. Um because I think even even
accelerate. Um because I think even even if even if you want to like uh assume that some non-nudgeable portion of people who are
using Fable and 56 today are going to switch to like Timmy K3 because it's cheaper and can do their workloads. I
think that is like completely overwhelmed by the people who still haven't like seriously tried this technology. you know, the people who've
technology. you know, the people who've like kind of tried it a little bit but are every day like discovering new use cases, like new cool things, high ROI things they can do with the models. I
think all those people are going to be using like 56 soul or fable 5 as a default to like unlock these new use cases. And that is like you're just not
cases. And that is like you're just not going to see like a a uh ARR slowdown, AR growth slowdown uh because like that is going to continue ramping so fast.
Yeah, I think we're in agreement on that. Think about how many people there
that. Think about how many people there are left to subscribe to this podcast and follow semi analysis man.
Dude, it's crazy to me like so I went to ICML last week. I went to AR Engineer the week before that. These are like you know nominally AI conferences, right? I
thought people would be pretty plugged in there. Like I would say 80 plus% of
in there. Like I would say 80 plus% of people had never heard of semi analysis before. I was like guys what what are we
before. I was like guys what what are we doing man? Like you claim to work in AI
doing man? Like you claim to work in AI but you haven't you haven't read you've never even heard of semi analysis. Like
we're still so early. It's
that's an ego check, Max. Come on, man.
Should calm it down.
I mean, maybe tootune our own horn a little bit.
Okay, man. I I I think we could keep talking about it this all day, but probably good to wrap here. Anything you
think was uh left upset, left unsaid?
Any burning questions?
Um I I just if the stock market crashes because all the investors, you know, have deep sea car one movement part two.
Uh buy the stocks, guys. Not investment
advice, though. Do your do your own diligence. Not investment advice.
diligence. Not investment advice.
Love it. Let's end it on that. Clip clip
out not saying anything about stocks and finish with me.
Do your own due diligence. You can do your own.
Good job, man. Okay.
Yeah. Cool.
Loading video analysis...