Opus 5.5 vs GPT-6 Sol vs Grok 4.7 - which is the best?
By No Code MBA
Summary
Topics Covered
- Cost is the real headline of this model cycle
- Cheap models unlock products that never made economic sense
- Frontier labs are intentionally pacing the most powerful models
- Pick the right tool, not the most powerful model
- Open-weight models are quietly squeezing the frontier
Full Transcript
Three new AI models were just released.
A new model from OpenAI, Anthropic, and Grock. By the end of the video, you're
Grock. By the end of the video, you're going to understand which models you should be using, which is the best, and which is the worst. I've brought on Will, the founder of Diad, to help
explain everything that you need to know about these models. Really excited to get into this video, and Will, thanks so much for joining and and sharing your knowledge with us. So this is a
benchmark um that we created for Diad is to create these realistic applications.
And maybe just the best way to show you what it is is actually show you like what um these new models created. So
let's actually just go play Opus 5.5.
We'll just play a few seconds. And what
it's doing is it's building basically a CRM app, right? A customer relationship management. You have different contacts.
management. You have different contacts.
You have different companies deals. And
you can basically see we're kind of going through the apps that each of these models built. Um I won't play through the whole video but I think you get a sense of these are realistic applications accepting authentication
creating different users. Um and if you look at GBT6 soul you know again they're all getting similar prompts slightly different look and feel but basically similar functionality and they all look
really nice you know like when you look at like Opus 5.5 GPD6 Soul these look really nice. This is the part that
really nice. This is the part that actually surprised me quite a bit was when I looked at GPT6 Luna it looks really good. Um, I'll show you the
really good. Um, I'll show you the benchmark score in just a minute, but if you look at it, you know, like honestly, I couldn't tell you that this was like a cheaper model than the than the other two demos I just showed you. There is a
couple functionalities that it didn't get quite right. But I think what's really kind of interesting to note here is if you look at the benchmark score here, right, you look at like GPT6
actually is the highest scorer right now. It's at 94%.
now. It's at 94%.
closely trailing by is GPD6 Astra, but then the new model Cloud Opus 5.5 also really strong at 93%.
Um, but this is the part that actually really surprised me. GPT6 Luna is at 87%.
And why that's so shocking to me is that it costs a dollar to build and GP6 Soul costs $10 to build. So that means GB6 Luna is 10 times cheaper and Opus 5.5 is
actually more than GP26 Soul. So it's
40% more than GPD6 soul. So now again you see Luna is 14 times cheaper. But
even Opus 5.5 you see here it's actually much cheaper than Opus 5 and much cheaper than Cloud Fable 5.1 which is definitely the most expensive. Fable 5.1
was like $47. So you see both of these labs are really getting into this cost competition, right? It's like these
competition, right? It's like these models are all so good now. They're
trying to drive things even cheaper and more affordable. And I think this is a
more affordable. And I think this is a really interesting new dynamic with this latest batch of models.
Got it. So this is showing like for example, you're on Fable 5.1 right now.
It costs $46.90 versus like Luna which cost $1 to to build. So 46 times more expensive
but only very slightly uh better in or like more you know qu higher quality result basically.
Yeah. Yeah. Totally. Everything is all bunched up here that it actually might be helpful if I show you one more model.
Um, let me show you GPD 5.6 Luna.
So 5.6 Luna, as you can imagine, was very affordable. Um, the Luna models are
very affordable. Um, the Luna models are very affordable, but you see it's way lower in score is at 60%. And this kind of matched my experience, too, which is, you know, obviously everyone loves the price of Luna. It's terrific, but it
wasn't really powerful enough to be a daily workh horses for a lot of different applications. But now this is
different applications. But now this is the part like I think is going to be interesting is that can you actually just use GPT6 Luna for a big chunk of your coding tasks. I think this is something that will be really
interesting to dig into because again it's getting so good and this is not to say of course there are going to be benchmarks where you know soul opus are crushing um Luna that for sure is the
case and you can see a whole bunch in the blog post. We'll look at some of them later, but I'm just talking about like the kind of things that I think people are going to be building a lot of times in just like everyday coding, right? You're probably not trying to
right? You're probably not trying to solve the latest frontier math problem.
You're just trying to build like like a SAS application that's useful for you.
And I think this is where like that gap is quickly narrowing. Everything is just getting so high up here, right? Part of
me is wondering like when I first ran this actually GP6 and Luna scored really low. Um it wasn't because they had a
low. Um it wasn't because they had a bug. It was because there was a
bug. It was because there was a ambiguation where my prompt wasn't very clear where like should you run this bootstrap process in development or in production process. So this was actually
production process. So this was actually like not the model's fault at all. This
is just an actual like benchmark bug.
And I think that's showing you how good these models are. It's like they're just getting right to the tippity top and all the differences that you're going to see are going to be in the like much harder benchmarks um that are out there. And
the question is do those benchmarks matter for your use case or not?
Right? Because if you if your benchmark for for this wasn't just like a CRM app, but it was a CRM app with 50 features that was like very complicated and you
needed it to to oneshot it or how far it could go with one shot, it would maybe we would start to see a bigger difference between uh Astra and Luna potentially.
U not not even be the best way to code anyway. But
anyway. But yeah, like it might be possible like we need to create like a diet bench v2 because everything's getting so good.
This is the part like I've been to it.
So each of these SAS applications actually is like three prompts, right?
It was sort of like there's like a basic one and then you get more advanced like multi-tenencies like oh you could invite someone by email. So I was trying to make it hard and harder but the thing is again the model is getting so good like even though this benchmark was hard like
a few months ago now it's just like everything is just crushing it. So I
think that just speaks to the like overall progress of these models.
Maybe if I could shift gears a little bit, there's of course other benchmarks.
A really interesting one is Cursor Bench from of course cursor. Um the one thing that's interesting here is that you'll see it's cursor bench 4.0. So they had a previous one is like 3.2 or something and now this is actually like way
harder. So everything got lower scores
harder. So everything got lower scores across the board. um they launched Grock 4.7 that looked pretty good but then Opus 5.5 again just kind of like crushes on this benchmark when you look at the
price you know for a given like $3 it's just giving by far the highest performance here so this is interesting and I think we can talk more about this later in terms of maybe uh some of the
nuance of what these models are good at and and not good at but based on this benchmark and it it seems like there'd be no reason to use Fable anymore right now and you would just want to use Opus
because it's as good if not better and cheaper from what I'm from from what the benchmarks are showing.
Yeah. And and this is what sort of Anthropic was saying in their blog post is like you know just use Opus 5.5. It's
cheaper. It's better. It's as they say it's as good as Fable 5.1. If you look at X some people are saying it's even better than Sable. I'm not sure but it's definitely Fable tier. And yeah, you're totally right. Like at this point you're
totally right. Like at this point you're just going to save way more money and get awesome performance with Opus 5.5.
Cool. Why don't we jump in real quick to each of the release notes because I think this will give us some more uh context about each of the models and then we can talk more about uh maybe
what people are are saying are are uh impressions of each model and I think giving people some more specifics on you know which model they should actually use. So hopefully you know by the end of
use. So hopefully you know by the end of the video they'll have some sense of of okay should I should use this or I shouldn't use this. But yeah, so so this is the the Opus uh yeah launch post it looks like.
Exactly. Um so this is what they start off with, right? It's like performs at Fable 5.1 and it costs 40% less to run than Opus 5. And I think these two things are the most important things,
right? It's it's stronger than Opus 5
right? It's it's stronger than Opus 5 Fable level and it's way cheaper. Um
we'll get into the pricing difference over here. Um the thing that like is
over here. Um the thing that like is really easy to miss is that their cash reads drops by 60%. Um and you don't think of cash reads a lot but it's
actually like when we use diet it's a huge chunk of your cost. is something
like even like you know half or even more than half your constant just caches because every time you're calling a tool you're doing like another turn in your agent loop you're constantly doing all this cash it's all the code that it read
beforehand old instructions and so making that way cheaper is a really big reason why this is like 40% cheaper than Opus 5 interestingly so like on the diet
benchmark we basically got exactly like 40% cheaper um from opus 5.5 to opus 5 so I think what they state here is is is what we're also seeing um out in the
real world. So, I think this is going to
real world. So, I think this is going to be awesome for everyone who's using cloud is that now you have a cheaper model that's also really smart. One of
the most interesting things about Opus 5.5 is that they've really improved the communication style. This was like a big
communication style. This was like a big complaint like all the claudisms we talked about in a previous video.
Everyone was complaining on social media and I found it just it's very grading, right? Like I would have to ask Claude
right? Like I would have to ask Claude like okay like what does this mean? like
can you just state this in like more simple terms explain it like in M5 and it seems like they've made a lot of improvements here you can kind of see there's like a few different examples of
course this is cherrypicked but I think what you sort of see is like it's just less weird like writing and formatting and it's just a little bit more straightforward in terms of like how
it's explaining things like I think they're trying to approach like you know how would a regular coworker explain things rather than this like very esoteric um explanations and I I think this is going to be a very welcome explanations for so far from the social
media vibes. People are definitely
media vibes. People are definitely liking 5.5. So yeah,
liking 5.5. So yeah, and no more m dashes based on the examples that they're that they're showing, you know. I guess so. And M dash is
you know. I guess so. And M dash is funny because you can always prompt it out, but it's always the weird ticks where it's like not this, but that, you know, and I think there's just all these weird phrasings and it seems like
they sort of just cleaned it up and you don't have to do these weird prompting or tricks to get to talk normally. So, I
think this is the one I'm very interested just playing around with and seeing if it feels better because I think this is the kind of thing it doesn't show up in benchmarks, but it just it honestly just makes you tired, right? It just makes it hard to work
right? It just makes it hard to work with AI hours a day when it's talking in a very strange form of English.
Yeah. And also, if you're using it for marketing tasks and not coding, then like that's where this stuff can become super helpful and and beneficial.
Yeah, for sure. For sure. And and I think this
for sure. For sure. And and I think this might be like an interesting angle is like this like non-coding just like creative writing or like marketing blog posts. I think that's the part
blog posts. I think that's the part where it's like yeah I don't have a good sense but I I I I would find it interesting like it might not be like you need the most powerful expensive models just to have good English like it
could be like I've heard Gemini is also like really nice and it's pretty affordable. So
affordable. So yeah. No, totally. And I I think that
yeah. No, totally. And I I think that there's I'm I can't remember the name of it, but I saw some AI lab that was like focused on writing and and yeah, I I wouldn't be surprised if we see more
specialized models like that where like right it doesn't need to be super powerful where it can build an app that's insane uh and very powerful the
way that like Opus is, but um I don't know why it couldn't be a good writer necessarily. So
necessarily. So yeah. Yeah. May maybe next time we'll
yeah. Yeah. May maybe next time we'll we'll take a look at some of those like riding benchmarks and and dig into that.
It's definitely interesting.
Yeah. Yeah. Cool. All right. So then uh GPT6 Soul. So you're saying people are a
GPT6 Soul. So you're saying people are a little more disappointed with this compared to the Opus launch.
Yeah. I think this is where like what are people looking for, right? And I
think it's like for a new launch.
Opening is hyping it up. They're saying
we have something awesome. And if you're looking for like the most intelligent model, like this is not it, right? I
think like right off the bat, it's like Opus 5.5 is crushing a benchmarks. I
don't think anybody's going to say GPD6 soul is is better. Like there is an interesting question like Astra is their most expensive, is their most intelligent model. Is Astra or Opus
intelligent model. Is Astra or Opus better? Right? Like I I'm not sure. I
better? Right? Like I I'm not sure. I
think we there needs to be more benchmarks on that. But what's
interesting to me is Opus is actually cheaper than Astra um just from the token pricing, but Soul is way cheaper than Opus. So now I think it's like
than Opus. So now I think it's like people are going to are now competing a lot on prices. And this to me is the real headline. It's basically right off
real headline. It's basically right off the bat everything is 50% um cheaper and whether you're using Chad GPD subscription. So like we actually added
subscription. So like we actually added a way recently where you can use your subscription in Diad you get way more usage out of it. But if you're paying API token price, you're going to feel it too. It's going to feel way nicer,
too. It's going to feel way nicer, right? Like just imagine sending twice
right? Like just imagine sending twice as many messages um with either of these models. And I think this to me is
models. And I think this to me is actually going to be like maybe like the bigger game changer, right? Is like
getting a lot more usage. Like people
are talking about cost all the time when I talk to users. This is like the number one thing. So I think having GPD6 just
one thing. So I think having GPD6 just be like an across the bat. Like if you told me a few months ago, right, like 5T six soul is going to be half the price, I'd be like, "Oh, hallelujah. This is
like the biggest news ever, right? So, I
think people like shouldn't sleep on how important price is. But, you know, enough of that. Let's get into some of the benchmarks. What you kind of see
the benchmarks. What you kind of see with soul is it's not really more intelligent than 5.6 soul, right? Like
if you look at this kind of like yellow color, I guess like the sun and the darker yellow, which is 5.6. Like what
you see is like everything is shifted to the left. It's cheaper, but it's not
the left. It's cheaper, but it's not really like way more intelligent. Uh and
they're also benchmarking against um Opus as as you can see um and Fable which are way more to the right way more expensive against similar performance.
So this is what open is basically saying is hey come use us we're way cheaper.
They even have this little table here.
It's like hey Astra is four times more expensive but Opus 5 is 11 times more expensive and FL fable iPhone is nine times more expensive. So, this is where you start seeing like they're really
just focusing on the cost here. And I'm
curious where 5.5 lands, right? I
imagine it's going to be way better than Opus 5, but if you just look at this price difference, it it's huge. And I
think this is where OpenI is really cracking a lead here.
And and from what I'm see, so Terra is no longer a part of any of this like is essentially even from uh it had a good run like like
three weeks. Um, but uh I feel like
three weeks. Um, but uh I feel like Terra, if I remember the pricing was around what Soul 6 is now. Is that is that right? And that's kind of how
that right? And that's kind of how they're That's which is like you said, it's crazy that now you're getting like better. It's Yeah. Everything is getting
better. It's Yeah. Everything is getting smarter and cheaper at the same time essentially. Exactly. Yeah. Yeah.
essentially. Exactly. Yeah. Yeah.
Is that the $210? Yeah, I think that's the same price. And then Luna to me, I feel like Luna is the big story out of this because Luna is now getting it's
almost like Luna is I I don't know what the benchmarks are like or maybe how you would think about it, but it seems like Luna might be at better than what Terra was for like massively cheaper than what Terra was essentially.
Yeah.
Is that how you would maybe think about it?
I think that seems fair because like even in 5.6 six Luna. Um it was already like super competitive with Terra, right? It was just like if you do like
right? It was just like if you do like max reasoning on Luna again, it's so cheap that max reasoning is like not that bad.
Um it was getting comparable level ter and now like you see like GP6 Luna is even stronger, right? Like just like per dollar or reasoning level you see improvements.
Um so I think this is where it's like yeah like GP6 Luna is better than 5.6 terra.
Um which to me is just like fascinating.
I think like personally I'm excited to play on Luna like there's all these use cases where it was like AI is kind of expensive like I don't know how we would make it work um justify the price and now it's like with six Luna is like okay
now something is really capable it's very affordable it's like I think there's going to be a whole bunch of products right like not even talking about coding like automations those kind of things I think that's going to be super interesting super exciting I think
that's going to actually make a big difference um in the long run is just these super affordable models that you can use for cases Yeah. So it's like okay the takeaway from Opus is okay now
there's a model that's basically better than Fable or as good as Fable but significantly cheaper. uh it seems like
significantly cheaper. uh it seems like with these open AI models it's more like okay there hasn't been a massive increase in in not Astra is still their best model if you are doing the most
complicated things but the price has come down significantly on the others and then I I do think like you said Luna opening up more use cases of like all these any like almost like um uh
openclaw type things like I would imagine uh if you have something like Luna that can like uh do it for so much cheaper Then it it it makes it affordable for people to to
use it in a way that like maybe they would spend they'd be happy to spend 20 cents but don't want to spend $2 per task basically.
Yeah. Yeah. Exactly. Like when things are 10 times cheaper, it's a totally different feel. What's also interesting
different feel. What's also interesting is like next week there's going to be openi dev day and they even hyping it up. They say there's going to be a ton
up. They say there's going to be a ton of stuff. So what I would expect next
of stuff. So what I would expect next week is there's probably going to be like some really powerful model like maybe it's like 6.1 Astra. I have no idea. Um
there's a lot of speculation but like we know that OpenAI and also Anthropic they have more powerful models but now there's this whole discussion of pacing the frontier and maybe this like the interesting part is that like there's just a little bit worry like okay these
models are getting more and more powerful are they going to like break out of their sandboxes we saw that in OPI hugging phase there's like other incidents from different labs so now I think it's like the labs are sort of shifting gears and they're saying like
okay we're not going to release the like very very most powerful model we have all these concerns about cyber security and like other typ risks. Now, let's
actually focus on creating things that like give a lot more value for a lot more intelligence for the dollar. And
this to me is actually like a really nice trend because frankly it's like I'm not working on the most like groundbreaking stuff. It was
like, you know, oftentimes I just want something cheaper and if you give me something twice as cheap, I'd actually sometimes like that more than twice as expensive. Um, often times I'm reaching
expensive. Um, often times I'm reaching for a model because it's like but this is actually like just way faster and I don't want to sit around for 10 minutes um with a more powerful model. So I think this is going to take
model. So I think this is going to take into like a lot more calculation now.
It's like it's not just you need to like most you know like powerful model ever in the world. It's like you just need to use the right tool for your task and it's going to be super effective.
Yeah. Yeah. I like to think about models as like it's almost like a triangle. You
have like cost, speed, and quality. And
quality, right, for whether it's because they're pacing the frontier or it's not there's only so much more they can go up or uh but it seems like right now, right, cost and speed are where we're
going to see more and more advancements which are going to open up a ton more use cases as well. I agree. So, Grock
also, actually, do you want to talk about Grock? I don't know if we we want
about Grock? I don't know if we we want to or not. Um, okay. May maybe we'll talk for a minute. Like so Grock is sort of an interesting release where it didn't get a ton of hype. It was it
was released earlier this week. Um if
you look at the benchmark performance like it's perfectly respectable. Um but
again you've got Opus 5.5 that's better for each dollar. So I think this is the part where like you know how good is it?
When you ran the benchmarks it also wasn't like crazy strong either. Um, so
I think this is the part where like my impression from Grock 4.7 was that it was an okay release, right? Like the
price stayed the same at 4.6. Uh, it
seems like on some benchmarks it got better. For this one, it didn't improve.
better. For this one, it didn't improve.
So yeah, in my opinion, it's like you've got cheaper models that are really good.
You've got more powerful models that are more expensive. And Grock is in this
more expensive. And Grock is in this weird middle land that I think it's like maybe it's like their next release is going to really catch up to the frontier. Um,
frontier. Um, yeah.
So, we'll see.
Yeah. So, it's like we're still waiting on Grock 5, which in theory is gonna gonna be something that that is at that level, but right now, unless you have a subscription that
includes Grock already, like there might not be a reason to go out there and choose Grock over over something else.
Yeah, I I I think so. I think unless you have some like real affinity rock, like maybe it's like tone or style or you have a subscription, you probably don't need to run out to go use it. Um,
use it. Um, but again, like this is the thing is like these models are all so good that like six months ago they would have been awesome, but now it's just like the landscape. This is changing every week
landscape. This is changing every week and now you've got really good options.
One option that actually like didn't talk about a lot but is worth mentioning is Deepseek V4.1 Flash. Um, and it's a little bit slept on, right? like V4
DeepSeek got quite a bit of attention.
V4.1 Flash actually replaces V4 Flash and V4 Pro. So, they say it's better across the board than V4 Pro, but with the flash pricing. So, that's really interesting. Um, and you can see this is
interesting. Um, and you can see this is the cheapest build by far. Um, it's even cheaper than Luna, which is sort of absurd. But the one thing that is kind
absurd. But the one thing that is kind of interesting though is when I actually watched, you know, the app that it was building, it just didn't look as nice. Um, and you know, granted like looks are not the
most important thing, but like this is kind of plain like this just was like a tier below the other applications that were being built by the other models.
Maybe the prompting could be better or whatnot. So, I think there is a
whatnot. So, I think there is a difference between DeepSeek and the other models, but again, it's like you could always have it make it look better. It's
getting the core functionality right.
And again, when you just look at this trend of like where the openweight models are going, I think this is what's interesting. I think this is actually
interesting. I think this is actually what's driving a lot of this cost competition is that anthropic and openi they see all these openweight labs creating really impressive ones and maybe they're not the same tier yet but
they're putting this pressure I think a really good healthy pressure and now it's like okay opus 5 is getting cheap 5.5 is getting cheaper now we get like soul luna um getting way cheaper so I think this is the nice part is that
we're getting more choices better intelligence for the dollar yeah it's really interesting where even if the frontier models always have like the most powerful,
maybe pe the the open weight models could really take over everything below that. If if they're cheaper and just as
that. If if they're cheaper and just as good, people are going to generally go for the the cheaper one. Uh at least not everyone. I mean, it depends on how
everyone. I mean, it depends on how people are using AI, but if you're a company and you're you're paying per API and and you um you can get the same value or better value, then that's where
people are going to go. But like you said, that's where the labs don't want you to do that. So that's why we're getting like Luna, the new Luna model, which is uh really good as well. Yeah. So
yeah. So even if you don't use opening models, you're actually benefiting from either way. So this life is awesome.
either way. So this life is awesome.
Right. Right.
Cool. Um all right. This is really helpful. I think that there's some
helpful. I think that there's some really interesting takeaways here of like how I'm going to think about using these these new models. anything else
will you want to cover or if people want to go to DIAD um yeah how can they do that?
Yeah, so maybe just a quick plug for DAD if you go to dy a.sh SH you can download it's AI app builder it's like a local open- source alternative to lovable replet um makes super easy build
applications and the big thing is we let you use any model right so all these models that we talked about you can use India today you can also use your custom models even run models locally one of the big things that we're doing is
actually like letting you use your chatb subscription so you can like way lower your costs by offloading you know the heavy load inference with your subscriptions we're working on that for cloud code now so I think that's going
to really interesting. Um, people should check it out. And I think overall in general, like we want to give people choices and I think as you see it's like choices is a really good thing. It
drives innovation and it just helps the customer at the end of the day. So yeah.
Cool. Cool. Awesome. Yeah. Everyone
check out Diad and Well, thanks so much for coming on and helping to explain this to the audience. And if you like this, leave a comment. Let us know what you think. If there's anything you want
you think. If there's anything you want us to cover in the future, uh, let us know. And again, this is a pretty new
know. And again, this is a pretty new format for us. We did one for the Fable 5.1 release.
Loading video analysis...