OpenAI should be scared of this one
Read full transcript 24 segments
-
It's hard to ignore just how well It's hard to ignore just how well Anthropic's been doing recently. From Anthropic's been doing recently. From Anthropic's been doing recently. From Fable 5.1 being my favorite code model Fable 5.1 being my favorite code model Fable 5.1 being my favorite code model ever released to Opus 5.5 taking over ever released to Opus 5.5 taking over ever released to Opus 5.5 taking over pretty much everyone's workflows, my own pretty much everyone's workflows, my own pretty much everyone's workflows, my own included. I haven't chosen Fable for any included. I haven't chosen Fable for any included. I haven't chosen Fable for any tasks for over a week now, which is just tasks for over a week now, which is just tasks for over a week now, which is just like mindblowingly crazy to now with yet like mindblowingly crazy to now with yet like mindblowingly crazy to now with yet another release, and this one seems another release, and this one seems another release, and this one seems strategically targeted to screw over strategically targeted to screw over strategically targeted to screw over OpenAI as much as possible. Sonnet is OpenAI as much as possible. Sonnet is OpenAI as much as possible. Sonnet is back and 5.5 is looking a hell of a lot back and 5.5 is looking a hell of a lot back and 5.5 is looking a hell of a lot more promising than Sonnet 5, which if more promising than Sonnet 5, which if more promising than Sonnet 5, which if you don't remember was one of my least you don't remember was one of my least you don't remember was one of my least favorite model releases like ever. favorite model releases like ever. favorite model releases like ever. Sonnet 5.5 is a model I did not expect Sonnet 5.5 is a model I did not expect Sonnet 5.5 is a model I did not expect to care about. I legitimately thought I to care about. I legitimately thought I to care about. I legitimately thought I would just kind of skip over this one would just kind of skip over this one would just kind of skip over this one and not talk about it because I've been and not talk about it because I've been and not talk about it because I've been so impressed with Opus's price and value so impressed with Opus's price and value so impressed with Opus's price and value and just everything I've been doing with and just everything I've been doing with and just everything I've been doing with it. But I was wrong. Sonnet 5.5 is an it. But I was wrong. Sonnet 5.5 is an it. But I was wrong. Sonnet 5.5 is an incredible model. I am actually very incredible model. I am actually very incredible model. I am actually very very impressed with it, but not in the very impressed with it, but not in the very impressed with it, but not in the ways you might expect cuz when you look ways you might expect cuz when you look ways you might expect cuz when you look at the numbers, it doesn't seem that at the numbers, it doesn't seem that at the numbers, it doesn't seem that impressive. On the artificial analysis impressive. On the artificial analysis impressive. On the artificial analysis intelligence index, it comes out as more intelligence index, it comes out as more intelligence index, it comes out as more expensive and dumber than Opus 5.5 at expensive and dumber than Opus 5.5 at expensive and dumber than Opus 5.5 at every single level. So, why am I so fond every single level. So, why am I so fond every single level. So, why am I so fond of Sonnet 5.5? It's going to be hard to of Sonnet 5.5? It's going to be hard to of Sonnet 5.5? It's going to be hard to justify, but I think I can do it well.
-
justify, but I think I can do it well. justify, but I think I can do it well. The simple way of putting it is there's The simple way of putting it is there's The simple way of putting it is there's finally a cheapish model from Anthropic finally a cheapish model from Anthropic finally a cheapish model from Anthropic that makes sense in the modern era, that makes sense in the modern era, that makes sense in the modern era, especially when you compare it to especially when you compare it to especially when you compare it to offerings from OpenAI that have offerings from OpenAI that have offerings from OpenAI that have historically been closer to this price historically been closer to this price historically been closer to this price range. This model almost feels squarely range. This model almost feels squarely range. This model almost feels squarely targeted at OpenAI, specifically at GPT6 targeted at OpenAI, specifically at GPT6 targeted at OpenAI, specifically at GPT6 Soul. There's a reason I'm filming this Soul. There's a reason I'm filming this Soul. There's a reason I'm filming this video and not a GPT6 Soul video, and video and not a GPT6 Soul video, and video and not a GPT6 Soul video, and it's not because Soul is so impressive. it's not because Soul is so impressive. it's not because Soul is so impressive. Trust me on that. I'll explain what I Trust me on that. I'll explain what I Trust me on that. I'll explain what I mean and more after a real quick word mean and more after a real quick word mean and more after a real quick word from today's sponsor. Staying on top of from today's sponsor. Staying on top of from today's sponsor. Staying on top of what's happening in your codebase has what's happening in your codebase has what's happening in your codebase has never been harder to do. AI has made never been harder to do. AI has made never been harder to do. AI has made contributing to most codebases way contributing to most codebases way contributing to most codebases way easier, but it's made keeping track of easier, but it's made keeping track of easier, but it's made keeping track of what's going on in them way harder. I what's going on in them way harder. I what's going on in them way harder. I can't tell you how many times I had some can't tell you how many times I had some can't tell you how many times I had some weird issue in a codebase. I was like, weird issue in a codebase. I was like, weird issue in a codebase. I was like, "What the hell happened here? Who would "What the hell happened here? Who would "What the hell happened here? Who would ever have merged this?" And then went to ever have merged this?" And then went to ever have merged this?" And then went to the notes and saw that it was actually the notes and saw that it was actually the notes and saw that it was actually changes I made with my agent that I changes I made with my agent that I changes I made with my agent that I merged because I was blindly trusting merged because I was blindly trusting merged because I was blindly trusting the AI reviewers. Today's sponsor is one the AI reviewers. Today's sponsor is one the AI reviewers. Today's sponsor is one of those reviewers. It's Code Rabbit. of those reviewers. It's Code Rabbit. of those reviewers. It's Code Rabbit. But I'm not here to talk about how good But I'm not here to talk about how good But I'm not here to talk about how good Code Rabbit's reviews are. Spoiler, Code Rabbit's reviews are. Spoiler, Code Rabbit's reviews are. Spoiler, they're very good. I want to talk about they're very good. I want to talk about they're very good. I want to talk about a new feature they introduced which has a new feature they introduced which has a new feature they introduced which has made actually keeping on top of the made actually keeping on top of the made actually keeping on top of the changes way easier. It's the Code Rabbit changes way easier. It's the Code Rabbit changes way easier. It's the Code Rabbit change stack. Now when you have Code change stack. Now when you have Code change stack. Now when you have Code Rabbit set up on a codebase, it will Rabbit set up on a codebase, it will Rabbit set up on a codebase, it will give you this simple button that you can give you this simple button that you can give you this simple button that you can click even signed out to see what click even signed out to see what click even signed out to see what changed in the PR in a way that makes changed in the PR in a way that makes changed in the PR in a way that makes way more sense. Instead of a pile of way more sense. Instead of a pile of way more sense. Instead of a pile of alphabetically listed files, you get a alphabetically listed files, you get a alphabetically listed files, you get a description of what changed and why.
-
description of what changed and why. description of what changed and why. Instead of a jumbled mess of commits Instead of a jumbled mess of commits Instead of a jumbled mess of commits that nobody's going to click through, that nobody's going to click through, that nobody's going to click through, you get this stack of the different you get this stack of the different you get this stack of the different things the PR changes with descriptions things the PR changes with descriptions things the PR changes with descriptions of what is changing and why. And if you of what is changing and why. And if you of what is changing and why. And if you look at this, you can see pretty look at this, you can see pretty look at this, you can see pretty immediately how much easier it is to immediately how much easier it is to immediately how much easier it is to read through the changes that were made read through the changes that were made read through the changes that were made in this poll request. It even summarizes in this poll request. It even summarizes in this poll request. It even summarizes these different parts so you can these different parts so you can these different parts so you can understand what's going on in each of understand what's going on in each of understand what's going on in each of them. In this part, it's the detection them. In this part, it's the detection them. In this part, it's the detection for the package install and making for the package install and making for the package install and making changes to how T3 Code is persisted on changes to how T3 Code is persisted on changes to how T3 Code is persisted on your machine. Here is where we fixed the your machine. Here is where we fixed the your machine. Here is where we fixed the update logic that was breaking for a lot update logic that was breaking for a lot update logic that was breaking for a lot of users. There's even a security blast of users. There's even a security blast of users. There's even a security blast radius diagram that makes it way easier radius diagram that makes it way easier radius diagram that makes it way easier to see what changes are risky in a given to see what changes are risky in a given to see what changes are risky in a given PR. I also really like the activity PR. I also really like the activity PR. I also really like the activity view, which is a super quick way to see view, which is a super quick way to see view, which is a super quick way to see what's changing in a PR to catch up what's changing in a PR to catch up what's changing in a PR to catch up without having to scroll through the without having to scroll through the without having to scroll through the mess that is the thread view in a given mess that is the thread view in a given mess that is the thread view in a given poll request on GitHub. Don't let SLOP poll request on GitHub. Don't let SLOP poll request on GitHub. Don't let SLOP take away your understanding of your take away your understanding of your take away your understanding of your codebase. Fight back at soy.link/code codebase. Fight back at soy.link/code codebase. Fight back at soy.link/code rabbit. As I was saying, the benchmarks rabbit. As I was saying, the benchmarks rabbit. As I was saying, the benchmarks don't necessarily look great depending don't necessarily look great depending don't necessarily look great depending on how you frame them, but I want to on how you frame them, but I want to on how you frame them, but I want to start with what Anthropic said with start with what Anthropic said with start with what Anthropic said with their official announcement post. This their official announcement post. This their official announcement post. This is the Claude Sonnet 5.5 announcement. is the Claude Sonnet 5.5 announcement. is the Claude Sonnet 5.5 announcement. Introducing Sonnet 5.5, the second model Introducing Sonnet 5.5, the second model Introducing Sonnet 5.5, the second model in the 5.5 family. The clear upgrade in the 5.5 family. The clear upgrade in the 5.5 family. The clear upgrade over Sonnet 5. It's 30% faster and up to over Sonnet 5. It's 30% faster and up to over Sonnet 5. It's 30% faster and up to 30% cheaper for most work. I don't love 30% cheaper for most work. I don't love 30% cheaper for most work. I don't love comparing this model to Sonnet 5 because comparing this model to Sonnet 5 because comparing this model to Sonnet 5 because on one hand, it's not going to showcase on one hand, it's not going to showcase on one hand, it's not going to showcase the benefits that well. Like the speed the benefits that well. Like the speed the benefits that well. Like the speed difference and the cost difference isn't difference and the cost difference isn't difference and the cost difference isn't that big a deal. Sonnet was relatively that big a deal. Sonnet was relatively that big a deal. Sonnet was relatively cheap. But also because Sonnet 5 was a cheap. But also because Sonnet 5 was a cheap. But also because Sonnet 5 was a garbage tier model and is not a good garbage tier model and is not a good garbage tier model and is not a good target to compare against. I said the target to compare against. I said the target to compare against. I said the same thing with the Opus 5.5 coverage.
-
same thing with the Opus 5.5 coverage. same thing with the Opus 5.5 coverage. It made no sense to compare it to Opus It made no sense to compare it to Opus It made no sense to compare it to Opus 5, which is why I'm thankful they 5, which is why I'm thankful they 5, which is why I'm thankful they compared it to Fable so much. But here compared it to Fable so much. But here compared it to Fable so much. But here again with Sonnet because Sonnet's the again with Sonnet because Sonnet's the again with Sonnet because Sonnet's the smallest that they're working on right smallest that they're working on right smallest that they're working on right now. Then there's Opus, then there's now. Then there's Opus, then there's now. Then there's Opus, then there's Fable. Haiku is eventually going to Fable. Haiku is eventually going to Fable. Haiku is eventually going to happen, but uh we'll see. They said it's happen, but uh we'll see. They said it's happen, but uh we'll see. They said it's coming in the next few weeks. I didn't coming in the next few weeks. I didn't coming in the next few weeks. I didn't think Sonnet would be here so quick, and think Sonnet would be here so quick, and think Sonnet would be here so quick, and honestly, I'm pretty impressed with it. honestly, I'm pretty impressed with it. honestly, I'm pretty impressed with it. Sonet 5.5 is faster, lower cost as a Sonet 5.5 is faster, lower cost as a Sonet 5.5 is faster, lower cost as a compliment to Opus 5.5, where Opus is compliment to Opus 5.5, where Opus is compliment to Opus 5.5, where Opus is built for complex work requiring careful built for complex work requiring careful built for complex work requiring careful judgment. Sonet 5.5 is strongest at well judgment. Sonet 5.5 is strongest at well judgment. Sonet 5.5 is strongest at well scoped everyday tasks. I will say it scoped everyday tasks. I will say it scoped everyday tasks. I will say it actually goes quite a bit further. Make actually goes quite a bit further. Make actually goes quite a bit further. Make sure you stay tuned for the fish slop sure you stay tuned for the fish slop sure you stay tuned for the fish slop demo. I I'm very impressed with it. demo. I I'm very impressed with it. demo. I I'm very impressed with it. That's all I'll say for now. It's also That's all I'll say for now. It's also That's all I'll say for now. It's also good for fixing bugs and creating good for fixing bugs and creating good for fixing bugs and creating polished documents, slides, and polished documents, slides, and polished documents, slides, and spreadsheets. I don't know how many of spreadsheets. I don't know how many of spreadsheets. I don't know how many of y'all are actually creating polished y'all are actually creating polished y'all are actually creating polished documents, slides, and spreadsheets with documents, slides, and spreadsheets with documents, slides, and spreadsheets with your models, but if you are, let me know your models, but if you are, let me know your models, but if you are, let me know in the comments. I am actually curious. in the comments. I am actually curious. in the comments. I am actually curious. And when you're on the way there, if you And when you're on the way there, if you And when you're on the way there, if you could hit that little red button, it could hit that little red button, it could hit that little red button, it does help us out a bunch. A lot of y'all does help us out a bunch. A lot of y'all does help us out a bunch. A lot of y'all aren't subscribed, and you'd be amazed aren't subscribed, and you'd be amazed aren't subscribed, and you'd be amazed at how much it helps the channel. And at how much it helps the channel. And at how much it helps the channel. And it's also a great way to keep up to it's also a great way to keep up to it's also a great way to keep up to date. and seeing that uh OpenAI dev day date. and seeing that uh OpenAI dev day date. and seeing that uh OpenAI dev day is in not very much time. In fact, it is in not very much time. In fact, it is in not very much time. In fact, it may have already happened is where may have already happened is where may have already happened is where you're going to want to be to keep up you're going to want to be to keep up you're going to want to be to keep up with how OpenAI is fighting back. And with how OpenAI is fighting back. And with how OpenAI is fighting back. And trust me, they're going to be fighting trust me, they're going to be fighting trust me, they're going to be fighting back. Anthropic also claims this model back. Anthropic also claims this model back. Anthropic also claims this model has a sharp eye for design. This is an has a sharp eye for design. This is an has a sharp eye for design. This is an interesting call out because uh yeah, interesting call out because uh yeah, interesting call out because uh yeah, I'll show you the designs in a bit. It's I'll show you the designs in a bit. It's I'll show you the designs in a bit. It's not quite what I was hoping for there.
-
not quite what I was hoping for there. not quite what I was hoping for there. And then they call out Haiku 5.5 which And then they call out Haiku 5.5 which And then they call out Haiku 5.5 which is going to be built for high volume and is going to be built for high volume and is going to be built for high volume and costsensitive applications will be costsensitive applications will be costsensitive applications will be joining Claude 5.5 family in the coming joining Claude 5.5 family in the coming joining Claude 5.5 family in the coming weeks. So how does Sonnet 5.5 improve weeks. So how does Sonnet 5.5 improve weeks. So how does Sonnet 5.5 improve over Sonnet 5? First off it gets the over Sonnet 5? First off it gets the over Sonnet 5? First off it gets the best score on Terminal Bench 4 to date best score on Terminal Bench 4 to date best score on Terminal Bench 4 to date which is kind of nuts cuz it's a Sonnet which is kind of nuts cuz it's a Sonnet which is kind of nuts cuz it's a Sonnet model but it's a huge jump. It got a model but it's a huge jump. It got a model but it's a huge jump. It got a 70.6% 70.6% 70.6% where previously it got 10.3. Like this where previously it got 10.3. Like this where previously it got 10.3. Like this just looks silly seeing sonnet 5.5 as just looks silly seeing sonnet 5.5 as just looks silly seeing sonnet 5.5 as the highest score ever in terminal the highest score ever in terminal the highest score ever in terminal bench. On one hand it makes me skeptical bench. On one hand it makes me skeptical bench. On one hand it makes me skeptical of terminal bench. On the other I'm just of terminal bench. On the other I'm just of terminal bench. On the other I'm just kind of skeptical of benches at this kind of skeptical of benches at this kind of skeptical of benches at this point. It's really hard to measure the point. It's really hard to measure the point. It's really hard to measure the capabilities and the ones I have been capabilities and the ones I have been capabilities and the ones I have been doing on my own have on one hand felt doing on my own have on one hand felt doing on my own have on one hand felt even sillier but on the other hand have even sillier but on the other hand have even sillier but on the other hand have better reflected my vibes when I use better reflected my vibes when I use better reflected my vibes when I use these things. So cover my benches in a these things. So cover my benches in a these things. So cover my benches in a bit cuz I was also pretty amused by the bit cuz I was also pretty amused by the bit cuz I was also pretty amused by the scores. So it's seven times better a scores. So it's seven times better a scores. So it's seven times better a score for terminal bench. which means score for terminal bench. which means score for terminal bench. which means it's like actually viable for code. It it's like actually viable for code. It it's like actually viable for code. It scored slightly below Opus 5.5 on GDP scored slightly below Opus 5.5 on GDP scored slightly below Opus 5.5 on GDP Val, which is a bench I don't really Val, which is a bench I don't really Val, which is a bench I don't really care that much about. And it's strong on care that much about. And it's strong on care that much about. And it's strong on long horizon work and image long horizon work and image long horizon work and image understanding. It's the first Sonnet understanding. It's the first Sonnet understanding. It's the first Sonnet model to beat Pokémon Reding and working model to beat Pokémon Reding and working model to beat Pokémon Reding and working only from screenshots. Feel like a lot only from screenshots. Feel like a lot only from screenshots. Feel like a lot of models can do this now, but it is of models can do this now, but it is of models can do this now, but it is cool to see these like cheaper tier cool to see these like cheaper tier cool to see these like cheaper tier models being able to do a long running models being able to do a long running models being able to do a long running task like playing a game. Seems like task like playing a game. Seems like task like playing a game. Seems like Anthropics had a couple like crazy Anthropics had a couple like crazy Anthropics had a couple like crazy internal unlocks around their RL stuff internal unlocks around their RL stuff internal unlocks around their RL stuff recently because previously any model recently because previously any model recently because previously any model that wasn't their biggest and best. They that wasn't their biggest and best. They that wasn't their biggest and best. They didn't really get the right vibe is all didn't really get the right vibe is all didn't really get the right vibe is all I can say. Like it's didn't feel like I can say. Like it's didn't feel like I can say. Like it's didn't feel like they actually understood the work they
-
they actually understood the work they they actually understood the work they were being given. It more felt like they were being given. It more felt like they were being given. It more felt like they were robots trying to operate in a very were robots trying to operate in a very were robots trying to operate in a very specific set of things. Now these just specific set of things. Now these just specific set of things. Now these just feel like slightly dumber, way faster feel like slightly dumber, way faster feel like slightly dumber, way faster versions of Fable. And I did actually versions of Fable. And I did actually versions of Fable. And I did actually get some of that vibe from Sonnet, which get some of that vibe from Sonnet, which get some of that vibe from Sonnet, which is still crazy to me. It really feels is still crazy to me. It really feels is still crazy to me. It really feels like they saw everybody else distilling like they saw everybody else distilling like they saw everybody else distilling anthropic models, got jealous, and anthropic models, got jealous, and anthropic models, got jealous, and decided to do it themselves. And turns decided to do it themselves. And turns decided to do it themselves. And turns out they're good at it now that they've out they're good at it now that they've out they're good at it now that they've been doing it more heavily. Another been doing it more heavily. Another been doing it more heavily. Another important piece, collaboration. Not like important piece, collaboration. Not like important piece, collaboration. Not like how well do our agents collaborate with how well do our agents collaborate with how well do our agents collaborate with each other. It's more about how well can each other. It's more about how well can each other. It's more about how well can they write and format their outputs to they write and format their outputs to they write and format their outputs to us. They have been taking on the like us. They have been taking on the like us. They have been taking on the like clawisms and all of the horrible Claudes clawisms and all of the horrible Claudes clawisms and all of the horrible Claudes stuff that everyone hates and doing stuff that everyone hates and doing stuff that everyone hates and doing everything they can to get rid of it. everything they can to get rid of it. everything they can to get rid of it. And as a result, Sonic 5.5 writes way And as a result, Sonic 5.5 writes way And as a result, Sonic 5.5 writes way more clearly than their previous more clearly than their previous more clearly than their previous generation of models. It feels like a generation of models. It feels like a generation of models. It feels like a better partner for collaboration than better partner for collaboration than better partner for collaboration than Sonet 5. The speed helps too, although I Sonet 5. The speed helps too, although I Sonet 5. The speed helps too, although I think they pushed the speed bit a little think they pushed the speed bit a little think they pushed the speed bit a little too hard here. The numbers are not as too hard here. The numbers are not as too hard here. The numbers are not as impressive as they make it sound. And impressive as they make it sound. And impressive as they make it sound. And now we have price. Sonet 5.5 is priced now we have price. Sonet 5.5 is priced now we have price. Sonet 5.5 is priced the same as Sonnet 5 at $2 per million the same as Sonnet 5 at $2 per million the same as Sonnet 5 at $2 per million in and $10 per million out, as well as in and $10 per million out, as well as in and $10 per million out, as well as 20 cents per million tokens for cash 20 cents per million tokens for cash 20 cents per million tokens for cash reads. But it typically needs far fewer reads. But it typically needs far fewer reads. But it typically needs far fewer tokens to do the same work. It's up to tokens to do the same work. It's up to tokens to do the same work. It's up to 30% less per task than its predecessor.
-
30% less per task than its predecessor. 30% less per task than its predecessor. Man, are there some rough edges to this Man, are there some rough edges to this Man, are there some rough edges to this statement. I I'm not going to call it statement. I I'm not going to call it statement. I I'm not going to call it outright a lie, but it is intentionally outright a lie, but it is intentionally outright a lie, but it is intentionally dancing around some important pieces. in dancing around some important pieces. in dancing around some important pieces. in particular, MAX and cash token reads. As particular, MAX and cash token reads. As particular, MAX and cash token reads. As I've talked about extensively, cash I've talked about extensively, cash I've talked about extensively, cash token reads and cash writes tend to be token reads and cash writes tend to be token reads and cash writes tend to be the majority of costs for our day-to-day the majority of costs for our day-to-day the majority of costs for our day-to-day agent code work nowadays, which is why agent code work nowadays, which is why agent code work nowadays, which is why the 20 cents per million token cash read the 20 cents per million token cash read the 20 cents per million token cash read number here, it's a little bit number here, it's a little bit number here, it's a little bit concerning because all of the other concerning because all of the other concerning because all of the other models that Anthropics put out recently, models that Anthropics put out recently, models that Anthropics put out recently, the two Fable 51 and Opus 55, got huge the two Fable 51 and Opus 55, got huge the two Fable 51 and Opus 55, got huge discounts in cash read costs. This model discounts in cash read costs. This model discounts in cash read costs. This model didn't. When you look at the pricing didn't. When you look at the pricing didn't. When you look at the pricing chart at the bottom here, you'll chart at the bottom here, you'll chart at the bottom here, you'll understand why I'm concerned. If you understand why I'm concerned. If you understand why I'm concerned. If you look at input token costs, Sonic 55 is look at input token costs, Sonic 55 is look at input token costs, Sonic 55 is $2 and Opus is $4. So, it's half the $2 and Opus is $4. So, it's half the $2 and Opus is $4. So, it's half the price. Same with output tokens, $10 to price. Same with output tokens, $10 to price. Same with output tokens, $10 to 20. And even the cash rate costs is 20. And even the cash rate costs is 20. And even the cash rate costs is roughly the same factor here with $ 250 roughly the same factor here with $ 250 roughly the same factor here with $ 250 for cash rights for Cloud Sonic 55 and for cash rights for Cloud Sonic 55 and for cash rights for Cloud Sonic 55 and $5 for Opus 55. So, why am I so upset? $5 for Opus 55. So, why am I so upset? $5 for Opus 55. So, why am I so upset? The first row cash reads. Previously, I The first row cash reads. Previously, I The first row cash reads. Previously, I was saying that cash reads are a very was saying that cash reads are a very was saying that cash reads are a very small percentage of my usage. That was small percentage of my usage. That was small percentage of my usage. That was the case for Fable 5.1 because Fable 5.1 the case for Fable 5.1 because Fable 5.1 the case for Fable 5.1 because Fable 5.1 dropped the cash read cost by 90%.
-
dropped the cash read cost by 90%. dropped the cash read cost by 90%. Meanwhile, Opus also dropped it, but Meanwhile, Opus also dropped it, but Meanwhile, Opus also dropped it, but only around 60%. Usually, cash reads are only around 60%. Usually, cash reads are only around 60%. Usually, cash reads are 90% off. So, if it's $4 per million 90% off. So, if it's $4 per million 90% off. So, if it's $4 per million input tokens, it'll be 40 per million input tokens, it'll be 40 per million input tokens, it'll be 40 per million cashed input tokens. It's $2 per million cashed input tokens. It's $2 per million cashed input tokens. It's $2 per million input tokens, it's 20 cents per cache input tokens, it's 20 cents per cache input tokens, it's 20 cents per cache million input tokens. Pardon me for million input tokens. Pardon me for million input tokens. Pardon me for using the Google AI summary here, but using the Google AI summary here, but using the Google AI summary here, but every website reporting on this pricing every website reporting on this pricing every website reporting on this pricing is garbage. In particular, Anthropics. I is garbage. In particular, Anthropics. I is garbage. In particular, Anthropics. I don't know why they don't have a good don't know why they don't have a good don't know why they don't have a good like model API dashboard like OpenAI like model API dashboard like OpenAI like model API dashboard like OpenAI does. I hope that they can throw up a does. I hope that they can throw up a does. I hope that they can throw up a prompt quickly to build that for them prompt quickly to build that for them prompt quickly to build that for them cuz this is garbage. Fable 5.1 cost $10 cuz this is garbage. Fable 5.1 cost $10 cuz this is garbage. Fable 5.1 cost $10 per million in and $50 per mill out per million in and $50 per mill out per million in and $50 per mill out which is double the cost of Opus 5.5. which is double the cost of Opus 5.5. which is double the cost of Opus 5.5. However, cash reads are 25 per million However, cash reads are 25 per million However, cash reads are 25 per million in. Remember all these other numbers for in. Remember all these other numbers for in. Remember all these other numbers for Fable 5.1 are double but for cash reads Fable 5.1 are double but for cash reads Fable 5.1 are double but for cash reads only 25% more expensive. Everything else only 25% more expensive. Everything else only 25% more expensive. Everything else 2x cash read 25%. So sonnet to opus 2x cash read 25%. So sonnet to opus 2x cash read 25%. So sonnet to opus doubles price for cash rights for input doubles price for cash rights for input doubles price for cash rights for input and output tokens. Opus to fable doubles and output tokens. Opus to fable doubles and output tokens. Opus to fable doubles again. But the cash read cost doesn't again. But the cash read cost doesn't again. But the cash read cost doesn't change between sonnet and opus and it change between sonnet and opus and it change between sonnet and opus and it only goes up 25% for fable. What this only goes up 25% for fable. What this only goes up 25% for fable. What this means is cash read becomes a more and means is cash read becomes a more and means is cash read becomes a more and more prominent part of your cost when more prominent part of your cost when more prominent part of your cost when you go down the model ti where it rounds you go down the model ti where it rounds you go down the model ti where it rounds out to under 5% with fable. It's closer out to under 5% with fable. It's closer out to under 5% with fable. It's closer to 20% with Opus and it's closer to like to 20% with Opus and it's closer to like to 20% with Opus and it's closer to like 50 plus with Sonnet 5.5 because that 50 plus with Sonnet 5.5 because that 50 plus with Sonnet 5.5 because that number is disproportionately large when number is disproportionately large when number is disproportionately large when you look at the other numbers. It would you look at the other numbers. It would you look at the other numbers. It would have been really nice if they could have have been really nice if they could have have been really nice if they could have knocked cash read down to like 15 cents knocked cash read down to like 15 cents knocked cash read down to like 15 cents or god forbid 10 cents. That would have or god forbid 10 cents. That would have or god forbid 10 cents. That would have been insane. But they didn't. Which
-
been insane. But they didn't. Which been insane. But they didn't. Which means that the costs for using this means that the costs for using this means that the costs for using this model don't end up being as much cheaper model don't end up being as much cheaper model don't end up being as much cheaper as you might think when you look at as you might think when you look at as you might think when you look at these numbers because normal input these numbers because normal input these numbers because normal input tokens are barely touched with agentic tokens are barely touched with agentic tokens are barely touched with agentic work because we tend to read from cash work because we tend to read from cash work because we tend to read from cash and output tokens are a very small and output tokens are a very small and output tokens are a very small portion of the overall costs. I feel portion of the overall costs. I feel portion of the overall costs. I feel obligated to call this all out because obligated to call this all out because obligated to call this all out because I've looked at the numbers for doing I've looked at the numbers for doing I've looked at the numbers for doing like forlike work with opus versus like forlike work with opus versus like forlike work with opus versus sonnet for real code tasks and sonnet sonnet for real code tasks and sonnet sonnet for real code tasks and sonnet consistently comes out as if not more consistently comes out as if not more consistently comes out as if not more expensive than opus does. In fact, with expensive than opus does. In fact, with expensive than opus does. In fact, with the artificial analysis intelligence the artificial analysis intelligence the artificial analysis intelligence index, Sonnet 5.5 cost roughly the same index, Sonnet 5.5 cost roughly the same index, Sonnet 5.5 cost roughly the same as Fable 5.1 for realworld tasks. as Fable 5.1 for realworld tasks. as Fable 5.1 for realworld tasks. Flashbang warning since I know a handful Flashbang warning since I know a handful Flashbang warning since I know a handful of you like those. We're going to of you like those. We're going to of you like those. We're going to artificial analysis. Yeah, this is artificial analysis. Yeah, this is artificial analysis. Yeah, this is insane. Sonnet 5.5 is neck andneck with insane. Sonnet 5.5 is neck andneck with insane. Sonnet 5.5 is neck andneck with Fable 5.1 for the most expensive run Fable 5.1 for the most expensive run Fable 5.1 for the most expensive run they've ever done on artificial they've ever done on artificial they've ever done on artificial analysis. That should kill any reason to analysis. That should kill any reason to analysis. That should kill any reason to use this model entirely, right? Kind of. use this model entirely, right? Kind of. use this model entirely, right? Kind of. It does kill one thing. Give you a hint. It does kill one thing. Give you a hint. It does kill one thing. Give you a hint. It's one particular word you can see It's one particular word you can see It's one particular word you can see here. starts with M and ends in max. here. starts with M and ends in max. here. starts with M and ends in max. Don't use it. I'm trying to explain why Don't use it. I'm trying to explain why Don't use it. I'm trying to explain why I hate max mode so much, and I don't I hate max mode so much, and I don't I hate max mode so much, and I don't want to make this whole video about it, want to make this whole video about it, want to make this whole video about it, but I'll do my best to explain here.
-
but I'll do my best to explain here. but I'll do my best to explain here. Reasoning levels aren't really levels. Reasoning levels aren't really levels. Reasoning levels aren't really levels. You're not saying you should reason this You're not saying you should reason this You're not saying you should reason this much. They're budgets. They are allowing much. They're budgets. They are allowing much. They're budgets. They are allowing the model to reason up to a certain the model to reason up to a certain the model to reason up to a certain amount. So, when you set low, you're amount. So, when you set low, you're amount. So, when you set low, you're saying, "I want you to do this with a saying, "I want you to do this with a saying, "I want you to do this with a small amount of reasoning tokens." When small amount of reasoning tokens." When small amount of reasoning tokens." When you set medium, high, or x high, you're you set medium, high, or x high, you're you set medium, high, or x high, you're saying, "I'm okay with you using more saying, "I'm okay with you using more saying, "I'm okay with you using more reasoning tokens up to a certain point." reasoning tokens up to a certain point." reasoning tokens up to a certain point." But you're not necessarily increasing But you're not necessarily increasing But you're not necessarily increasing the amount of tokens used. I've had the amount of tokens used. I've had the amount of tokens used. I've had benchmarks where the gap between low and benchmarks where the gap between low and benchmarks where the gap between low and X high was like 5 to 8% tokens. Like X high was like 5 to 8% tokens. Like X high was like 5 to 8% tokens. Like it's not a big difference if the tasks it's not a big difference if the tasks it's not a big difference if the tasks are simple. But if the tasks are are simple. But if the tasks are are simple. But if the tasks are complex, then having the extra budget complex, then having the extra budget complex, then having the extra budget can absolutely help. So what is the can absolutely help. So what is the can absolutely help. So what is the issue with max? The issue with max is issue with max? The issue with max is issue with max? The issue with max is that it's not really increasing the that it's not really increasing the that it's not really increasing the ceiling for how many reason tokens are ceiling for how many reason tokens are ceiling for how many reason tokens are allowed to be used. It's more so allowed to be used. It's more so allowed to be used. It's more so increasing the floor for how many have increasing the floor for how many have increasing the floor for how many have to be used. It's effectively telling the to be used. It's effectively telling the to be used. It's effectively telling the model it's not done until it does a model it's not done until it does a model it's not done until it does a certain number of reasoning tokens. And certain number of reasoning tokens. And certain number of reasoning tokens. And the result is that I've had benchmarks the result is that I've had benchmarks the result is that I've had benchmarks where from low to X high you get a 5% where from low to X high you get a 5% where from low to X high you get a 5% increase in token usage. And from X high increase in token usage. And from X high increase in token usage. And from X high to max you get, and I'm not to max you get, and I'm not to max you get, and I'm not exaggerating, a 1,500% exaggerating, a 1,500% exaggerating, a 1,500% increase in token usage. Literally 15 increase in token usage. Literally 15 increase in token usage. Literally 15 times more tokens for that one little times more tokens for that one little times more tokens for that one little bump at the end because you're no longer bump at the end because you're no longer bump at the end because you're no longer letting it be done with easy tasks. And letting it be done with easy tasks. And letting it be done with easy tasks. And this can often actually hurt performance this can often actually hurt performance this can often actually hurt performance because if the model is told to keep because if the model is told to keep because if the model is told to keep overthinking the thing, it's going to overthinking the thing, it's going to overthinking the thing, it's going to overthink the hell out of it. It's going overthink the hell out of it. It's going overthink the hell out of it. It's going to start second guessing itself and is to start second guessing itself and is to start second guessing itself and is going to start getting wrong answers going to start getting wrong answers going to start getting wrong answers because you're forcing it to think too because you're forcing it to think too because you're forcing it to think too much. It worked more like this where it much. It worked more like this where it much. It worked more like this where it slightly increased the floor. I still
-
slightly increased the floor. I still slightly increased the floor. I still wouldn't recommend it, but it'd be fine. wouldn't recommend it, but it'd be fine. wouldn't recommend it, but it'd be fine. But it's not. It's forcing the model to But it's not. It's forcing the model to But it's not. It's forcing the model to do too much reasoning. And I can prove do too much reasoning. And I can prove do too much reasoning. And I can prove this very easily. So on 5.5, max this very easily. So on 5.5, max this very easily. So on 5.5, max reasoning effort was $7.60. reasoning effort was $7.60. reasoning effort was $7.60. If we compare to something like, I don't If we compare to something like, I don't If we compare to something like, I don't know, Astron Max, $3.26. know, Astron Max, $3.26. know, Astron Max, $3.26. So, Sonnet was two times more expensive So, Sonnet was two times more expensive So, Sonnet was two times more expensive than Astra for the same bench. That than Astra for the same bench. That than Astra for the same bench. That sounds insane until you realize Sonnet sounds insane until you realize Sonnet sounds insane until you realize Sonnet 5.5 can also be run on, I don't know, 5.5 can also be run on, I don't know, 5.5 can also be run on, I don't know, Xi, high, god forbid, medium. That looks Xi, high, god forbid, medium. That looks Xi, high, god forbid, medium. That looks a little less bad, right? XI is still a a little less bad, right? XI is still a a little less bad, right? XI is still a bit more expensive than I would like, bit more expensive than I would like, bit more expensive than I would like, but it's pretty close to Astra's price. but it's pretty close to Astra's price. but it's pretty close to Astra's price. It's cheaper, but still more than I It's cheaper, but still more than I It's cheaper, but still more than I would want. But once you go down to high would want. But once you go down to high would want. But once you go down to high and medium, you have crazy low prices, and medium, you have crazy low prices, and medium, you have crazy low prices, $18 and 59 respectively for running the $18 and 59 respectively for running the $18 and 59 respectively for running the same benches. Realistically speaking same benches. Realistically speaking same benches. Realistically speaking though, if we go look at the though, if we go look at the though, if we go look at the intelligence scores when you bump down intelligence scores when you bump down intelligence scores when you bump down these reasoning efforts, sure X high now these reasoning efforts, sure X high now these reasoning efforts, sure X high now is roughly Astra level, but Sonnet 55 is roughly Astra level, but Sonnet 55 is roughly Astra level, but Sonnet 55 was roughly three points higher than was roughly three points higher than was roughly three points higher than Astra. Do we actually think Sonnet's Astra. Do we actually think Sonnet's Astra. Do we actually think Sonnet's going to be that much better than Astra?
-
going to be that much better than Astra? going to be that much better than Astra? Okay, that I feel bad asking that Okay, that I feel bad asking that Okay, that I feel bad asking that because realistically speaking, Sonnet because realistically speaking, Sonnet because realistically speaking, Sonnet does have fewer dumb spikes that I get does have fewer dumb spikes that I get does have fewer dumb spikes that I get so frustrated with. So, I actually so frustrated with. So, I actually so frustrated with. So, I actually personally would take Sonnet over Astra personally would take Sonnet over Astra personally would take Sonnet over Astra for my day-to-day code work. Call me for my day-to-day code work. Call me for my day-to-day code work. Call me insane. I would just use Opus, but yeah, insane. I would just use Opus, but yeah, insane. I would just use Opus, but yeah, wanted to call this out because people wanted to call this out because people wanted to call this out because people are looking too closely at the max are looking too closely at the max are looking too closely at the max reasoning efforts and I really don't reasoning efforts and I really don't reasoning efforts and I really don't think anyone should use them. Like, I think anyone should use them. Like, I think anyone should use them. Like, I have not been shown a good enough use have not been shown a good enough use have not been shown a good enough use case for why max effort makes sense for case for why max effort makes sense for case for why max effort makes sense for most users. Pretend it doesn't exist. most users. Pretend it doesn't exist. most users. Pretend it doesn't exist. It'll make your life much easier. Back It'll make your life much easier. Back It'll make your life much easier. Back to the benches quick before I start to the benches quick before I start to the benches quick before I start covering the actual intricacies of using covering the actual intricacies of using covering the actual intricacies of using the model. In particular, its speed, the model. In particular, its speed, the model. In particular, its speed, which I don't love the way it's been which I don't love the way it's been which I don't love the way it's been reported on so far. Funny coming back reported on so far. Funny coming back reported on so far. Funny coming back here for the last rant because as you here for the last rant because as you here for the last rant because as you can see on some benches like Frontier can see on some benches like Frontier can see on some benches like Frontier Code, they put two scores in because XI Code, they put two scores in because XI Code, they put two scores in because XI scored better than Max did. Yes, really. scored better than Max did. Yes, really. scored better than Max did. Yes, really. Max put it below Soul, but XI put it Max put it below Soul, but XI put it Max put it below Soul, but XI put it above Soul. Very interesting. And GBD6 above Soul. Very interesting. And GBD6 above Soul. Very interesting. And GBD6 Soul is kind of just DOA, isn't it? I've Soul is kind of just DOA, isn't it? I've Soul is kind of just DOA, isn't it? I've barely had any reason to talk about it barely had any reason to talk about it barely had any reason to talk about it and haven't put it in videos for a and haven't put it in videos for a and haven't put it in videos for a reason. It's just not that impressive to reason. It's just not that impressive to reason. It's just not that impressive to me. Hopeful, fingers crossed, we'll get me. Hopeful, fingers crossed, we'll get me. Hopeful, fingers crossed, we'll get something better in the near future.
-
something better in the near future. something better in the near future. Everything else here looks pretty good. Everything else here looks pretty good. Everything else here looks pretty good. Computer use, it's scoring much better Computer use, it's scoring much better Computer use, it's scoring much better than it did before. Still neck andneck than it did before. Still neck andneck than it did before. Still neck andneck with Opus 5.5. I wish we had the numbers with Opus 5.5. I wish we had the numbers with Opus 5.5. I wish we had the numbers for OS World 2.1 for GPT6, not just for OS World 2.1 for GPT6, not just for OS World 2.1 for GPT6, not just Soul, but Astra because I still Soul, but Astra because I still Soul, but Astra because I still personally use Astra almost exclusively personally use Astra almost exclusively personally use Astra almost exclusively for computer use. Occasionally a review for computer use. Occasionally a review for computer use. Occasionally a review here and there, but even then iffy. I here and there, but even then iffy. I here and there, but even then iffy. I think it's kind of insane of them to put think it's kind of insane of them to put think it's kind of insane of them to put this chart as the first chart showing this chart as the first chart showing this chart as the first chart showing performance of the model inside of their performance of the model inside of their performance of the model inside of their reporting. First off, it kind of shows reporting. First off, it kind of shows reporting. First off, it kind of shows terminal bench isn't the greatest terminal bench isn't the greatest terminal bench isn't the greatest measure because Opus 5.5 dipped with measure because Opus 5.5 dipped with measure because Opus 5.5 dipped with max, but it also much worse shows that max, but it also much worse shows that max, but it also much worse shows that Sonnet 5.5 is more expensive than Opus Sonnet 5.5 is more expensive than Opus Sonnet 5.5 is more expensive than Opus on Max. It somehow is outperforming on Max. It somehow is outperforming on Max. It somehow is outperforming Opus. This is the type of chart that you Opus. This is the type of chart that you Opus. This is the type of chart that you would see and think something went would see and think something went would see and think something went wrong. Not the type of chart that you wrong. Not the type of chart that you wrong. Not the type of chart that you would publish is the first chart in the would publish is the first chart in the would publish is the first chart in the blog post, but sure kind of sad that the blog post, but sure kind of sad that the blog post, but sure kind of sad that the only place it is outperforming opus for only place it is outperforming opus for only place it is outperforming opus for the cost is in the max reasoning effort, the cost is in the max reasoning effort, the cost is in the max reasoning effort, which again you shouldn't use. Over to which again you shouldn't use. Over to which again you shouldn't use. Over to Frontier Code, you can see the same Frontier Code, you can see the same Frontier Code, you can see the same pattern where it plummets on max, but pattern where it plummets on max, but pattern where it plummets on max, but does decently on X high and high as does decently on X high and high as does decently on X high and high as well. Medium and low are pretty big well. Medium and low are pretty big well. Medium and low are pretty big drops. And similar to how I feel about drops. And similar to how I feel about drops. And similar to how I feel about not using max, I don't think you should not using max, I don't think you should not using max, I don't think you should use any anthropic model on low right use any anthropic model on low right use any anthropic model on low right now. They actually stopped providing the now. They actually stopped providing the now. They actually stopped providing the option to do no reasoning on Opus and option to do no reasoning on Opus and option to do no reasoning on Opus and I'm assuming now on Sonnet as well.
-
I'm assuming now on Sonnet as well. I'm assuming now on Sonnet as well. Previously there was a no reasoning Previously there was a no reasoning Previously there was a no reasoning option where it would just start option where it would just start option where it would just start responding immediately like an instant responding immediately like an instant responding immediately like an instant mode. They don't ship that anymore mode. They don't ship that anymore mode. They don't ship that anymore because these models suck without because these models suck without because these models suck without reasoning. And if they're given too reasoning. And if they're given too reasoning. And if they're given too little budget to reason, they will suck little budget to reason, they will suck little budget to reason, they will suck even harder as they have proven to in even harder as they have proven to in even harder as they have proven to in benches like this. So generally avoid benches like this. So generally avoid benches like this. So generally avoid low, avoid max. And for the most part, low, avoid max. And for the most part, low, avoid max. And for the most part, medium's fineish. Medium is a lot medium's fineish. Medium is a lot medium's fineish. Medium is a lot stronger with Opus than it is with stronger with Opus than it is with stronger with Opus than it is with Sonnet. So keep that in mind. Cursor Sonnet. So keep that in mind. Cursor Sonnet. So keep that in mind. Cursor bench was fun. I can look at the bench was fun. I can look at the bench was fun. I can look at the official numbers for it here where official numbers for it here where official numbers for it here where Sonnet did outperform with high kind of Sonnet did outperform with high kind of Sonnet did outperform with high kind of and it does complement the Opus 5.5 and it does complement the Opus 5.5 and it does complement the Opus 5.5 curve relatively well. But again, the curve relatively well. But again, the curve relatively well. But again, the drop from medium opus to low is just too drop from medium opus to low is just too drop from medium opus to low is just too big. And Opus is already such a big. And Opus is already such a big. And Opus is already such a surprisingly good value that it's hard surprisingly good value that it's hard surprisingly good value that it's hard for me to justify using much else. I for me to justify using much else. I for me to justify using much else. I also can't help but notice that all of also can't help but notice that all of also can't help but notice that all of these are score to cost and none of them these are score to cost and none of them these are score to cost and none of them are score to number of tokens. There's a are score to number of tokens. There's a are score to number of tokens. There's a reason for that. If we hop back over to reason for that. If we hop back over to reason for that. If we hop back over to Cursors Bench and click tokens, you'll Cursors Bench and click tokens, you'll Cursors Bench and click tokens, you'll see that Sonnet does have yet another see that Sonnet does have yet another see that Sonnet does have yet another greatest of all time score, their token greatest of all time score, their token greatest of all time score, their token usage in Cursor Bench, where they did usage in Cursor Bench, where they did usage in Cursor Bench, where they did 271,920 271,920 271,920 tokens per task, putting it ahead of tokens per task, putting it ahead of tokens per task, putting it ahead of even Opus at 218K and Gemini 38 Flash at even Opus at 218K and Gemini 38 Flash at even Opus at 218K and Gemini 38 Flash at 162K. It's almost double the number of 162K. It's almost double the number of 162K. It's almost double the number of tokens to Gemini 38 at Flash, which is tokens to Gemini 38 at Flash, which is tokens to Gemini 38 at Flash, which is insane. It is more than 5x the number of insane. It is more than 5x the number of insane. It is more than 5x the number of tokens that GPT56 soul and also six tokens that GPT56 soul and also six tokens that GPT56 soul and also six Astra used. So keep that in mind. This Astra used. So keep that in mind. This Astra used. So keep that in mind. This is not a token efficient model, which is not a token efficient model, which is not a token efficient model, which should be made up for by the speed, should be made up for by the speed, should be made up for by the speed, right? Cuz it's so fast. More bad news
-
right? Cuz it's so fast. More bad news right? Cuz it's so fast. More bad news sadly. According to Open Router, the TPS sadly. According to Open Router, the TPS sadly. According to Open Router, the TPS for using Sonnet 5.5 through anthropic for using Sonnet 5.5 through anthropic for using Sonnet 5.5 through anthropic is around 94 tokens per second, which is around 94 tokens per second, which is around 94 tokens per second, which sounds insane when you're used to sounds insane when you're used to sounds insane when you're used to something like Astra going at 30. They something like Astra going at 30. They something like Astra going at 30. They also report Opus 5.5 at around 70. For also report Opus 5.5 at around 70. For also report Opus 5.5 at around 70. For what it's worth, for my numbers using what it's worth, for my numbers using what it's worth, for my numbers using things through the official things through the official things through the official subscriptions, I have seen around 100 subscriptions, I have seen around 100 subscriptions, I have seen around 100 TPS on average for Opus 5.5 and around TPS on average for Opus 5.5 and around TPS on average for Opus 5.5 and around 150 for Sonnet 5.5. So, it is faster, 150 for Sonnet 5.5. So, it is faster, 150 for Sonnet 5.5. So, it is faster, but it's also not very token efficient. but it's also not very token efficient. but it's also not very token efficient. And result here is that when I used it And result here is that when I used it And result here is that when I used it for something like fish slop, it ended for something like fish slop, it ended for something like fish slop, it ended up taking quite a bit longer than Opus up taking quite a bit longer than Opus up taking quite a bit longer than Opus did. Sonnet took 43 minutes to finish did. Sonnet took 43 minutes to finish did. Sonnet took 43 minutes to finish and Opus only took 36. And when you and Opus only took 36. And when you and Opus only took 36. And when you measure how much time they spent measure how much time they spent measure how much time they spent generating tokens, it's an even bigger generating tokens, it's an even bigger generating tokens, it's an even bigger gap from 39 minutes to 27 minutes gap from 39 minutes to 27 minutes gap from 39 minutes to 27 minutes because again, this model uses a lot of because again, this model uses a lot of because again, this model uses a lot of tokens. Let's take a quick look at Witch tokens. Let's take a quick look at Witch tokens. Let's take a quick look at Witch AI, the design showcase made by Dra that AI, the design showcase made by Dra that AI, the design showcase made by Dra that has been very useful to see the has been very useful to see the has been very useful to see the capabilities of these new models when capabilities of these new models when capabilities of these new models when they drop. Reference have also opened up they drop. Reference have also opened up they drop. Reference have also opened up Fables and Opus' runs. Let's take a look Fables and Opus' runs. Let's take a look Fables and Opus' runs. Let's take a look at the first set generation. This one is at the first set generation. This one is at the first set generation. This one is interesting. It turned the page into interesting. It turned the page into interesting. It turned the page into like a fake app with the little things like a fake app with the little things like a fake app with the little things on the side to feel more like what the on the side to feel more like what the on the side to feel more like what the product might feel like. It's unique product might feel like. It's unique product might feel like. It's unique style and that's something I've seen the style and that's something I've seen the style and that's something I've seen the other models do. Not bad, but a bit other models do. Not bad, but a bit other models do. Not bad, but a bit opinionated. Let's see what else we got opinionated. Let's see what else we got opinionated. Let's see what else we got here.
-
This one's a bit rough. And it has some This one's a bit rough. And it has some issues with the scroll, too, where it issues with the scroll, too, where it issues with the scroll, too, where it just like rotated from today to last just like rotated from today to last just like rotated from today to last month back to today. These little lines month back to today. These little lines month back to today. These little lines in the background aren't my favorite in the background aren't my favorite in the background aren't my favorite thing. They hurt the readability. Don't thing. They hurt the readability. Don't thing. They hurt the readability. Don't love it. Next, we have this card design, love it. Next, we have this card design, love it. Next, we have this card design, and it's hideous. We got this highlighty and it's hideous. We got this highlighty and it's hideous. We got this highlighty one. Pretty boring. Don't love it. The one. Pretty boring. Don't love it. The one. Pretty boring. Don't love it. The animations are awful. Then here we have animations are awful. Then here we have animations are awful. Then here we have a weird hierarchy of like a geologic. I a weird hierarchy of like a geologic. I a weird hierarchy of like a geologic. I don't know the term for this type of don't know the term for this type of don't know the term for this type of like cross-section, but yeah, not great like cross-section, but yeah, not great like cross-section, but yeah, not great compared to Opus. Definitely a little worse, but compared Definitely a little worse, but compared to Fable, to Fable, to Fable, significantly worse. I still think Fable significantly worse. I still think Fable significantly worse. I still think Fable 5.1 is the best overall design model in 5.1 is the best overall design model in 5.1 is the best overall design model in particular for nicel lookinging particular for nicel lookinging particular for nicel lookinging front-end designs for your marketing front-end designs for your marketing front-end designs for your marketing site and whatnot, but I have found Opus site and whatnot, but I have found Opus site and whatnot, but I have found Opus to be relatively steerable towards good to be relatively steerable towards good to be relatively steerable towards good designs and it follows instructions designs and it follows instructions designs and it follows instructions around design much better. I've not around design much better. I've not around design much better. I've not pushed the limits of Sonnet for design pushed the limits of Sonnet for design pushed the limits of Sonnet for design personally very much, but from all of personally very much, but from all of personally very much, but from all of the demos I've seen here, not very the demos I've seen here, not very the demos I've seen here, not very impressed with it front end impressed with it front end impressed with it front end capabilities. still way ahead of capabilities. still way ahead of capabilities. still way ahead of anything OpenAI and especially anything anything OpenAI and especially anything anything OpenAI and especially anything that XAI has, but not the model I'd that XAI has, but not the model I'd that XAI has, but not the model I'd reach to for front end, especially since reach to for front end, especially since reach to for front end, especially since in like forl like work, Opus often ends in like forl like work, Opus often ends in like forl like work, Opus often ends up being around the same price. If you up being around the same price. If you up being around the same price. If you remember the intro of this video, you remember the intro of this video, you remember the intro of this video, you might be a bit confused at this point might be a bit confused at this point might be a bit confused at this point because it doesn't seem like this model because it doesn't seem like this model because it doesn't seem like this model is all that great. It's bad at front is all that great. It's bad at front is all that great. It's bad at front end. It uses too many tokens. It's end. It uses too many tokens. It's end. It uses too many tokens. It's faster, but actually slower because of faster, but actually slower because of faster, but actually slower because of the token differences, and it doesn't the token differences, and it doesn't the token differences, and it doesn't seem meaningfully smarter or cheaper
-
seem meaningfully smarter or cheaper seem meaningfully smarter or cheaper than Opus in day-to-day work. So, what than Opus in day-to-day work. So, what than Opus in day-to-day work. So, what the hell do I like this model so much the hell do I like this model so much the hell do I like this model so much for? Well, to be frank, I don't think for? Well, to be frank, I don't think for? Well, to be frank, I don't think this is a model that you or I should be this is a model that you or I should be this is a model that you or I should be selecting. If you are presented options selecting. If you are presented options selecting. If you are presented options inside of Claude Code between Fable, inside of Claude Code between Fable, inside of Claude Code between Fable, Opus, and Sonnet, I really don't think Opus, and Sonnet, I really don't think Opus, and Sonnet, I really don't think many people should be picking Sonnet. I many people should be picking Sonnet. I many people should be picking Sonnet. I can already tell how the comment section can already tell how the comment section can already tell how the comment section is going to look after I said that. is going to look after I said that. is going to look after I said that. Well, Theo, not everyone can afford Well, Theo, not everyone can afford Well, Theo, not everyone can afford Opus. Sonnet's more expensive. Shut the Opus. Sonnet's more expensive. Shut the Opus. Sonnet's more expensive. Shut the hell up. Seriously, I I don't want to hell up. Seriously, I I don't want to hell up. Seriously, I I don't want to have that argument today. We're talking have that argument today. We're talking have that argument today. We're talking about how these things operate in our about how these things operate in our about how these things operate in our actual realworld usage. It's why I'm actual realworld usage. It's why I'm actual realworld usage. It's why I'm excited to say for a handful of types of excited to say for a handful of types of excited to say for a handful of types of tasks, Sonnet does prove to be tasks, Sonnet does prove to be tasks, Sonnet does prove to be meaningfully more cost effective than meaningfully more cost effective than meaningfully more cost effective than Opus. Not necessarily the types of tasks Opus. Not necessarily the types of tasks Opus. Not necessarily the types of tasks I would send a model after, but I would send a model after, but I would send a model after, but absolutely the types of tasks that Opus absolutely the types of tasks that Opus absolutely the types of tasks that Opus and Fable would. Its strengths come as a and Fable would. Its strengths come as a and Fable would. Its strengths come as a tool for our other models to use. We tool for our other models to use. We tool for our other models to use. We shouldn't be calling Sonnet directly. It shouldn't be calling Sonnet directly. It shouldn't be calling Sonnet directly. It should be used as one of many things should be used as one of many things should be used as one of many things that Opus or Fable will orchestrate when that Opus or Fable will orchestrate when that Opus or Fable will orchestrate when it's trying to break up real complex it's trying to break up real complex it's trying to break up real complex work. I kind of made a bench for this work. I kind of made a bench for this work. I kind of made a bench for this for my Grock review where I was trying for my Grock review where I was trying for my Grock review where I was trying to find the strengths of Grock 4.7. And to find the strengths of Grock 4.7. And to find the strengths of Grock 4.7. And I did it by having all of these models I did it by having all of these models I did it by having all of these models do a really big deep audit of T3 code, do a really big deep audit of T3 code, do a really big deep audit of T3 code, specifically this giant orchestrator specifically this giant orchestrator specifically this giant orchestrator v2pr that's been iterated on for far too v2pr that's been iterated on for far too v2pr that's been iterated on for far too long. It's hundreds of thousands of long. It's hundreds of thousands of long. It's hundreds of thousands of lines of code. I gave all of these lines of code. I gave all of these lines of code. I gave all of these models a pretty detailed prompt asking models a pretty detailed prompt asking models a pretty detailed prompt asking them to break up all the things being them to break up all the things being them to break up all the things being added to orchestrator v2 to propose a added to orchestrator v2 to propose a added to orchestrator v2 to propose a strategy where we can get these things strategy where we can get these things strategy where we can get these things landed into main sooner rather than landed into main sooner rather than landed into main sooner rather than doing it all through the single giant doing it all through the single giant doing it all through the single giant PR. And I found that at the time Astra
-
PR. And I found that at the time Astra PR. And I found that at the time Astra was by far the best at doing these was by far the best at doing these was by far the best at doing these breakdowns. Fable was not great at it. breakdowns. Fable was not great at it. breakdowns. Fable was not great at it. Grock actually was outperforming Fable. Grock actually was outperforming Fable. Grock actually was outperforming Fable. And then Opus, Sonnet, and Soul didn't And then Opus, Sonnet, and Soul didn't And then Opus, Sonnet, and Soul didn't do particularly well. I have overhauled do particularly well. I have overhauled do particularly well. I have overhauled this since and updated the way that it this since and updated the way that it this since and updated the way that it is judged. And those changes have is judged. And those changes have is judged. And those changes have brought Fable up to a meaningfully brought Fable up to a meaningfully brought Fable up to a meaningfully higher score. Still below Opus and still higher score. Still below Opus and still higher score. Still below Opus and still far below Astro, which by far had the far below Astro, which by far had the far below Astro, which by far had the best plan on how to do these things. I best plan on how to do these things. I best plan on how to do these things. I still find that Astra is just like still find that Astra is just like still find that Astra is just like uniquely willing to dig into the details uniquely willing to dig into the details uniquely willing to dig into the details for things. So, it performed really well for things. So, it performed really well for things. So, it performed really well here. The thing that came as a massive here. The thing that came as a massive here. The thing that came as a massive surprise to me was Sonnet's performance surprise to me was Sonnet's performance surprise to me was Sonnet's performance where it ended up being around half the where it ended up being around half the where it ended up being around half the price of Opus. you know what it should price of Opus. you know what it should price of Opus. you know what it should be and also performing slightly better be and also performing slightly better be and also performing slightly better than Opus did according again to my than Opus did according again to my than Opus did according again to my automated judging panel that goes automated judging panel that goes automated judging panel that goes through the changes and proposals to through the changes and proposals to through the changes and proposals to make decisions on various axes. So make decisions on various axes. So make decisions on various axes. So remember this isn't traditional code remember this isn't traditional code remember this isn't traditional code work. This is a deep dive type task work. This is a deep dive type task work. This is a deep dive type task where the model has to go through a where the model has to go through a where the model has to go through a large code base and figure out what can large code base and figure out what can large code base and figure out what can be changed in it and how to explain it be changed in it and how to explain it be changed in it and how to explain it to someone else. This is the type of to someone else. This is the type of to someone else. This is the type of task that requires going through a ton task that requires going through a ton task that requires going through a ton of different things to come back with of different things to come back with of different things to come back with good information. It's not testing the good information. It's not testing the good information. It's not testing the coding capabilities in a traditional coding capabilities in a traditional coding capabilities in a traditional sense. It's much more analytical, I sense. It's much more analytical, I sense. It's much more analytical, I guess, where it's like trying to make guess, where it's like trying to make guess, where it's like trying to make good architectural decisions and good architectural decisions and good architectural decisions and comprehend what's going on in a comprehend what's going on in a comprehend what's going on in a codebase. So, the cost per point here is codebase. So, the cost per point here is codebase. So, the cost per point here is insane. It is like the best value I've insane. It is like the best value I've insane. It is like the best value I've seen in this bench by far. More seen in this bench by far. More seen in this bench by far. More importantly, in my opinion, is how much importantly, in my opinion, is how much importantly, in my opinion, is how much time it took because it ended up only time it took because it ended up only time it took because it ended up only taking around 5 minutes to do all of
-
taking around 5 minutes to do all of taking around 5 minutes to do all of this work. And like, yeah, sure, Soul this work. And like, yeah, sure, Soul this work. And like, yeah, sure, Soul was able to do it even faster, but at was able to do it even faster, but at was able to do it even faster, but at like half the score. Opus took almost like half the score. Opus took almost like half the score. Opus took almost twice as long, and Aster took almost twice as long, and Aster took almost twice as long, and Aster took almost three times as long. So, Sonnet as a three times as long. So, Sonnet as a three times as long. So, Sonnet as a tool that your agents can call on to do tool that your agents can call on to do tool that your agents can call on to do this type of research to help it plan this type of research to help it plan this type of research to help it plan and scope work. If you give Opus or and scope work. If you give Opus or and scope work. If you give Opus or Fable a big task and they want to Fable a big task and they want to Fable a big task and they want to analyze the codebase before starting, analyze the codebase before starting, analyze the codebase before starting, they can now call on Sonnet to do that they can now call on Sonnet to do that they can now call on Sonnet to do that and get results that they're more than and get results that they're more than and get results that they're more than happy with. This is huge and it's one of happy with. This is huge and it's one of happy with. This is huge and it's one of the biggest strengths and I hope others the biggest strengths and I hope others the biggest strengths and I hope others start to make benchmarks like this start to make benchmarks like this start to make benchmarks like this because it is such a good way to see because it is such a good way to see because it is such a good way to see these capabilities and to see the these capabilities and to see the these capabilities and to see the strength of what Sonnet is introducing strength of what Sonnet is introducing strength of what Sonnet is introducing here. Little peering behind the curtain here. Little peering behind the curtain here. Little peering behind the curtain here. As you can guess, I have been here. As you can guess, I have been here. As you can guess, I have been testing a ton of models over the last testing a ton of models over the last testing a ton of models over the last few weeks honestly and it's been chaos. few weeks honestly and it's been chaos. few weeks honestly and it's been chaos. One of my bigger tests is that I've been One of my bigger tests is that I've been One of my bigger tests is that I've been working on rewriting all of TypeScript working on rewriting all of TypeScript working on rewriting all of TypeScript like the compiler for the language in like the compiler for the language in like the compiler for the language in Rust. There's already a Go rewrite by Rust. There's already a Go rewrite by Rust. There's already a Go rewrite by Microsoft that's really good, but I Microsoft that's really good, but I Microsoft that's really good, but I wanted to see how much further I could wanted to see how much further I could wanted to see how much further I could push and also potentially get it working push and also potentially get it working push and also potentially get it working in WASM. This was meant to be a Hail in WASM. This was meant to be a Hail in WASM. This was meant to be a Hail Mary project that would never happen, Mary project that would never happen, Mary project that would never happen, but Opus has fully unblocked and has it but Opus has fully unblocked and has it but Opus has fully unblocked and has it going really, really far. But one of the going really, really far. But one of the going really, really far. But one of the things I noticed is that it left over a things I noticed is that it left over a things I noticed is that it left over a ton of the slop that Astra and Soul had ton of the slop that Astra and Soul had ton of the slop that Astra and Soul had made, like millions of lines of it, and made, like millions of lines of it, and made, like millions of lines of it, and it wasn't getting rid of it. So, I told it wasn't getting rid of it. So, I told it wasn't getting rid of it. So, I told it to halt and go clean up all of the it to halt and go clean up all of the it to halt and go clean up all of the legacy slop. And then I had a thread legacy slop. And then I had a thread legacy slop. And then I had a thread over here in T3 Code where I'm over here in T3 Code where I'm over here in T3 Code where I'm consistently monitoring changes as they consistently monitoring changes as they consistently monitoring changes as they come in. And this thread is one where I come in. And this thread is one where I come in. And this thread is one where I found all these legacy things that need found all these legacy things that need found all these legacy things that need to be deleted. So I asked for an update.
-
to be deleted. So I asked for an update. to be deleted. So I asked for an update. How about now? We should have lots of How about now? We should have lots of How about now? We should have lots of legacy stuff cleared out. And I legacy stuff cleared out. And I legacy stuff cleared out. And I scrolled. None of this is talking about scrolled. None of this is talking about scrolled. None of this is talking about how much legacy stuff was deleted. And I how much legacy stuff was deleted. And I how much legacy stuff was deleted. And I was really confused. I was like, what was really confused. I was like, what was really confused. I was like, what the hell? We delete a bunch of code. the hell? We delete a bunch of code. the hell? We delete a bunch of code. Aren't you telling me about it? Is this Aren't you telling me about it? Is this Aren't you telling me about it? Is this a regression in Sonnet 5.5's behavior? a regression in Sonnet 5.5's behavior? a regression in Sonnet 5.5's behavior? Maybe it's not as good as Opus at like Maybe it's not as good as Opus at like Maybe it's not as good as Opus at like understanding my intent. Nope. I was in understanding my intent. Nope. I was in understanding my intent. Nope. I was in the wrong thread. This is one where I the wrong thread. This is one where I the wrong thread. This is one where I was always talking about numbers. So was always talking about numbers. So was always talking about numbers. So this is actually a really good thing is this is actually a really good thing is this is actually a really good thing is that it was able to make what I would that it was able to make what I would that it was able to make what I would consider a pretty logical decision based consider a pretty logical decision based consider a pretty logical decision based on what the contents of the thread are. on what the contents of the thread are. on what the contents of the thread are. It was continuing to operate the way the It was continuing to operate the way the It was continuing to operate the way the thread had instead of trying to figure thread had instead of trying to figure thread had instead of trying to figure out what the hell I meant when I said out what the hell I meant when I said out what the hell I meant when I said should have lots of legacy stuff cleared should have lots of legacy stuff cleared should have lots of legacy stuff cleared out. It just did I would argue the right out. It just did I would argue the right out. It just did I would argue the right thing here which is nice. I was about to thing here which is nice. I was about to thing here which is nice. I was about to crash out about it not understanding my crash out about it not understanding my crash out about it not understanding my intent but it totally does. I asked it intent but it totally does. I asked it intent but it totally does. I asked it how much code has been deleted. It's how much code has been deleted. It's how much code has been deleted. It's about 1.6 millions lines of code have about 1.6 millions lines of code have about 1.6 millions lines of code have been deleted since I started filming cuz been deleted since I started filming cuz been deleted since I started filming cuz I just kicked this job off before I just kicked this job off before I just kicked this job off before filming. Funny enough. So yeah, the real filming. Funny enough. So yeah, the real filming. Funny enough. So yeah, the real reason I came here was to grab my fish reason I came here was to grab my fish reason I came here was to grab my fish slot thread. Both to see how everything slot thread. Both to see how everything slot thread. Both to see how everything came out in terms of cost, but more came out in terms of cost, but more came out in terms of cost, but more importantly to show you guys the actual importantly to show you guys the actual importantly to show you guys the actual demo. Ended up costing roughly the same demo. Ended up costing roughly the same demo. Ended up costing roughly the same as the Opus one did, but sadly the Opus as the Opus one did, but sadly the Opus as the Opus one did, but sadly the Opus one did actually use a Fable 5.1 sub one did actually use a Fable 5.1 sub one did actually use a Fable 5.1 sub agent at some point. So it's score isn't agent at some point. So it's score isn't agent at some point. So it's score isn't necessarily the most accurate. I guess I necessarily the most accurate. I guess I necessarily the most accurate. I guess I do have to rerun Opus 5.5 on fish slop do have to rerun Opus 5.5 on fish slop do have to rerun Opus 5.5 on fish slop at some point to get even better numbers at some point to get even better numbers at some point to get even better numbers there. But the input token difference is there. But the input token difference is there. But the input token difference is insane. Sonnet 55 used way more, like insane. Sonnet 55 used way more, like insane. Sonnet 55 used way more, like more than three and a half times more more than three and a half times more more than three and a half times more input tokens than Opus did, and the input tokens than Opus did, and the input tokens than Opus did, and the resulting cost was actually quite resulting cost was actually quite resulting cost was actually quite similar. Although, I would guess if I
-
similar. Although, I would guess if I similar. Although, I would guess if I had not accidentally spawned the Fable had not accidentally spawned the Fable had not accidentally spawned the Fable sub agents, Opus would have even been sub agents, Opus would have even been sub agents, Opus would have even been cheaper. How are the results? It's one cheaper. How are the results? It's one cheaper. How are the results? It's one of the best I've ever seen. There are of the best I've ever seen. There are of the best I've ever seen. There are some ways where it is worse. Like some some ways where it is worse. Like some some ways where it is worse. Like some of the models are just not as good as of the models are just not as good as of the models are just not as good as the models that we were getting with the models that we were getting with the models that we were getting with Opus, but it is still without question Opus, but it is still without question Opus, but it is still without question one of the best I've ever seen. I would one of the best I've ever seen. I would one of the best I've ever seen. I would argue in some ways it is better than the argue in some ways it is better than the argue in some ways it is better than the Opus one and in all ways it is better Opus one and in all ways it is better Opus one and in all ways it is better than everything I've gotten out of Astra than everything I've gotten out of Astra than everything I've gotten out of Astra and obviously out of every other lab. I and obviously out of every other lab. I and obviously out of every other lab. I remember just like two or three months remember just like two or three months remember just like two or three months ago being so impressed that Kim K3 could ago being so impressed that Kim K3 could ago being so impressed that Kim K3 could use Blender at all and here I am now use Blender at all and here I am now use Blender at all and here I am now with like a bunch of actual like 3D with like a bunch of actual like 3D with like a bunch of actual like 3D models that you can tell what they are models that you can tell what they are models that you can tell what they are properly. It's not even like bad. Like properly. It's not even like bad. Like properly. It's not even like bad. Like these fish are kind of cute and have these fish are kind of cute and have these fish are kind of cute and have their own like little cute unique design their own like little cute unique design their own like little cute unique design to them. The coral, the seaweed, all of to them. The coral, the seaweed, all of to them. The coral, the seaweed, all of this is like decent. I just realized this is like decent. I just realized this is like decent. I just realized none of the sound is coming through. So, none of the sound is coming through. So, none of the sound is coming through. So, let me figure out why that is quick. One let me figure out why that is quick. One let me figure out why that is quick. One sec. a few audio issues later, but now sec. a few audio issues later, but now sec. a few audio issues later, but now at the very least I should actually be at the very least I should actually be at the very least I should actually be able to hear and you can hopefully too. able to hear and you can hopefully too. able to hear and you can hopefully too. One of the things I I can't help but One of the things I I can't help but One of the things I I can't help but notice it's like all the anthropic notice it's like all the anthropic notice it's like all the anthropic models have this, but Sonnet especially models have this, but Sonnet especially models have this, but Sonnet especially does. It just feels good to play. Like does. It just feels good to play. Like does. It just feels good to play. Like you're seeing a 30fps YouTube video of you're seeing a 30fps YouTube video of you're seeing a 30fps YouTube video of this. I might have to start uploading this. I might have to start uploading this. I might have to start uploading these for y'all to try though because these for y'all to try though because these for y'all to try though because this is running at a buttery smooth 120 this is running at a buttery smooth 120 this is running at a buttery smooth 120 fps on my laptop and like all the fps on my laptop and like all the fps on my laptop and like all the movement and all the mechanics actually movement and all the mechanics actually movement and all the mechanics actually feel pretty solid and balanced.
-
I actually love the animation for the I actually love the animation for the fish, too. It's adorable. Little uh fish, too. It's adorable. Little uh fish, too. It's adorable. Little uh pelican there. The craziest thing is the pelican there. The craziest thing is the pelican there. The craziest thing is the aliens. aliens. aliens. Where are they? It says that there's an Where are they? It says that there's an Where are they? It says that there's an attack. Where is he? Yeah, there he is. attack. Where is he? Yeah, there he is. attack. Where is he? Yeah, there he is. The alien is the best looking I've seen The alien is the best looking I've seen The alien is the best looking I've seen by far in any of these demos. Oh, how by far in any of these demos. Oh, how by far in any of these demos. Oh, how crazy is it that like I threw a prompt crazy is it that like I threw a prompt crazy is it that like I threw a prompt at a model and then 40 minutes later for at a model and then 40 minutes later for at a model and then 40 minutes later for a few dollars it spit this out? Like if a few dollars it spit this out? Like if a few dollars it spit this out? Like if I had paid API prices would have been I had paid API prices would have been I had paid API prices would have been $16. I've paid more than $16 for games $16. I've paid more than $16 for games $16. I've paid more than $16 for games worse than this. Shamefully enough. I worse than this. Shamefully enough. I worse than this. Shamefully enough. I was a kid at some point. I bought crappy was a kid at some point. I bought crappy was a kid at some point. I bought crappy PlayStation games that were made after PlayStation games that were made after PlayStation games that were made after like the movies that I was watching. like the movies that I was watching. like the movies that I was watching. We've all been there, right? But like We've all been there, right? But like We've all been there, right? But like god damn, for the price of a skin in god damn, for the price of a skin in god damn, for the price of a skin in Fortnite, you can make your own game. Fortnite, you can make your own game. Fortnite, you can make your own game. And that's assuming you're paying the And that's assuming you're paying the And that's assuming you're paying the API prices. Obviously, this is super API prices. Obviously, this is super API prices. Obviously, this is super super cheap if you're doing it over your super cheap if you're doing it over your super cheap if you're doing it over your subscription, which uh by the way, cuz I subscription, which uh by the way, cuz I subscription, which uh by the way, cuz I I know a lot of people are always I know a lot of people are always I know a lot of people are always curious what the numbers look like curious what the numbers look like curious what the numbers look like there. I did actually run some cost there. I did actually run some cost there. I did actually run some cost breakdowns for my own usage across my breakdowns for my own usage across my breakdowns for my own usage across my now six cloud accounts. And from my now six cloud accounts. And from my now six cloud accounts. And from my rough analysis of my real accounts that rough analysis of my real accounts that rough analysis of my real accounts that I emptied, which I killed three accounts I emptied, which I killed three accounts I emptied, which I killed three accounts in the past few days, you get around in the past few days, you get around in the past few days, you get around $2,300 $2,300 $2,300 a week of usage on the $200 plan. That a week of usage on the $200 plan. That a week of usage on the $200 plan. That puts you at almost 10 grand a month for puts you at almost 10 grand a month for puts you at almost 10 grand a month for 30 days with your Cloud Code sub on the 30 days with your Cloud Code sub on the 30 days with your Cloud Code sub on the $200 plan. That's going to stretch you $200 plan. That's going to stretch you $200 plan. That's going to stretch you pretty far with these models. While it pretty far with these models. While it pretty far with these models. While it might not look like Sonnet's going to might not look like Sonnet's going to might not look like Sonnet's going to get you much further per dollar than get you much further per dollar than get you much further per dollar than Opus, that is the case for general work.
-
Opus, that is the case for general work. Opus, that is the case for general work. If you're having it as the model you If you're having it as the model you If you're having it as the model you select to use for things you do, sure. select to use for things you do, sure. select to use for things you do, sure. But if you have it as a tool, Opus calls But if you have it as a tool, Opus calls But if you have it as a tool, Opus calls to do things like deep dives into your to do things like deep dives into your to do things like deep dives into your codebase to find specific behaviors or codebase to find specific behaviors or codebase to find specific behaviors or characteristics or confirming a hunch it characteristics or confirming a hunch it characteristics or confirming a hunch it has about some thing that you found in has about some thing that you found in has about some thing that you found in an API or all of those types of an API or all of those types of an API or all of those types of investigatory things. It is really cheap investigatory things. It is really cheap investigatory things. It is really cheap and surprisingly fast, too. I would bet and surprisingly fast, too. I would bet and surprisingly fast, too. I would bet that if you set Opus up properly to use that if you set Opus up properly to use that if you set Opus up properly to use Sonnet at the right times for these sub Sonnet at the right times for these sub Sonnet at the right times for these sub agents that the result will be Opus agents that the result will be Opus agents that the result will be Opus feeling way faster and getting real work feeling way faster and getting real work feeling way faster and getting real work done for slightly cheaper. That sounds done for slightly cheaper. That sounds done for slightly cheaper. That sounds like a pretty good deal to me. So, while like a pretty good deal to me. So, while like a pretty good deal to me. So, while I don't think you should actually use I don't think you should actually use I don't think you should actually use Sonnet 5.5 yourself, I think it is an Sonnet 5.5 yourself, I think it is an Sonnet 5.5 yourself, I think it is an incredible addition to this new family incredible addition to this new family incredible addition to this new family of models, and I'm excited to see how of models, and I'm excited to see how of models, and I'm excited to see how Opus chooses to wield it. I will not be Opus chooses to wield it. I will not be Opus chooses to wield it. I will not be using this model much going forward using this model much going forward using this model much going forward myself, but I do hope that my Opus is myself, but I do hope that my Opus is myself, but I do hope that my Opus is able to because there is clearly value able to because there is clearly value able to because there is clearly value here. While not in like front-end here. While not in like front-end here. While not in like front-end building work or day-to-day coding building work or day-to-day coding building work or day-to-day coding tasks, there is a ton of value in how tasks, there is a ton of value in how tasks, there is a ton of value in how this model can be utilized to confirm this model can be utilized to confirm this model can be utilized to confirm things in real world code bases. And I things in real world code bases. And I things in real world code bases. And I would assume for real world documents, would assume for real world documents, would assume for real world documents, office, and all that type of stuff, too.
-
office, and all that type of stuff, too. office, and all that type of stuff, too. That's not what you guys are here for. That's not what you guys are here for. That's not what you guys are here for. You're here for the code. And I'm hyped You're here for the code. And I'm hyped You're here for the code. And I'm hyped to say this model is useful as long as to say this model is useful as long as to say this model is useful as long as you're not the one sending it prompts. you're not the one sending it prompts. you're not the one sending it prompts. But godamn, OpenAI really needs to But godamn, OpenAI really needs to But godamn, OpenAI really needs to respond to this. They have lost in all respond to this. They have lost in all respond to this. They have lost in all of the places they are strongest. They of the places they are strongest. They of the places they are strongest. They don't have the smartest model. They don't have the smartest model. They don't have the smartest model. They don't have the most efficient model. don't have the most efficient model. don't have the most efficient model. They don't have the cheapest model. They They don't have the cheapest model. They They don't have the cheapest model. They don't have the best model for calling don't have the best model for calling don't have the best model for calling other models. It's rough for OpenAI other models. It's rough for OpenAI other models. It's rough for OpenAI right now. And I can't wait to see how right now. And I can't wait to see how right now. And I can't wait to see how they catch up because right now it just they catch up because right now it just they catch up because right now it just doesn't feel like a good value. The $200 doesn't feel like a good value. The $200 doesn't feel like a good value. The $200 codeex plan gets me nowhere near as much codeex plan gets me nowhere near as much codeex plan gets me nowhere near as much usage and nowhere near as much realworld usage and nowhere near as much realworld usage and nowhere near as much realworld code as I'm getting out of my $200 Cloud code as I'm getting out of my $200 Cloud code as I'm getting out of my $200 Cloud subs. So take that as you will. Fingers subs. So take that as you will. Fingers subs. So take that as you will. Fingers crossed we got some fun announcements crossed we got some fun announcements crossed we got some fun announcements coming. I have a feeling that things are coming. I have a feeling that things are coming. I have a feeling that things are about to speed up, not slow down. about to speed up, not slow down. about to speed up, not slow down. Hopefully this was a useful breakdown. Hopefully this was a useful breakdown. Hopefully this was a useful breakdown. You can better know how to wield this You can better know how to wield this You can better know how to wield this model. And until next time, peace nerds.
Summary
Anthropic's Sonnet 5.5 is presented as a strategically priced model challenging OpenAI's offerings, particularly its GPT-6 Soul. Despite lower benchmark scores, Sonnet 5.5's affordability and effectiveness in modern workflows make it a compelling alternative to more expensive models. The takeaway is that cost-effective AI models can still deliver significant value and competitive advantages.