← Back
Aaron Zisk September 22, 2026 14m

M5 Ultra… Apple Wasn’t Messing Around

Read full transcript 12 segments
  1. Well, the M5 Ultra is finally here, and Well, the M5 Ultra is finally here, and it's right here on my desk, right next it's right here on my desk, right next it's right here on my desk, right next to the M3 Ultra. Apple says it has local to the M3 Ultra. Apple says it has local to the M3 Ultra. Apple says it has local AI performance up to 4.3 times faster AI performance up to 4.3 times faster AI performance up to 4.3 times faster than the previous generation. Two times than the previous generation. Two times than the previous generation. Two times faster storage, 1.3 times faster CPU faster storage, 1.3 times faster CPU faster storage, 1.3 times faster CPU speeds. Sounds pretty good. But is that speeds. Sounds pretty good. But is that speeds. Sounds pretty good. But is that real or is that just marketing? Well, real or is that just marketing? Well, real or is that just marketing? Well, according to the spec sheet, the new according to the spec sheet, the new according to the spec sheet, the new memory bandwidth is 1.2 terabytes per memory bandwidth is 1.2 terabytes per memory bandwidth is 1.2 terabytes per second. I'm going to test that out, of second. I'm going to test that out, of second. I'm going to test that out, of course. M5 Ultra is what I have here. course. M5 Ultra is what I have here. course. M5 Ultra is what I have here. That one starts at $54.99. And if you That one starts at $54.99. And if you That one starts at $54.99. And if you configure it the way I have it here, configure it the way I have it here, configure it the way I have it here, you're going to need to bump that GPU to you're going to need to bump that GPU to you're going to need to bump that GPU to 80 cores. You're going to need to bump 80 cores. You're going to need to bump 80 cores. You're going to need to bump the memory to 256 GB. 512 is coming next the memory to 256 GB. 512 is coming next the memory to 256 GB. 512 is coming next month. And you're going to need to bump month. And you're going to need to bump month. And you're going to need to bump the storage to 8 TB. I mean, you don't the storage to 8 TB. I mean, you don't the storage to 8 TB. I mean, you don't need to, but that's what I have here. need to, but that's what I have here. need to, but that's what I have here. So, I thought I'd give you the price, So, I thought I'd give you the price, So, I thought I'd give you the price, which comes to 14,299. which comes to 14,299. which comes to 14,299. Not only do you get a performance bump Not only do you get a performance bump Not only do you get a performance bump apparently, but you also apparently get apparently, but you also apparently get apparently, but you also apparently get a price bump quite a bit. Apple loaned a price bump quite a bit. Apple loaned a price bump quite a bit. Apple loaned me this one. It's nice to be recognized me this one. It's nice to be recognized me this one. It's nice to be recognized after covering them for many years and after covering them for many years and after covering them for many years and buying my own. I bought this one though.

  2. buying my own. I bought this one though. buying my own. I bought this one though. And by the way, the reason both of these And by the way, the reason both of these And by the way, the reason both of these machines are on my desk right now is machines are on my desk right now is machines are on my desk right now is because of you. You folks watching this because of you. You folks watching this because of you. You folks watching this channel and supporting me allows me to channel and supporting me allows me to channel and supporting me allows me to get real hardware to do these real tests get real hardware to do these real tests get real hardware to do these real tests for you instead of just reading the for you instead of just reading the for you instead of just reading the numbers off a spec sheet. So, thank you. numbers off a spec sheet. So, thank you. numbers off a spec sheet. So, thank you. Look at all those cores. We got 36 of Look at all those cores. We got 36 of Look at all those cores. We got 36 of them. By the way, there's no efficiency them. By the way, there's no efficiency them. By the way, there's no efficiency cores in the Mac Studio here. We have 24 cores in the Mac Studio here. We have 24 cores in the Mac Studio here. We have 24 performance cores and 12 super cores. performance cores and 12 super cores. performance cores and 12 super cores. Now, to keep this fair, both machines Now, to keep this fair, both machines Now, to keep this fair, both machines ran the exact same builds of Llama CPP ran the exact same builds of Llama CPP ran the exact same builds of Llama CPP and MLX, and I use the exact same model and MLX, and I use the exact same model and MLX, and I use the exact same model files on both. By the way, Llama CPP is files on both. By the way, Llama CPP is files on both. By the way, Llama CPP is the cross-platform library that works on the cross-platform library that works on the cross-platform library that works on pretty much all the systems, not only pretty much all the systems, not only pretty much all the systems, not only Apple, and is very popular. MLX is Apple, and is very popular. MLX is Apple, and is very popular. MLX is Apple's own stack for running AI. Apple's own stack for running AI. Apple's own stack for running AI. Typically, MLX has better performance, Typically, MLX has better performance, Typically, MLX has better performance, but not always. but not always. but not always. Now, if you're buying this machine just Now, if you're buying this machine just Now, if you're buying this machine just for developer tasks, let me tell you for developer tasks, let me tell you for developer tasks, let me tell you something, it's overkill. Here's a something, it's overkill. Here's a something, it's overkill. Here's a speedometer score. 47 is the M3 Ultra speedometer score. 47 is the M3 Ultra speedometer score. 47 is the M3 Ultra score. We got 60.2 on the M5 Ultra. I score. We got 60.2 on the M5 Ultra. I score. We got 60.2 on the M5 Ultra. I just did a M6 mini review. And yeah, it just did a M6 mini review. And yeah, it just did a M6 mini review. And yeah, it destroys that score by quite a lot. Now, destroys that score by quite a lot. Now, destroys that score by quite a lot. Now, this is the M5 chip. That's the M6 chip.

  3. this is the M5 chip. That's the M6 chip. this is the M5 chip. That's the M6 chip. So, a single core of that is going to be So, a single core of that is going to be So, a single core of that is going to be much faster in the M6 family than the M5 much faster in the M6 family than the M5 much faster in the M6 family than the M5 family. All right, quick CPU test for family. All right, quick CPU test for family. All right, quick CPU test for devs. Since this thing has 36 cores, devs. Since this thing has 36 cores, devs. Since this thing has 36 cores, let's use all of them. Here, I've got a let's use all of them. Here, I've got a let's use all of them. Here, I've got a large.net net compilation with a 100,000 large.net net compilation with a 100,000 large.net net compilation with a 100,000 name spaces and classes. And I ran the name spaces and classes. And I ran the name spaces and classes. And I ran the compilation. We've got 71.7 seconds on compilation. We've got 71.7 seconds on compilation. We've got 71.7 seconds on the M3 Ultra, 52.4 the M3 Ultra, 52.4 the M3 Ultra, 52.4 on I think that's the fastest I've ever on I think that's the fastest I've ever on I think that's the fastest I've ever seen it on the M5 Ultra. That's a seen it on the M5 Ultra. That's a seen it on the M5 Ultra. That's a compilation. What about uh Python, which compilation. What about uh Python, which compilation. What about uh Python, which is an interpreted test. So, let's run is an interpreted test. So, let's run is an interpreted test. So, let's run that. And boom. All right. Now, this is that. And boom. All right. Now, this is that. And boom. All right. Now, this is a Mandler test, which is basically doing a Mandler test, which is basically doing a Mandler test, which is basically doing the Mandlero algorithm in Python. Oh. Oh the Mandlero algorithm in Python. Oh. Oh the Mandlero algorithm in Python. Oh. Oh my gosh. What? Uh, this is two times my gosh. What? Uh, this is two times my gosh. What? Uh, this is two times faster. Wow. Okay. 10.25 seconds on the faster. Wow. Okay. 10.25 seconds on the faster. Wow. Okay. 10.25 seconds on the M3 Ultra. 5.8 seconds. This is crazy. M3 Ultra. 5.8 seconds. This is crazy. M3 Ultra. 5.8 seconds. This is crazy. I've never seen it that fast. Shh. It's I've never seen it that fast. Shh. It's I've never seen it that fast. Shh. It's okay. It'll get the job done. Storage is okay. It'll get the job done. Storage is okay. It'll get the job done. Storage is supposed to be twice as fast. And I just supposed to be twice as fast. And I just supposed to be twice as fast. And I just ran Amorphous Disc Mark M3 ultra ran Amorphous Disc Mark M3 ultra ran Amorphous Disc Mark M3 ultra sequential speed. So, this is like sequential speed. So, this is like sequential speed. So, this is like copying a large file back and forth.

  4. copying a large file back and forth. copying a large file back and forth. 6,500 megabytes per second. So, 6.5 GB 6,500 megabytes per second. So, 6.5 GB 6,500 megabytes per second. So, 6.5 GB per second for read and 2.7 for right. per second for read and 2.7 for right. per second for read and 2.7 for right. Wow. The M5 Ultra 14,888 Wow. The M5 Ultra 14,888 Wow. The M5 Ultra 14,888 read and almost 20,000 in right. Wow, read and almost 20,000 in right. Wow, read and almost 20,000 in right. Wow, that is crazy. That's for sequential. that is crazy. That's for sequential. that is crazy. That's for sequential. Now, for compilation, for developer Now, for compilation, for developer Now, for compilation, for developer related things, you're going to want to related things, you're going to want to related things, you're going to want to look at the random numbers, too, because look at the random numbers, too, because look at the random numbers, too, because you're copying lots of little files back you're copying lots of little files back you're copying lots of little files back and forth. But that's actually more than and forth. But that's actually more than and forth. But that's actually more than two times faster also. So yeah, two two times faster also. So yeah, two two times faster also. So yeah, two times faster is a conservative label by times faster is a conservative label by times faster is a conservative label by marketing team. It's actually faster marketing team. It's actually faster marketing team. It's actually faster than that. This matters for AI as well. than that. This matters for AI as well. than that. This matters for AI as well. When you're loading a 140 GB model, it's When you're loading a 140 GB model, it's When you're loading a 140 GB model, it's going to read it off the disc much going to read it off the disc much going to read it off the disc much faster than this one. Editor Alex here. faster than this one. Editor Alex here. faster than this one. Editor Alex here. This number turned out to be way too low This number turned out to be way too low This number turned out to be way too low and I noticed it later. This is actually and I noticed it later. This is actually and I noticed it later. This is actually rerunning it. So the right speed looks a rerunning it. So the right speed looks a rerunning it. So the right speed looks a lot better. And this is on an 8 terbte lot better. And this is on an 8 terbte lot better. And this is on an 8 terbte drive by the way. So both of them are 8 drive by the way. So both of them are 8 drive by the way. So both of them are 8 terabytes. All right, let's switch to terabytes. All right, let's switch to terabytes. All right, let's switch to local LLM talk. For local LLMs, there's local LLM talk. For local LLMs, there's local LLM talk. For local LLMs, there's really two numbers I care about. PROM really two numbers I care about. PROM really two numbers I care about. PROM processing or PP as sometimes it's processing or PP as sometimes it's processing or PP as sometimes it's referred to as and token generation or referred to as and token generation or referred to as and token generation or TG. These are terms that I didn't TG. These are terms that I didn't TG. These are terms that I didn't invent. This is from Llama CPP world.

  5. invent. This is from Llama CPP world. invent. This is from Llama CPP world. They like to have fun over there. Okay. They like to have fun over there. Okay. They like to have fun over there. Okay. PP prompt processing is the model PP prompt processing is the model PP prompt processing is the model reading your prompt and it leans on reading your prompt and it leans on reading your prompt and it leans on compute or the GPU course. So the more compute or the GPU course. So the more compute or the GPU course. So the more cores you have, the faster those cores, cores you have, the faster those cores, cores you have, the faster those cores, the better. And token generation, TG, the better. And token generation, TG, the better. And token generation, TG, that's the model writing the answer. And that's the model writing the answer. And that's the model writing the answer. And that one leans on memory bandwidth. So that one leans on memory bandwidth. So that one leans on memory bandwidth. So both of those things were improved in both of those things were improved in both of those things were improved in the M5 Ultra. But it's not just an the M5 Ultra. But it's not just an the M5 Ultra. But it's not just an incremental improvement. It's not a incremental improvement. It's not a incremental improvement. It's not a small improvement because in the M3 small improvement because in the M3 small improvement because in the M3 Ultra, in the M3 generation, the GPU Ultra, in the M3 generation, the GPU Ultra, in the M3 generation, the GPU cores were just GPU cores. They were cores were just GPU cores. They were cores were just GPU cores. They were just regular old cores. But now with the just regular old cores. But now with the just regular old cores. But now with the M5 generation, we have for every GPU M5 generation, we have for every GPU M5 generation, we have for every GPU core, there's a neural accelerator that core, there's a neural accelerator that core, there's a neural accelerator that lives inside of it. That's dedicated lives inside of it. That's dedicated lives inside of it. That's dedicated matrix multiplication or AI hardware matrix multiplication or AI hardware matrix multiplication or AI hardware inside of every single core. Now, check inside of every single core. Now, check inside of every single core. Now, check this out. On the M3 Ultra, Llama CPP this out. On the M3 Ultra, Llama CPP this out. On the M3 Ultra, Llama CPP starts up and says has tensor equals starts up and says has tensor equals starts up and says has tensor equals false. Same exact build on the M5 Ultra, false. Same exact build on the M5 Ultra, false. Same exact build on the M5 Ultra, has tensor is true. So, Llama CPP has tensor is true. So, Llama CPP has tensor is true. So, Llama CPP supports metal. It's actually one of the supports metal. It's actually one of the supports metal. It's actually one of the first things Llama CPP supported. And first things Llama CPP supported. And first things Llama CPP supported. And now it has specific path that the now it has specific path that the now it has specific path that the software takes for the M5 family of software takes for the M5 family of software takes for the M5 family of chips to take advantage of those neural chips to take advantage of those neural chips to take advantage of those neural accelerators. Quick word from the accelerators. Quick word from the accelerators. Quick word from the sponsor and we'll take a look at memory sponsor and we'll take a look at memory sponsor and we'll take a look at memory bandwidth. Every developer account bandwidth. Every developer account bandwidth. Every developer account leaves another piece of your information leaves another piece of your information leaves another piece of your information online. Data brokers collect those online. Data brokers collect those online. Data brokers collect those pieces, build a profile, and then sell pieces, build a profile, and then sell pieces, build a profile, and then sell it. Making your data harder to find it. Making your data harder to find it. Making your data harder to find makes you harder to target. Privacy laws makes you harder to target. Privacy laws makes you harder to target. Privacy laws in many countries let you tell data in many countries let you tell data in many countries let you tell data brokers to remove the information they brokers to remove the information they brokers to remove the information they have about you. The hard part is finding have about you. The hard part is finding have about you. The hard part is finding all those brokers, then submitting every all those brokers, then submitting every all those brokers, then submitting every request and checking again when your

  6. request and checking again when your request and checking again when your information eventually comes back. information eventually comes back. information eventually comes back. Incogn handles those requests Incogn handles those requests Incogn handles those requests automatically and then keeps checking up automatically and then keeps checking up automatically and then keeps checking up and following up on them. My dashboard and following up on them. My dashboard and following up on them. My dashboard showed over 231 databases with my showed over 231 databases with my showed over 231 databases with my details in them. Most of those are details in them. Most of those are details in them. Most of those are already gone. And if I spot a specific already gone. And if I spot a specific already gone. And if I spot a specific page with my data on it, I send Incogn page with my data on it, I send Incogn page with my data on it, I send Incogn the URL through custom removals and they the URL through custom removals and they the URL through custom removals and they take it from there. Basically, find the take it from there. Basically, find the take it from there. Basically, find the data, remove it, and keep it removed. data, remove it, and keep it removed. data, remove it, and keep it removed. For some added reassurance, Deoid For some added reassurance, Deoid For some added reassurance, Deoid verified Incogn's data removal verified Incogn's data removal verified Incogn's data removal processes. Now, there are limits. It's processes. Now, there are limits. It's processes. Now, there are limits. It's not going to be deleting government not going to be deleting government not going to be deleting government records or social posts, but it can records or social posts, but it can records or social posts, but it can tackle data brokers and eligible pages. tackle data brokers and eligible pages. tackle data brokers and eligible pages. And you can try it with a 30-day money And you can try it with a 30-day money And you can try it with a 30-day money back guarantee. So, make your personal back guarantee. So, make your personal back guarantee. So, make your personal information harder to find with Incogn. information harder to find with Incogn. information harder to find with Incogn. Go to incogni.com/alexiscant Go to incogni.com/alexiscant Go to incogni.com/alexiscant and use code alexiscin for 60% off an and use code alexiscin for 60% off an and use code alexiscin for 60% off an annual plan. And now back to the video. annual plan. And now back to the video. annual plan. And now back to the video. So 1.2 TB per second memory bandwidth on So 1.2 TB per second memory bandwidth on So 1.2 TB per second memory bandwidth on the M5 Ultra. I wanted to measure it. the M5 Ultra. I wanted to measure it. the M5 Ultra. I wanted to measure it. There's a good old benchmark for that There's a good old benchmark for that There's a good old benchmark for that called stream sustained memory bandwidth called stream sustained memory bandwidth called stream sustained memory bandwidth and high performance computers. It's and high performance computers. It's and high performance computers. It's old, but it still works. So I pulled it old, but it still works. So I pulled it old, but it still works. So I pulled it down. You can do the same. You can down. You can do the same. You can down. You can do the same. You can compile it. And boom. Now the M3 Ultra compile it. And boom. Now the M3 Ultra compile it. And boom. Now the M3 Ultra 819 GB per second is what Apple claims.

  7. 819 GB per second is what Apple claims. 819 GB per second is what Apple claims. We're getting 312.9 We're getting 312.9 We're getting 312.9 gigabytes per second for the Triad and gigabytes per second for the Triad and gigabytes per second for the Triad and 583.6 583.6 583.6 gigabytes per second on the M5 Ultra. gigabytes per second on the M5 Ultra. gigabytes per second on the M5 Ultra. So, it is quite a lot higher than the M3 So, it is quite a lot higher than the M3 So, it is quite a lot higher than the M3 Ultra. And the reason for the lower Ultra. And the reason for the lower Ultra. And the reason for the lower number is because stream is actually a number is because stream is actually a number is because stream is actually a CPU related test. So, it's running on CPU related test. So, it's running on CPU related test. So, it's running on the CPU memory bandwidth. So, let's the CPU memory bandwidth. So, let's the CPU memory bandwidth. So, let's measure the memory bandwidth from the measure the memory bandwidth from the measure the memory bandwidth from the GPU side, which is what matters for AI. GPU side, which is what matters for AI. GPU side, which is what matters for AI. Memory bandwidth micro benchmark. And Memory bandwidth micro benchmark. And Memory bandwidth micro benchmark. And let's go. That's better. 724 gigabytes let's go. That's better. 724 gigabytes let's go. That's better. 724 gigabytes per second on the M3 Ultra. Apple says per second on the M3 Ultra. Apple says per second on the M3 Ultra. Apple says 819. So that's about 89 90%. And the M5 819. So that's about 89 90%. And the M5 819. So that's about 89 90%. And the M5 Ultra 1,039 GB per second out of 1.2 TB Ultra 1,039 GB per second out of 1.2 TB Ultra 1,039 GB per second out of 1.2 TB that's about 8586%. So that's 1.4 times that's about 8586%. So that's 1.4 times that's about 8586%. So that's 1.4 times the actual measured memory bandwidth the actual measured memory bandwidth the actual measured memory bandwidth over the M3 Ultra. Does that mean that over the M3 Ultra. Does that mean that over the M3 Ultra. Does that mean that token generation is going to be 1.4 token generation is going to be 1.4 token generation is going to be 1.4 times faster? All right, let's start big. Deepseek V4 All right, let's start big. Deepseek V4 Flash is 284 billion parameters. By the Flash is 284 billion parameters. By the Flash is 284 billion parameters. By the way, this is a pretty big model and it's way, this is a pretty big model and it's way, this is a pretty big model and it's pretty popular now. And we're getting 37 pretty popular now. And we're getting 37 pretty popular now. And we're getting 37 tokens per second on the M3 Ultra, 53 on tokens per second on the M3 Ultra, 53 on tokens per second on the M3 Ultra, 53 on the M5 Ultra. And that is about one and the M5 Ultra. And that is about one and the M5 Ultra. And that is about one and a half times faster. Kind of lines up, a half times faster. Kind of lines up, a half times faster. Kind of lines up, right? So that's token generation.

  8. right? So that's token generation. right? So that's token generation. That's memory bandwidth bound. Remember That's memory bandwidth bound. Remember That's memory bandwidth bound. Remember that. What about prompt processing? That that. What about prompt processing? That that. What about prompt processing? That thing that the M3 generation was so good thing that the M3 generation was so good thing that the M3 generation was so good at token generation. I'm saying at token generation. I'm saying at token generation. I'm saying generation a lot. It had really good generation a lot. It had really good generation a lot. It had really good memory bandwidth. It was impressive for memory bandwidth. It was impressive for memory bandwidth. It was impressive for a box like this, but prompt processing a box like this, but prompt processing a box like this, but prompt processing was always the issue of the M3s and the was always the issue of the M3s and the was always the issue of the M3s and the M4s until we got those neural M4s until we got those neural M4s until we got those neural accelerators. So, what about prime accelerators. So, what about prime accelerators. So, what about prime processing now? 483 tokens per second on processing now? 483 tokens per second on processing now? 483 tokens per second on the M3 Ultra and 1485 the M3 Ultra and 1485 the M3 Ultra and 1485 on the M5 Ultra. That's three times. on the M5 Ultra. That's three times. on the M5 Ultra. That's three times. Now, we're talking, but Apple said four. Now, we're talking, but Apple said four. Now, we're talking, but Apple said four. Now, different models behave Now, different models behave Now, different models behave differently. And you might remember this differently. And you might remember this differently. And you might remember this one from the M5 Max video that I did a one from the M5 Max video that I did a one from the M5 Max video that I did a few months ago. And at that time, the M3 few months ago. And at that time, the M3 few months ago. And at that time, the M3 Ultra was king at 82 tokens per second. Ultra was king at 82 tokens per second. Ultra was king at 82 tokens per second. That's GPTOSS120B, That's GPTOSS120B, That's GPTOSS120B, 120 billion parameter model. Now, that 120 billion parameter model. Now, that 120 billion parameter model. Now, that was in LM Studio, not Pure Lama CPP. was in LM Studio, not Pure Lama CPP. was in LM Studio, not Pure Lama CPP. Now, we're getting 90. I tested it today Now, we're getting 90. I tested it today Now, we're getting 90. I tested it today again. Now, we're getting 90 tokens per again. Now, we're getting 90 tokens per again. Now, we're getting 90 tokens per second. And on the M5 Ultra, 130 tokens second. And on the M5 Ultra, 130 tokens second. And on the M5 Ultra, 130 tokens per second. That's fast for a 120 per second. That's fast for a 120 per second. That's fast for a 120 billion parameter model. This is not billion parameter model. This is not billion parameter model. This is not Quen 34B like I usually would show on Quen 34B like I usually would show on Quen 34B like I usually would show on the channel, but we're going to skip the channel, but we're going to skip the channel, but we're going to skip that today. Prompt processing goes from that today. Prompt processing goes from that today. Prompt processing goes from 1,300 to about 3,200 tokens per second.

  9. 1,300 to about 3,200 tokens per second. 1,300 to about 3,200 tokens per second. Now, watch this. With 64,000 tokens Now, watch this. With 64,000 tokens Now, watch this. With 64,000 tokens already in context, I gave it 32,000 already in context, I gave it 32,000 already in context, I gave it 32,000 more. That took 80 seconds on the M3 more. That took 80 seconds on the M3 more. That took 80 seconds on the M3 Ultra and 56 seconds on the M5 Ultra. Ultra and 56 seconds on the M5 Ultra. Ultra and 56 seconds on the M5 Ultra. So, that's 1.4 times. Now, Deepsec and GPT OSS are both models Now, Deepsec and GPT OSS are both models or mixture of experts. With those kind or mixture of experts. With those kind or mixture of experts. With those kind of models, each token only goes through of models, each token only goes through of models, each token only goes through a few small experts. Most of it sits a few small experts. Most of it sits a few small experts. Most of it sits idle. In a dense model, it's every idle. In a dense model, it's every idle. In a dense model, it's every weight, every token. So, I added Quen weight, every token. So, I added Quen weight, every token. So, I added Quen 3.827B, 3.827B, 3.827B, a pretty popular model. Not as many a pretty popular model. Not as many a pretty popular model. Not as many parameters, but it is dense. So, it's parameters, but it is dense. So, it's parameters, but it is dense. So, it's heavy, right? Density is the same thing. heavy, right? Density is the same thing. heavy, right? Density is the same thing. No, I took science class. cuz I know No, I took science class. cuz I know No, I took science class. cuz I know what density is. Token generation on the what density is. Token generation on the what density is. Token generation on the M3 Ultra, 37 tokens per second, 54 M3 Ultra, 37 tokens per second, 54 M3 Ultra, 37 tokens per second, 54 tokens per second on the M5 Ultra. So tokens per second on the M5 Ultra. So tokens per second on the M5 Ultra. So that's 1.47 times. Basically, we're that's 1.47 times. Basically, we're that's 1.47 times. Basically, we're getting the same kind of uh scale on all getting the same kind of uh scale on all getting the same kind of uh scale on all kinds of models here. Token generation kinds of models here. Token generation kinds of models here. Token generation is basically following Apple's spec.

  10. is basically following Apple's spec. is basically following Apple's spec. Now, prompt processing 430 on the M3 Now, prompt processing 430 on the M3 Now, prompt processing 430 on the M3 Ultra and on the M5 Ultra 1,800. Holy Ultra and on the M5 Ultra 1,800. Holy Ultra and on the M5 Ultra 1,800. Holy cow, that's kind of like four times, cow, that's kind of like four times, cow, that's kind of like four times, right? It's actually more than four right? It's actually more than four right? It's actually more than four times. I also gave them both a 14,000 times. I also gave them both a 14,000 times. I also gave them both a 14,000 token prompt, equivalent to about a token prompt, equivalent to about a token prompt, equivalent to about a handful of source files if you're doing handful of source files if you're doing handful of source files if you're doing code, just to see if it falls apart if code, just to see if it falls apart if code, just to see if it falls apart if I'm doing a real code base instead of I'm doing a real code base instead of I'm doing a real code base instead of just a benchmark. And I timed the prompt just a benchmark. And I timed the prompt just a benchmark. And I timed the prompt processing. That's 33 seconds on the M3 processing. That's 33 seconds on the M3 processing. That's 33 seconds on the M3 Ultra and 8 seconds on the M5 Ultra. Ultra and 8 seconds on the M5 Ultra. Ultra and 8 seconds on the M5 Ultra. Apple did say it's up to four times Apple did say it's up to four times Apple did say it's up to four times faster prompt processing and I got 4.2. faster prompt processing and I got 4.2. faster prompt processing and I got 4.2. So, Apple again was not lying. They did So, Apple again was not lying. They did So, Apple again was not lying. They did it again. Now, don't get me wrong, I'm impressed. Now, don't get me wrong, I'm impressed. But before you get your credit card out, But before you get your credit card out, But before you get your credit card out, and it is going to be a big credit card and it is going to be a big credit card and it is going to be a big credit card bill, that four times needs a long bill, that four times needs a long bill, that four times needs a long prompt. At about 1,700 tokens, it's 3.4 prompt. At about 1,700 tokens, it's 3.4 prompt. At about 1,700 tokens, it's 3.4 times. It takes about 4,500 tokens to times. It takes about 4,500 tokens to times. It takes about 4,500 tokens to hit the four time scale, which shouldn't hit the four time scale, which shouldn't hit the four time scale, which shouldn't be a problem if you're using this be a problem if you're using this be a problem if you're using this machine professionally. However, on a machine professionally. However, on a machine professionally. However, on a short prompt around 350 tokens or so, short prompt around 350 tokens or so, short prompt around 350 tokens or so, it's only 1.6 times. It's still an it's only 1.6 times. It's still an it's only 1.6 times. It's still an improvement, but if you're mostly improvement, but if you're mostly improvement, but if you're mostly chatting, then you won't really feel the chatting, then you won't really feel the chatting, then you won't really feel the improvement. If you're feeding it a few improvement. If you're feeding it a few improvement. If you're feeding it a few files, then you will. It's also not much files, then you will. It's also not much files, then you will. It's also not much more efficient. During generation, the more efficient. During generation, the more efficient. During generation, the chip reported about 45 watts versus 36 chip reported about 45 watts versus 36 chip reported about 45 watts versus 36 on the M3 Ultra. And that's GPU plus on the M3 Ultra. And that's GPU plus on the M3 Ultra. And that's GPU plus CPU, not wall power. So, tokens per watt CPU, not wall power. So, tokens per watt CPU, not wall power. So, tokens per watt only went up to about 12%. This is only went up to about 12%. This is only went up to about 12%. This is basically the machine idling 12 11 or 12

  11. basically the machine idling 12 11 or 12 basically the machine idling 12 11 or 12 watts for the M3 Ultra and about eight watts for the M3 Ultra and about eight watts for the M3 Ultra and about eight or nine up to 10 watts on the M5 Ultra. or nine up to 10 watts on the M5 Ultra. or nine up to 10 watts on the M5 Ultra. However, during heavy intense GPU work, However, during heavy intense GPU work, However, during heavy intense GPU work, I got it pegging right now. Uh that's a I got it pegging right now. Uh that's a I got it pegging right now. Uh that's a GPU history chart right now on both of GPU history chart right now on both of GPU history chart right now on both of them all the way to 100% usage. And the them all the way to 100% usage. And the them all the way to 100% usage. And the M3 Ultra is hitting almost 200 watts. M3 Ultra is hitting almost 200 watts. M3 Ultra is hitting almost 200 watts. The M5 over 400 watts. That is crazy. The M5 over 400 watts. That is crazy. The M5 over 400 watts. That is crazy. Wow. And that's pretty stable there. Wow. And that's pretty stable there. Wow. And that's pretty stable there. Wow, look at that. So, I started hearing Wow, look at that. So, I started hearing Wow, look at that. So, I started hearing uh the M5 Ultra first. Is definitely uh the M5 Ultra first. Is definitely uh the M5 Ultra first. Is definitely much much more audible. much much more audible. much much more audible. You can probably hear it from my You can probably hear it from my You can probably hear it from my microphone. microphone. microphone. That's all M5 Ultra right now. I'm not even hearing the M3 Ultra. I'm I'm not even hearing the M3 Ultra. I'm hearing the M5 Ultra. And look at the hearing the M5 Ultra. And look at the hearing the M5 Ultra. And look at the temperatures. You can just tell from the temperatures. You can just tell from the temperatures. You can just tell from the bright orange of the M5 Ultra that it's bright orange of the M5 Ultra that it's bright orange of the M5 Ultra that it's heating up quite a bit more. So, I'm heating up quite a bit more. So, I'm heating up quite a bit more. So, I'm about 43 44 degrees on the M5 Ultra about 43 44 degrees on the M5 Ultra about 43 44 degrees on the M5 Ultra and and and considerably lower on the M3 Ultra. And considerably lower on the M3 Ultra. And considerably lower on the M3 Ultra. And then there's memory. Everything I ran then there's memory. Everything I ran then there's memory. Everything I ran today fits in 256 gigs. Ooh. Ooh, that's today fits in 256 gigs. Ooh. Ooh, that's today fits in 256 gigs. Ooh. Ooh, that's so much toastier than the M3s were. Wow.

  12. so much toastier than the M3s were. Wow. so much toastier than the M3s were. Wow. Okay, that gets toasty. Wow. Now, Okay, that gets toasty. Wow. Now, Okay, that gets toasty. Wow. Now, there's some bigger models out there, there's some bigger models out there, there's some bigger models out there, and I showed plenty of examples of uh my and I showed plenty of examples of uh my and I showed plenty of examples of uh my tower of 512 GB Mac Studios with the M3 tower of 512 GB Mac Studios with the M3 tower of 512 GB Mac Studios with the M3 Ultra, including clustering, which I'll Ultra, including clustering, which I'll Ultra, including clustering, which I'll come back to later on, in Mac Studio come back to later on, in Mac Studio come back to later on, in Mac Studio format and in uh Mac Mini format as format and in uh Mac Mini format as format and in uh Mac Mini format as well. So, stay tuned for that. well. So, stay tuned for that. well. So, stay tuned for that. Obviously, the 512 GB models of M5 Ultra Obviously, the 512 GB models of M5 Ultra Obviously, the 512 GB models of M5 Ultra are going to be very expensive, but are going to be very expensive, but are going to be very expensive, but they're going to be able to run much they're going to be able to run much they're going to be able to run much larger models, too, and faster. So larger models, too, and faster. So larger models, too, and faster. So bottom line, token generation one and a bottom line, token generation one and a bottom line, token generation one and a half times faster. Nice upgrade. Prompt half times faster. Nice upgrade. Prompt half times faster. Nice upgrade. Prompt processing two to four times faster. For processing two to four times faster. For processing two to four times faster. For coding agents, you will feel that. So coding agents, you will feel that. So coding agents, you will feel that. So this was a first look. I've got more this was a first look. I've got more this was a first look. I've got more testing to do, of course, just to give testing to do, of course, just to give testing to do, of course, just to give you a little preview with eight requests you a little preview with eight requests you a little preview with eight requests at once. This thing puts out 134 tokens at once. This thing puts out 134 tokens at once. This thing puts out 134 tokens per second against M3 Ultra 75. I'm also per second against M3 Ultra 75. I'm also per second against M3 Ultra 75. I'm also going to dive a little deeper into MLX going to dive a little deeper into MLX going to dive a little deeper into MLX versus Llama CPP and 8bit models. So versus Llama CPP and 8bit models. So versus Llama CPP and 8bit models. So stay tuned for that. Make sure you stay tuned for that. Make sure you stay tuned for that. Make sure you subscribe so you don't miss it. And if subscribe so you don't miss it. And if subscribe so you don't miss it. And if you want to check out the M5 Max and its you want to check out the M5 Max and its you want to check out the M5 Max and its capabilities, that video is right over capabilities, that video is right over capabilities, that video is right over here. Thanks for watching and I'll see here. Thanks for watching and I'll see here. Thanks for watching and I'll see you in the next one.

Summary

The discussion centers on the new M5 Ultra chip, comparing its performance to the M3 Ultra with references to AI speed, storage, CPU, and memory bandwidth. The practical takeaway is that while the M5 Ultra offers significant performance gains, particularly in AI tasks, it comes at a considerably higher price point, making it overkill for general developer tasks.

View original episode ↗