The M6 Mac mini Had Me Checking My Numbers
Read full transcript 15 segments
-
So, this is the M6 [music] Mac Mini. So, this is the M6 [music] Mac Mini. This is the M4. Kind of hard to tell the This is the M4. Kind of hard to tell the This is the M4. Kind of hard to tell the difference, right? And the M6 Mac Mini difference, right? And the M6 Mac Mini difference, right? And the M6 Mac Mini [music] is supposed to make the M4 [music] is supposed to make the M4 [music] is supposed to make the M4 obsolete. Naturally, the first thing I obsolete. Naturally, the first thing I obsolete. Naturally, the first thing I did was run Geekbench. Hold on. The M6 did was run Geekbench. Hold on. The M6 did was run Geekbench. Hold on. The M6 scored lower than the M5 in this MacBook scored lower than the M5 in this MacBook scored lower than the M5 in this MacBook Pro. Pro. Pro. 4,085 versus 4190. Let's see. When the 4,085 versus 4190. Let's see. When the 4,085 versus 4190. Let's see. When the M5s came out, Geekbench [music] 6 was M5s came out, Geekbench [music] 6 was M5s came out, Geekbench [music] 6 was the norm. Now, we have Geekbench 7. In the norm. Now, we have Geekbench 7. In the norm. Now, we have Geekbench 7. In version 7, all the scores are lower. So, version 7, all the scores are lower. So, version 7, all the scores are lower. So, I reran the M5 number in version 7. And I reran the M5 number in version 7. And I reran the M5 number in version 7. And look at that, 3715 on the M5. Not that look at that, 3715 on the M5. Not that look at that, 3715 on the M5. Not that it matters because the M5 is not it matters because the M5 is not it matters because the M5 is not available in a Mac [music] Mini. We went available in a Mac [music] Mini. We went available in a Mac [music] Mini. We went from M4 to M6. All right, Geekbench is from M4 to M6. All right, Geekbench is from M4 to M6. All right, Geekbench is one thing. We'll get into actual tests one thing. We'll get into actual tests one thing. We'll get into actual tests in a minute. And it makes me wonder, is in a minute. And it makes me wonder, is in a minute. And it makes me wonder, is the M6 a [music] big enough leap over the M6 a [music] big enough leap over the M6 a [music] big enough leap over the M4 to really matter? In other words, the M4 to really matter? In other words, the M4 to really matter? In other words, if you already have the M4, do you care? if you already have the M4, do you care? if you already have the M4, do you care? I'll do some dev tests as well as local I'll do some dev tests as well as local I'll do some dev tests as well as local AI. This is just my first look here. I'm AI. This is just my first look here. I'm AI. This is just my first look here. I'm going to be digging a lot deeper into going to be digging a lot deeper into going to be digging a lot deeper into all of it in videos coming up, including all of it in videos coming up, including all of it in videos coming up, including that M6, the M5 [music] that M6, the M5 [music] that M6, the M5 [music] Pro, building a load cluster, and the Pro, building a load cluster, and the Pro, building a load cluster, and the big boy, the M5 Ultra. Oh, yeah.
-
So, what's changed in the chip? Because So, what's changed in the chip? Because [music] it's actually way more than a [music] it's actually way more than a [music] it's actually way more than a spec bump. Starting with the CPU, for spec bump. Starting with the CPU, for spec bump. Starting with the CPU, for the first time ever, Apple is mixing the first time ever, Apple is mixing the first time ever, Apple is mixing three different kinds of cores. There's three different kinds of cores. There's three different kinds of cores. There's two big super cores. That's the new name two big super cores. That's the new name two big super cores. That's the new name for their [music] fastest cores. Four for their [music] fastest cores. Four for their [music] fastest cores. Four performance cores in the middle and six performance cores in the middle and six performance cores in the middle and six efficiency cores. You can even see that efficiency cores. You can even see that efficiency cores. You can even see that an activity monitor. That's a lot of an activity monitor. That's a lot of an activity monitor. That's a lot of cores. And that [music] translates cores. And that [music] translates cores. And that [music] translates directly into compilations and developer directly into compilations and developer directly into compilations and developer related activities later on. Then related activities later on. Then related activities later on. Then there's a GPU. This is also the first there's a GPU. This is also the first there's a GPU. This is also the first base Mac to get a 12 core GPU. The M4 base Mac to get a 12 core GPU. The M4 base Mac to get a 12 core GPU. The M4 only had 10. [music] And here's one that only had 10. [music] And here's one that only had 10. [music] And here's one that matters most for everything we're about matters most for everything we're about matters most for everything we're about to do. Every single one of those GPU to do. Every single one of those GPU to do. Every single one of those GPU cores now has [music] a neural cores now has [music] a neural cores now has [music] a neural accelerator built in. This was new in accelerator built in. This was new in accelerator built in. This was new in the M5. It's not new period to Apple, the M5. It's not new period to Apple, the M5. It's not new period to Apple, but it was not in the M4. That's the but it was not in the M4. That's the but it was not in the M4. That's the hardware that runs AI fast. [music] And hardware that runs AI fast. [music] And hardware that runs AI fast. [music] And you'd think that they were done there, you'd think that they were done there, you'd think that they were done there, but they're not. They added another but they're not. They added another but they're not. They added another neural engine. Now, there's two of them. neural engine. Now, there's two of them. neural engine. Now, there's two of them. Now, that was computational stuff. Now, that was computational stuff. Now, that was computational stuff. There's also a few other things that got There's also a few other things that got There's also a few other things that got bumped in [music] the chip. There's a bumped in [music] the chip. There's a bumped in [music] the chip. There's a new Apple N1 wireless chip. So, you get new Apple N1 wireless chip. So, you get new Apple N1 wireless chip. So, you get Wi-Fi 7 and Bluetooth 6. The Ethernet Wi-Fi 7 and Bluetooth 6. The Ethernet Wi-Fi 7 and Bluetooth 6. The Ethernet jack doubled in size. I mean, it's the jack doubled in size. I mean, it's the jack doubled in size. I mean, it's the same size. It's just doubled in speed, I same size. It's just doubled in speed, I same size. It's just doubled in speed, I should say. But unfortunately, there's should say. But unfortunately, there's should say. But unfortunately, there's one thing that did not change. Uh, one thing that did not change. Uh, one thing that did not change. Uh, Thunderbolt 4, not Thunderbolt 5. They Thunderbolt 4, not Thunderbolt 5. They Thunderbolt 4, not Thunderbolt 5. They saved five for the Pro model. So, if saved five for the Pro model. So, if saved five for the Pro model. So, if you're moving a lot of data, that one's you're moving a lot of data, that one's you're moving a lot of data, that one's going to matter.
-
going to matter. going to matter. >> This is a good test for [music] >> This is a good test for [music] >> This is a good test for [music] developers as well as end users. This is developers as well as end users. This is developers as well as end users. This is speedometer. It measures browser speed. speedometer. It measures browser speed. speedometer. It measures browser speed. Basically, how fast the web feels. does Basically, how fast the web feels. does Basically, how fast the web feels. does a bunch of todo applications in a bunch of todo applications in a bunch of todo applications in different frameworks like React, different frameworks like React, different frameworks like React, Angular, and [music] pure JavaScript and Angular, and [music] pure JavaScript and Angular, and [music] pure JavaScript and TypeScript and all that. What the TypeScript and all that. What the TypeScript and all that. What the Oh my gosh. What? Okay, I thought we Oh my gosh. What? Okay, I thought we Oh my gosh. What? Okay, I thought we were going to be like 60s something, but were going to be like 60s something, but were going to be like 60s something, but 73. This is the highest I've ever seen 73. This is the highest I've ever seen 73. This is the highest I've ever seen it. This is the first time I ran it, and it. This is the first time I ran it, and it. This is the first time I ran it, and I'm genu I'm actually surprised here. I'm genu I'm actually surprised here. I'm genu I'm actually surprised here. This is crazy. Let's move on to This is crazy. Let's move on to This is crazy. Let's move on to something heavy. Python [music] something heavy. Python [music] something heavy. Python [music] time. Python main 16,000. And this time. Python main 16,000. And this time. Python main 16,000. And this basically does the Mandelroad algorithm, basically does the Mandelroad algorithm, basically does the Mandelroad algorithm, the fractals [music] patterns, but in the fractals [music] patterns, but in the fractals [music] patterns, but in Python. Boom. You can find this on Python. Boom. You can find this on Python. Boom. You can find this on benchmarks game. This one goes crazy. It benchmarks game. This one goes crazy. It benchmarks game. This one goes crazy. It uses up all the cores available. There's uses up all the cores available. There's uses up all the cores available. There's 12 here. So, I'm expecting that to win 12 here. So, I'm expecting that to win 12 here. So, I'm expecting that to win over the one with the M4 where there's over the one with the M4 where there's over the one with the M4 where there's only 10 cores. It's already finished at only 10 cores. It's already finished at only 10 cores. It's already finished at 21.5. 21.5. 21.5. Oh my gosh, this is crazy. 32.8. Oh my gosh, this is crazy. 32.8. Oh my gosh, this is crazy. 32.8. That is a huge difference. I'm going to That is a huge difference. I'm going to That is a huge difference. I'm going to rerun this with a higher number just so rerun this with a higher number just so rerun this with a higher number just so that you can see what's going on with that you can see what's going on with that you can see what's going on with those cores. Oh.
-
those cores. Oh. those cores. Oh. >> [laughter] >> [laughter] >> [laughter] >> I'm hearing them. Look at those soldiers >> I'm hearing them. Look at those soldiers >> I'm hearing them. Look at those soldiers marching on. Those are green soldiers marching on. Those are green soldiers marching on. Those are green soldiers using every single core. That's the using every single core. That's the using every single core. That's the activity, by the way. 10 [music] here, activity, by the way. 10 [music] here, activity, by the way. 10 [music] here, 12 here. There's the power utilization 12 here. There's the power utilization 12 here. There's the power utilization right there. This is fully active CPU right there. This is fully active CPU right there. This is fully active CPU cores, not GPU, CPU. So, [music] we're cores, not GPU, CPU. So, [music] we're cores, not GPU, CPU. So, [music] we're at 40 watts usage on the M4 and about 48 at 40 watts usage on the M4 and about 48 at 40 watts usage on the M4 and about 48 on the M6. It's pretty stable. Just for on the M6. It's pretty stable. Just for on the M6. It's pretty stable. Just for reference, the M4 Pro Mac Mini got 23.2 reference, the M4 Pro Mac Mini got 23.2 reference, the M4 Pro Mac Mini got 23.2 seconds. EM6 BTM4 Pro. Now, that was a seconds. EM6 BTM4 Pro. Now, that was a seconds. EM6 BTM4 Pro. Now, that was a Python interpreted test. Here is one Python interpreted test. Here is one Python interpreted test. Here is one that's a compilation. We're using Xcode that's a compilation. We're using Xcode that's a compilation. We're using Xcode benchmark here. It's up on GitHub. I can benchmark here. It's up on GitHub. I can benchmark here. It's up on GitHub. I can link to this down below. I'm going to link to this down below. I'm going to link to this down below. I'm going to start it off. Boom. Now, while it's start it off. Boom. Now, while it's start it off. Boom. Now, while it's building, this is basically a framework building, this is basically a framework building, this is basically a framework that includes 76 popular [music] Coco that includes 76 popular [music] Coco that includes 76 popular [music] Coco libraries and their dependencies. So, libraries and their dependencies. So, libraries and their dependencies. So, it's pretty substantial. If you're a it's pretty substantial. If you're a it's pretty substantial. If you're a developer who needs to compile their developer who needs to compile their developer who needs to compile their code and wait for it, it's pretty large code and wait for it, it's pretty large code and wait for it, it's pretty large project several times a day. This will project several times a day. This will project several times a day. This will matter to you. There we go. Build matter to you. There we go. Build matter to you. There we go. Build succeeded. 114 seconds for this bottom succeeded. 114 seconds for this bottom succeeded. 114 seconds for this bottom one, which is M6. The M4 [music] is one, which is M6. The M4 [music] is one, which is M6. The M4 [music] is still working at it. Here is the M5 Max.
-
still working at it. Here is the M5 Max. still working at it. Here is the M5 Max. 84 seconds. The M1 Ultra Max Studio got 84 seconds. The M1 Ultra Max Studio got 84 seconds. The M1 Ultra Max Studio got [music] 112 seconds, which is on par [music] 112 seconds, which is on par [music] 112 seconds, which is on par with the M6. Yeah, think about that. with the M6. Yeah, think about that. with the M6. Yeah, think about that. These are all older Xcode and Mac OS These are all older Xcode and Mac OS These are all older Xcode and Mac OS results, so the numbers don't exactly results, so the numbers don't exactly results, so the numbers don't exactly match, but it's pretty close. Finally, match, but it's pretty close. Finally, match, but it's pretty close. Finally, the M4 finished at 172 seconds. the M4 finished at 172 seconds. the M4 finished at 172 seconds. >> [music] >> [music] >> [music] >> All right, storage. Storage matters also >> All right, storage. Storage matters also >> All right, storage. Storage matters also for developers because if you're doing for developers because if you're doing for developers because if you're doing compilations, it's a lot of reading and compilations, it's a lot of reading and compilations, it's a lot of reading and writing of small files. Also, if you're writing of small files. Also, if you're writing of small files. Also, if you're copying large language model files, copying large language model files, copying large language model files, you're going to want to look at this you're going to want to look at this you're going to want to look at this number right here, which is the number right here, which is the number right here, which is the sequential read and write. They're both sequential read and write. They're both sequential read and write. They're both pretty fast here, but one is pretty fast here, but one is pretty fast here, but one is considerably faster than the other one. considerably faster than the other one. considerably faster than the other one. And what matters more for us developers And what matters more for us developers And what matters more for us developers is these random read and writes, which is these random read and writes, which is these random read and writes, which is the bottom two rows. And wow, random is the bottom two rows. And wow, random is the bottom two rows. And wow, random 4K QD64. Much faster on the M6. This 4K QD64. Much faster on the M6. This 4K QD64. Much faster on the M6. This one, the [music] QD1, eh, it's kind of a one, the [music] QD1, eh, it's kind of a one, the [music] QD1, eh, it's kind of a draw. Overall, a big improvement in the draw. Overall, a big improvement in the draw. Overall, a big improvement in the SSD speeds. All right, red light. Just a SSD speeds. All right, red light. Just a SSD speeds. All right, red light. Just a quick little break to go to my favorite quick little break to go to my favorite quick little break to go to my favorite store.
-
store. store. You guys know I've been doing a lot of You guys know I've been doing a lot of You guys know I've been doing a lot of AI on the channel, and MicroEnter has a AI on the channel, and MicroEnter has a AI on the channel, and MicroEnter has a lot of hardware for that. If you're lot of hardware for that. If you're lot of hardware for that. If you're building AI workstations, which we've building AI workstations, which we've building AI workstations, which we've done before, by the way, or you just done before, by the way, or you just done before, by the way, or you just want something, a device that'll run want something, a device that'll run want something, a device that'll run models locally. But today I'm here for models locally. But today I'm here for models locally. But today I'm here for the max. >> Oh, there's Dan. Hi, Dan. >> Oh, there's Dan. Hi, Dan. >> How's it going? >> How's it going? >> How's it going? >> Hey, >> Hey, >> Hey, >> how's everyone doing? >> how's everyone doing? >> how's everyone doing? >> They've got a whole section of the store >> They've got a whole section of the store >> They've got a whole section of the store dedicated to Apple. You got your dedicated to Apple. You got your dedicated to Apple. You got your MacBooks, Mac Minis, Mac Studios. You MacBooks, Mac Minis, Mac Studios. You MacBooks, Mac Minis, Mac Studios. You have no idea how many times I needed a have no idea how many times I needed a have no idea how many times I needed a random dongle or adapter, right? And random dongle or adapter, right? And random dongle or adapter, right? And even Thunderbolt cable. I can just stop even Thunderbolt cable. I can just stop even Thunderbolt cable. I can just stop in the store and grab one right here. in the store and grab one right here. in the store and grab one right here. And they also have a Dan in here. And they also have a Dan in here. And they also have a Dan in here. sometimes depends on the day. sometimes depends on the day. sometimes depends on the day. >> Dan has been in many videos before and >> Dan has been in many videos before and >> Dan has been in many videos before and with the M6 and M5 Pro Mac Minis coming with the M6 and M5 Pro Mac Minis coming with the M6 and M5 Pro Mac Minis coming out and the M5 Ultra Mac Studios, out and the M5 Ultra Mac Studios, out and the M5 Ultra Mac Studios, MicroEnter is going to be carrying them MicroEnter is going to be carrying them MicroEnter is going to be carrying them as soon as they become available. I'll as soon as they become available. I'll as soon as they become available. I'll leave links in the description to their leave links in the description to their leave links in the description to their AI section and the new Mac lineup. And AI section and the new Mac lineup. And AI section and the new Mac lineup. And by the way, if you're in the Austin by the way, if you're in the Austin by the way, if you're in the Austin area, they got a good deal where you can area, they got a good deal where you can area, they got a good deal where you can get 128 GB flash drive for free. Just get 128 GB flash drive for free. Just get 128 GB flash drive for free. Just check out the link below. Thanks to check out the link below. Thanks to check out the link below. Thanks to MicroEnter for sponsoring this video.
-
MicroEnter for sponsoring this video. MicroEnter for sponsoring this video. Now, let's get back to those Macs. Now, let's get back to those Macs. Now, let's get back to those Macs. Now, the M6 is Apple's first 2nm chip, 2 Now, the M6 is Apple's first 2nm chip, 2 Now, the M6 is Apple's first 2nm chip, 2 nm. The M4 is 3 nm. Did I just say 26 nm. The M4 is 3 nm. Did I just say 26 nm. The M4 is 3 nm. Did I just say 26 34? That's confusing. The new one has 34? That's confusing. The new one has 34? That's confusing. The new one has smaller, newer transistors. Okay, that smaller, newer transistors. Okay, that smaller, newer transistors. Okay, that they leak less power and run cooler. You they leak less power and run cooler. You they leak less power and run cooler. You might have heard that before, but what might have heard that before, but what might have heard that before, but what does that actually mean? Well, you get does that actually mean? Well, you get does that actually mean? Well, you get more speed for the same amount of power more speed for the same amount of power more speed for the same amount of power used. Bolt minis sitting at idle pull used. Bolt minis sitting at idle pull used. Bolt minis sitting at idle pull the same tiny amount of power. So, they the same tiny amount of power. So, they the same tiny amount of power. So, they have the same floor, but the M6 does a have the same floor, but the M6 does a have the same floor, but the M6 does a lot more when you push it. You already lot more when you push it. You already lot more when you push it. You already saw one example of that earlier. And saw one example of that earlier. And saw one example of that earlier. And there's one more number that matters for there's one more number that matters for there's one more number that matters for AI. I want to show you this before we AI. I want to show you this before we AI. I want to show you this before we get into some of the [music] tests. This get into some of the [music] tests. This get into some of the [music] tests. This is memory bandwidth right here. is memory bandwidth right here. is memory bandwidth right here. Basically, it's how fast the chip can Basically, it's how fast the chip can Basically, it's how fast the chip can move data around. The M4 could do about move data around. The M4 could do about move data around. The M4 could do about 120 GB per second. The M6 is different 120 GB per second. The M6 is different 120 GB per second. The M6 is different depending on which model you get. So, depending on which model you get. So, depending on which model you get. So, you get 153, but also 170. That's pretty you get 153, but also 170. That's pretty you get 153, but also 170. That's pretty confusing, right? How do you get 170? I confusing, right? How do you get 170? I confusing, right? How do you get 170? I want more. Well, if you go down here, want more. Well, if you go down here, want more. Well, if you go down here, you'll see that if you upgrade to 32 you'll see that if you upgrade to 32 you'll see that if you upgrade to 32 gigabytes of memory, that's the only gigabytes of memory, that's the only gigabytes of memory, that's the only case where you would get 170. Otherwise, case where you would get 170. Otherwise, case where you would get 170. Otherwise, if you have 16 GB, then you're getting if you have 16 GB, then you're getting if you have 16 GB, then you're getting 153. But those are just Apple numbers.
-
153. But those are just Apple numbers. 153. But those are just Apple numbers. How do we actually measure it? Well, How do we actually measure it? Well, How do we actually measure it? Well, there's a stream benchmark that's been there's a stream benchmark that's been there's a stream benchmark that's been around for many, many years, and [music] around for many, many years, and [music] around for many, many years, and [music] it's designed to measure memory it's designed to measure memory it's designed to measure memory bandwidth. You can check it out. It's up bandwidth. You can check it out. It's up bandwidth. You can check it out. It's up on GitHub. So, [music] I pulled it down. on GitHub. So, [music] I pulled it down. on GitHub. So, [music] I pulled it down. I built it. It's a C program. If you I built it. It's a C program. If you I built it. It's a C program. If you don't know how to do it, just ask Claude don't know how to do it, just ask Claude don't know how to do it, just ask Claude Code to do it for you. Boom. And it's Code to do it for you. Boom. And it's Code to do it for you. Boom. And it's very fast. So, we got 112 GB per second very fast. So, we got 112 GB per second very fast. So, we got 112 GB per second for the copy operation on the M4, which for the copy operation on the M4, which for the copy operation on the M4, which is actually pretty close. It's about 8 is actually pretty close. It's about 8 is actually pretty close. It's about 8 GB off from official Apple numbers. And GB off from official Apple numbers. And GB off from official Apple numbers. And we got 143 down here, which is quite a we got 143 down here, which is quite a we got 143 down here, which is quite a bit off cuz this machine [music] is bit off cuz this machine [music] is bit off cuz this machine [music] is supposed to get 170. We'll talk about supposed to get 170. We'll talk about supposed to get 170. We'll talk about that shortly, but I'm measuring 143. that shortly, but I'm measuring 143. that shortly, but I'm measuring 143. Let's do it again cuz it's [music] so Let's do it again cuz it's [music] so Let's do it again cuz it's [music] so fast. 144.7 this time. Yeah, that's fast. 144.7 this time. Yeah, that's fast. 144.7 this time. Yeah, that's going to play a direct role in AI. going to play a direct role in AI. going to play a direct role in AI. [music] [music] [music] Speaking of which, now the fun part. Speaking of which, now the fun part. Speaking of which, now the fun part. Local AI. I mean, it's all fun parts. Local AI. I mean, it's all fun parts. Local AI. I mean, it's all fun parts. Depends what you're doing, right? And Depends what you're doing, right? And Depends what you're doing, right? And this is exactly the kind of machine this is exactly the kind of machine this is exactly the kind of machine that'll let you do all those things. But that'll let you do all those things. But that'll let you do all those things. But Apple is really shifting hard towards Apple is really shifting hard towards Apple is really shifting hard towards AI, you know, whatever you want to call AI, you know, whatever you want to call AI, you know, whatever you want to call it, Apple intelligence or the plain old it, Apple intelligence or the plain old it, Apple intelligence or the plain old artificial kind. The hardware is being artificial kind. The hardware is being artificial kind. The hardware is being built built built to run it locally, to run it locally, to run it locally, especially this guy. That's for another especially this guy. That's for another especially this guy. That's for another video. Running local AI really comes video. Running local AI really comes video. Running local AI really comes down to two things. Compute, which is down to two things. Compute, which is down to two things. Compute, which is how strong the GPU and its neural engine how strong the GPU and its neural engine how strong the GPU and its neural engine is. That's the prompt processing stage is. That's the prompt processing stage is. That's the prompt processing stage or how fast it reads your prompt and or how fast it reads your prompt and or how fast it reads your prompt and decides what to do and calculate stuff.
-
decides what to do and calculate stuff. decides what to do and calculate stuff. And then there's the memory bandwidth or And then there's the memory bandwidth or And then there's the memory bandwidth or how fast it writes, which is what moves how fast it writes, which is what moves how fast it writes, which is what moves the data when you're generating text or the data when you're generating text or the data when you're generating text or images or even video. And this little images or even video. And this little images or even video. And this little mini leans on both. Also, you might see mini leans on both. Also, you might see mini leans on both. Also, you might see two different memory related numbers two different memory related numbers two different memory related numbers show up. Bandwidth is how fast the show up. Bandwidth is how fast the show up. Bandwidth is how fast the memory moves data. You can think of it memory moves data. You can think of it memory moves data. You can think of it as the width of the pipe. Wider the as the width of the pipe. Wider the as the width of the pipe. Wider the pipe, the faster the water flows. pipe, the faster the water flows. pipe, the faster the water flows. Capacity is how much memory you've got. Capacity is how much memory you've got. Capacity is how much memory you've got. This is the size of the tank that feeds This is the size of the tank that feeds This is the size of the tank that feeds into the pipe. I'm thirsty. So speed decides how fast the model runs So speed decides how fast the model runs and capacity decides whether your model and capacity decides whether your model and capacity decides whether your model runs at all. If there's too many runs at all. If there's too many runs at all. If there's too many billions of parameters, then it might billions of parameters, then it might billions of parameters, then it might not even fit in the amount of memory you not even fit in the amount of memory you not even fit in the amount of memory you have in there. Move the glass away so we have in there. Move the glass away so we have in there. Move the glass away so we don't spill it on our new Mac minis and don't spill it on our new Mac minis and don't spill it on our new Mac minis and load up a model. All right, I'm running load up a model. All right, I'm running load up a model. All right, I'm running Llama CPP here because it's a very Llama CPP here because it's a very Llama CPP here because it's a very popular tool, but we'll get into LM popular tool, but we'll get into LM popular tool, but we'll get into LM Studio in a bit. Just want to get some Studio in a bit. Just want to get some Studio in a bit. Just want to get some basic numbers here for prompt processing basic numbers here for prompt processing basic numbers here for prompt processing and token generation, also abbreviated and token generation, also abbreviated and token generation, also abbreviated as PP and TG. Don't make that joke.
-
as PP and TG. Don't make that joke. as PP and TG. Don't make that joke. Don't do it. I know you want to. Current Don't do it. I know you want to. Current Don't do it. I know you want to. Current 3.59 3.59 3.59 billion parameter model. This is the billion parameter model. This is the billion parameter model. This is the [music] Q4K [music] Q4K [music] Q4K quant. Okay, it's a GGUF model. I know quant. Okay, it's a GGUF model. I know quant. Okay, it's a GGUF model. I know it's a lot of acronyms and words that I it's a lot of acronyms and words that I it's a lot of acronyms and words that I just threw at you, but I go into a lot just threw at you, but I go into a lot just threw at you, but I go into a lot more detail in other videos. The reason more detail in other videos. The reason more detail in other videos. The reason I'm running a 9 billion parameter model I'm running a 9 billion parameter model I'm running a 9 billion parameter model here is because it's a good fit. You here is because it's a good fit. You here is because it's a good fit. You always want to fit the model to your always want to fit the model to your always want to fit the model to your hardware. And we've got a big difference hardware. And we've got a big difference hardware. And we've got a big difference here, [music] especially for prompt here, [music] especially for prompt here, [music] especially for prompt processing. 210 tokens per second on the processing. 210 tokens per second on the processing. 210 tokens per second on the M4 compared to 742 M4 compared to 742 M4 compared to 742 on the M6. Those are the improvements to on the M6. Those are the improvements to on the M6. Those are the improvements to [music] the GPU, especially with those [music] the GPU, especially with those [music] the GPU, especially with those neural accelerators and token neural accelerators and token neural accelerators and token generation. That's the memory bandwidth generation. That's the memory bandwidth generation. That's the memory bandwidth part. 18 tokens per second on the M4, part. 18 tokens per second on the M4, part. 18 tokens per second on the M4, 26.9 tokens per second on the M6. I 26.9 tokens per second on the M6. I 26.9 tokens per second on the M6. I wanted to show you this in Llama Bench, wanted to show you this in Llama Bench, wanted to show you this in Llama Bench, but also in Llama Beni because Llama but also in Llama Beni because Llama but also in Llama Beni because Llama Beni uses a more realistic approach Beni uses a more realistic approach Beni uses a more realistic approach where it actually talks to the server where it actually talks to the server where it actually talks to the server over the [music] network, which is how over the [music] network, which is how over the [music] network, which is how you're most likely going to be using you're most likely going to be using you're most likely going to be using something like this. So, with Llama something like this. So, with Llama something like this. So, with Llama Beni, we first launch a server and then Beni, we first launch a server and then Beni, we first launch a server and then it runs through it several times to get it runs through it several times to get it runs through it several times to get an average score. Llama is open [music] an average score. Llama is open [music] an average score. Llama is open [music] source by Yuger. You can check it out on source by Yuger. You can check it out on source by Yuger. You can check it out on GitHub. I'll link to it down below. And GitHub. I'll link to it down below. And GitHub. I'll link to it down below. And also, it runs against any kind of also, it runs against any kind of also, it runs against any kind of server, PLLM, SG Lang, not only Llama server, PLLM, SG Lang, not only Llama server, PLLM, SG Lang, not only Llama CPP. Oh wow. Okay, this also shows CPP. Oh wow. Okay, this also shows CPP. Oh wow. Okay, this also shows another piece of information that's very another piece of information that's very another piece of information that's very important. Time to first token. And wow, important. Time to first token. And wow, important. Time to first token. And wow, look at this time to first token. 2.5 look at this time to first token. 2.5 look at this time to first token. 2.5 [music] [music] [music] seconds on the M4, 721 milliseconds,
-
seconds on the M4, 721 milliseconds, seconds on the M4, 721 milliseconds, less than a second on the M6. Now the PP less than a second on the M6. Now the PP less than a second on the M6. Now the PP and the TG numbers are a little bit and the TG numbers are a little bit and the TG numbers are a little bit lower here because it's using HTTP to lower here because it's using HTTP to lower here because it's using HTTP to communicate to the server. [music] So PP communicate to the server. [music] So PP communicate to the server. [music] So PP is 198 on the M4, TG17 versus the M6 is 198 on the M4, TG17 versus the M6 is 198 on the M4, TG17 versus the M6 which got PP 711 and TG26. [music] which got PP 711 and TG26. [music] which got PP 711 and TG26. [music] But what does that actually look like But what does that actually look like But what does that actually look like when you're seeing text generated? Well, when you're seeing text generated? Well, when you're seeing text generated? Well, this is what it looks like. Look how this is what it looks like. Look how this is what it looks like. Look how much it's already written compared to much it's already written compared to much it's already written compared to this one. [music] Half a screen of this this one. [music] Half a screen of this this one. [music] Half a screen of this one versus a full screen of that one. one versus a full screen of that one. one versus a full screen of that one. That's how we'll scientifically measure That's how we'll scientifically measure That's how we'll scientifically measure this stuff. This is a thinking model. So this stuff. This is a thinking model. So this stuff. This is a thinking model. So it's doing a lot [music] of thinking. it's doing a lot [music] of thinking. it's doing a lot [music] of thinking. But no matter this is all token But no matter this is all token But no matter this is all token generation. So whether it's thinking generation. So whether it's thinking generation. So whether it's thinking tokens or actual tokens, it's doing it tokens or actual tokens, it's doing it tokens or actual tokens, it's doing it at the same speed. Okay, it was thinking at the same speed. Okay, it was thinking at the same speed. Okay, it was thinking too much, so I had to terminate it. Here too much, so I had to terminate it. Here too much, so I had to terminate it. Here are the numbers. The prompt was 45.9 are the numbers. The prompt was 45.9 are the numbers. The prompt was 45.9 tokens per second. Generation 17.4 on tokens per second. Generation 17.4 on tokens per second. Generation 17.4 on the M4. On the M6, 71 tokens per second the M4. On the M6, 71 tokens per second the M4. On the M6, 71 tokens per second for the prompt [music] and 26 tokens per for the prompt [music] and 26 tokens per for the prompt [music] and 26 tokens per second for the generation. Clearly, the second for the generation. Clearly, the second for the generation. Clearly, the M6 is faster both at reading your prompt M6 is faster both at reading your prompt M6 is faster both at reading your prompt and writing the reply. That newer and writing the reply. That newer and writing the reply. That newer architecture and the extra bandwidth are architecture and the extra bandwidth are architecture and the extra bandwidth are doing that real work here. All right, doing that real work here. All right, doing that real work here. All right, text is one thing. Let's make it draw text is one thing. Let's make it draw text is one thing. Let's make it draw something. And boom. This is the Flux something. And boom. This is the Flux something. And boom. This is the Flux one Chanel model. Same exact prompt on one Chanel model. Same exact prompt on one Chanel model. Same exact prompt on both of these. You can already see a big both of these. You can already see a big both of these. You can already see a big difference here in the estimated time.
-
difference here in the estimated time. difference here in the estimated time. This one says 33. This one says 33. This one says 33. It already popped up with a preview It already popped up with a preview It already popped up with a preview image. Wow. This one is estimating it to image. Wow. This one is estimating it to image. Wow. This one is estimating it to be much longer. And it's [music] done. be much longer. And it's [music] done. be much longer. And it's [music] done. 35 seconds to make this image over here 35 seconds to make this image over here 35 seconds to make this image over here on the M6. That one is still sampling. on the M6. That one is still sampling. on the M6. That one is still sampling. Now, this model fits comfortably [music] Now, this model fits comfortably [music] Now, this model fits comfortably [music] in 16 gigabytes of memory. So, we don't in 16 gigabytes of memory. So, we don't in 16 gigabytes of memory. So, we don't have an issue with that. What we're have an issue with that. What we're have an issue with that. What we're seeing here is that memory bandwidth at seeing here is that memory bandwidth at seeing here is that memory bandwidth at work. Come on. Come on. Come on. work. Come on. Come on. Come on. work. Come on. Come on. Come on. Yes. 1 minute 34 seconds. That is a big Yes. 1 minute 34 seconds. That is a big Yes. 1 minute 34 seconds. That is a big big difference. That model is loaded on big difference. That model is loaded on big difference. That model is loaded on both of these machines right now. There both of these machines right now. There both of these machines right now. There is no memory pressure. Everything is in is no memory pressure. Everything is in is no memory pressure. Everything is in the green. I'm going to kick it up a the green. I'm going to kick it up a the green. I'm going to kick it up a notch here. [laughter] notch here. [laughter] notch here. [laughter] I'm looping a local model a 100 times, I'm looping a local model a 100 times, I'm looping a local model a 100 times, thousand tokens each, and watching this thousand tokens each, and watching this thousand tokens each, and watching this thing generate. Mactop is showing 34 thing generate. Mactop is showing 34 thing generate. Mactop is showing 34 watts used on the chip and about 31 watts used on the chip and about 31 watts used on the chip and about 31 used. Well, it's actually bouncing used. Well, it's actually bouncing used. Well, it's actually bouncing around between Yeah, it's about 30. I'd around between Yeah, it's about 30. I'd around between Yeah, it's about 30. I'd say 30 31 on the M6. Maximum power was say 30 31 on the M6. Maximum power was say 30 31 on the M6. Maximum power was 38 on the M6, 44 on the M4. So, even 38 on the M6, 44 on the M4. So, even 38 on the M6, 44 on the M4. So, even though it's faster, it's using less though it's faster, it's using less though it's faster, it's using less power. And this is at 100% GPU power. And this is at 100% GPU power. And this is at 100% GPU utilization. This is where that 2nmter utilization. This is where that 2nmter utilization. This is where that 2nmter chip is really supposed to shine here.
-
chip is really supposed to shine here. chip is really supposed to shine here. Not just fast, but fast and cool. Check Not just fast, but fast and cool. Check Not just fast, but fast and cool. Check this out. You can even see the [music] this out. You can even see the [music] this out. You can even see the [music] difference in the color here in the difference in the color here in the difference in the color here in the thermal camera. Just take a look here. thermal camera. Just take a look here. thermal camera. Just take a look here. 37.6 37.6 37.6 37.7 on the M4. And we're at 35.9 37.7 on the M4. And we're at 35.9 37.7 on the M4. And we're at 35.9 33. It's a couple of degrees lower. 33. It's a couple of degrees lower. 33. It's a couple of degrees lower. Okay, I'm not [music] going to nitpick Okay, I'm not [music] going to nitpick Okay, I'm not [music] going to nitpick here, but it is a little bit lower. Wo, here, but it is a little bit lower. Wo, here, but it is a little bit lower. Wo, look at this guy. Same temperature as look at this guy. Same temperature as look at this guy. Same temperature as the M6. Who is that guy? I don't know. the M6. Who is that guy? I don't know. the M6. Who is that guy? I don't know. So, remember that dual neural engine I So, remember that dual neural engine I So, remember that dual neural engine I talked to you about earlier? There's talked to you about earlier? There's talked to you about earlier? There's definitely some good opportunities going definitely some good opportunities going definitely some good opportunities going on there along with new stuff like on there along with new stuff like on there along with new stuff like native FP8 support. That's floatingoint native FP8 support. That's floatingoint native FP8 support. That's floatingoint 8. Definitely needs some proper testing, 8. Definitely needs some proper testing, 8. Definitely needs some proper testing, so I'll talk about that in another so I'll talk about that in another so I'll talk about that in another video. Make sure you subscribe so you video. Make sure you subscribe so you video. Make sure you subscribe so you don't miss that. One of the most popular don't miss that. One of the most popular don't miss that. One of the most popular models and very capable right now is models and very capable right now is models and very capable right now is Quen 3.827B. Quen 3.827B. Quen 3.827B. By the time you watch this video, By the time you watch this video, By the time you watch this video, there's going to be something new out. there's going to be something new out. there's going to be something new out. But right now, and this model is just But right now, and this model is just But right now, and this model is just over 16 GB in size. It won't work on a over 16 GB in size. It won't work on a over 16 GB in size. It won't work on a machine that has 16 gigabytes. Whether machine that has 16 gigabytes. Whether machine that has 16 gigabytes. Whether it's M4 or M6, it's going to say it's M4 or M6, it's going to say it's M4 or M6, it's going to say something like likely too large. You can something like likely too large. You can something like likely too large. You can try it. You can download it. You can try it. You can download it. You can try it. You can download it. You can loosen up LM Studios, but you're going loosen up LM Studios, but you're going loosen up LM Studios, but you're going to run into issues. It's just not going to run into issues. It's just not going to run into issues. It's just not going to work for you. However, 32 GB, [music] to work for you. However, 32 GB, [music] to work for you. However, 32 GB, [music] no problem. Check it out. And I'm going no problem. Check it out. And I'm going no problem. Check it out. And I'm going to bump up that context all the way to to bump up that context all the way to to bump up that context all the way to 262,000. It's going to load it up. No 262,000. It's going to load it up. No 262,000. It's going to load it up. No problem. Look at that memory going up.
-
problem. Look at that memory going up. problem. Look at that memory going up. 23 gigabytes used. Still plenty left. 23 gigabytes used. Still plenty left. 23 gigabytes used. Still plenty left. Hi, how can I help you today? And it's Hi, how can I help you today? And it's Hi, how can I help you today? And it's not super fast. 9.7 tokens a second. But not super fast. 9.7 tokens a second. But not super fast. 9.7 tokens a second. But it did it and did it smoothly. No it did it and did it smoothly. No it did it and did it smoothly. No complaints. The 16 GB model that I have complaints. The 16 GB model that I have complaints. The 16 GB model that I have here is not going to do that. But the here is not going to do that. But the here is not going to do that. But the kicker is that the M4 Mac Minis are not kicker is that the M4 Mac Minis are not kicker is that the M4 Mac Minis are not going to give you that 32 GB option. going to give you that 32 GB option. going to give you that 32 GB option. They only go up to 24. The new M6 Mac They only go up to 24. The new M6 Mac They only go up to 24. The new M6 Mac minis go up to [music] 32. For local AI, minis go up to [music] 32. For local AI, minis go up to [music] 32. For local AI, the 32 GB upgrade isn't a nice to have. the 32 GB upgrade isn't a nice to have. the 32 GB upgrade isn't a nice to have. It's the difference between running a 27 It's the difference between running a 27 It's the difference between running a 27 billion parameter model [music] or not billion parameter model [music] or not billion parameter model [music] or not at all. So, how much Mac Mini do you at all. So, how much Mac Mini do you at all. So, how much Mac Mini do you actually need? For most people, the base actually need? For most people, the base actually need? For most people, the base M6 Mac Mini with one smart upgrade is an M6 Mac Mini with one smart upgrade is an M6 Mac Mini with one smart upgrade is an incredible value. And that 256 GB one, I incredible value. And that 256 GB one, I incredible value. And that 256 GB one, I would not get that. Especially because would not get that. Especially because would not get that. Especially because you don't get Thunderbolt 5 on there. you don't get Thunderbolt 5 on there. you don't get Thunderbolt 5 on there. It's going to be very slow transferring It's going to be very slow transferring It's going to be very slow transferring data back and forth. Not very slow, but data back and forth. Not very slow, but data back and forth. Not very slow, but slower than the current capabilities out slower than the current capabilities out slower than the current capabilities out there. So, 512 GB for storage or a there. So, 512 GB for storage or a there. So, 512 GB for storage or a terabyte is even better. That would be a terabyte is even better. That would be a terabyte is even better. That would be a smart upgrade for casual users, for AI smart upgrade for casual users, for AI smart upgrade for casual users, for AI users. 32 GB of memory. And the price, users. 32 GB of memory. And the price, users. 32 GB of memory. And the price, starting price is $8.99, which is more starting price is $8.99, which is more starting price is $8.99, which is more expensive than the M4s came when they expensive than the M4s came when they expensive than the M4s came when they came out, but yeah, I think it's kind of came out, but yeah, I think it's kind of came out, but yeah, I think it's kind of a fair price for that chip. Right now, a fair price for that chip. Right now, a fair price for that chip. Right now, you're still getting that 256 GBTE you're still getting that 256 GBTE you're still getting that 256 GBTE storage for that price. But if you can storage for that price. But if you can storage for that price. But if you can handle external storage and offloading handle external storage and offloading handle external storage and offloading your data, then for one of the fastest your data, then for one of the fastest your data, then for one of the fastest chips out there right now on a machine
-
chips out there right now on a machine chips out there right now on a machine this small, that is a really good deal. this small, that is a really good deal. this small, that is a really good deal. The M5 Pro runs you about $1,700 and up, The M5 Pro runs you about $1,700 and up, The M5 Pro runs you about $1,700 and up, twice as much. Is it twice as much twice as much. Is it twice as much twice as much. Is it twice as much power? Uh, I don't know. We'll have to power? Uh, I don't know. We'll have to power? Uh, I don't know. We'll have to find out. Stay tuned. But if you're find out. Stay tuned. But if you're find out. Stay tuned. But if you're getting just one of these, then you're getting just one of these, then you're getting just one of these, then you're probably making money off of your work probably making money off of your work probably making money off of your work and using it for pro things. So, this and using it for pro things. So, this and using it for pro things. So, this was my first look. Hope you liked it. was my first look. Hope you liked it. was my first look. Hope you liked it. We're going to dig deeper, of course, We're going to dig deeper, of course, We're going to dig deeper, of course, that uh M5 Pro cluster that I'm building that uh M5 Pro cluster that I'm building that uh M5 Pro cluster that I'm building and the M5 Ultra. Keep an eye out for and the M5 Ultra. Keep an eye out for and the M5 Ultra. Keep an eye out for those videos. Thanks for watching. I'll those videos. Thanks for watching. I'll those videos. Thanks for watching. I'll see you next time.
Summary
This analysis compares the M4 and M6 Mac Mini chips, initially using Geekbench to assess performance differences. Key features of the M6 include a new CPU core configuration with "super cores," a 12-core GPU, dual neural engines for AI acceleration, and updated wireless connectivity. The practical takeaway is that the M6 represents a significant upgrade over the M4, particularly for developer tasks and AI workloads, rather than just a minor spec bump.