Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
Read full transcript 11 segments
-
Good afternoon everybody. Um Good afternoon everybody. Um I'm Satan Shu from Kov going to be I'm Satan Shu from Kov going to be I'm Satan Shu from Kov going to be talking about vertical uh mobility. It's talking about vertical uh mobility. It's talking about vertical uh mobility. It's a quite a fancy topic uh the title that a quite a fancy topic uh the title that a quite a fancy topic uh the title that we came up with but basically going to we came up with but basically going to we came up with but basically going to be talking about the inference platform be talking about the inference platform be talking about the inference platform that we have at Corv that we are that we have at Corv that we are that we have at Corv that we are building to serve small to big models building to serve small to big models building to serve small to big models and various different types of and various different types of and various different types of workloads. workloads. workloads. Uh a quick intro about me. Um I joined Uh a quick intro about me. Um I joined Uh a quick intro about me. Um I joined core just about four months back. Uh core just about four months back. Uh core just about four months back. Uh leading all of inference over there and leading all of inference over there and leading all of inference over there and um before this I was managing everything um before this I was managing everything um before this I was managing everything at AWS Anapuna labs for uh training and at AWS Anapuna labs for uh training and at AWS Anapuna labs for uh training and before that inference and training at before that inference and training at before that inference and training at SANOVA. So quite a bit of experience in SANOVA. So quite a bit of experience in SANOVA. So quite a bit of experience in uh in this particular space. uh in this particular space. uh in this particular space. Um what I will the way I'll be taking Um what I will the way I'll be taking Um what I will the way I'll be taking you through is explaining to you the you through is explaining to you the you through is explaining to you the consumption models that we have and from consumption models that we have and from consumption models that we have and from that how we have derived what the that how we have derived what the that how we have derived what the platform should look like so that we do platform should look like so that we do platform should look like so that we do not need to keep changing the platform not need to keep changing the platform not need to keep changing the platform and we keep making enhancements in the and we keep making enhancements in the and we keep making enhancements in the platform that we have for serving platform that we have for serving platform that we have for serving inference and how and why the why inference and how and why the why inference and how and why the why performance plays such an important role performance plays such an important role performance plays such an important role over there. Uh I think a little bit of over there. Uh I think a little bit of over there. Uh I think a little bit of this might be common with the previous this might be common with the previous this might be common with the previous topic that was discussed over here.
-
topic that was discussed over here. topic that was discussed over here. [snorts] [snorts] [snorts] Um consumption models. So we have uh at Um consumption models. So we have uh at Um consumption models. So we have uh at large two biggest consumption models. large two biggest consumption models. large two biggest consumption models. One is a serverless which is where the One is a serverless which is where the One is a serverless which is where the customers can come in, consumers can customers can come in, consumers can customers can come in, consumers can come in uh do not need to worry about come in uh do not need to worry about come in uh do not need to worry about managing the hardware themselves. Do not managing the hardware themselves. Do not managing the hardware themselves. Do not need to worry about managing the need to worry about managing the need to worry about managing the clusters orchestration anything at all. clusters orchestration anything at all. clusters orchestration anything at all. There's API, there's UI, you come in, There's API, there's UI, you come in, There's API, there's UI, you come in, you pay per token uh and you get your you pay per token uh and you get your you pay per token uh and you get your model served. Um biggest thing over here model served. Um biggest thing over here model served. Um biggest thing over here is that uh the type of models uh that we is that uh the type of models uh that we is that uh the type of models uh that we serve in the catalog that is the breadth serve in the catalog that is the breadth serve in the catalog that is the breadth of models that the customer will be able of models that the customer will be able of models that the customer will be able to go through. Uh I'll talk about to go through. Uh I'll talk about to go through. Uh I'll talk about dedicated and then I'll come back to dedicated and then I'll come back to dedicated and then I'll come back to serverless because there is one thing serverless because there is one thing serverless because there is one thing unique on the serverless side. Dedicated unique on the serverless side. Dedicated unique on the serverless side. Dedicated inference service that we provide is inference service that we provide is inference service that we provide is more for customers who want to know more for customers who want to know more for customers who want to know exactly what hardware they are going to exactly what hardware they are going to exactly what hardware they are going to be using and running on. Uh but the be using and running on. Uh but the be using and running on. Uh but the model deployment also depends on them. model deployment also depends on them. model deployment also depends on them. So they use our service, they use our So they use our service, they use our So they use our service, they use our orchestration layers. Uh but the model orchestration layers. Uh but the model orchestration layers. Uh but the model deployment depends on them. Model deployment depends on them. Model deployment depends on them. Model performance also depends on them as long performance also depends on them as long performance also depends on them as long as we provide in the platform the as we provide in the platform the as we provide in the platform the capability and the knobs to serve those capability and the knobs to serve those capability and the knobs to serve those features. Coming back to serverless, uh features. Coming back to serverless, uh features. Coming back to serverless, uh one of the interesting pieces over here one of the interesting pieces over here one of the interesting pieces over here is um typically serverless models are is um typically serverless models are is um typically serverless models are very are noisy neighbor problems where very are noisy neighbor problems where very are noisy neighbor problems where uh if let's say everyone is banging on uh if let's say everyone is banging on uh if let's say everyone is banging on the exact same model then you might be the exact same model then you might be the exact same model then you might be timing out quite a bit depending on how timing out quite a bit depending on how timing out quite a bit depending on how much capacity I have behind it. So much capacity I have behind it. So much capacity I have behind it. So another feature that we have on the another feature that we have on the another feature that we have on the serverless side is what we calling serverless side is what we calling serverless side is what we calling provisioned throughput. So as a
-
provisioned throughput. So as a provisioned throughput. So as a customer, if you know your traffic customer, if you know your traffic customer, if you know your traffic profile and if you can let us know about profile and if you can let us know about profile and if you can let us know about that, we can carve it out specifically that, we can carve it out specifically that, we can carve it out specifically for you behind the scenes. You still do for you behind the scenes. You still do for you behind the scenes. You still do not need to worry about what hardware it not need to worry about what hardware it not need to worry about what hardware it is exactly running on as long as your is exactly running on as long as your is exactly running on as long as your throughput, your SLAs's are maintained. throughput, your SLAs's are maintained. throughput, your SLAs's are maintained. So that is another one on the serverless So that is another one on the serverless So that is another one on the serverless side and that is still charged by by the side and that is still charged by by the side and that is still charged by by the token but you know that you're not token but you know that you're not token but you know that you're not running into the noise enabled problem running into the noise enabled problem running into the noise enabled problem over there. over there. over there. Um let me take a quick stab at few Um let me take a quick stab at few Um let me take a quick stab at few different types of workloads workload different types of workloads workload different types of workloads workload shapes that uh we have um that we are shapes that uh we have um that we are shapes that uh we have um that we are seeing and the ratio between these is seeing and the ratio between these is seeing and the ratio between these is like continuously changing though like continuously changing though like continuously changing though agentic is like really high up there. agentic is like really high up there. agentic is like really high up there. Uh, agentic and chat kind of very Uh, agentic and chat kind of very Uh, agentic and chat kind of very similar. Super high on the input similar. Super high on the input similar. Super high on the input sequence lengths, very low on the output sequence lengths, very low on the output sequence lengths, very low on the output sequence lengths typically. But the sequence lengths typically. But the sequence lengths typically. But the biggest difference between agentic and biggest difference between agentic and biggest difference between agentic and chat being the fact that the multi-turns chat being the fact that the multi-turns chat being the fact that the multi-turns in agentic are super low latency versus in agentic are super low latency versus in agentic are super low latency versus in chats because when you get the in chats because when you get the in chats because when you get the response as a user, you have to read the response as a user, you have to read the response as a user, you have to read the answer and then you respond to it. So answer and then you respond to it. So answer and then you respond to it. So there are there are differences over there are there are differences over there are there are differences over there and that big difference ultimately there and that big difference ultimately there and that big difference ultimately converts into something uh related to converts into something uh related to converts into something uh related to the KV cache management uh but these are the KV cache management uh but these are the KV cache management uh but these are these two are both real time and another these two are both real time and another these two are both real time and another real-time workload is your voice and real-time workload is your voice and real-time workload is your voice and videos uh which are study streaming and videos uh which are study streaming and videos uh which are study streaming and super latency sensitive on the agent super latency sensitive on the agent super latency sensitive on the agent agent and chat side um largely the agent and chat side um largely the agent and chat side um largely the requirements are from throughput point requirements are from throughput point requirements are from throughput point of view not so much from latency but
-
of view not so much from latency but of view not so much from latency but real-time voice and videos are real-time voice and videos are real-time voice and videos are absolutely totally latency sensitive. absolutely totally latency sensitive. absolutely totally latency sensitive. Coming to batch, batch is where the Coming to batch, batch is where the Coming to batch, batch is where the SLAs's are like super loose. They run SLAs's are like super loose. They run SLAs's are like super loose. They run into like way like seconds and minutes. into like way like seconds and minutes. into like way like seconds and minutes. Sometimes for some customers actually Sometimes for some customers actually Sometimes for some customers actually even in hours. They're like I'll just even in hours. They're like I'll just even in hours. They're like I'll just throw give me 10 to 12 hours of workload throw give me 10 to 12 hours of workload throw give me 10 to 12 hours of workload capability and uh I'll throw whatever I capability and uh I'll throw whatever I capability and uh I'll throw whatever I can process it whenever you can. uh can process it whenever you can. uh can process it whenever you can. uh these batch workloads um the way they these batch workloads um the way they these batch workloads um the way they come into the picture over here in um come into the picture over here in um come into the picture over here in um deciding uh sorry being the requirement deciding uh sorry being the requirement deciding uh sorry being the requirement for some of the design choices that we for some of the design choices that we for some of the design choices that we make in the stack. Imagine these uh four make in the stack. Imagine these uh four make in the stack. Imagine these uh four different types of workload shapes um in different types of workload shapes um in different types of workload shapes um in the time dimension you have to play the the time dimension you have to play the the time dimension you have to play the game of tetris on how you can fit it in game of tetris on how you can fit it in game of tetris on how you can fit it in to utilize the underlying infrastructure to utilize the underlying infrastructure to utilize the underlying infrastructure the most. I'll give a high level on how our stack I'll give a high level on how our stack is shaped right now. Um, and I'll walk is shaped right now. Um, and I'll walk is shaped right now. Um, and I'll walk you through a bit of a request flow over you through a bit of a request flow over you through a bit of a request flow over here. So for both serverless and here. So for both serverless and here. So for both serverless and dedicated, if you look at the right hand dedicated, if you look at the right hand dedicated, if you look at the right hand side of the screen, you'll see that on side of the screen, you'll see that on side of the screen, you'll see that on the platform side we you'll go to the the platform side we you'll go to the the platform side we you'll go to the control plane to uh to have your uh control plane to uh to have your uh control plane to uh to have your uh authorizations, your rate limitings and authorizations, your rate limitings and authorizations, your rate limitings and your usage being tracked etc. so that we your usage being tracked etc. so that we your usage being tracked etc. so that we can be built accordingly and uh can be built accordingly and uh can be built accordingly and uh observability so that we can make sure observability so that we can make sure observability so that we can make sure that we are not violating the SLAs's that we are not violating the SLAs's that we are not violating the SLAs's that have been signed right uh that have been signed right uh that have been signed right uh underlying on the platform I've shown at underlying on the platform I've shown at underlying on the platform I've shown at a super high level that we have these a super high level that we have these a super high level that we have these different um inference engines VLMs SG
-
different um inference engines VLMs SG different um inference engines VLMs SG langs and tensorl but there are quite langs and tensorl but there are quite langs and tensorl but there are quite there's quite a few quite a bit of there's quite a few quite a bit of there's quite a few quite a bit of detail over here that I'll touch upon detail over here that I'll touch upon detail over here that I'll touch upon uh and underlying that what I'm trying uh and underlying that what I'm trying uh and underlying that what I'm trying to show over here in green is um various to show over here in green is um various to show over here in green is um various different pieces of hardware. So it's different pieces of hardware. So it's different pieces of hardware. So it's not that um so the platform needs to be not that um so the platform needs to be not that um so the platform needs to be capable enough of share of having the capable enough of share of having the capable enough of share of having the workload getting distributed across workload getting distributed across workload getting distributed across various different generations of these various different generations of these various different generations of these GPUs [snorts] uh specifically Nvidia GPUs [snorts] uh specifically Nvidia GPUs [snorts] uh specifically Nvidia GPUs that we use right so let's take up GPUs that we use right so let's take up GPUs that we use right so let's take up a let's take a few examples uh over here a let's take a few examples uh over here a let's take a few examples uh over here um let's say the request originates from um let's say the request originates from um let's say the request originates from the client side through apps or the client side through apps or the client side through apps or notebooks any of those or through the notebooks any of those or through the notebooks any of those or through the agents right it hits the gateway once it agents right it hits the gateway once it agents right it hits the gateway once it hits the gateway then uh like I hits the gateway then uh like I hits the gateway then uh like I mentioned on the control plane goes mentioned on the control plane goes mentioned on the control plane goes through authentication etc etc etc but through authentication etc etc etc but through authentication etc etc etc but then um comes either the serverless or then um comes either the serverless or then um comes either the serverless or dedicated so in the case of serverless dedicated so in the case of serverless dedicated so in the case of serverless it'll be paper token so that the token it'll be paper token so that the token it'll be paper token so that the token usage would be monitored over here not usage would be monitored over here not usage would be monitored over here not the exact tokens but just the token the exact tokens but just the token the exact tokens but just the token usage because we maintain ZDR zero data usage because we maintain ZDR zero data usage because we maintain ZDR zero data retention policies uh it is multi retention policies uh it is multi retention policies uh it is multi depending on the multi-tenency or the depending on the multi-tenency or the depending on the multi-tenency or the provisioned uh if it is provisioned then provisioned uh if it is provisioned then provisioned uh if it is provisioned then we know underlying for the router it we know underlying for the router it we know underlying for the router it needs to go in and target the uh the needs to go in and target the uh the needs to go in and target the uh the explicit deployments for the provision explicit deployments for the provision explicit deployments for the provision throughput uh customers. For the throughput uh customers. For the throughput uh customers. For the multi-end customers, there are separate multi-end customers, there are separate multi-end customers, there are separate deployments.
-
deployments. deployments. Router over here specifically uh the the Router over here specifically uh the the Router over here specifically uh the the router is very important since um the router is very important since um the router is very important since um the router is responsible for making KV router is responsible for making KV router is responsible for making KV cache aware routing choices. Why is it cache aware routing choices. Why is it cache aware routing choices. Why is it important? Because like I mentioned when important? Because like I mentioned when important? Because like I mentioned when we were discussing the workload uh we were discussing the workload uh we were discussing the workload uh profiles profiles profiles uh the agentic use cases are typically uh the agentic use cases are typically uh the agentic use cases are typically super heavy on the input sequence super heavy on the input sequence super heavy on the input sequence lengths and bulk of the input sequence lengths and bulk of the input sequence lengths and bulk of the input sequence length about 80 to 90% depending on length about 80 to 90% depending on length about 80 to 90% depending on which company it is depending on the which company it is depending on the which company it is depending on the customers 80 to 90% of it is the same customers 80 to 90% of it is the same customers 80 to 90% of it is the same for various different requests. So there for various different requests. So there for various different requests. So there is no point in going in and recomputing is no point in going in and recomputing is no point in going in and recomputing the prefill or redoing the prefill for the prefill or redoing the prefill for the prefill or redoing the prefill for that prefill is supercomputebound very that prefill is supercomputebound very that prefill is supercomputebound very expensive that's why as much as you can expensive that's why as much as you can expensive that's why as much as you can hit the cache more you can save which is hit the cache more you can save which is hit the cache more you can save which is why if you look at the token pricing why if you look at the token pricing why if you look at the token pricing anywhere uh there's a specific price for anywhere uh there's a specific price for anywhere uh there's a specific price for input tokens and there's a way cheaper input tokens and there's a way cheaper input tokens and there's a way cheaper price for the cash input tokens so price for the cash input tokens so price for the cash input tokens so caching becomes like really important caching becomes like really important caching becomes like really important over here underlying uh The underlying over here underlying uh The underlying over here underlying uh The underlying how you want to split the hardware is how you want to split the hardware is how you want to split the hardware is totally dependent on the choice in the totally dependent on the choice in the totally dependent on the choice in the platform and we provide the capability platform and we provide the capability platform and we provide the capability to do either either do a pre-filled to do either either do a pre-filled to do either either do a pre-filled decode disagregation if the use case decode disagregation if the use case decode disagregation if the use case desires it or do not do it because desires it or do not do it because desires it or do not do it because prefilled decode disagregation is not uh prefilled decode disagregation is not uh prefilled decode disagregation is not uh cheap for every type of use case. Uh cheap for every type of use case. Uh cheap for every type of use case. Uh let's take another request flow over let's take another request flow over let's take another request flow over here. Let's see if uh when it was a here. Let's see if uh when it was a here. Let's see if uh when it was a dedicated customer then what will
-
dedicated customer then what will dedicated customer then what will happen? uh dedicated customer again will happen? uh dedicated customer again will happen? uh dedicated customer again will go through the gateways that have been go through the gateways that have been go through the gateways that have been set up for them with proper isolations. set up for them with proper isolations. set up for them with proper isolations. Um billing is not based on tokens. Um billing is not based on tokens. Um billing is not based on tokens. Billing is based on usage of per GPU per Billing is based on usage of per GPU per Billing is based on usage of per GPU per hour. Uh it's a private gateway so that hour. Uh it's a private gateway so that hour. Uh it's a private gateway so that uh there is noisy enable problem. No one uh there is noisy enable problem. No one uh there is noisy enable problem. No one else can get in. uh same router logic else can get in. uh same router logic else can get in. uh same router logic over here so that if there are uh cache over here so that if there are uh cache over here so that if there are uh cache heavy if there are requests which are heavy if there are requests which are heavy if there are requests which are very similar then it's it hits the cache very similar then it's it hits the cache very similar then it's it hits the cache most and depending on the deployment most and depending on the deployment most and depending on the deployment that the customer makes in their on that the customer makes in their on that the customer makes in their on their dedicated GPUs their dedicated GPUs their dedicated GPUs uh they can decide if they want to do uh they can decide if they want to do uh they can decide if they want to do prefill decode disagregation or not they prefill decode disagregation or not they prefill decode disagregation or not they can decide which which engine to use VLM can decide which which engine to use VLM can decide which which engine to use VLM or S lang or tenslm and given the bulk or S lang or tenslm and given the bulk or S lang or tenslm and given the bulk of capacity that the customer has of capacity that the customer has of capacity that the customer has reserved they can decide if they want to reserved they can decide if they want to reserved they can decide if they want to have just one deployment uh with the have just one deployment uh with the have just one deployment uh with the ability to scale through the whole ability to scale through the whole ability to scale through the whole cluster or they want to have multiple cluster or they want to have multiple cluster or they want to have multiple different models, multiple different different models, multiple different different models, multiple different deployments. Um deployments. Um deployments. Um one thing that I do want to mention one thing that I do want to mention one thing that I do want to mention about the router over here um is the about the router over here um is the about the router over here um is the fact that heterogeneous capacity across fact that heterogeneous capacity across fact that heterogeneous capacity across different zones and regions is different zones and regions is different zones and regions is supported. It is a little it's quite a supported. It is a little it's quite a supported. It is a little it's quite a bit of a hard problem to load balance bit of a hard problem to load balance bit of a hard problem to load balance across that. So the priority order that across that. So the priority order that across that. So the priority order that we typically take is uh first KV cache we typically take is uh first KV cache we typically take is uh first KV cache locality and then the least loaded locality and then the least loaded locality and then the least loaded fallback.
-
fallback. fallback. Um Um Um that's that um another request flow that that's that um another request flow that that's that um another request flow that I want to go over here which um might be I want to go over here which um might be I want to go over here which um might be a little hard to see from the diagram is a little hard to see from the diagram is a little hard to see from the diagram is I want to take the batch workflow. For I want to take the batch workflow. For I want to take the batch workflow. For the batch workflow what we would the batch workflow what we would the batch workflow what we would actually do is the underlying capacity actually do is the underlying capacity actually do is the underlying capacity that the customer has let's say the same that the customer has let's say the same that the customer has let's say the same dedicated inference customer uh during dedicated inference customer uh during dedicated inference customer uh during US daytime they're running their US daytime they're running their US daytime they're running their real-time workloads and from evening to real-time workloads and from evening to real-time workloads and from evening to night they want to run batch workloads night they want to run batch workloads night they want to run batch workloads the same capacity after time can be the same capacity after time can be the same capacity after time can be scheduled to run the batch workloads. So scheduled to run the batch workloads. So scheduled to run the batch workloads. So we provide the capability in the API to we provide the capability in the API to we provide the capability in the API to tell when to scale up and when to scale tell when to scale up and when to scale tell when to scale up and when to scale down and as per schedule if we can if down and as per schedule if we can if down and as per schedule if we can if they tell us that we have to scale down they tell us that we have to scale down they tell us that we have to scale down we will scale down and open it up for we will scale down and open it up for we will scale down and open it up for batch processing through the night. I think I've spoken quite a bit about uh I think I've spoken quite a bit about uh optimizations on the KV cache side but I optimizations on the KV cache side but I optimizations on the KV cache side but I do want to repeat a little bit because do want to repeat a little bit because do want to repeat a little bit because this is one of the most interesting this is one of the most interesting this is one of the most interesting pieces. Um it if we can hit on the cache pieces. Um it if we can hit on the cache pieces. Um it if we can hit on the cache more you can you will avoid the cost of more you can you will avoid the cost of more you can you will avoid the cost of prefill which is the most expensive prefill which is the most expensive prefill which is the most expensive piece over here. Um piece over here. Um piece over here. Um reusing the cache across multiple reusing the cache across multiple reusing the cache across multiple different turns in your agentic different turns in your agentic different turns in your agentic workloads between turns also there is workloads between turns also there is workloads between turns also there is lot of um similar prefill uh that comes lot of um similar prefill uh that comes lot of um similar prefill uh that comes in in the input sequence length. Um, in in the input sequence length. Um, in in the input sequence length. Um, think about the chat workloads which is think about the chat workloads which is think about the chat workloads which is where offloading AV cache also becomes where offloading AV cache also becomes where offloading AV cache also becomes extremely important because with cache
-
extremely important because with cache extremely important because with cache with with the chat workloads we have a with with the chat workloads we have a with with the chat workloads we have a lot of latency between different between lot of latency between different between lot of latency between different between multiple turns that we as users put in multiple turns that we as users put in multiple turns that we as users put in but uh if we completely evict whatever but uh if we completely evict whatever but uh if we completely evict whatever we had in our particular conversation we had in our particular conversation we had in our particular conversation then the next time we ask a question in then the next time we ask a question in then the next time we ask a question in that same chat it's going to take a that same chat it's going to take a that same chat it's going to take a little bit longer. Uh so instead of little bit longer. Uh so instead of little bit longer. Uh so instead of actually completely evicting and redoing actually completely evicting and redoing actually completely evicting and redoing the prefill again what the the the prefill again what the the the prefill again what the the techniques being used are uh maybe using techniques being used are uh maybe using techniques being used are uh maybe using u we we are using our own but externally u we we are using our own but externally u we we are using our own but externally we know about LM cache and moon cakes. we know about LM cache and moon cakes. we know about LM cache and moon cakes. Um what we do is we will offload the KV Um what we do is we will offload the KV Um what we do is we will offload the KV cache to a high bandwidth storage so cache to a high bandwidth storage so cache to a high bandwidth storage so that we can store a lot of these that we can store a lot of these that we can store a lot of these prefills prefills prefills uh such that whenever the accompanying uh such that whenever the accompanying uh such that whenever the accompanying request comes for that particular request comes for that particular request comes for that particular conversation it can be loaded in right conversation it can be loaded in right conversation it can be loaded in right away into the HPM. Um on the performance liver I would I Um on the performance liver I would I just want to mention a few performance just want to mention a few performance just want to mention a few performance livers. We've discussed the PD livers. We've discussed the PD livers. We've discussed the PD disagulative decoding are others and how to carefully decoding are others and how to carefully choose the parallelization degrees and choose the parallelization degrees and choose the parallelization degrees and the strategies that is actually very the strategies that is actually very the strategies that is actually very important. Uh two of the biggest livers important. Uh two of the biggest livers important. Uh two of the biggest livers that we've been working with are that we've been working with are that we've been working with are quantization to NVFP4 and spec.
-
quantization to NVFP4 and spec. quantization to NVFP4 and spec. uh we do provide capability where if the uh we do provide capability where if the uh we do provide capability where if the customer has their data set and they customer has their data set and they customer has their data set and they want us to train in speculators for want us to train in speculators for want us to train in speculators for their data sets for better acceptance their data sets for better acceptance their data sets for better acceptance lens which will ultimately make the lens which will ultimately make the lens which will ultimately make the output throughput significantly higher. output throughput significantly higher. output throughput significantly higher. Uh we do have that as well. So but that Uh we do have that as well. So but that Uh we do have that as well. So but that happens async. We uh we get the data happens async. We uh we get the data happens async. We uh we get the data async we train the speculators async and async we train the speculators async and async we train the speculators async and then we deploy the speculators into the then we deploy the speculators into the then we deploy the speculators into the customer deployments uh if that's what customer deployments uh if that's what customer deployments uh if that's what they wanted. uh you see three they wanted. uh you see three they wanted. uh you see three screenshots over here. I have posted screenshots over here. I have posted screenshots over here. I have posted them from uh the last one month one them from uh the last one month one them from uh the last one month one month's worth of work that uh some of us month's worth of work that uh some of us month's worth of work that uh some of us in my team have done. You can see we in my team have done. You can see we in my team have done. You can see we came quickly on top of the leaderboard came quickly on top of the leaderboard came quickly on top of the leaderboard on Kimmy 2.6 2.7 and um those are those on Kimmy 2.6 2.7 and um those are those on Kimmy 2.6 2.7 and um those are those are from artificial analysis and going are from artificial analysis and going are from artificial analysis and going back to the session before this can we back to the session before this can we back to the session before this can we trust that that's why for GLM I have the trust that that's why for GLM I have the trust that that's why for GLM I have the results from open router. So artificial results from open router. So artificial results from open router. So artificial analysis when they run benchmarks analysis when they run benchmarks analysis when they run benchmarks they're running very specific workloads they're running very specific workloads they're running very specific workloads open router is actual user traffic uh open router is actual user traffic uh open router is actual user traffic uh and you can see uh on the open router and you can see uh on the open router and you can see uh on the open router side weights and biases so the branding side weights and biases so the branding side weights and biases so the branding is different but weights and bias is is different but weights and bias is is different but weights and bias is basically k we bought weights and biases basically k we bought weights and biases basically k we bought weights and biases about a year back uh you can see the about a year back uh you can see the about a year back uh you can see the speed over here that we have from our speed over here that we have from our speed over here that we have from our deployment is pretty close to what deployment is pretty close to what deployment is pretty close to what fireworks is providing us fireworks fast fireworks is providing us fireworks fast fireworks is providing us fireworks fast right but underlying techniques that we right but underlying techniques that we right but underlying techniques that we are using is what I want to emphasize are using is what I want to emphasize are using is what I want to emphasize the most over here uh for for the most over here uh for for the most over here uh for for performance optimization that becomes performance optimization that becomes performance optimization that becomes critical because ultimately what you critical because ultimately what you critical because ultimately what you want to serve to the customer what we want to serve to the customer what we want to serve to the customer what we want to serve to the customer is uh want to serve to the customer is uh want to serve to the customer is uh price performance benefit
-
quick recap um single platform is what quick recap um single platform is what I've been trying to emphasize is what I've been trying to emphasize is what I've been trying to emphasize is what I've tried to show uh two different I've tried to show uh two different I've tried to show uh two different consumption models serverless and consumption models serverless and consumption models serverless and dedicated for customers and within dedicated for customers and within dedicated for customers and within serverless I describe two different serverless I describe two different serverless I describe two different consumption models as well pay as you go consumption models as well pay as you go consumption models as well pay as you go and provision vision throughput if you and provision vision throughput if you and provision vision throughput if you care about that and ultimately care about that and ultimately care about that and ultimately compounding the gains uh through compounding the gains uh through compounding the gains uh through performance optimizations uh in the performance optimizations uh in the performance optimizations uh in the stack. stack. stack. That's all. Thank you folks.
Summary
This tech talk focuses on the inference platform at Corv, designed to serve a variety of models and workloads. The presenter, with experience from AWS and SANOVA, discusses consumption models like serverless and dedicated inference, emphasizing the importance of performance and platform flexibility. The key takeaway is understanding these consumption models to build a robust and adaptable inference platform.