**From Caching to Batching to Flex  How to optimize AI system for production - Google DeepMind — Transcript & Summary | SozAI**
Source: https://sozai.app/transcript/caching-batching-flex-optimize-ai-production/

Learn how Google DeepMind optimizes AI systems for production using caching, batching, and Flex to reduce costs and improve scalability.

## Key Takeaways

- Caching is essential to reduce inference costs by reusing previously computed context.
- Batch API allows efficient handling of large volumes of requests with acceptable latency trade-offs.
- Flex pricing introduces composability in AI system optimization, enabling tailored cost-performance balance.
- AI workloads vary in latency sensitivity, requiring different optimization strategies within the same system.
- Continuous innovation, such as blockwise caching, is needed to further enhance scalability and cost-efficiency.

## What the video covers

- The session focuses on making AI more affordable and scalable for production environments.
- Discussion centers around three optimization pillars: caching, batch API, and Flex pricing models.
- Caching reduces costs by avoiding repeated computation of previously processed context.
- Batch API enables high throughput and low latency by batching multiple requests, suitable for background operations.
- Flex pricing offers a new approach to balance criticality, reliability, and cost with composable trade-offs.
- Challenges include handling repeated context in multi-turn conversations and agentic systems.
- Latency mismatches across different AI workloads require flexible optimization strategies.
- Google’s Gemini API incorporates these pillars and offers significant cost discounts (up to 90% for caching).
- The talk references similar approaches by Anthropic and OpenAI for context and API design.
- Future concepts like blockwise caching could further improve caching efficiency and user experience.

## Chapters

1. 00:00 Introduction to AI affordability and scaling
2. 01:44 Challenges of repeated context in agentic systems
3. 03:10 Latency differences and use case considerations
4. 04:50 Pricing models and cumulative discounts
5. 06:25 Batch API for high throughput and background processing
6. 08:07 Caching mechanisms and cost reductions
7. 11:31 Flex pricing: balancing criticality, reliability, and cost
8. 13:10 Future concepts: blockwise caching and demos

Answers

## Questions about this video

What are the main methods to optimize AI system costs discussed in the video?

The video discusses three main optimization methods: caching to avoid repeated computations, batch API for handling multiple requests efficiently, and Flex pricing to balance cost and reliability.

How does caching help reduce AI inference costs?

Caching stores previously processed context so the system does not recompute it, resulting in up to 90% cost reduction on repeated content.

What is the Flex pricing model introduced by Google DeepMind?

Flex is a new pricing approach that allows composable trade-offs between criticality, reliability, and cost, providing more flexibility than standard or cached pricing.

## Full Transcript — Download SRT & Markdown

00:00

Speaker A

Hi everyone, and welcome. Uh, we have a session about how to make AI more affordable at scale, and as you guys are all building at scale, this is a perfect moment to talk about it.

00:12

Speaker A

it. Um I'm trying to um keep it quite conceptual like and then also refer to the tools that we have. We are both working in Gemini API and then Patrick will take it over and make it very hands-on and tangible with demos and

00:27

Speaker A

Um, I'm trying to keep it quite conceptual and then also refer to the tools that we have. We are both working in Gemini API, and then Patrick will take it over and make it very hands-on and tangible with demos and best practice cases.

00:42

Speaker A

model development has greatly accelerated over the past couple of days both days both in the um the Gemini core modeling as well as the free model space. However, the models are getting more efficient and effective. they are um cheaper in price for what the

00:59

Speaker A

So let's get started just for context setting, and I know this is focused around Gemini, but you can put that like you can put any other naming convention and any other big AI lab on this. We see that the speed of model development has greatly accelerated over the past couple of days, both days, both in the Gemini core modeling as well as the free model space.

01:15

Speaker A

optimize um this is not a exhaustive lift list of challenges that we have but those are the ones that I um selfishly focus on and that I built product for and API designs and features for as well. So, one of the biggest things and

01:30

Speaker A

However, the models are getting more efficient and effective. They are cheaper in price for what the performance is. We still have quite a lot of challenges when it comes to actually scaling because the inference cost doesn't scale with the value just linearly. In order to build sustainable models and sustainable systems, we need to have other ways to optimize.

01:44

Speaker A

repeated context. So as we move into more agentic systems, as we move into like multi-turn conversations and uh the context that we are sending and the and the content that we're sending gets longer and longer. So we want to rem the

01:58

Speaker A

This is not an exhaustive list of challenges that we have, but those are the ones that I selfishly focus on and that I built product for and API designs and features for as well. So, one of the biggest things, and I think I'm not sure if you guys have seen the mainstage presentation, is around context.

02:06

Speaker A

Well, we all started with just regular standard interactive traffic or interactive serving which means you send you get back. But as we kind of broaden horizon of different use cases that we cover, we see that latencies are different. They differ from use case to

02:23

Speaker A

So, I'm trying to kind of brush over that a little bit faster because you guys are all very knowledgeable around caching, but repeated context. So as we move into more agentic systems, as we move into like multi-turn conversations, the context that we are sending and the content that we're sending gets longer and longer.

02:34

Speaker A

million times and you will hear it a million times over, agents. So as you're building background agents, as you're building long long running operations, uh latency becomes less of a concern and you're willing to maybe have that as a

02:47

Speaker A

So we want the system to remember more things, and there's a lot of repeat sending. Then the second one is latency mismatch.

02:55

Speaker A

So basically right now we don't have a lot of tools to kind of do these tradeoffs in different settings in the same like in the same kind of AI um um workflow or AI system. So you basically have one way of

03:10

Speaker A

Well, we all started with just regular standard interactive traffic or interactive serving, which means you send, you get back. But as we kind of broaden the horizon of different use cases that we cover, we see that latencies are different. They differ from use case to use case.

03:21

Speaker A

make optimization a more composable solution for everyone so that they can all choose the right setting. So this is kind of the area of challenges that we have. Okay, I'm seeing my time. Um so these are the three kind of pillars I

03:35

Speaker A

You will have something where it's critical you get a response right away. But you will also have AI workloads and systems that can absolutely run in the background. And I'm pretty sure you heard the word a million times, and you will hear it a million times over: agents.

03:49

Speaker A

implement these pillars, but we definitely have it too. So caching um caching is kind of the solution to repeat content sending. So you stop paying for your what you already sent and what you already computed. Batch API is a way to have high throughput um with

04:06

Speaker A

So as you're building background agents, as you're building long, long running operations, latency becomes less of a concern, and you're willing to maybe have that as a trade-off for pricing. So that is another big kind of challenge field that I'm trying to tackle with the things that I offer.

04:18

Speaker A

run in the background. And flex is something that we launched rather new. However, it's been conceptually around is um moving the needle on criticality and basically uh and on reliability but also at a different cost uh at a

04:37

Speaker A

And then all or nothing. So basically right now, we don't have a lot of tools to kind of do these tradeoffs in different settings in the same AI workflow or AI system. So you basically have one way of interacting with an AI model, and that is either standard or it's cached or non-cached, but it's not very composable.

04:50

Speaker A

regular uh interactive pricing. So that will make a cumulative like you can't accumulate them all together but if you accumulate two together and I will get to that you get around 95% discount and I think that becomes more viable as you

05:03

Speaker A

So I am looking at the challenge of how to make optimization a more composable solution for everyone so that they can all choose the right setting. So this is kind of the area of challenges that we have.

05:19

Speaker A

back from a KV cache and then only compute new tokens that the system sees and that's basically the principle of caching and I know there was a lot around caching already said so I'm going to be moving on to just ex um explain

05:35

Speaker A

Okay, I'm seeing my time. So these are the three kind of pillars I chose to discuss. They relate to things that we're doing in the Gemini API, and they relate to things that everyone else is doing too.

05:54

Speaker A

It's a way of uh it's created via the client SDK. So it's a little bit harder to integrate and do. You have to set that up. You have to create the cache.

06:03

Speaker A

I will also reference Anthropic and OpenAI with regard to specific ways of how they implement these pillars, but we definitely have it too. So caching, caching is kind of the solution to repeat content sending. So you stop paying for what you already sent and what you already computed.

06:17

Speaker A

libraries will return agents. So, this is wherever you have big chunks that you want to make sure gets cached.

06:25

Speaker A

Batch API is a way to have high throughput with low latency kind of use cases covered. So have like, stop paying for the speed that you don't need, but also the ability to batch up a bunch of requests into one and then have that run in the background.

06:39

Speaker A

This is basically zero coding change. Um for us it's Gemini that ret that um that detects repeated um prefix automatically and then basically um once and hits the cache um and for each of the cash hits you get a 90% discount. However, really

06:57

Speaker A

And Flex is something that we launched rather new. However, it's been conceptually around. It is moving the needle on criticality and basically on reliability but also at a different cost, at a different cost point or price point.

07:09

Speaker A

So there's definitely a lot of optimization that can be done uh for implicit caching and I think a lot of the previously talked about um methods of um or best practices in caching kind of goes to kind of uh to increase

07:24

Speaker A

And all of these, you can see that we have like also how we currently discount. So we have caching at 90%, and then batch as well as flex at 50% of regular interactive pricing.

07:36

Speaker A

There's a new concept that is currently as far as I know not released and and not in use with any of the big um providers but definitely something that uh we are looking at conceptually in the industry and that would be a huge

07:51

Speaker A

So that will make a cumulative—like you can't accumulate them all together, but if you accumulate two together, and I will get to that, you get around 95% discount, and I think that becomes more viable as you guys scale.

08:07

Speaker A

don't have that. You cache arbitrary segments and then basically have a like you have a lookup um to kind of look up the actual segments. And this can be incredibly [snorts] powerful, especially when it comes to things like coding agents where you have

08:22

Speaker A

I'm not sure how much I'm going to go into context caching, but basically computing attention is quite expensive for long context. What we're doing is we basically save a precomputed attention state and then we load that back from a KV cache and then only compute new tokens that the system sees, and that's basically the principle of caching.

08:35

Speaker A

and match different documents all kind of um uh kind of cached into separate blocks and then you can so like you can uh retrieve them and create your own like uh context very easily very composable and also for agents tools

08:50

Speaker A

I know there was a lot around caching already said, so I'm going to be moving on to just explain the core principles very high level. We have explicit caching and implicit caching, and different labs and different AI providers provide different versions of this.

09:02

Speaker A

pitch that I'm trying to get through and if technology improves this can be very very interesting um for future development I also just wanted to show something that you guys haven't heard about caching um before we move on okay

09:17

Speaker A

So explicit caching, for example, is provided by Anthropic as well as by us. It's a way of—it's created via the client SDK. So it's a little bit harder to integrate and do. You have to set that up. You have to create the cache.

09:34

Speaker A

file. You submit it um then it gets cued and it waits for max 24 hours to have uh off peak um capacity and then basically get it done. And then new we launched this new now conceptually new is web

09:49

Speaker A

However, for example, we guarantee that it comes at the 90% cost reduction, that anything that you explicitly cache will hit that cache or will at least get you the cost reduction.

10:04

Speaker A

definitely look at batch especially for large sums like where you look for like big chunks of of data that you need to process like eval um other workloads.

10:18

Speaker A

And that obviously is for large static libraries or return agents. So this is wherever you have big chunks that you want to make sure gets cached.

10:31

Speaker A

launch both priority and flex. Flex is really like best effort reliability. So we are kind of like we are um working with criticality in the background and just make it less critical. So if the model like is used a lot and we are like

10:44

Speaker A

And then implicit caching is basically what OpenAI runs with. However, everyone starts to offer this as well. I think OpenAI does like standard two hours with the opt-in mechanism of doing it up to 24 hours.

10:58

Speaker A

discount. So if you are getting like flex um replies then the discount is if you put in um con in your config the service tier flex and it hits it then you get the discount. This is definitely not for real time critical um workloads

11:14

Speaker A

This is basically zero coding change. For us, it's Gemini that detects repeated prefix automatically and then basically once and hits the cache. And for each of the cache hits, you get a 90% discount.

11:31

Speaker A

like the SDK integration is minimal for flex. Um batch has its own endpoint. Both come at the same discount. Um the SLAs's and SLOs's are slightly different with 24 hours and then best efforts but probably minutes. Um and one is for

11:46

Speaker A

However, really important, cache hits are not guaranteed, that discount is not guaranteed, and you will see it with different providers. You have different kind of cache hit rates also kind of reliant on the use case that you have.

12:05

Speaker A

how you can use the different tools. Explicit hashing is a great tool to be used either with flex uh or batch. I would say reflex even more than batch because I have to tell you guys um if you're using explicit caching plus batch

12:20

Speaker A

So there's definitely a lot of optimization that can be done for implicit caching, and I think a lot of the previously talked about methods or best practices in caching kind of go to increase implicit cache hit rate.

12:28

Speaker A

So I'm not trying to hide that. So I would say if you're using it with flex then you get the full reduction. If you're using it with batch you have to make sure that the cash is lifelong enough um for that actually to kind of

12:40

Speaker A

That's based for the dynamic user content and just interactive without having to set up databases around it.

12:56

Speaker A

sure that you're using um whatever the tier service tier is or batch bin depending on your latency. Really think about it what is necessary is what is very critical and what has more time to um process. and then think about how to

13:10

Speaker A

There's a new concept that is currently, as far as I know, not released and not in use with any of the big providers, but definitely something that we are looking at conceptually in the industry, and that would be a huge unblocking for caching, which is blockwise caching.

13:23

Speaker A

Thank you. Yeah. Hi everyone again. I'm Patrick. I have the honor of doing the fun part and doing the live demo and we're going to see all of this now in action and look at one concrete example.

13:34

Speaker A

So right now, as some of you heard, prefix caching is very constrained. You need to use very static content for it to actually have a cache hit. With blockwise, you don't have that. You cache arbitrary segments and then basically have a lookup to kind of look up the actual segments.

13:41

Speaker A

So here we have different workloads uh on our platform where we can leverage all the different techniques. For example, let's say uh every at the end of the day overnight or every weekend we want to embed the student home

13:54

Speaker A

And this can be incredibly powerful, especially when it comes to things like coding agents where you have entire repositories that you can build blocks around and then only update the block in the cache where you actually make the updates to the document or the code.

14:11

Speaker A

can combine it with caching. And then last but not least for example let's say we want to generate a personalized learning road map um that the ch student wants to generate there. We still want to get this relatively fast, but it's

14:23

Speaker A

So that is one really cool use case. Another one is when you like—

14:31

Speaker A

Um, I prepared all of this and I will also uh uploaded or is already upload uploaded to GitHub and I will share the link later. But basically, I have uh three different scripts and I want to quickly walk over this and show you how

14:47

Speaker A

to apply these techniques. Um, this is using the Google Chennai Python SDK. Who is who is already building with the with the Gemini API and Python SDK?

14:58

Speaker A

Not that many. So, some of you might see this for the first time. Um, but yeah, it's it's the Python SDK. You can simply install this and then import the package. And it's one of the easiest way to work with the Gemini models. Um, and

15:11

Speaker A

there is some code to get some nice terminal output. We can ignore this but basically this is the first um example to generate uh embeddings via the batch API and basically what we need to do is we need to store the requests in a JSON

15:29

Speaker A

or JSON lines file. So here for each line we have one request or one text that we want to embed and step one is to upload the file. So we call client files upload and then once it's uploaded we

15:46

Speaker A

start the batch shop. So we call client batches create embeddings. And this is our latest model Gemini embedding 2.

15:53

Speaker A

Maybe you've heard of it. We launched it just a few weeks ago. It's a multimodal embedding. So it can actually do much more than just text. Um and then we wait. So let's actually jump to the uh to the terminal and start the script. Uh

16:07

Speaker A

Python. And it's the first one. And now it's pulling uh the info from the files. And then we see it uploaded the file. And now it submitted the batch job. And now we're waiting. So it says here job state

16:24

Speaker A

pending. If you can see this. Um and yeah, we we guarantee that this is coming back within 24 hours. In practice, it's actually often much faster. So maybe we are lucky and we get this while I'm doing the demo. Uh we can

16:39

Speaker A

check back in a in a minute. But maybe while this is still running, I can give you a quick tip. So here we're kind of every 10 seconds we're pulling for the job state. But we could actually just

16:52

Speaker A

remove this and instead Lucia mentioned this. We can configure a web hook hook. So we can say client web hooks.create create and there you can subscribe to different chops um uh to different events. So for example here you can

17:08

Speaker A

subscribe to the batch job succeeded and then define a URL and then you get a notification to this URL uh once this is finished. Uh and then you need all of don't need all of this pulling. Uh yeah

17:24

Speaker A

and this is the first example. Let's quickly check if yeah no it's still running. So this is out of my control.

17:29

Speaker A

But I'm not sure when it's coming back. Uh but let's check out the second example. So second example, we want to use caching. Uh so here we are um kind of assuming that the student brings a large textbook uh to the chat and I have

17:46

Speaker A

one book that I downloaded uh the mysteries of beekeeping explained and then the student uh wants to ask different questions about this um this book. And this is a nice example where we can leverage caching. And to do that,

18:01

Speaker A

uh, we're loading the system prompts together with the book as a cache. And for this, we're configuring client caches create. And then here we can also specify the time to live. And then once we have the cache later, then when we

18:18

Speaker A

generate the chat responses, client models generate content. Here we simply apply the cache. And then it's uh using the cache. So let's start the second um second script. Then I have to speed up a little bit. Uh so here it's estimating

18:39

Speaker A

the size. You see it has around 150,000 tokens. So Gemini can can handle this pretty well. And now it created the cache. We get a key. We also see the expiration time. And now we're simulating the first student question.

18:54

Speaker A

and we get the first response and I did a second one and after that I do some calculations. We also heard in the or maybe you heard in the other uh talk that it's important to have uh to

19:07

Speaker A

monitor your system to like calculate the the cash rate and the pricing. Um so I'm doing this in this utility class which you can check out later on on GitHub. Uh yeah, let's give it a few seconds and then hopefully we should see

19:26

Speaker A

the comparison or the number of of cache tokens. If the live demo gods are with me and I have 30 more seconds, so I need to speed up if the live demo gods are Oh, yeah.

19:42

Speaker A

Okay, we get it actually. Nice. Yeah. So, we have the input tokens 150,000. The cash tokens pretty close. So this is an extreme example where we load a lot of context into the cache. So it's actually 99.9%.

19:56

Speaker A

And then the output tokens. Output tokens are a little bit more expensive than input tokens. So if you do the calculations under the hood, you will see here we get around 87% cost savings.

20:07

Speaker A

And I actually prepared a second script where you can try the same with implicit caching. So there you don't even have to configure all this uh explicit caching.

20:16

Speaker A

And then you can compare how the cache rate uh cach it rate is here. And then I have one more um script but I'm running out of time. But basically to apply the flex tier you simply configure it uh

20:30

Speaker A

here uh when sending the prompt client models generate content. You apply the service to your flex and then it's also good practice to configure a timeout um so you don't get a timeout. And then you also maybe want to implement um retries

20:45

Speaker A

with an exponential backoff and then maybe a fallback um to fall back to the standard tier. And then here I'm also wrapping this into a thread so it's not blocking the main um the main thread. Um let's start this as well very quickly. It

21:06

Speaker A

should come back relatively um relatively quickly and then I'll wrap it up. And yeah, here you see here we have attempt one. Now it's trying the flex here. Sometimes it might make might fail. Then you can do the the retries. Uh but here we get it.

21:26

Speaker A

And then here you see it's applying the flex routing. And I did some like cost calculations. And if we scale this up to, for example, 1,000 requests in this example, we save uh $9.

21:39

Speaker A

And yeah, that's it. That's how to apply it in action. I do have some best practices, but some of them I already mentioned and I will add them just to the GitHub repository. So feel free to grab the uh the link with this QR code.

21:52

Speaker A

And yeah, thank you so much for your attention. I hope you learned something new. And we're also around uh or I'm around for questions. Uh yeah, thank you. Amazing. [applause] Thank you very much. Um, we have time for maybe one or two quick

22:13

Speaker A

questions if anybody has any. Sorry. Thanks. Uh, I enjoyed the end of that. I'm sorry it's a bit late, but I was curious um how you see caching developing over the next two or three years uh in terms of um the cost and the

22:36

Speaker A

capability. Is it going to become cheaper uh versus the costs of the models? like is is the are fresh tokens or rather are cash tokens going to become cheaper in compar in comparison to fresh tokens uh over time or or do

22:51

Speaker A

you think it'll still be around 90%. So I guess like [laughter] I'm I'm hardressed to say they're going to become cheaper than like 90% discounted. I'm not sure if we can like how much margin we still have to go even

23:05

Speaker A

cheaper. But what I think like the biggest areas are is like increasing cash hit rates like even for implicit not just for explicit and then coming up with more um sophisticated solutions of how it's done so that the user doesn't

23:18

Speaker A

have to make sure that they really like you know like the the blockwise is kind of like the area where I see like loads of improvements from user experience perspective on how to build it and how to make sure we refine it but like

23:30

Speaker A

purely cost I'm not sure I'm not sure how that developed. What's the limiting factor?

23:36

Speaker A

What's the limiting factor? Like for um for fresh tokens, it's uh basically the energy to like move memory, right? But for um for cache tokens, you're not doing that you're not doing those calculations you need to do. So like why

23:52

Speaker A

is it uh what's the limiting factor on there's I mean one for one for example that I know um is um the uh the tal like the the balance between like um general kind of like um serving optimization and

24:10

Speaker A

routing and like how to like split into different um like uh chips in when when kind of computing and then on the other hand like having cache where you like that needs to be at the same location than the compute. to like kind of like

24:24

Speaker A

optimizing just the general serving and optimizing specifically for caching and cache storage like that's one of the balances to keep for example and then also like how to extend TTL like time like you know like to extend the lifetime of the uh of the

24:40

Speaker A

cache and make it more durable while at the same time like optimizing cost of actual durable storage space like also like the type of chips that we use in order to store like all of that is like all of those are optimization talent

24:53

Speaker A

challenges. But yeah, I see like the we definitely like those are the areas to explore like the user experience side as well as the kind of internal like how we're going to do it text tech side and Alan.

25:05

Speaker A

Amazing. Thank you very much. Yes. Thank you. [applause] There we go.

Topics: AI optimization caching batch API Flex pricing Google DeepMind Gemini API AI scalability inference cost multi-turn conversations agentic systems


---
This is the markdown twin of https://sozai.app/transcript/caching-batching-flex-optimize-ai-production/ — the same content, without the markup.
Published by SozAI (https://sozai.app). Reuse and quotation are allowed with attribution and a link back.
Machine-readable index: https://sozai.app/llms.txt · data API: https://sozai.app/api/
