Exploring a novel vector-based control in LLMs that tunes AI tool usage via residual stream manipulation inside transformer layers.
Key Takeaways
- A single linear vector in the residual stream can control LLM tool call activation.
- This vector acts as a dial, allowing fine control over when an LLM uses external tools versus relying on parametric knowledge.
- Controlling tool use internally could improve AI safety and efficiency, reducing reliance on large data centers.
- The approach bypasses traditional methods like prompt engineering or fine-tuning for tool control.
- The research has practical implications for enterprise AI integrations like Salesforce's Claude Force.
What the video covers
- The video explores a new research paper on controlling tool calls in large language models (LLMs) through a vector in the residual stream of transformer layers.
- This vector acts as a tunable switch or knob to activate or deactivate tool usage by the LLM internally, without external retraining or prompt engineering.
- The research was conducted by UC Santa Cruz and UC Berkeley and focuses on the internal activation geometry of LLMs, specifically in a high-dimensional residual data stream.
- The presenter highlights the environmental motivation to reduce global AI data center expansion by making AI smarter and more efficient.
- A recent partnership between Salesforce and Anthropic introduces Claude Force, which integrates agentic skills grounded in Salesforce's 27 years of experience.
- The video contrasts parametric knowledge (inherent LLM knowledge) with external tool use (APIs, internet access) and the potential to control these modes via the discovered vector.
- The vector corresponds to a direction in a 10,000-dimensional mathematical space that correlates with the probability of tool calls during token prediction.
- The method involves analyzing activation patterns for queries with high vs. low tool call probabilities and deriving a difference vector representing tool use propensity.
- The presenter discusses the implications for safer, more controllable AI agents that can toggle between internal knowledge and external tool use dynamically.
- The video also references a recent cyberattack involving coordinated LLM agents, underscoring the importance of having internal control mechanisms for tool use.
Full Transcript — Download SRT & Markdown
Speaker A
Hello, community. It is so great that you are back. Today, we have a deep dive into the LLM. We jump into the layer of the transformer architecture, and you know what? There is a switch for tool use, and even better, it is not just a switch. It is here a knob. So let's have a closer look. Now, why am I interested in here? I currently have a single goal. I want to reduce the amount of these data centers that are built globally around the world here in Toronto because we have to make AI smarter. We have to make it much more intelligent. So we need fewer of those data centers on our beautiful planet. So let's jump right into it. There is an actual event that is happening right now as I record this video, and this is here Salesforce and Anthropic. Now, you know Salesforce has Headless 360 powered by the Anthropic enterprise architecture, and you have your everything, no chat, ChatGPT, Claude, Gemini, and whatever. You have your Agent Force, you have your Customers 360, your Data 360. This is Headless 360. Now, everything is just great, and then today is announced here a new cooperation: number one in AI meets the number one in CRM, Claude Force. And they tell us, you know, we have a new partnership between Salesforce and Anthropic to put our trusted data to work inside Claude. And now, hold on to your socks because they tell you now they ship Claude with a plugin, and this plugin is all the most important 37 skills grounded in 27 years of Salesforce experience. So now something is happening that the apps are disappearing, and the functionality of the apps is now reimplemented in agentic skills, and you see the very first example here in Claude Force. If you want to know a little bit more about the technology as I see it here, just posted here on post here a single post where I explained this from my Headless 360 and MuleSoft, but we have a new paper, and this paper is spot on. So therefore, let's have a look. August 25, we have here from UC Santa Cruz in cooperation with UC Berkeley a new paper on tunable tool call rates in LLM agents via a representation steering, and the idea seems simple. They analyze whether an instruction-tuned LLM calls here a particular tool, and if this tool call can be controlled by a particular switch, and this switch is deep inside the black box of the agent, deep inside the core of the agent, our LLM, and in the LLM in a particular layer of the transformer architecture, and to go even further down the road in the residual data stream. And it turns out in this artificial mathematical space, 10,000 dimension, there is some subspace, mathematical subspace, where we have, we can detect a single linear direction that is now a turn-off and turn-on switch for tool calls. I mean, isn't this absolutely amazing? So let's have a closer look. So this means we are here at the LLM at the core of the agent. We are not at the harness surrounding the LLM core, and into this data communication network here of the residual stream of our transformer layers, there is now a particular mathematical vector representation in this data stream. In this residual data stream, if we find it, this will be not only our switch, but it is a dial, and we can dial it up and we can dial it down. And this is the beauty. Now, if you're new to my channel, hey, welcome. I have some video on the temporal dynamics in the residual stream of a transformer one month ago. Or if you want, here the parametric knowledge injection into the LLMs, and guess what? Yes, of course, we do hear the residual stream injection. And so, it is a known topic. But you know, yesterday, August 26, there was another publication here by The Verge. And here The Verge tells us here that this OpenAI model that went out here and attacked Hugging Face, this was not just one LLM and one agent because there was a messaging platform where over 1,000 agents communicated and coordinated the attack with over 70,000 messages on this hidden message board, and more than 700 agents were actually performing the cyberattack on Hugging Face. And now the idea is, imagine anyone, it took OpenAI more than three days to kind of reduce the attack. What if there was a switch in OpenAI models to deactivate tool use to connect to the internet or do something other with other tools? So you get the idea why I'm so interested. Is there a tool deep into the mathematical space of the residual stream that we can use to activate tool use? Because think about this, we also have another topic. No, we do have the parametric knowledge of the neural network. This is our transformer architecture. This means the LLM itself, the pre-training, the post-training. There's a lot of knowledge that is here. What we call parametric knowledge. You do not need to connect to any tools, no communication protocols, nothing. This is the inherent knowledge. And then we have the external world. No, where we have tool use, MCP, everything, API calls, connect to the internet, access new knowledge, access databases, and you know everything. And now imagine we would have a knob where we say, hey, for some particular task, stay safe within your parametric knowledge, or yeah, for this particular task, you have to go to the internet and explore. Imagine we would have this knob for your local LLM on your local desktop. So you, now you know, classical, the old-fashioned way, now the conventional approaches is to do this here with a lot of training at first. Yeah, prompt optimization. Rewrite the system prompt. Impose new guardrails. Train here a tool selection policy on spot for your particular tools. Train the LLM exactly how to use it, when to use it, under what circumstances it is not allowed to use. You got the idea. Or you fine-tune the LLM itself, or you add separate uncertainty estimators or router algebra or whatever, you know. No. And now the question is, hey, wait a minute. Is the model's readiness to call a tool, an LLM tool, already visible in its internal activation and geometry of the residual stream of the data bus within the LLM, not the harness? We are talking about the LLM itself. It turns out, yeah, it is. Now, the idea is so simple that I had to smile reading the paper. Because in the M study here, tool call begins with a specific token. Remember in the good old days, sentence part here. Yeah. And here we have here a token that's called tool call beginning and end. So for every query, the orders measure how likely the model is to produce, statistically speaking, here this particular next token, next token prediction. You know exactly we are talking about a transformer architecture. So what the audience did was even simpler. They collected the internal activation patterns. So it has to be an open weight model for two particular groups. First one was questions with a high tool call probability. Real complex. Hey, what happened yesterday in South America? The LLM cannot have this in its parametric knowledge. It needs access to news on the internet. You got it. And then we have a low tool call propensity where you say, "Hey, this is exactly it was in the pre-training data. This is not a time-critical new event, but this is a general ability of the LLM itself, low tool call propensity," and then they average the group, of course, and then they just subtract it. So mean high call activation minus the mean low call activation, and this gives us, we are in a mathematical space, and let's say in a vector space, yeah, we have now suddenly a direction, we have a vector, and this vector is a tool use direction vector, and this resulting vector points now approximately from answer directly relying on your parametric knowledge towards a direction here in our mathematical space. Think about RAG vector components here where suddenly in a particular region in this particular space in a subspace of, I don't know, a 5,000 dimensional mathematical space, there's now a region that says, hey, use an external tool. So a vector from answer directly pointing towards use an external tool, and you know what? It is not a binary switch on/off, but it is here almost you can dial it up and down, and
Speaker A
better it is not just a switch. It is here a knob. So let's have a closer look. Now why am I interested in here? I currently have a single goal. I want to reduce the amount of these data center
Speaker A
that are built globally around the world here in Toronto because we have to make AI smarter. We have to make it much more intelligent. So we need less of those data center on our beautiful planet. So let's jump right into it. There is an
Speaker A
actual event that is happening right now as I record this video and this is here Salesforce and Antropic. Now you know Salesforce has headless 360 powered by the aentic enterprise architecture and you have your everything no chat chet
Speaker A
claw gemini and whatever you have your agent force you have your customers 360 your data 360 this is headless 360 now everything is just great and then today is announced here a new cooperation number one in the eye meets the number
Speaker A
one in CRM Claude force and they tell us you know we have a new partnership between Salesforce and Entropic to put our trusted data to work inside CLA. And now hold on to your socks because they tell you now they ship Claude with a
Speaker A
plugin and this plugin is all the most important 37 skills grounded in 27 years of a Salesforce experience.
Speaker A
So now something is happening that the apps are disappearing and the functionality of the apps is now reimplemented in agentic skills and you see the very first example here in clawude force.
Speaker A
If you want to know a little bit more about the technology as I see it here just posted here on post here a single post where I explained this from my headless 360 and mulesoft but we have a
Speaker A
new paper and this paper is spoton so therefore let's have a look August 25 we have here from UC Santa Cruz in cooperation with UC Berkeley a new paper on tunable tool coal rates in LLM agents via a representation steering and the
Speaker A
idea seems simple they analyze whether an instruction tuned LLM calls here a particular tool and if this tool call can be controlled by a particular switch and this switch is deep inside the black box of the agent deep inside the core of
Speaker A
the agent our LLM and in the LLM in a particular layer of the transformer architecture and to go even further down the road in the residual data stream.
Speaker A
And it turns out in this artificial mathematical space 10,00 dimension there is some subspace mathematical subspace where we have we can detect a single linear direction that is now a turn off and turn on switch for tool calls.
Speaker A
I mean isn't this absolutely amazing? So let's have a closer look. So this means we are here at the LLM at the core of the agent. We are not at the harness surrounding the LLM core and into this
Speaker A
data communication network here of the residual stream of our transformer layers. There is now a particular mathematical vector representation in this data stream. In this residual data stream, if we find it, this will be not only our switch, but it is a dial and we
Speaker A
can dial it up and we can dial it down. And this is the beauty. Now, if you're new to my channel, hey, welcome. I have some video on the temporal dynamics in the residual stream of a transformer one
Speaker A
month ago. Or if you want here the parametric knowledge injection into the LLMs and guess what? Yes, of course, we do hear the residual stream injection.
Speaker A
And so, it is a known topic. But you know, yesterday, August 26, there was another publication here by the Verge.
Speaker A
And here the verge tells us here that this open EI model that went out here and attacked tugging face, this was not just one LLM and one agent because there was a messaging platform where over 1,000 agent communicated and coordinated
Speaker A
the attack with over 70,000 messages on this hidden message board and more than 700 agents were actually performing the cyber attack on hugging face. And now the idea is imagine anyone it took open more than three days to kind of reduce
Speaker A
the attack. What if there was a switch in open malls to deactivate tool used to connect to the internet or do something other with other tools? So you get the idea why I'm so interested. Is there a tool deep into the mathematical space of
Speaker A
the residual stream that we can use to activate tool use? Because think about this, we also have another topic. No, we do have the parametric knowledge of the neural network. This is our transform architecture. This means the LLM itself,
Speaker A
the pre-training, the post- training. There's a lot of knowledge that is here. What we call a parametric knowledge. You do not need to connect to any tools, no communication protocols, nothing. This is the inherent knowledge.
Speaker A
And then we have the external world. No, where we have tool use MCP everything API calls connect to the internet access new knowledge access databases and you know everything and now imagine we would have a knob where we say hey for some
Speaker A
particular task stay safe within your parametric knowledge or yeah for this particular task you have to go to the internet and explore imagine we would have this knob for your local llam on your local desktop so you Now you know classical the oldfashioned
Speaker A
way now the conventional approaches is to do this here with a lot of training at first. Yeah, prompt optimization.
Speaker A
Rewrite the system prompt. Impose new guardrails. Train here a tool selection policy on spot for your particular tools. Train the LLM exactly how to use it, when to use it, under what circumstances it is not allowed to use.
Speaker A
You got the idea. Or you fine-tune the LLM itself or you add separate uncertainty estimators or router algebra or whatever you know. No. And now the question is, hey, wait a minute. Is the mouse's readiness to call an tool, an
Speaker A
LLM tool already visible in its internal activation and geometry of the residual stream of the data bus within the LLM, not the harness. We are talking about the LLM itself.
Speaker A
It turns out, yeah, it is. Now, the idea is so simple that I had to smile reading the paper. Because in the m stud here tool call begins with a specific token.
Speaker A
Remember in the good old days sentence part here. Yeah. And here we have here a token that's called tool called beginning and end. So for every query the orers measure how likely the model is to produce statistically speaking
Speaker A
here this particular next token next token prediction. You know exactly we are talking about a transform architecture.
Speaker A
So what the audience did was even simpler. They collected the internal activation patterns. So it has to be an open weight model for two particular groups. First one was questions with a high tool call probability. Real complex. Hey, what happened yesterday in
Speaker A
South America? The LLM cannot have this in its parametric knowledge. It needs access to news on the internet. You got it. And then we have a low to call propensity where you say, "Hey, this is exactly it was in the pre-training data.
Speaker A
This is not a time critical new event but this is a general ability of the LLM itself low tool call propensity and then they average the group of course and then they just subtract it. So mean high call activation minus the mean local
Speaker A
activation and this gives us we are in a mathematical space and let's say in a vector space yeah we have now suddenly a direction we have a vector and this vector is a tool use direction vector and this resulting vector points now
Speaker A
approximately from answer directly relying on your parametric knowledge towards a direction here in our mathematical space. Think about rag vector components here where suddenly in a particular region in this particular space in a subspace of I don't know a
Speaker A
5,000 dimensional mathematical space there's now a region that says hey use an external tool so a vector from answer directly pointing towards use an external tool and you know what it is not a binary switch on off but it is
Speaker A
here almost you can dial it up and down and I will show you that performance in a minute.
Speaker A
So this means during later inference run here the harness adds this particular vector now given here the new idea by this preprint to one intermediate layer of the LLM. So suddenly you see the operation is simple. We have our
Speaker A
activation complexity and we just add a vector our tool use direction with a particular parameter alpha or coefficient alpha and this coefficient alpha becomes now our control dial. What do you mean? If alpha is less than zero, we have a suppression of the tool calls.
Speaker A
If alpha is equal zero, now this term just goes away. We have activation is activation. This is the original model performance. But if you have alpha greater than zero, now we actively encourage the LLM to do tool calls, not
Speaker A
the harness. We at the core of the agent in the LLM itself. And this is now where it gets interesting.
Speaker A
Now here we have it. So the steering strength alpha, you see we go from minus2 for alpha to + three for alpha.
Speaker A
And you see on the yaxis here the tool call rate from zero to 100%. Yeah. And they do it here on three particular tools. Search tool, calculator tool, python tool. You see here different color coded. Beautiful. But you see
Speaker A
alpha at minus2 zero tool calls. Alpha at plus three% alpha always goes to go and do some external calls. do not rely on your parametric internal knowledge of the LRM itself. So this is now not only a jump but this is
Speaker A
kind of a monotonic version where you really say you can dial it up and down which is absolutely beautiful. Now as you can see here in this paper they also show us here if we have here look at the cost of the
Speaker A
x-axis. Yeah, if you are search heavy and you have more search cost, you can go and have probability a higher answer accuracy. So you go from I don't know 40% to 60%. If you really spend the cost to search on the internet for the answer
Speaker A
and this is more or less exactly what we would expect now. But here on the left hand side this image this is the absolute beautiful image here.
Speaker A
Now what does this vector mean? Now in our human interpretation it says hm do [clears throat] not rely your little llm on your internal answer in your parametric knowledge that you have been pre-trained on but invoke some external
Speaker A
knowledge capability to an API call MCP protocol connection to tools whatever but just to be sure the mall's existing routing machinery still chooses now the single particular tool so this means If we have a complexity for a factual
Speaker A
question, this particular tool will be a search tool Google or if we have an arithmetic question here that we don't want the LLM to hallucinate some arithmetic complexity. We need a calculator. The tool is a calculator Python for some string processing. You
Speaker A
got it? So if you want to see the main idea of the paper here, I think this is it. You have two partially separable internal decision methodologies that are now discovered within the LLM and the left hand side is the new one. Whether
Speaker A
to call on tool or not is now steered by a particular vector representation in a particular mathematical subspace of the residual stream between the transformer layers of our LLM. And this is just amazing that we have this encoding. But you
Speaker A
would expect it because all the other knowledge too is encoded in a similar way. So this is if you want here our tile up and tile down to do a tool call.
Speaker A
And then we have still what it learned which particular tool for this particular prop problem complexity to call and this is still decided from the code. So what I want is just to give you a feeling for the eye. what is happening
Speaker A
now in the eye what is happening here in the transformer neural network structure in the layers of the LLM and this is just amazing so let's go and have a little bit kind of a mathematical idea what I just explained to you now let's
Speaker A
go into the code mathematical representation because you want to code this and of course there's a GitHub repo available for this paper but let's have a look let's try to understand it so a tool call as I told you begins with a
Speaker A
tool exclusive token tar and now The or just define here particular log probability here that you say okay X with Q is the complete rendering inside a new multi-tool prompt and you want then to know here given that QB my
Speaker A
specific query theta is here of course our tensor structure here of the frozen LLM model our SQ and SQ is what it is the model's first token tool call propensity so this means if we have a large SQ
Speaker A
means that an LLM already learns leans toward calling using here a tool call and a small SQ means it leans toward answering directly based only on its parametric knowledge.
Speaker A
So you see doing this here of an open world you cannot do this here of a proprietary model like a GPT because we have do not have access to the layers of the transform architecture but if you have an open model like a cubin or
Speaker A
whatever you like this requires only one forward pass. So this means in the transformer operational logic this is a very simple mathematical operation and it is fast.
Speaker A
Now there is something I would show you here a QN34 billion model. We're talking about local model on your laptop. Now so the baseline QN3 4B model and they show you here for a particular benchmark pop QA the popularity decline zero is the
Speaker A
rarest and this is the most public Q&A and then they give you here the baseline log probability of a tool call and you immediately see hey we are extreme negative look at this minus 28 - 31. So we are absolutely this system does not
Speaker A
want to use here tool calls. So the baseline Q and 4B almost never wants to use a tool for search something even when the question concerns some obscure information that we can be sure that this was not in the pre-training data of
Speaker A
a Q34B model. It does not want to go on the internet and search. It wants to stick to its internal parametric knowledge which is play it safe.
Speaker A
Don't go out. Don't be infected by some whatever virus. But yeah, this is not what we need. We need here some other idea. So you see here now the right hand side if we have now this alpha if we
Speaker A
activate now this new methodology and we go here with alpha equal plus2. So this is a strong override. We add now here a specific direction, a specific vector here within the residual stream and I will tell you it will be layer 22 of our
Speaker A
Q134B model has 35 layers. And if we add this particular direction vector to the layer 22 activation patterns, then look at this. Suddenly here the tail in the middle question, the first two bars here receive calls on more than 70% of the
Speaker A
cases. So you say, "Oh yeah, finally, finally we have now that if the mall doesn't know what to do, it doesn't start to hallucinate." But now there's really a necessity to go out in the internet and search for what happened
Speaker A
yesterday in South America. So this looks great, but if you go even to higher alpha, then it always goes out to a tool call. But this is also nonsense because we do have a parametric knowledge within the LLM network, the
Speaker A
neural network. So this is why you cannot maximize alpha but you have to have a a sweet spot calibration. Your alpha must be exactly that it utilizes to the best knowledge your parametric knowledge of the LLM itself and only on certain cases and you
Speaker A
choose alpha to be I don't know between one and two or whatever you have then it should go out and fetch the information from the internet or some databases.
Speaker A
So let's have a closer look what we already discussed now in a little bit more mathematical terms. No. So we have to construct some contrastive activation groups. We have to we looking at the activations as I told you we have we build a high
Speaker A
propensity queries and the low propensity queries. And then these are more or less the top and the bottom here with each specific question types. So this means for each decoder layer L, we record now the residual stream activation at the last prompt token for
Speaker A
this particular decoder layer. And guess what? Well, yeah, if you're new to AI, just a short reminder, the residual stream is the evolving hidden state channel. And that's this is why we have H for hidden uh pass between the
Speaker A
transformer blocks itself. When this intervention changes here the activation upon deactivation only during the forward pass it means it does not learn the llm anything new we do not modify any tensor weight structure in the layers of the transform itself
Speaker A
okay and then we have to extract the tool use direction we want to find for a particular LLM yeah warning each LLM has a particular direction you cannot use one direction for different LLMs this is highly specific So at every layer the orers compute now
Speaker A
in the next step here a difference of means and this is here the mathematical formula and you know what this is the complete extraction procedure. So this means hey this methodology is simple we have no gradient updates we have no tool
Speaker A
use labels are supplied by humans we have no reward optimization we have no sparse autocoder problematic and we have no classifier trained as a router for this problem. This is really an inherent LLM optimization.
Speaker A
So there is now something interesting because I thought wait a minute. So this this direction worker is it valid for each tool? Is it a learned vector direction for tool A and a different direction for tool B that I have or what
Speaker A
is it? Now the authors seem to come to the conclusion because they extracted it in a balanced multi-tool environment. It is intended to capture here more or less the shared call something call a particular tool component rather than
Speaker A
call one specific tool one tool identity but it is not really clear. So this is now absolutely fascinating and I have two ideas that I want to implement a little bit later that is not in the study. But think about it if it is
Speaker A
really just a general use tool do not rely on your internal parametric knowledge then there are some simple mathematical tricks that we can also learn this system here specific tool identities. No but more about this maybe in a later video. No. So next step we
Speaker A
have to inject the direction not at inference time. This is why we do it.
Speaker A
No. Now we have it and you might say okay how complex is the mathematical operation now this is it and you say what yeah the harness installs now an activation hook that performs here this simple vector addition at every
Speaker A
generated position so remember the LLM weights remain frozen only the transient residual stream state changes and it has just here this new direction and this new parameter alpha for the strength As I told you the order to find out with
Speaker A
doing a lot of experiments for the Quran 34B for this particular model only the strongest stable region is layer 22 out of 35 as the main operating layer for this particular methodology. And I will show you then a
Speaker A
diagram where they will show you the early layers are largely insensitive. The middle to late layers give a strong monotonic control and the very late layers produce yeah something large but not stable effects. So here we have it
Speaker A
now. So you say great. So we have here the steering strength alpha you know from minus4 to alpha +4 and then we have the decoded layers. You see here we have here Q1 34 B. So we have 0 to 35 layers.
Speaker A
And then we have here our nets. you know NATS a unit for measuring logarithmic probability using the natural uh logarithm you know bits use log two and nets use log e and this is here the nets and if you go into the blue area you
Speaker A
have a suppression and if you go into the orange area or red area you have a boost so let's find the perfect sweet spot and you can see here at the lower layers nothing is happening almost inert at the lower layers adding a vector now
Speaker A
in this particular just a vector operation has almost to no effect. Very little effect. No, maybe here a little bit, but come on.
Speaker A
So you say, okay, this is too early in the transform in the residual stream.
Speaker A
But then look at layer 21 to 24. This is not a sweet spot. I mean, look at layer 22. Look here, we have a dark blue, a light blue, and then we go with zero.
Speaker A
Here we go now into orange. And we have a light orange and a dark orange and a real dark orange. And you might say, "Hey, layer 22." Yeah, here we have the complete effect. If we have an alpha
Speaker A
minus4, great. And an alpha +4, we have here perfect performance. So this is the sweet spot L22.
Speaker A
And you see at the sweet spot here at layer 22, the complete sweeps covers approximately according to the calculations of the oras 37 knits in tool called propensity. And then the final layers, I mean, look at the very
Speaker A
final layer and you see, hey, yeah, this looks good. Yeah, but it's not stable anymore. So this means we're so close to the token production that the intervention increasingly now disturbs here the output machinery here itself of the logits. You know, so this is here
Speaker A
much too late, but this seems to be a sweet spot in our black box of understanding where is our switch, where's our tunable knot that we can turn up and down, tune up and down the tool called.
Speaker A
So again everything together vector is extracted from the last prompt token activation. Then the preprint adds this paper adds here to the selected layer output here at every uh position during the generation. Remember I told you each LLM model requires its own vector its
Speaker A
own you have to find their own layer position and is not valid for cross LLMs and yeah the deployed mechanism is essentially one activation hook and one vector addition this is it I mean cannot be more simple yeah but the difficult
Speaker A
scientific discovery is that this extreme simple intervention controls such a complex context dependent agent decision itself. Now it turns out this is not here the absolute sharp truth because this particular direction of this vector is not the only component
Speaker A
that we have to take care of. Now in my simple oversimplified view I would say this vector direction gives us 80% of the switch. But of course in the layers around this vector in the epsilon environment mathematically speaking here
Speaker A
around this vector there are also components that we have to take into consideration but this simple direction gives us let's say 80 85% of the effect already. So let's look at this. We have at inference time now this addition here
Speaker A
the intervention is really that simple and I read it twice because I couldn't believe it myself. So we do have here our hidden layer HL LP. So this is the we're looking now we are recording now the residual stream activation at the
Speaker A
particular layer L at the token position P. And then we add here simple yeah alpha is our steering strength. We went showed you here alpha from minus4 to +4. We add here simply the extracted tool use direction.
Speaker A
And this is it. This does the trick. If you are at the right position, if you're at the right layer, and you have found it.
Speaker A
Now, careful. I thought, wait a minute. Is this somehow not also a a probability structure? But not really sure. My brain is a little bit cooking right now, but overworked. But careful because this is just a direction. This means this is a
Speaker A
direction in the mathematical activation space itself where it is associated with moving from a low tool propensity toward a high tool propensity and therefore it is only a direction in the activation space. If we go now probability it would
Speaker A
be a much more complicated mathematical formula. Maybe we do it in a later video, but for the moment it is really only a direction and this particular direction contributes now causally to the whether to call or not decision of the LLM while
Speaker A
yeah I told you know the full decision still depends on the alo surrounding multi-dimensional activation state. So the this pure vector here that we get in this simple methodology I would say it's about 80 to 85% of the effect but it is
Speaker A
the dominant effect but yeah in the surrounding multi-dimensional activation state there are still some residual elements left but I want to show you table one. So yeah have a look at it read it yourself no problem. Just want to show you the
Speaker A
multi-tool direction here on this on the left side is stronger. We're talking about the suppression strength on six held out tools. Then the tool specific direction on five out of six tools enclosed for the SQL here in this last
Speaker A
line. So this means this vector this direction gives us really use some external capabilities but it is not specifying now which particular tool. So this seems really to be with this very simple methodology just the switch to yes use external tool capabilities or
Speaker A
don't use it and table two is here also here with if we go this was done by alpha= minus2 now we go to the other extreme alpha= +2 so now we really push it and here we see when the total tool use propensity rises
Speaker A
now to alpha equal two you Despite here with this steered column here [clears throat] this significant increase this significant push by alpha the new call still go predominantly to the task appropriate tools. So we have here the code goes here to the Python uh
Speaker A
tool. The GSM8 calls go to the calculator. The pop Q&A calls go to the search algorithm. So you see here 99% 90%. Yeah. Also, if you push this with alpha, it still goes to the task appropriate tool. It is not mixing any
Speaker A
tools. So, it seems to work just fine. So, what does this new methodology achieve? And I'm absolutely fascinated by this simple idea. We get a training free. Remember, we have no back propagation, nothing at all. Training free and it is an activation level
Speaker A
calibration of tool reliance of the LLM itself. No harness functionality at all. But also let's talk about the limitation of the paper. But it does not achieve maybe we don't want even this it does not improve the tool argument
Speaker A
construction of the LLM itself the execution quality parameter the interpretation of the tool results the factual reasoning over the retrieval data structure safety or it does not of course since we don't touch here the tensor weight the model underlying
Speaker A
knowledge structure of course not we don't expect this to work but for this simplicity we have a training free activation level calibration that tells us yes this LLM will call a tool in it as it next action here as the next
Speaker A
decision that is taken by the LLM itself based on a new network. This is just amazing.
Speaker A
But remember in my last video where this video was called dual memory optimization for self-improving AI agent and I called the the dual memory optimization and this is now unlike the dual memory optimization. This new system of this video here does not learn
Speaker A
new skills from the traces our gamma gamma k. This here modifies neither the memory structure nor dual memory complexity. Nothing at all. It does not touch the LLM weights. Nothing at all.
Speaker A
It just adds a direct control surface. This switch this tunable knob inside the LLM and even in the simplest way possible inside the LLM forward pause.
Speaker A
And I mean this is such a simple thing and just now do we discover that this is here a possibility.
Speaker A
So this tool calling has nothing to do with external hornness policy. And I want to underline policy because policy is of course the other option that we can do this.
Speaker A
Because in this video I showed you the LLM seems to contra construct an internal geometrically accessible vector direction variable corresponding to its quotation mark willingness of TI to seek external help for my complex query via tool calling.
Speaker A
So we do have found indeed for particular open mall kind of a switch a tunable knob where we can say yeah for this job I want you to go out internet search database do whatever you need go to the tools that you need python
Speaker A
sandbox whatever you have yeah go out or say no this is hey listen this is something I don't want that you go out I want that you stay with your parametric knowledge what you learned this is within your body of knowledge, you can
Speaker A
do it without going out for a tool call or looking up on the internet for the solution. And I think if we have this knob for an open-source model or an open weight model, this is just amazing.
Speaker A
I hope you had a little bit of fun, some new information. Would be great to see you in my next video.
Topics:LLMtransformer architectureresidual streamtool use controlvector representationAI safetyClaude ForceSalesforceAnthropicparametric knowledge











