**Kimi K3 + GLM-5.3: Self-Improvement (RSI) Unlocked — Transcript & Summary | SozAI**
Source: https://sozai.app/transcript/kimi-k3-glm-53-self-improvement-rsi-unlocked/

Explore the RSI agent combining Kimi K3 and GLM-5.3 for recursive self-improvement with verified memory, outperforming proprietary models.

## Key Takeaways

- RSI agent leverages frozen LLMs with optimized memory for domain-specific adaptation.
- Multi-agent system enables continuous learning and verification in new environments.
- Open-weight models like GLM 5.3 can outperform larger proprietary models through recursive self-improvement.
- Memory optimization is key to autonomous, nonparametric adaptation without retraining the LLM.
- The approach separates general intelligence from environment-specific knowledge for better specialization.

## What the video covers

- Introduction to RSI agent: a recursive self-improving AI with a frozen LLM backbone and verified memory.
- Combination of open-weight models Kimi K3 and GLM 5.3 to outperform proprietary models like GBD6 Astra and Fable 5.1.
- Focus on memory optimization rather than changing the LLM tensor weights.
- RSI agent uses a multi-agent framework with curriculum, actor, and verify agents for autonomous environment exploration.
- The agent specializes in domain-specific knowledge by learning causal relationships through experiments in unfamiliar environments.
- Separation of general reasoning capabilities from environment-specific operational knowledge to reduce agent failure.
- Memory is autonomously constructed and optimized through experimental science, enabling recursive self-improvement.
- The learning loop involves hypothesis generation, experimentation, observation, verification, and memory revision.
- Broad and deep exploration modes enhance environmental coverage and resolve uncertainties respectively.
- The RSI agent’s performance improves without tensor weight updates, highlighting the importance of memory and agent design.

## Chapters

1. 00:00 Introduction to RSI Agent and Memory Optimization
2. 01:10 Performance Comparison with Proprietary Models
3. 02:22 Memory Optimization and Frozen LLM Concept
4. 03:52 Scientific Understanding of RSI Agent Adaptation
5. 05:09 Separation of General Reasoning and Environment-Specific Knowledge
6. 06:28 Multi-Agent Learning Loop: Curriculum, Actor, and Verify Agents
7. 07:32 Knowledge Retention and Verification Process
8. 08:45 Autonomous Experimental Science and Memory Construction
9. 09:52 Memory Optimization and Recursive Self-Improvement
10. 11:00 Exploration Modes and Benchmark Performance

Answers

## Questions about this video

What is an RSI agent as described in this video?

An RSI agent is a recursive self-improving AI that uses a frozen large language model backbone combined with a verified memory system to autonomously learn and adapt to new environments without retraining the model weights.

How does the RSI agent outperform proprietary models like Fable 5.1 or Astra?

It outperforms these models by optimizing an external memory component through a multi-agent framework that experiments, verifies, and retains environment-specific knowledge, enabling specialization without changing the frozen LLM weights.

What roles do the curriculum, actor, and verify agents play in the RSI framework?

The curriculum agent generates tasks to explore uncertainties, the actor agent performs experiments and retains lessons learned, and the verify agent independently examines outcomes to ensure knowledge validity before updating memory.

## Full Transcript — Download SRT & Markdown

00:01

Speaker A

Hello community, welcome back. Today we talk about an RSI agent. Yes, a recursive self-improving agent. We have a frozen LLM in the backbone where it has the central intelligence. And we're going to work with a verified memory.

00:16

Speaker A

And you know what is special about it? We will combine a Kim K3 and a GLM 5.3.

00:21

Speaker A

So open weight model to really outperform a GBD6 Astro model. We just need a little bit of a recursive self-improving methodology, a new methodology on purely memory optimization. So let's have a look. This is our paper here. This is by University

00:38

Speaker A

of California, San Diego and University of Illinois, Chicago. And those authors, as you can see here, are also working here in the startup. Do we support startups? Absolutely. Therefore, I will show you this and the study is called

00:52

Speaker A

RSI agent autonomous exploration for recursive self-improvement in new environments, published September 14, 2026. So, let's have a look just to make you a little bit more interested in this. Look at this. We have really that on particular benchmarks here like

01:10

Speaker A

agents last exam. We have now, we build now, and I will show you how to do it, an RSI agent that is really outperforming a GPD6 Astra Claude Opus 5 GBD 5.6, beautiful, and we do this with open

01:23

Speaker A

weight model and you know this is what I like to show you on this channel. So let's start. What is the general problem that we're going to solve? Oh, you know, digital agents, we applied it here and those agents have to adapt to new

01:37

Speaker A

environments. Let's say finance, physics, chemistry, whatever. Now those interfaces, the tools, and all the failure modes are not fully captured here by a classical pre-trained LLM like, I don't know, an Astro model or a sole model by OpenAI.

01:53

Speaker A

Of course not, those are general models here, and if you have a specific environment, finance, medical, whatever, you have to have some specialization. So let's look at this. The authors tell us, hey, we introduce now a new RSI agent

02:09

Speaker A

here and this is a training-free multi-agent framework, multi-agent for recursive self-improvement of the AI agent through an autonomous memory construction.

02:22

Speaker A

Now you won't say, okay, so the memory since we have frozen weight is in the harness and nothing else, nothing else. We will just optimize here the memory. All the other elements in the harness will stay the same, so this is nice. Let's have

02:37

Speaker A

a look at this RSI, the recursive self-improvement. What we have, this agent coordinates now three other agents: the curriculum agents, the actor agent, and the verify agents to continuously explore this new environment, this new physics, finance, medical, biochemical

02:55

Speaker A

environment and retain some new environment-specific knowledge that this agent learned interacting with this specific environment. And this knowledge will include reusable causal relationships between all the actions, all the conditions, and all the consequences as observed observation. So

03:16

Speaker A

great. What is our storyline? And the storyline is simple: that an agent enters an unfamiliar environment, let's say physics. It has to invent some experiments and it will learn some causal rules of this environment doing here this experiment in this environment

03:35

Speaker A

and doing so with time it will surpass here the stronger proprietary models like fable 5.1 or Astra. So this is it: specialization in particular domains with particular knowledge. You have to train those models on these particular

03:52

Speaker A

domains and then they outperform here the basic proprietary models. Now you can have, of course, a scientific understanding and the scientific understanding is that RSI agent performs here autonomous target condition nonparametric adaptation. This means the LLM tensor weights remain frozen and the

04:13

Speaker A

learning occurs through an evolving external memory constructed from verified interaction with this particular domain environment.

04:22

Speaker A

So let's look at the difference. Why is this happening? Why suddenly we have an open weight model and we do a little bit of mirror optimization and we outperform here the big players? Well, a strong model here like a fable 5.1 may

04:35

Speaker A

understand Python, beautiful, have a perfect planning, perfect geometry, perfect software interfaces in general. I don't know how many trillion free trainable parameters, but it is lacking here the local knowledge required here for a very particular environment. Let's say you have a clinical and medical

04:52

Speaker A

environment. No, so it lacks here which application APIs actually work, which action changes which state variables, which export procedure preserves here an edible artifact, which apparent success is rejected by the real evaluator in the environment, which constraints appear

05:09

Speaker A

only in unusual cases. So what it does, take a step back, think about it. It separates a general reasoning capability from this environment-specific operational knowledge. So therefore our new RSI agent assumes that much of the agent failure comes from the second category,

05:29

Speaker A

this environment-specific knowledge, this interlink. So therefore instead of changing now the foundation model from fable 5.0 to fable 5.1, it lets now the system experimentally construct a local model of the environment in an open weight LLM that is free to download. What a

05:50

Speaker A

nice idea. So let's do this. Now of course you say, hey, where is my mathematical equation? Here it is. The basic execution equation is conventional. There's nothing specific.

06:03

Speaker A

We just optimize a memory component in a harness. So this is all there is. So your specific action is done with a frozen policy pi data. But we act, we have now something that is here the query. Then h is our history at a

06:17

Speaker A

particular time step t and our persistent memory, of course, m. And then, yeah, we want to come up with a new action that is specific for this domain environment.

06:28

Speaker A

So if you want, the new methodology divides now with our three agents here this learning loop. Yes, guess what, it's looping among now three particular roles. At first, as I showed you in the image, a curriculum agent. This agent decides what

06:43

Speaker A

uncertainty should be investigated next. It is in an unknown or partially known environment, so there are open questions and the agent, their curriculum agent, is now generating practice tasks that will expose post the missing prerequisite, the difficult variance, the possible

07:00

Speaker A

failure condition or unknown rules that are valid in this particular medical environment or physical environment.

07:08

Speaker A

Then we have the actor. The actor performs, guess what, the task using now executable Python patch programs, whatever you have, owns now the durable memory and decides what lessons to retain. What have we learned from the interaction doing now

07:23

Speaker A

our experiments in this environment? We have some insights, some analysis of the observation. So what have we learned?

07:32

Speaker A

What knowledge should we retain? What knowledge should we put into our persistent memory? And then we need a verify agent. The verify agent examines the produced artifact and observed environment state, does not see the actor as private reasoning of the memory producing here

07:48

Speaker A

the probability that both agents simply repeat the same unsupported explanation. This is where you need a verifier agent that is not within the curriculum agent.

07:59

Speaker A

So we have a simple loop. We have a hypothesis. This is in the new environment that I want to test. We do an experiment for this hypothesis. We have with a particular action that we set in the environment an environmental outcome

08:15

Speaker A

that we observe. This observation goes into a verification agent. The verification agent says, hmm, I learned something or this is what I think this could mean. So we do have now a new knowledge learned. This new knowledge learned will go in the memory. This

08:32

Speaker A

means we have a memory revision and from this new memory, guess what, the machine can come up with a new hypothesis, recursive self-improvement.

08:45

Speaker A

So it means memory now is constructed or optimized or modified through a form of an autonomous experimental science.

08:55

Speaker A

Yeah. Yeah. Yeah. Rather than summarizing just some trajectories, we have an agent that is really doing some experiments in this environment to explore the environment to ex...

09:07

Speaker A

Let's think the environment is a new planet. You have no idea what's happening over there. So you do experiments workflow.

09:15

Speaker A

If you're rather processoriented here we go. So we start a frozen open source mall here on Kimmy or GLM 5.3 enters an unfamiliar software environment. Its general intelligence is insufficient because it lacks the local causal knowledge that is valid only in

09:32

Speaker A

this particular environment. So therefore it invents a curriculum of experiments. Parallel exploration maps now the environment focused experiment expose hidden constraints and independent verifier aent determines what actually happened and interprets the results. The agent converts those outcomes into insights into knowledge

09:52

Speaker A

into stored during the memory optimization. The memory of the LLM remains frozen. So this means we only optimize here the memory here in the harness. And then the same m re-enters now the environment with lessons learned with a different

10:09

Speaker A

more effective policy what to do in this particular environment. And then we ask whether this deserves to be called recursive self-improvement. And I will have a very particular view at the very end of this video and I will explain

10:22

Speaker A

why. So what is specific and beautiful and special in this publication? The authors go with two time scales and I love it.

10:31

Speaker A

So let's have a look. We have at first a broad recursive self exploration. What it does it maps the environment in a very broad style. So the curriculum agent proposes several related project which execute of course in parallel. You

10:46

Speaker A

can launch multiple probes into the atmosphere of an unknown planet from the same memory snapshot that in the memory you have the understanding what is going on what is the world model of this new planet. So this project explore

11:00

Speaker A

different operation formats for case workflow but the results are verified independently of course and memory updates are then applied sequentially so that the contradictory lessons can be reconciled before the next exploration wave a broad recursive self exploration.

11:17

Speaker A

Then the second stage is deep we have now understood to the broad scale now we have to have a deep dive into one particular segment. So a deep recursive self exploration and now we have chosen a particular target a particular

11:33

Speaker A

mountain range on this planet or a particular location of a lake or an ocean or whatever.

11:40

Speaker A

So this means the eye system attempts. Now this target with its knowledge identifies the remaining bottleneck because doing the experiment we still find some inconsistencies or some new observation generates now a focus practice task updates the memory based

11:57

Speaker A

on the new observation from the environment and then attempts the target again. So you see a recursive cycle.

12:05

Speaker A

If you want to see this here in table we have the broad exploration mode the deep exploration mode scientific function of the broad is increased environmental coverage in total and the deepest then resolve uncertainty around a specific target where the we hope that the LLM is

12:22

Speaker A

intelligent enough to identify a important representative target and then examine this with some particular experiments.

12:33

Speaker A

The orus call this here an aentic causal discovery process. Beautiful. So what does it mean? The llam tenzo weights are frozen. We do not optimize here the core of the agent to llm. We go here with an external memory structure here and here

12:51

Speaker A

in this memory structure we have now the construction of experiment on a particular planet on a particular software environment. You got the idea.

13:01

Speaker A

This experiments are coded are executed the results are observed and we have a verification. Hey, does this provide new knowledge? Does it change our view of this world model? Does this contain a component of a new world update? The

13:16

Speaker A

verify if yes comes back into this hornness memory. If not, throw it away. Wait a causal discovery. But careful the phrase is directionally reasonable but scientifically it is a little bit too general. No because why? The system performs intervention and records the

13:36

Speaker A

action condition outcome relationship. Beautiful but it is more causal than passive trajectory summarization. However, it does not provide the following facts. It does not provide an explicit causal graph representation. It does not provide a formal intervention variable a set of variables. It does not

13:55

Speaker A

identify assumption. It does not control for co confounding. It does not some counterfactual estimation. And you got the idea.

14:05

Speaker A

Or if you want to have it in very simple words, we have an interactive causal hypothesis formation with some empirical checking through some experiments that are run by our agent. So this means we have here if you want a discovery of

14:20

Speaker A

operational cause and effect rules in this new environment which is great. I love it. Yeah. If you want to see this here on an example here of a cat system. Here we have the broad recursive self exploration and the deep

14:34

Speaker A

self exploration and memory use. Here you go. So you see the RSI agent converts the interaction with an unfamiliar environment into a durable procedural knowledge that we know and then tests whether the same frozen model performs better when that particular

14:53

Speaker A

procedural knowledge just learned is now reused in the next experiment. It's a loop. Let's go to the numerical result and you see here for a particular benchmark OS world 2.0 Hero h you say what is this?

15:09

Speaker A

So we have spreadsheet repair audio editing web representation here and some segmentation. So why suddenly we have here this differentiation. So let's look what it is. The first one here the gray one is without RSI. This is no

15:23

Speaker A

exploration and no persistent memory. Then the orange one is without BRS. This is deep exploration only. The next one here in blue is the broad exploration only. And then the in green the full RSI is then brought followed by the deep

15:39

Speaker A

exploration. Now you see interesting for different task you have a different performance indicator but in general if you do the green stuff the full RSI you have more or less the best performance but you also have a detail where the

15:54

Speaker A

deeper result of this graph you can see that if you do the broad exploration this contributes in general more than a deep exploration alone. Of course, if you have a general understanding of a new exploration of a new planet, it is

16:13

Speaker A

better to have a broad understanding compared to that you just have one or two detailed deep explorations.

16:21

Speaker A

So, this provides some evidence of a complimentary search scale. No, the broad exploration learns the environment's vocabulary, the general rules, not the general atmospheric dynamics. You got it. And the deep exploration learns the exception and the boundary condition that are needed for

16:38

Speaker A

the specific target chosen by the LLM. Now of course here you have it now for the different models. So at first we start with the open source model. Then we have the closed source model. You have your GPD 5.6 soul, your clo fable

16:51

Speaker A

5, your GPD Astra. And you see this new RSI agent with opensource model outperforms all the other stuff. Great.

17:01

Speaker A

Let's have a closer look. If they say RSI agent and they are not really precise here in this table. So let me be clear. This RSI agent has a clear configuration. We have here a GLM 5.3 open weight M as the actor agent and the

17:17

Speaker A

other two agents are guess what a KI K3. Now the auto fire this combination of a GLM 5.3 with a Kim K3 provides here the best benchmark results. But you also see a line before the last line that is here

17:31

Speaker A

without RSI. So this means they use here exactly the same GLM and Kim K3 act verify a harness configuration but in this particular case they disable the autonomous exploration and the persistent memory. So absolute interesting to see the performance of

17:48

Speaker A

the different model and the outperformance here of this combination of open weight models. So therefore if you want to be absolute precise what you see here that a multi-agent opensource system here our RSI agent the very last line in this table combining now the GLM

18:05

Speaker A

5.3 agent and the Kim K3 double agent surpasses now a pure GPD6 Astra propriatory model that we have no access to and that we have to pay to every moment we use it on a reported partial credit here the very first column on

18:22

Speaker A

this particular benchmark of OS as world 2.0 great what are the inside here the model AI structure becomes substantially better without a tensor weight update so it is not the model it is the agent but what really is

18:41

Speaker A

important you see but what evolves is its experimental constructed memory only the memory is optimized not yet the architecture or architecture that performs performs this improving. So you see you don't have to modify the harness here the components multicomponents

19:00

Speaker A

changes. No, if you just perform here or optimize the memory representation with this simple methodology, you can outperform here a proprietary model like GP6 Astra.

19:14

Speaker A

Of course, as always, we do have limitations. So when you go with a strong interpretation of recursive self-improvement, it would require in my notation that the system improvement mechanism to be part of what the system can modify. But you know that RSI agent

19:32

Speaker A

here as presented in this particular study does not generally rewrite the architecture, the active verify topology, the verification procedure, the exploration protocol, the broad then deep control structure or the harness code that manages [clears throat] here the memory update. Now so we have really

19:49

Speaker A

an absolute if you want rock solid architecture that we do not modify. We just have an update here of what is learned of this agent interacting now in this new environment. And this is why I think this is such a beautiful simple

20:05

Speaker A

study with such beautiful results. So is here the RSI agent here the strong RSI in real scientific terminology? H maybe not. I think RSI a join is therefore best classified here in a more scientific precise way as a bounded

20:23

Speaker A

anchored self-improvement through a recursive accumulated memory representation doing here experiment in this particular domain environment it is its recussion is in the feedback process no new memory changes the next experiment which changes the next memory you see some simple sort some simple

20:47

Speaker A

ideas how to optimize here the performance of AI agents and therefore done with open weight models. I just love it. Have a look at this paper. I hope to see you in my next

Topics: RSI agent recursive self-improvement Kimi K3 GLM 5.3 frozen LLM memory optimization multi-agent framework autonomous learning domain specialization open weight model


---
This is the markdown twin of https://sozai.app/transcript/kimi-k3-glm-53-self-improvement-rsi-unlocked/ — the same content, without the markup.
Published by SozAI (https://sozai.app). Reuse and quotation are allowed with attribution and a link back.
Machine-readable index: https://sozai.app/llms.txt · data API: https://sozai.app/api/
