Skip to content

The Death of Monolithic AI: The Scaffold & Harness Paradigm

Explores the shift from monolithic AI to scaffold and harness paradigms, analyzing recent research on evolving AI architectures and memory systems.

Ask about this video. Answers come from its transcript only — with the timestamp, so you can check them.

Generated from the transcript and can be wrong — check the timestamp.

Key Takeaways

  • AI is evolving from monolithic large language models to scaffolded architectures with harness elements for better optimization.
  • Structured representations like knowledge graphs and skill memories are crucial for efficient AI reasoning and learning.
  • Theoretical physics concepts provide useful metaphors and mathematical tools to understand AI learning processes.
  • Balancing exploration and exploitation via gated mechanisms is key to self-evolving AI systems.
  • Non-parametric scaffolds and architectural topology optimization are emerging as important directions in AI research.

What the video covers

  • The video discusses a new paradigm in AI moving away from monolithic models towards scaffold and harness structures that optimize AI architectures.
  • It highlights recent scientific papers from 2026 focusing on neuro-symbolic frameworks, self-evolving skill agents, and ontology-aware multi-agent systems.
  • The presenter connects these papers to concepts like belief states, knowledge graphs, and memory hierarchies in transformer architectures.
  • Mathematical analogies from theoretical physics, such as wave function collapse and phase transitions, are used to explain AI learning dynamics.
  • The importance of structured knowledge graphs and skill banks as computational lattices for efficient reasoning is emphasized.
  • The video explores how AI systems balance exploration and exploitation through gated mechanisms and reward distributions.
  • It introduces the idea of a harness system as an architectural scaffold that guides AI learning and tool acquisition.
  • The presenter shares insights on how these evolving AI systems use non-parametric scaffolds to improve reliability and performance.
  • The video includes a discussion on numerical simulations to understand AI optimization and architectural topology.
  • Overall, it offers a theoretical and practical perspective on the future of AI system design beyond traditional transformer models.

Answers

Questions about this video

What is the main theme of the video?

The video explores the transition from monolithic AI models to scaffold and harness paradigms that incorporate structured memory and evolving architectures for improved AI performance.

How does the video relate AI concepts to theoretical physics?

It uses metaphors like wave function collapse and phase transitions to explain how AI systems condense unstructured data into structured knowledge and optimize learning through critical boundaries.

What role do knowledge graphs and skill memories play in the discussed AI systems?

Knowledge graphs and skill memories act as computational lattices or scaffolds that structure AI reasoning, enabling efficient memory updates and dynamic balancing of exploration and exploitation.

Full Transcript — Download SRT & Markdown

00:01
Speaker A
[laughter] Hello, community. So great that you are back. I promise you today will be really a strange video day. So, let's have a look.
00:10
Speaker A
Can you imagine what I want to show you with this image? Think about it.
00:16
Speaker A
If not, I help you a little bit. So, I think at the core we have our LLM, beautiful, great, and we have here a harness structure with all our harness elements. And today I had a special idea because I thought, you know what? We
00:29
Speaker A
are operating here on mathematical optimization methodologies that have a direct effect here on the scaffolding itself that we built.
00:37
Speaker A
Think about a geometry. Think about a reasoning space, a complexity that you have in an N-dimensional space.
00:43
Speaker A
And this geometry, what do we do? We build a scaffold on this geometry. We build a little computational lattice structure so that we are able to have a numerical calculation within our AI mathematical operations.
00:57
Speaker A
So, the first paper that I wanted to show you today, I'm going to have a little peek behind the curtain how I operate here on my channel. You know, this is here a paper that, reading here several hundred papers here, I thought,
01:08
Speaker A
hey, this is nice, no? July 31st, 2026, here Georgia Institute of Technology. This is here a neuro-symbolic fast-slow thinking framework for LLM agents under Markov, and I said, this is great. Look, they have here a knowledge graph to
01:25
Speaker A
present the belief state. And I thought, this is a video that has two elements from my last videos.
01:31
Speaker A
Because in one of my last videos, when we talked about MEET as the new transformer architecture, I showed you that we have a local block memory in the transformer layer, but we also have a hyper memory block here.
01:42
Speaker A
And I think this is here a beautiful mathematical coherence if we go to this fast and slow thinking framework here on a memory level.
01:52
Speaker A
But you know, the second optimum would be this is here a real representation of the belief state, and they really take this belief state from the video where I showed you here AI is learning whom to watch and how to influence here with the
02:05
Speaker A
theory of mind and the cost approximation. Here we have a real example that the belief state is now materializing as a simple knowledge graph. And I saw this is great, so I can now connect three videos and you have
02:19
Speaker A
here really a real-world example also July 31st. Unfortunately, I continue to read and this is here another paper here. This is University of Chinese Academy of Science and Institute of Automation, Chinese Academy of Science and Peking University and Tsinghua University and a lot of
02:34
Speaker A
other beautiful highly intelligent people, and they talk about here a self-play skill evolution. So, they go for skill memories and self-evolving AI system also July 31st, 2026. And I saw this is absolutely beautiful because this is here a skill
02:53
Speaker A
augmented agent which makes use of procedural memory and evolving state of a tool-augmented search self-play system.
03:02
Speaker A
This is something I would love to show you. And I continue to read and I said, "Wow, look at this article." This is here from Zhejiang University and they have here also July 31st, 2026 here an ontology-aware self-evolving agent for
03:18
Speaker A
and now the most important word, a scientific tool acquisition for our multi-agent system. And I said, "Look at this." This is about ontologies, so this is driven by an evolving memory of skills, of experiences, and an ontologized tool
03:35
Speaker A
graph structure. And it distills here some generalizable knowledge from contrastive trajectories during the accumulation, while during the inference it formulates some active requests and utilizes here a gate to dynamically balance here exploration exploitation. And I saw this is great.
03:52
Speaker A
This is a paper I would like to show you. And of course, if you continue reading, this would be from Georgia Institute of Technology and Kuaishou Technology of China, and this is about a self-evolving. Yeah, you have it in the title.
04:05
Speaker A
Abandoned routed agentic harness system, a new harness system where I thought, yeah, absolutely also July 31st, 2026.
04:13
Speaker A
This is the paper I would love to make a video for my audience. And then I thought, wait, wait a minute, there's another paper.
04:21
Speaker A
I don't have to mention that it's July 31st, 2026. This is here by Hong Kong University of Science and Technology and Bytedance, and they go here for an analytical memory here multimodal, but they go here also with a rec system.
04:34
Speaker A
So this is a framework that supports you the retrieval that we know from rec and the analytic memory building. And I thought, wow, this would be the paper for a video.
04:46
Speaker A
And then something [clears throat] happened. Then something happened because I thought, so which is it now?
04:51
Speaker A
And normally I would converge on one video and I would make you the video on one single paper.
04:57
Speaker A
But today is different. Today suddenly my brain activated itself and said, wait a minute.
05:04
Speaker A
Do you see the pattern? Now four of these five preprints have, let's call it, a scientific similarity, but this would be not precise and it is not a similarity.
05:16
Speaker A
It is a scientific pattern. And this is here the main sword of today that changed the day for me and I just want to bring you along with my journey.
05:27
Speaker A
So let's talk about my feelings. And as a theoretical physicist, you know when measuring your quantum system you must collapse your wave function to extract your physical observable, and with this feeling I started to think about the similarities of those papers.
05:41
Speaker A
Think about it. In a standard AI agent, a belief state probability distribution or they talked about in my last or my last two papers.
05:49
Speaker A
Think, what is it? It is a smeared wave function, no? Because we have unstructured text histories that contains some immense redundant entropy.
05:58
Speaker A
Those humans like to talk forever. My goodness, imagine a YouTube video with this, no?
06:03
Speaker A
But the first paper here forces here this wave function collapse in my idea, in my mind here, by condensing the raw text into a structured knowledge graph manifold.
06:16
Speaker A
Do you feel it? Do you understand the pattern? Do you see it? Raw text, just some blah blah blah, in a structured knowledge graph manifold.
06:25
Speaker A
So, if you want the first mathematical breakthrough of the first paper in their slow thinking module, is it updates here this particular manifold with a particular algorithm. This is a twisted sequential Monte Carlo simulation.
06:36
Speaker A
Beautiful, if you are interested, have a look at the first paper. It's really nice.
06:43
Speaker A
And then there was this second sort of in my brain, you know, phase transition.
06:47
Speaker A
We have a lot of phase transitions in theoretical physics, I mean, even in experimental physics, no? But you know that phase transitions occur only at critical boundaries, no?
06:55
Speaker A
Water, vapor, fluid, frozen ice. There's certain temperatures that are connected with this, no? Now, look at this paper. They talk about here bidirectional loop makes here the task generation and the skill memory co-evolve.
07:10
Speaker A
Am I said okay? Hmm, think about this. They use here an asymmetric self-play mechanism in their paper. If you read the paper, please do. And the agent then generates the task dynamically, but the challenge is reward space. What is the
07:26
Speaker A
main element where the AI learns is now mathematically forced into a strict bell-shaped boundary condition.
07:34
Speaker A
Bell-shaped probability distribution, what a coincidence, no? It looks again like a wave structure, no, that we have in theoretical physics, no?
07:44
Speaker A
But this is now, as I told you, the reward space. But this is a very specific feedback probability distribution that is igniting here a specific learning pattern.
07:57
Speaker A
Now, they have a simple formula, as you see here, beautiful, no? And the reward peaks at some intermediate success rate and penalizes both the trivial cases. So, this is not where you have to face transition, no?
08:09
Speaker A
If it's trivial, the AI says, "Hey, I noticed this is boring." If the problem is unsolvable for a little AI machine, then it doesn't make sense. The machine is not able to solve it if you give it 10,000 years. So, this is also not close
08:21
Speaker A
to a phase transition. So, what we are looking for is, can we come i
08:34
Speaker A
And you see, this is now exactly because the proposal continuously pushes you to pose problem just beyond the solver's current ability. It goes in baby steps above the current level that the AI can solve and says, "Listen, I'll give you,
08:49
Speaker A
I don't know, some time, no? 5 milliseconds, 100 milliseconds. Try to solve it anyway." This has a very nice integration of skill distillation, memory receiving, and yeah, I thought this is this is this is something it rang a bell in my brain, you know?
09:07
Speaker A
And then we have another problem in physics no? Imagine exploring here a new physical system with the dimensionality of your Hamiltonian basis vectors, let's say in AI this could be here some actions here on your tools that you're integrating
09:22
Speaker A
your multi-agent system, is constantly expanding. So, you're working with a Hamiltonian basis vector expansion here in space.
09:31
Speaker A
How does now the agent optimize a trade-off between exploiting a known basis set of tools? Let's talk about action tools, not just a coincidence.
09:40
Speaker A
And how does it know when to stop this and explore and start exploring unknown basis vector?
09:47
Speaker A
Because it has no idea maybe those vectors are helpful and provide a path to tools that are not yet discovered to be included here in the reasoning complexity.
09:58
Speaker A
And here the third paper here, scientific tool agent evolution, restricts now the infinite action space here by bifurcating it into a simple simple two elements, no?
10:09
Speaker A
At first again, we have a known tool graph and you see the classification here from a discrete or continuous medium into a graph structure is here something that is imminent today in research, no? But then we go with an
10:23
Speaker A
unknown set. Now you know that the mathematical definition of a set is anything else than simple or trivial, no?
10:32
Speaker A
But do you discover the pattern? Do you see the pattern? Can you feel it?
10:38
Speaker A
No? No problem. Let's continue. So we are here with this beautiful study side tool agent evolution here. And then they show us here these particular steps, no? And when exploration of an unknown action succeeds, there is here in the in the quasi code here in the
10:53
Speaker A
pseudo code here, step 35 or line 35. And this dictates the agent it must execute now a market subroutine that is permanently crystallizing the tools on topology into well, onto the known manifold. And this is happening here if
11:11
Speaker A
you see in the accumulation progress from 0 to 25. Look how the performance jumps from a low 55% to 83%.
11:20
Speaker A
This is here RGPD 5.4 or Gemini 3 flash. You see 28 to 75%. Whenever you change this and you discover new like Hamiltonian basis vector in your description of your physical system or here in computer science here available
11:37
Speaker A
action tools, you do have that something changes in the underlying space geometry that that is the basis of your calculation. Like here crystallizing the tools ontology into the unknown unknown manifold.
11:53
Speaker A
Do you see it now? What happens when the optimization target is not an external environment, but the exact architectural topology of another machine learning matrix?
12:06
Speaker A
Simulated annealing over combinatorial code edits require now a both a scalar magnitude gradient and a semantic directional gradient computation, no?
12:16
Speaker A
And rather than blindly allowing now our LLM to walk around in a Markov chain just hello, here am I. Let's explore the code base alteration that they're able to perform.
12:26
Speaker A
The last paper, Rack-A-Horn, is here creates now what? It creates here a specific instrument of a decoupled dual feedback loop. If you read the paper, you will see it.
12:39
Speaker A
So, this is now the next paper that has the same pattern. If you have not found it, there's the simple solution. Look.
12:48
Speaker A
Across all the four domains that we looked at, the system kind of solve an identical optimization mathematical objective, no? Maximize the expected reward of our policy pi data.
13:02
Speaker A
And now comes the interesting point. All the all papers show us here that optimizing the neural weights, this is here the tensor structure of our neural network, optimizing the neural weights alone is strictly insufficient to escape a local
13:18
Speaker A
minima in heavily obscured high entropy spaces. And those high entropy spaces, I showed you four different spaces.
13:25
Speaker A
So, this mean by explicitly now, this is the solution, condensing the environment's complexity into some simpler, maybe symmetric or geometric invariants, we can apply a mathematical theorem to find an elegant solution.
13:43
Speaker A
And those geometric invariants either in the first case, it was a knowledge graph. In the second, it was a skill bank. In the third, it was a known ontology. And in the fourth, it was an experimental skill state itself.
13:58
Speaker A
Do you see the pattern now? The pattern is simply that these architectures freeze the systemic entropy of the system, allowing the parametric neural policies to operate efficiently across mathematically stabilized topologies.
14:15
Speaker A
We are looking for the deep deep underlying geometric optimal solutions. This is the main cause of an optimization theory. If we don't find the right mathematical representation of our problem, of my problem with your physics, maybe you work in physics or
14:36
Speaker A
chemistry or finance or whatever, then the system will fail. So, now you understand maybe this image.
14:42
Speaker A
Now, at the core, we have our LLM, our frozen weight structures, beautiful. Now, we don't we don't actually learn out directly. We have not a full fine-tuning cycle. We have no lower since this fails here by complex structures. We have constructed here our
14:58
Speaker A
AI harness around the core of the agent. Now, and this harness now has also a beautiful possibility, an extreme power, to prepare and analyze the data stream before entering here, coming closer here to the core LLM.
15:14
Speaker A
And this wave that I ask AI to generate here is simply the geometry of the mathematical space of the underlying problem. Because whatever task you have, you have a simple mathematical solution.
15:30
Speaker A
But you also have an optimized mathematical representation of the task you're asking the system to perform.
15:38
Speaker A
Like here, now. This is here of the form of a particular wave function, let's say, in the simplest case, now. We cannot solve it really from a classical understanding here.
15:48
Speaker A
So, what we do in quantum theory, we put here this particular cover over it. We call it here a lattice, a particular computational quantum lattice structure, and then we can calculate only on the dots. We don't have the smooth surface,
16:03
Speaker A
but we have only computational dots, and we can jump from dot to dot in our numerical simulation. But the dots are real close, now. They are real cover over this manifold surface.
16:14
Speaker A
So, we are almost able to find here the correct solution. Since we're not able to find the correct analytical solution, the pure mathematical solution.
16:24
Speaker A
So, you see, we try to have your pure compute power and solve a problem that we do not understand and we do not understand the solution at all.
16:33
Speaker A
We just compute, compute, compute, put over here this lattice structure, and we hope from the result we can reflect on this and then try to understand the solution.
16:44
Speaker A
Now, you can have fun and reread your four papers now with this knowledge, now.
16:49
Speaker A
Think about it. The first paper is, you know, stops using here the raw text every time the AI takes an action. It passes here the observation, permanently adds it here to the structured knowledge graph. Now, this knowledge graph is now
17:01
Speaker A
the scaffold. This is exactly the example now. But how does it decide which path to explore without getting lost? As I told you, they use a twist a twisted sequential Monte Carlo algorithm. So, what it means, instead of a messy tree
17:14
Speaker A
search that would be By simplification, they track now a set of particles on this particular surface and they just look, hey, what is happening in the future? We can predict the future. We think as an AI I can predict the future.
17:26
Speaker A
So, therefore, we do have the future, so let's just play it out. Let's run a numerical simulation and see what is the result.
17:35
Speaker A
You can see this here, beautiful. Here you have your particle one, particle two, particle three in different futures and you calculate what is the right one, yeah?
17:44
Speaker A
Now the the idea of the simulation is simple, yeah? To evaluate each particle to see if it made physically progress to what our defined goal. If it did, it will survive. If it did not reach the goal, it is killed. And this is not maybe the
17:58
Speaker A
wrong way. It's a brutal way to approximate this, but it's a simple way because it aggressively filters out all the noise and all the things that are not working.
18:09
Speaker A
Now, the other paper here has a beautiful idea and I want to show you this.
18:16
Speaker A
They were not aware that I'm going to set up this crazy video here together with you, but they have something I can use and I think this is maybe what ignited my specific thought of today.
18:29
Speaker A
They are for different Alpha Worlds and Web Shop and Science World. They have some different benchmark, never mind about it, but they compared now GPT-5 on these different worlds.
18:39
Speaker A
And they have now three different methodology: history, belief, and here it is knowledge graph. So, let's have a look at this.
18:45
Speaker A
So, when forced to read the unstructured history, more or less the pure text, now our reflection model in this particular paper struggles heavily, now. You see here, yeah, there are two particular statistical parameters, never mind, the first one is the total detection error
19:00
Speaker A
and the second is the effective reliability ER, but you see, oh wow, we have really problems. It suffers a 95 TDE and achieves only here an effective reliability of 53%.
19:13
Speaker A
So, more or less this history of text is drowning in the background noise. The eye system is not clearly able to identify, hey, what is this text all about? What should I extract? What is the logical sequence of arguments to
19:27
Speaker A
extract from this text? This is just noise to me. Belief system. We talked about it in my last videos, no?
19:35
Speaker A
When we transition out to the belief summary of this particular model here, GPT-5, the TD drops slightly to 89 from 95, and our effective reliability climbs up to 55%. That's not That's an indicator, no?
19:50
Speaker A
So, the lossy compression removed some of the noise, but the discrimination is still weak. We have not found, no, the real strong signal here in this white noise or whatever you have, no?
20:01
Speaker A
And now let's see if we have here this new topology, this new knowledge graph in the context.
20:08
Speaker A
So, what we have? We have now a structured topological condensation of the environment, and this is our This is our knowledge graph.
20:17
Speaker A
And then we have that the errors collapse. Now, you might say, "Hmm, the data, 77 for TD, and here for effective reliability just jumps to 65% is not a massive jump." But look at the other one. So, yeah,
20:31
Speaker A
this is already a kind of an indication. Remember, we are not in theoretical physics, so we are not in real science.
20:38
Speaker A
We are here in computer science, and over there, yeah. Okay, they are not real theoretical physicists, they are computer scientists. So, okay, let's just say, "Okay, we accept here this result." But you see what I mean?
20:53
Speaker A
Now, reflect on this. This table is two that we just looked at, two different versions, shows us that an LLM itself cannot execute some reliable logical reflection over some raw unconstrained text streams, just some wave of text over text, thousands of pages.
21:09
Speaker A
Yeah, it can try to distill here and extract you the the meaning, but it's not really the best system.
21:17
Speaker A
But by crystallizing the environment state into the symbolic knowledge graph, suddenly the job is so much easier, no?
21:25
Speaker A
Because the first paper suppresses here the false positives, and this is all the hallucination by the I mile by the way, no? And allows you the reasoning agent to maintain a high purity trajectory that is not hallucinating towards the
21:38
Speaker A
global objective. Reduce the search space of all the vector representation of all this stream of text elements of text chunks here, and we condense everything and we build a symbolic knowledge graph. And this reduction of complexity is the solution. So, you see,
21:55
Speaker A
this is here the pattern. And this is exactly the coherence that I felt reading those papers, and I could not really immediately articulate to you.
22:08
Speaker A
Think about it. All these papers stop relying on the AI internal weight structure and our pipe data or some messy unstructured text histories the fluid to store the information yeah?
22:21
Speaker A
You know what we did? We built a harness around our LLMs. I mean, that's the reason for this, and this is the reason.
22:28
Speaker A
So, instead now this or the authors of these papers actively build now rigid non-parametric scaffolds.
22:37
Speaker A
And I've showed you one is a graph, one is an ontology, one is a skill bank. It doesn't really matter if you understand the underlying principle in order to freeze here what they have learned into a solid structure.
22:51
Speaker A
And this goes hand in hand with scaffold and harness. If you are still not convinced, look at the next paper, no?
23:02
Speaker A
How does a system learn best? In physics, dynamics happen at the phase boundaries, no? As I have already discussed. So, if any AI generates tasks that are too easy, it learns nothing. If the task is impossible, it just produces noise. We
23:14
Speaker A
talked about it. So, it must operate exactly at the boundary of its own competence.
23:21
Speaker A
It's not so easy to really determine this boundary for a black box AI system, no?
23:26
Speaker A
So, they build now a particular system. This is a prototype where the AI generates its own practice problems, but it forces you to generate it only to create problems that are just a little bit a baby step out of reach.
23:38
Speaker A
So, that you hope, give them enough time, the system is able to do it. You place it so close here that you say, "Oh, hey. Okay, I have two sigmas and then it's just a little bit outside no?"
23:50
Speaker A
And this is exactly what we already talked about. And look, it is more or less the same pattern of the deeper learning pattern, no? When the AI fails on one of these boundary tasks, it doesn't just throw the failure away, if you read the paper,
24:03
Speaker A
no? It extracts the lesson, beautiful, and solidifies here it into a skill card in an external memory bank.
24:12
Speaker A
Now, interestingly, authors show us, you will see that the total number of skills rise, experiment experiment experiment and then it sharply contracts, goes down, and then it stabilizes. What happened?
24:25
Speaker A
This is exactly the system decoupling and deleting useless skills. This is here like this phase transition, exactly like a crystal dissolving and reforming into a more stable lattice configuration.
24:40
Speaker A
So, another form like this can happen. And the last one again is again, using your tools already know versus gambling on unknown tools. I just rushed through this because you know this now.
24:54
Speaker A
They use here a particular gate mechanism. They use here the upper confidence bound algorithm, calculates a score based on the tools historical success plus an exploration bonus.
25:03
Speaker A
Is a simple mathematical formula. If you're familiar with it, great. If not, this is your particular policy that operate.
25:10
Speaker A
And yeah. Now, what is interesting is this sentence here in the paper, you know?
25:17
Speaker A
They capture here the agent's ability to express a task-specific tool need under an ontology-guided reasoning. And here we have it again.
25:28
Speaker A
So, you see there are different representation of this problem and the solution, but it is almost all time the same pattern.
25:37
Speaker A
Here, you have the gate control and ontology-guided retrieval. So, what are the main insight of all these papers?
25:46
Speaker A
Neural network architectures are great, but they're also terrible at maintaining a reliable memory of a long noisy timelines.
25:56
Speaker A
And the unifying breakthrough of these four preprints is that they stop forcing the neural network of the LLM to remember everything or the eye in the harness structure.
26:07
Speaker A
Instead, what they do, they build here, and this is the common pattern of these papers preprints a deterministic external scaffold.
26:16
Speaker A
This is either a graph, a skill based on ontology, you have seen it. The neural network is simply here only the engine.
26:24
Speaker A
The scaffold is if you want the transmission, that actually applies the power to the road, converting now the noisy exploration into permanent structured progress in an optimized system.
26:36
Speaker A
So, this brought me then to the main idea that I do not understand exactly what is the difference of a scaffold and a harness.
26:45
Speaker A
So, what is the scaffold? Looking at these four papers, a scaffold is, if you want the detection medium, this is the geometry, the underlying mathematical geometry of your problem.
26:56
Speaker A
This scaffold is the static or slowly evolving data structure. It is the geometric invariant that we're looking for. As a theoretical physicist, you know exactly we're looking here for invariant groups. We're looking for symmetries to be able to solve our
27:08
Speaker A
problems. Otherwise, if we will never find a way. So, we are looking here for the topological map or if you want to get this lattice structure.
27:18
Speaker A
And in the first paper, this scaffold is the knowledge graph. In the second paper, this is the skill bank. In the third scientific tool agent evolution, it is the known tool ontology that we operate. And with Rec Harness, it is the
27:30
Speaker A
experimental skill text manifold. So, this means this scaffold, just looking at these four papers, is now the structured medium that recon- that records the physical events.
27:43
Speaker A
It possesses a internal inherent mathematical geometry, but it does not possess logic. Or is it in a way that the logic is hidden in the geometric configuration?
27:59
Speaker A
So, it just sits there maintaining here the collapsed entropy of the complete system state.
28:07
Speaker A
That's also, what is now the harness? The harness in those four papers is nothing else than the experimental apparatus, the dynamics. This is, if you want, the algorithmic control loop that surrounds both the LLM and the scaffold.
28:20
Speaker A
The LLM here with our tensor weight structure and the scaffold with the optimized geometry, with the optimized mathematical representation with symmetry and rotation group and everything there. It governs here the flow of time, the allocation of budget, the evaluation of rewards, and it is
28:34
Speaker A
simply the statistical gating of the information. And I think this is here really what resonates with me. This is a harness.
28:43
Speaker A
The statistical gating of the information flow. In the first paper, you saw the harness is the twisted sequential Monte Carlo algorithm determine which particles live or die.
28:54
Speaker A
In the scientific tool agent evolution the harness is here the specific gate deciding whether to exploit known tools or to explore unknown ones. So this is exactly this what I call delicate equilibrium that we always have, no?
29:08
Speaker A
And in Rec Harness, I mean you literally have here the name harness here in the title, it is here the Thompson sampling rate where they decide which of the actual edits to test next.
29:21
Speaker A
So, if you want and yeah, I'm thinking about physics, I'm thinking about particle accelerator. So this is here my image. This is the way I think. The harness dictates how the active beam this is here the the LLM if you want.
29:34
Speaker A
This is here the pipe interacts with the detection medium and the detection medium is here the scaffold. The scaffold is here I mean it's just a coincidence that this is here a green net here with this area.
29:47
Speaker A
But this is here really the underlying geometry. The underlying geometric pattern that we want to explore, crush, detect, build up.
29:57
Speaker A
This is here why suddenly I'm quite satisfied for today with this definition. The harness dictates how the active beam the LLM interacts with the detection medium and this is here the scaffold the geometric medium.
30:12
Speaker A
There is another table that is absolutely fascinating in this papers and this is this one here on Caesar.
30:21
Speaker A
Now, they give us here and let's go with a Qwen 38 billion model here. They give us here yeah, different benchmark.
30:28
Speaker A
Forget about we just look at the last column the average one. Now, they give us here three options.
30:33
Speaker A
They say, "Look, this is just the raw LLM as a beam no memory. This is here just a motor spinning in a vacuum." And it achieve here on all the different benchmark here, 56.3%. This is here.
30:45
Speaker A
Now, the second line is Caesar off. What is it? The LLM after being trained via the harness itself, but with the external scaffold disconnected at test time, you see, okay, we got higher. We go from 56 to 58%. Great.
31:01
Speaker A
And then we have Caesar on. And this is now the LLM hooked up to the final fully populated scaffold full configuration, and it jumps now, I mean, it's just just one percentage point, but let's say it jumps now to 59%.
31:16
Speaker A
This is interesting. So, you can turn here kind of here in this experiment they show us, you can turn the scaffold on and off and get here better or worse results.
31:27
Speaker A
Of course, this is not really the perfect experiment for the argumentation, but I think it it's just a coincidence that you can see this, no?
31:35
Speaker A
First line is here just the raw LLM. Then you have the LLM with the harness, but without the scaffold. And only if we activate all three elements, you get the best performance indicator.
31:47
Speaker A
Just think about it. Maybe there is something to this sort. If we talk about Let's think about it.
31:55
Speaker A
Let's go to another another step forward. No, but I stop here. Come on. Let's use this.
32:01
Speaker A
What about this sentence? The future of AI, as proven by these four or five preprints, is building some elegant mathematically optimized harnesses that operate on mathematically optimized geometric scaffold. Would you agree with this sentence of mine?
32:17
Speaker A
What about this? The LLM is just the engine block. The harness is the chassis and the transmission and the steering wheel.
32:26
Speaker A
And of course, you can transplant a completely different LLM into your harness, no? And the system will still function because the true architect of the problem is mapped within the scaffold, not the weights.
32:40
Speaker A
Now, think about this. This is not a trivial sentence. Because they has at least three different levels going down if you want to analyze it. So, leave me some comment here.
32:53
Speaker A
And then I thought, okay, and I want to visualize this for you. And I had this prompt here and I said, listen, create another of our special design patterns here talking to my AI with a star gate motif, forget about it, and
33:04
Speaker A
depict the underlying complex geometry that we want to regularize, that we want to master with a computational architecture in order to morph it a general problem given to an AI into a geometric pattern that allows the harness configuration to find an elegant
33:19
Speaker A
and mathematical primitive solution, means it is solvable and all, utilizing now the symmetries and the rotation groups and the special mathematical constructs.
33:28
Speaker A
So, this is my way of prompting your my AI machine. I trained this AI machine here on my pattern, so you see, and say, "Hey, create the image." And I was really waiting for this image.
33:39
Speaker A
And here you have it. This is here how the AI understands here my crazy thoughts, but there's something to it. You see this?
33:48
Speaker A
Here we have our complex manifold where we try to to have here to tame here this complexity here with some regularized structured with some lattice, with some knowledge graph, with any graph structure at all.
34:03
Speaker A
And then here our harness element here is looking especially here for the geometry, for the underlying geometry of the problem.
34:11
Speaker A
And it is building here specific geometry detectors that if one of those geometries is detected, that this harness now knows an easy solution.
34:22
Speaker A
So, yeah, it is absolutely fascinating. It is just here a video to think about it, to go crazy if you have here the pros and the cons and you compare everything in your in your brain and you think, wait a minute, what exactly is
34:35
Speaker A
this? And there's no easy solution because you can ask any AI system. They are not really there yet. So, have fun, use your human brain, enjoy this, and please, if you have a breakthrough here, if you have some new ideas how we
34:52
Speaker A
can cope with this or how I can better explain this to my audience, I would highly appreciate a comment from you.
34:59
Speaker A
Or, I'll see you in my next video.
Topics:AI architecturescaffold paradigmharness systemlarge language modelsneuro-symbolic AIknowledge graphself-evolving agentstransformer memoryphase transition AIAI optimization

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →