Explores runtime graph repair inspired by human cognition, revealing flaws in AI benchmarks for physics and proposing a new adaptive knowledge exploration framework.
Key Takeaways
- Current AI benchmarks in physics are largely flawed due to AI evaluators' inability to recognize alternative correct solutions.
- Human expert evaluation is crucial for accurate assessment of AI performance in complex scientific tasks.
- The new runtime graph repair framework mimics human cognition by combining graph and text knowledge for adaptive reasoning.
- Bidirectional synergy between structured and unstructured data enhances AI's ability to navigate and reason over large knowledge bases.
- This approach could improve AI reliability in complex domains like physics, medicine, and finance.
What the video covers
- The video discusses a new approach to runtime graph repair inspired by human cognition, integrating structured graph and unstructured text knowledge.
- It highlights a study by Yale, Cambridge, and USC showing that 95.2% of AI benchmarks in physics are incorrectly evaluated due to AI evaluators' limited understanding.
- Benchmarks of models like Fable 5, GPT-5.6 Soul, and Gemini 3.1 Pro are analyzed, revealing non-comparable conditions and surprising performance similarities.
- The presenter critiques reliance on AI-only benchmarks and emphasizes the importance of human expert evaluation in complex scientific domains.
- A new paper from Beijing University and Chinese Academy of Sciences proposes a training-free framework for adaptive knowledge exploration inspired by human problem-solving.
- This framework establishes a bidirectional synergy between graph-based and text-based representations to improve reasoning and navigation in large knowledge bases.
- The cognitive cycle involves planning, exploration, reflection, and dynamic decision-making to repair and update knowledge graphs at runtime.
- The video explains the mathematical and algorithmic structure behind this approach, including state updates, policy decisions, and control transitions.
- The presenter discusses the implications for AI in complex fields like physics, medicine, and finance, questioning the trustworthiness of AI evaluators.
- The video concludes with a call for more robust, human-inspired AI evaluation and knowledge integration methods to advance AI research.
Chapters
- 00:00Introduction to Runtime Graph Repair and AI Research
- 01:27Flaws in AI Benchmarks for Physics Revealed by Human Experts
- 03:22Comparative Performance of AI Models on Public Benchmarks
- 05:19Critique of AI Benchmarking Practices and OpenAI Results
- 06:49Introduction to Human Cognition-Inspired Graph Reasoning Framework
- 08:35Bidirectional Synergy Between Graph and Text Knowledge
- 10:19Cognitive Cycle: Planning, Exploration, and Reflection in AI
- 12:10Mathematical Structure and Policy Decisions in Graph Repair
- 14:35Implications for AI Trustworthiness in Complex Domains
- 20:45Summary and Future Directions in AI Knowledge Integration
Full Transcript — Download SRT & Markdown
Speaker A
Hello community. Hey, it's so great that you are back. Let's have some fun today with artificial intelligence research.
Speaker A
We have a brand new topic, and today we're going to talk about runtime graph repair in multiple dimensions, of course. And I know what you're saying.
Speaker A
You say, "Hey, wait a minute. We just had a video where we had some procedural graph, some new procedural graph methodology. And now you will build on this." And I have to tell you, "No, not at all." I have a surprise for you
Speaker A
because today we're going to map, as the authors claim, human cognition, your science here. And we'll copy this onto any AI machine. Interested? So, let's start. But before we start, a short interrupt and I will tell you this study
Speaker A
here by Yale University, University of Cambridge, and University of Southern California. And you know what they did?
Speaker A
They said, "Listen, those benchmarks, these AI benchmarks we have, that are where the benchmark is done by another AI machine, we now employ a lot of human expert, of human specialist. And we do this for physics. We answer the
Speaker A
question, how good are the frontier models at physics?" So, a big thank you to all the people, as you see here, the authors that participated in this study. And you know what they found? They found that the benchmark of all AI benchmark that we
Speaker A
have on physics, on science, is 95.2% incorrect. Absolutely. So, almost all of the benchmark results we have on AI, on science, is incorrect. And you might say "Why?
Speaker A
All our AI model benchmarks are wrong? We cannot trust AI as an evaluator, especially in science? Absolutely." And they looked [clears throat] here at the performance of a Fable 5, of a GPT-5.6, of a Soul and Gemini 1.0
Speaker A
and they said, "You know what? The AI that produced your answer was evaluated by another AI machine." Now, it turns out that the AI machine that evaluated this did not understand that the answers produced by Fable 5 and Soul and Gemini,
Speaker A
were correct, just it was a different path. A different different methodology in the mathematical physics that they applied.
Speaker A
So, therefore, the AI machine that were the evaluator of those AI benchmarks, evaluated the answer as wrong, although it was a correct answer, but just a different approach to solve this.
Speaker A
So, Yale University, thank you for paying all these human expert to look at this and find out that this whole AI benchmark was 95.2% that the what the AI evaluator said is a failure to 95.2% were not failures at all, but just a
Speaker A
different way to solve equation, a different mathematical approach. But, the AI evaluator model was simply too stupid to understand that this is a different approach to their limited understanding of physics. So, therefore, all the benchmarks that we have on
Speaker A
physics are more or less incorrect. So much about AI and science. Now, if you say, "Okay, so what model to use now?" I told you here in this prompt, uh today is September 14th, that I will go now with GPT 5.6 Soul High, but I
Speaker A
told you also careful, Fable 5 and GPT 5.6 Soul were evaluated with the coding tools.
Speaker A
Gemini 3.1 Pro was not allowed to use. So, I would say this is a highly non-comparable situation, but okay, this is the data that we have. But, look at this. For the benchmark on public benchmark on public sources, you see
Speaker A
Fable 5 High has not 92.6% accuracy, GPT 5.6 Soul has 93.9, but you know what? The Gemini 3.1 Pro, the grandfather model, without tool use for coding mathematical physics, achieves absolutely the same performance like the best GPT-5.6 so
Speaker A
without tool use. So, this is absolutely impressive, but if you look now at expert ordered benchmark here, and this is here especially here the last one, the last line here, you have the critical point benchmark. And this is really complex
Speaker A
research using critical thinking in mathematical physics. So, this is as complex as it can get.
Speaker A
And here you see Fable 5 90% GPT-5.6 so 94% and of course now it hurts Gemini 1.5 Pro that it has no does not is not allowed here to have coding tools for the validation of its mathematical physics.
Speaker A
So, therefore at least we still have 68% so interesting, absolutely fascinating. So, at first notice that there is an not really comparable scenario, no?
Speaker A
But yeah, for the time if you have to use activated a GPT-5.6 so might be here really the reason to go. So, I'm desperately waiting for a Gemini 4 because if this continues I think a Gemini 4 if we have with no tool use on a
Speaker A
classical public benchmark here the same performance like we have in a GPT-5.6 so with tool use fingers crossed. Okay.
Speaker A
So, this is also one of the reason a lot of you ask me, "Why do you not present here in my videos here the benchmark result that are produced by OpenAI? All the other YouTubers show the the papers
Speaker A
by the marketing departments of OpenAI or the marketing department of Meta or Anthropic, but you always go to independent sources and you always go to university sources. Why do you don't trust here OpenAI?" And I think this is
Speaker A
a good question. So, let's look at a paper here by Yale University who paid humans to do the job. And of course, OpenAI, why should they pay human expert professor of physics to do the job if they have several hundred thousand GPUs
Speaker A
just behind? So, if you pay humans, you find out that all these physics benchmark were incorrect.
Speaker A
So, great. So, now to a simple question for you. Can we, can you trust the eye in complex topics like physics, medicine chemistry?
Speaker A
What about some complex job in finance? It turns out that the EI, EI machines that evaluated Soul and Fable did not even understand that Soul and Fable came up with correct solution, but the EI evaluator were simply too stupid,
Speaker A
the EI machine was itself too stupid, too limited in its understanding to understand that its evaluation job was absolutely incorrect.
Speaker A
So, what about medicine? So, let's start out today's video, and this was just here the starter, this was here the precognition, so let's go.
Speaker A
We talk now about the main cognition on graphs. And here we have it, this is Beijing University of Post and Telecommunication, some academy in Beijing, the Chinese Academy of Sciences, and the University of the Chinese Academy of Sciences, a lot of
Speaker A
unbelievable clever, beautiful, intelligent people, and they published September 11, 2026 a new paper. And they have now this idea, and they say, "You know what? If we have real complex knowledge bases, some large-scale knowledge graph, no? And text corpora for complex reasoning, it
Speaker A
is still challenging. Yes, of course, we have Rag and GraphRag and thousands of Rag versions, but you know what? We propose now a new Rag methodology, and we go now and re-map the human cognition on a graph structure."
Speaker A
And this is a human cognition-inspired training-free framework for adaptive knowledge exploration. So, if you want, and this is just a joke, of course, I say, "This is my rag on rag.
Speaker A
So, let's have a look. They have here a beautiful comparison of all the different exploration strategy for rag. You have here a simple beam structure reflect backtrack, then you have proactive planning, adaptive entity linking, text base pruning. And then you
Speaker A
have the new methodologies and guess what? The new methodologies integrating everything that we know from the other systems into this new system.
Speaker A
You might say, "Hey, great." So, let's have a look. Let's have a little bit of fun with rag.
Speaker A
So, discognition on graph integrates here all prior capabilities in a cognitive cycle establishing here, and this is important now, a bidirectional synergy where text knowledge actively bridges here um symbolic graph gaps, whole in the graph knowledge enabling now uh better
Speaker A
robust navigation in real global knowledge bases. So, what they do they establish here a bidirectional synergy between a structured graph representation of a complexity and an unstructured text representation of a complexity.
Speaker A
And they say, "Hm, how can we combine them that we have a synergy in both directions?" Of course, you have here with this beautifu
Speaker A
Great. Zero stars here watching because nobody really values these studies and says, "Wow, this is really of important." So, what are we talking about? We talking about it efficient retrieval of new data for our agent is of course the
Speaker A
core of any rag system, retrieval augmented generation. And we had in the old times here in the stone age of AI, we had vector rag, and then at the bronze age, we had graph rag here where we have knowledge graph for structural
Speaker A
guidance. And 7 months ago I did did a deep graph rag with [laughter] your PO and hierarchical reasoning and I forget graph rag we have now have new better system and we have memory graph rag and we have hierarchical reasoning
Speaker A
on graph rag a year ago. So you got we have a ton of methodologies. So what is this new system now doing?
Speaker A
It is progressively constructing now evidence chains by and now hold on to your socks synthesizing structural associations from knowledge graphs and semantic information from text corpora.
Speaker A
Now if you like mathematics you say hey that's cool. So now how do I code a synthesized structural association from a knowledge graph now? So let's talk about it.
Speaker A
So the authors are clear now. We draw inspiration from the cognitive processes of human problem solving. So when humans do not blindly towards your information instead humans engage in a dynamic cycle of planning exploration and reflection and you say hey that's great. So they I
Speaker A
suppose they they are those humans. So those humans proactively formulate investigation plans flexible synthesize information from diverse sources and then those humans I mean they must be quite clever now. Those humans continuously then reflect on accumulated evidence to adjust their strategy. So
Speaker A
this is here under quotation mark here from the authors of this paper. And therefore the authors tell us and now we will propose here the cognition on graph COG. It is a framework that instantiate this human-like reasoning process.
Speaker A
And I'm just amazed as a theoretical physicist that neuroscience can now map human-like reasoning processes into a continuous cognitive loop on an AI machine.
Speaker A
So let's do this. Now strictly between you and me now I have to tell you I don't think so now. I think that this cognition on graph does not train a cognitive model or discover here a cognitive algorithm at all.
Speaker A
It modules are implemented through a simply hand-based design prompt structure. I will show you this in a minute. And some deterministic runtime control. So, from the point of a theoretical physicist, this is not happening. But we are here
Speaker A
with the publication. We are here in the hype here with AI and mathematics. And of course, we map now cognition. I mean, come on, what's the problem with this?
Speaker A
Now, so let's compare this methodology to vector rag, to graph rag, to all the different rag system. And here at the last one, we see now, oh, beautiful, the key capabilities. Now, imagine we have now all the capabilities in the world.
Speaker A
Right. This is here the framework. And I just want to show you here, this is really from the artist the framework. So, you have a query.
Speaker A
And then you have a planning, what to do. This is an AI structure, no? Then you have your classic knowledge graph retrieval. Then you have your text retrieval. Let's go with Wikipedia.
Speaker A
Then you just have a synthesis, a summary, a retrieval as sufficiency. Then we have a reflection, and then adjust the planning, and you continue, and you have an answer generation. And come on, what can go wrong?
Speaker A
I want to show you here a real example and they provide in the annex here really here. So, you have a particular question, then you have your multiple turns. So, let's go through the cognition on graph reasoning process.
Speaker A
So, at first, for this particular question, you have here the planning phase. Then you have the exploration phase.
Speaker A
Notice, we have two different knowledge bases. We have a knowledge graph and Wikipedia itself.
Speaker A
And then we have a synthesis. Well, guess what? In the first turn, the synthesis is zero. So, we have a reflection state. So, this is it.
Speaker A
Planning exploration synthesis reflection. Then turn two. So, we recover now the planning. We have now some insight that the first planning was incorrect. Then again, we have the exploration phase.
Speaker A
Then we have maybe a key discovery from either the knowledge graph or the textual version of Wikipedia. Then we have the synthesis and reflection. And you might say, "And how we do this?" Guess what? Here the reasoning trace here is simply a
Speaker A
refraction proactive planning adaptive entity linking text entity utilization and this is it. And if you see the prompt template that they use, they provide everything in the annex. Thank you for providing this here. Table 19, you have here the prompt template for
Speaker A
the relation discovery module, the row context available relation, the task, the output format. Everything is given to you. And you see this is just prompt engineering.
Speaker A
But let's reflect on this. What is it? It is much more than you might think.
Speaker A
Because this cognition on graph, I do not like the term, is not primarily a new graph search algorithm.
Speaker A
It is actually, if you use this in runtime, it is an agentic runtime that converts some retrieval from an open-loop search procedure into a closed-loop process of epistemic control over two complementary knowledge representation. And this is yet a
Speaker A
knowledge graph and here the text corpora from Wikipedia. This has nothing to do anymore with the classical rack from years ago.
Speaker A
This is now a sequential navigation process through a partially observed knowledge space. So, just give you a feeling why.
Speaker A
For the preprint, they did some experiments. We go with their data. You have about close to 90 million processed Wiki data entities. You have about 1.4 billion processed edges in the graph representation. You have more than 5 billion words here in the textual body
Speaker A
here of Wikipedia. And you see you are dealing now with dimensions and unbelievable amount of words, 5 billion words. So, if you search this, you know, this is not going to work.
Speaker A
The system cannot exhaustively inspect the complete space. So, it needs now a particular policy, a search policy in this search space that is a mathematical space, for choosing now which small region to explore first, and then in the next step, and then in
Speaker A
the next step. So, you get the sequence. And therefore, I see this here, this cognition on graph, I don't like it, as a closed-loop controller, eh?
Speaker A
Now, the paper and the authors go with here the four stages. And I told you the four stages here officially are planning exploration synthesis and yeah, you got it, the authors are really proud that they have then the
Speaker A
reflection eh? The EI system reflects on its actions, where another EI reflect on the action of this first EI machine, and then we have some deep insights. Great.
Speaker A
Now, you know, as a theoretical physicist, I prefer a more precise system description, or at least an interpretation. So, therefore, let's be a little bit more precise than the published paper. We have a state, then we have an information gathering action
Speaker A
that is designed by the LLM, then we do have an observation that is returned from the environment here to the agent.
Speaker A
We have a state system-wise update procedure, and then we have a policy decision. This is it if you want to be a little bit more precise. Why? Because we have to code this. This means we have to define here the mathematical structure
Speaker A
for the internal state state S at a turn T, and this is simply represented now by this tuple, where you have here, of course, describing here the internal state here at a particular turn T, with Q, this is
Speaker A
the original question that I as a human user have to the EI machine, then we have the current investigation planning, then with NT, this is our notebook, where we store our verified facts that we have, C is here our candidate
Speaker A
entities, the pool of promising candidate entities that one want to explore next, and H of T is the look back, the history of previous plans and retrieved outcomes and retrieved arguments, and everything that happened in the past.
Speaker A
And guess what? Now, it is simple, because mathematically, what it is from this state equation, the agent now chooses here a particular exploration action. You know, this is always the same thing. We the LLM decides now on a
Speaker A
particular action based on its neural network. Beautiful. Consisting now of sub-queries and corresponding anchor entities. Guess where the anchor entities are going to be mapped to? Of course, the graph structure the graph structure.
Speaker A
And then then it receives simply two observations. An observation here from the knowledge graph and an observation here from the textual body, yeah?
Speaker A
They are synthesized into a compressed evidence state, after which and then then comes the nice thing. We have now a judgment variable, yeah?
Speaker A
And this determines now a control transition. So, think about it. We only have three options yeah?
Speaker A
The judgment variable, where you say, "Okay, great. We have another AI machine that now comes in and says, 'This is it.
Speaker A
This is sufficient. So, we have our answer.'" Beautiful. Done. Or this judgment machine comes in and says, "Hey, this is insufficient, but useful. So, continue along this particular direction in our reasoning trace.
Speaker A
Like, "Hey, this sounds interesting. Continue to reason. Maybe you will find something." Or we have now a judgment by the other AI machine, "This is insufficient and this is useless. So, we know, 'Hey, stop. Don't follow this path
Speaker A
anymore. Change the strategy. Go back maybe here.'" Great. And now, guess what? Yes, of course, the third case is here the innovation of this paper.
Speaker A
Because think about it. What was the problem? What is the current problem, right? Current state of the art retrieval agents like TOG, like think on graph, operate reactively. So, they use here a heuristic beam search. So, if they go
Speaker A
down a wrong path on this particular reasoning tree, they just keep digging until they hit a dead end, resulting in either a failed query or they hallucinate here a logical chain and the answer. Great. We don't want this.
Speaker A
Now, do I have to show you again this study where here the AI machine that judged here another outcome of another AI machine, that this is here exactly what we have here the study by Yale where we said,
Speaker A
"Yes, here we have here an AI machine, Fable 5, GPT, Solo, Gemini." And this was judged here by another AI machine.
Speaker A
And now, doing this by human researcher, we find out the 95% of the audited AI model failures were not model failures at all.
Speaker A
So, yeah. Whenever you have an AI that is judging work of another AI, I have now for me myself a simple advice, do not trust AI machines.
Speaker A
Basics, never ever. Mathematics, never ever. Anyone say, "Hey, wait a minute. I just remembered that just days ago there was a beautiful video here on this channel about topological intelligence, and we were talking about this same process." What a coincidence, huh?
Speaker A
And we were talking about this knowledge hole, this this areas here in our carrier space where suddenly there were no observation points.
Speaker A
Exactly, this is what we have what can happen now also on a graph's representation, we will also on a graph detect knowledge holes.
Speaker A
And, of course, with our detailed and a mathematical understanding of topology and homotopy classes here, you know exactly what we're going to talk about now.
Speaker A
Absolutely, we will use here the semantic complexity to find a bridging function to the graph representation.
Speaker A
One of my viewers asked me, "Hey, do you have some homework to do?" I mean, of course, if you ask me, simple. Just verify that this statement is correct.
Speaker A
Planning is a hypothesis formation. Second, graph exploration is progressive structural filtering. This means we work with entity linking, relation discovery, and fact pruning. Third, text exploration is more or less a hierarchical reading exercise with adaptive page selection, global
Speaker A
scheming, and detailed reading. And therefore, the real novelty of this complete paper is simply we have a graph representation and a text representation that form now a feedback system, a closed loop.
Speaker A
This is your homework. Show me the proof. Great. If not, we continue with the video.
Speaker A
So, just that you understand, a graph entity, what does the graph supply? Canonical identities, aliases, relations table, anchors for selecting documents, structural hints about what information should exist.
Speaker A
So, therefore, you want the graph organizes you the textual reading. On the other side, the text can expose an entity that is missing from the local graph neighborhood, where we have this whole structure not connected through the selected relations that we know. So,
Speaker A
we're out missing here out on relation, on edges, implicit in a statistical table or not mentioned in original questions.
Speaker A
So, if the graph representation is incomplete, we do not build here. We do not try to detect missing edges or missing data.
Speaker A
We just switch to another mathematical representation of knowledge. And the other mathematical representation of knowledge is a text-based representation of 5 billion Wikipedia words and their syntactic and semantic sequence.
Speaker A
So, you see, this can be quite interesting to combine. So, let's say the text has effectively constructed, if you want now here, a temporary bridge over a missing or inaccessible graph edge. It is not like three videos ago I showed you uh
Speaker A
methodology how to build missing edges between unknown elements in your test graph representation. Here we do a much simpler methodology.
Speaker A
We just say, "Hey, listen, I have a graph representation, but at the same time on the internet I have a text representation.
Speaker A
So, I now fuse them together and let the text build temporary bridges over my graph complexity.
Speaker A
But careful, this cog- cognition on graph does not permanently modify the knowledge graph. This bridge functionality exists only if you look here at the GitHub repo in the code in the agent's evolving search state representation.
Speaker A
Of course, we have to talk about memory. We have 5 billion words, you understand?
Speaker A
A memory optimization would be a nice thing, no? So, let's go. Three kind of memories. I just give you here the indication. Now, we have a notebook with the verified facts, and the functional role in our system is
Speaker A
this is a stable epistemic memory. The candidate pool is simply this is promising, but yeah, but we have to explore it, no? This is the search frontier memory. You know exactly, oh, we have to have here a deeper
Speaker A
exploration of this candidate pool. And then we have the interaction history. What happened in the past, previous plans, previous outcomes. If you want, this is the episodic strategic memory.
Speaker A
Or in very simple terms, the notebook answers, "Hey, what have I already established as an AI machine?" The candidate pool answers, "Hey, what might I investigate later as an AI machine?" And the interaction history answers, "Yeah, what already happened?"
Speaker A
Results. Now, I will tell you here the main result. With this new Regon Rego, whatever, cognition on graph methodology, with a Q and 3 8 billion mile, you achieve here a performance that exceeds here a 32 billion mile on
Speaker A
the old methodology. Across six different data set and with a particular selection of some complexities and benchmarks and yeah, here you have the numerical data.
Speaker A
If you want to compare this over 100 samples per data set, here you have a Q and 3 8B, a Q and 3 20 32B, and you have here of the COG here really outperforming here of course the direct
Speaker A
methodology but this I want to show you that you see here you get a difference now as 8 B does 50% overall average 32 B gives you 57% a deep sea version 3.2 gives you 62.6% so you have an idea how the systems
Speaker A
perform. But of course we are here mathematical theory and here almost said mathematical artificial intelligence about some basic insights so what is it on a structural level rack I mean the now this new rack this new idea that we want to try to build
Speaker A
this new rack is becoming an agent architecture now this is nothing to do with the vector space of all the ages now.
Speaker A
So this also means the distinction between reasoning and retrieval in a system architecture is more or less slowly dissolving.
Speaker A
Why? Because in this cognition of graph methodology the retrieval decision actively depends on the current hypothesis formulation the previous verified evidence the field strategy the unresolved constraint the candidate entities and an estimate of epistemic sufficiency.
Speaker A
So reasoning the classical term reasoning and the classical retrieval they become real real close here in architectural terms.
Speaker A
The representation translation becomes now the reasoning operation per se. So this means we have a graph and a text corpus and they are not merely just two different databases now because they impose two different inductive biases now.
Speaker A
Graph in the simplest way I can express this it has an identity plus the relation to the nodes and the edges and a discrete graph structure now.
Speaker A
A text has a context a description and some implicit relations. And of course we can find a mapping but do a mapping over 5 billion words.
Speaker A
I mean, you have to build a supercomputer for this, eh? So, therefore, the act of translating between two different representation might reveal that information is not easily accessible within either representation alone. So, this is where we have our bidirectional synergy
Speaker A
emerging. I also think that missing knowledge can be really repaired temporarily also at runtime.
Speaker A
Because this cognition on graph methodology does not need to reconstruct the entire incomplete knowledge graph.
Speaker A
Yes, we have a methodology, but in the paper, they just say, "Hey, listen, come on. Just discover here the textual bridge that is required for the current question here on the graph manifold, eh?" So, therefore, this is kind of a
Speaker A
solution that is query-conditioned temporary graph repair. And I think the last one is memory.
Speaker A
Everything that we can put external on external memory creates here this particular search state machine.
Speaker A
So, think about it. We have three different kinds of memory: notebook, candidate pool, and interaction history memory current and otherwise stateless LLM into a kind of a stateful search process.
Speaker A
So, is this an indication and this is just a speculation to a new AI architecture principle?
Speaker A
The LLM supplies the local semantic judgment and the harness preserves the global arithmetic state also in this rag on rag complexity.
Speaker A
Or if you wanted simple, the LLM decides whether an entity or a relation is useful in a graph representation, and then during runtime, the runtime determines which information persists and how it will change the future actions. So, here again, this duality
Speaker A
here of energetic structure between the core LLM and the outer control cycle of our harness.
Speaker A
Yeah, test time. If we have this really that we apply it here during runtime test time architecture, can it really substitute for model size? I showed you that we have here an 8 billion model and with a clever architecture, we come
Speaker A
close to the performance or even outperform a 32 billion model with an older architecture. Yes, of course, this is possible. It was demonstrated here.
Speaker A
But careful, this is a highly selected examples, selected spaces, benchmarks. So, for your knowledge, for your domain, for your capacity, you have to check this yourself.
Speaker A
But in general, I think we can say an 8 billion result suggested some capability attributed to the model size, so to that next 32 billion model, can instead be produced by a clever architecture when we think about the
Speaker A
decomposition of the complexity of the complete system, the memory optimization. I showed you three different layers of memory, the verification process. I showed you here my massive doubts on an AI machine verification, the adaptive retrievals bidirectional synergy that I showed you
Speaker A
at the beginning of the video, the repeated inference and the explicit recovery mechanisms that are also, of course, in place in this new methodology.
Speaker A
So, therefore, this is interesting. Different representation can have synergy, can have bidirectional synergy where the symbolic structure data, our graph representation or hypergraph representation, and the other representation that is a dense semantic text body, actively query and repair one another in
Speaker A
real time. I love the simplicity of the idea. I like the implementation of the idea.
Speaker A
The title that it is human cognition on graph structure, so human cognition on the eye machines.
Speaker A
Maybe not really, but read this paper. There are some beautiful technical insights. You can even go and develop your own mathematical representation of the ideas like I've shown you I started to show you. So, please because you need the
Speaker A
mathematics if you really want to write the code or you find a GitHub repo as I've shown you at the beginning of this video. Anyway, I hope you had a little bit of fun. There was some new information you have some new ideas that
Speaker A
you might try out here in the next days. Anyway, it would be great to see you in my next video.
Topics:runtime graph repairhuman cognitionAI benchmarksknowledge graphadaptive knowledge explorationphysics AI evaluationGPT-5.6Fable 5Gemini AIgraph reasoning











