Skip to content

Rich Sutton, The OaK Architecture: A Vision of SuperInt… — Transcript

Rich Sutton presents the OAK architecture, a vision for domain-general, experiential AI based on reinforcement learning and options.

Key Takeaways

  • AI development requires better reinforcement learning algorithms beyond current deep learning.
  • The OAK architecture proposes an experiential, domain-general approach to building intelligent agents.
  • Options are fundamental units consisting of policies and termination conditions for agent behavior.
  • Learning and knowledge acquisition should occur at runtime, not just during a training phase.
  • Understanding and modeling the mind conceptually simply is a central goal of AI research.

Summary

  • Rich Sutton introduces the OAK architecture as a vision for superintelligence derived from experience.
  • The talk emphasizes AI as a grand quest to understand intelligence and create powerful, domain-general agents.
  • Sutton critiques current learning algorithms as inadequate and stresses the need for better reinforcement learning methods.
  • OAK is based on the concepts of options (policy pairs) and knowledge learned at runtime rather than design time.
  • The architecture aims for domain generality, experiential learning from runtime, and open-ended sophistication in abstractions.
  • Sutton argues the path to strong AI runs through reinforcement learning, not solely through large language models.
  • The agent in OAK learns a high-level transition model enabling planning with larger jumps in the environment.
  • The design avoids embedding domain-specific knowledge, focusing instead on learning everything from experience.
  • Sutton highlights the importance of continual learning and the ability to discover hierarchical features and subproblems.
  • The talk situates OAK as a conceptual framework to understand minds and achieve the 'holy grail' of AI.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
[Applause] I had so much fun talking with so many of you here today. Um, and you know, we need a conference on reinforcement. It's clear. I wasn't sure, but now I think it's clear. This
00:16
Speaker A
was the right move. Um so I have prepared for you a talk uh on the oak architecture. It's a vision of super intelligence from experience. It really is my attempt to um address the issues at the center of of
00:38
Speaker A
was the right move. Um, so I have prepared for you a talk on the OAK architecture. It's a vision of superintelligence from experience. It really is my attempt to address the issues at the center of
00:53
Speaker A
Um so AI AI is a grand quest. Uh we're trying to understand how people work. We're trying to make people. We're trying to make ourselves powerful.
01:07
Speaker A
AI. And I want to start just by recognizing how difficult and important the task of AI is.
01:23
Speaker A
to the origin of life on the earth at least when we under when we when some part of the of the of the planet understands how it works and how it thinks and how it uh can be so
01:39
Speaker A
Um, so AI is a grand quest. We're trying to understand how people work. We're trying to make people. We're trying to make ourselves powerful.
01:58
Speaker A
just going to be good. Lots of people are worried about it. I think it's it's going to be good. It's an unalloyed good and I think the greatest advances are still ahead of us. It's a marathon.
02:10
Speaker A
Um, and this is a profound intellectual milestone. It's going to change everything. You know this, but it's good to take a moment and pause and recognize what we're doing is incredibly hard, incredibly important. As an intellectual milestone, I think it'll be comparable
02:25
Speaker A
The biggest bottleneck is strangely is we have inadequate learning algorithms. You may think we have our deep learning and we that's the one thing we know but I think it's not like that at all. I think it's it's more like we they are
02:38
Speaker A
to the origin of life on the earth, at least when some part of the planet understands how it works and how it thinks and how it can be so
02:45
Speaker A
Okay. So, and then myself uh I've tried to think deeply about this intelligence for half a century. Every day I'm sort of in the trenches designing algorithms, trying to design algorithms, seeking better algorithms for reinforcement learning, for learning
02:59
Speaker A
transformative to our planet. Okay. But it's also a continuation of things that we've always done, and it's just the next big step. So now myself, I
03:12
Speaker A
architecture called oak and I think it provides a line of sight towards towards our grand prize of understanding mind.
03:22
Speaker A
think this is just going to be good. Lots of people are worried about it. I think it's going to be good. It's an unalloyed good, and I think the greatest advances are still ahead of us. It's a marathon.
03:32
Speaker A
It comes from the idea of options and knowledge. Now, as many of you are very familiar with, an option is a pair.
03:40
Speaker A
I think, you know, good for this group that the path to full AI, strong AI, runs through reinforcement learning and not, I think, through things like non-experiential things like large language models.
03:52
Speaker A
a way of deciding to stop behaving in the Okay, so just those two, just a pair. And in oak the agent has lots of options and it's going to learn its knowledge will be about what happens when you follow the option. So in this
04:10
Speaker A
The biggest bottleneck is, strangely, we have inadequate learning algorithms. You may think we have deep learning, and that's the one thing we know, but I think it's not like that at all. I think our algorithms are very crude. They need to be better, and that is what we should be working on.
04:20
Speaker A
Um yeah so that's that's uh where the name is from. I think it's is a grand challenge and a grand quest. And so I show it like this that we are seeking the holy grail. The holy grail of AI.
04:36
Speaker A
Okay. So, and then myself, I've tried to think deeply about this intelligence for half a century. Every day I'm sort of in the trenches designing algorithms, trying to design algorithms, seeking better algorithms for reinforcement learning, for learning
04:56
Speaker A
more in a minute. But let me just put down my three main uh design goals. It should be domain general. It should be experential. That is the mind should grow from runtime experience not from a special training phase.
05:12
Speaker A
from experience. And I follow this Alberta plan for AI research, which you may know about. Mike and Patrick and I did it a couple years ago. And today I'm going to talk about the vision of an overall agent architecture, AI agent
05:28
Speaker A
computational resources. So those are the three main deterata and we'll talk about them. I guess first I want to establish some words. I want to talk about I'm going to talk about design time and runtime. Design time and
05:43
Speaker A
architecture called OAK, and I think it provides a line of sight towards our grand prize of understanding mind.
05:56
Speaker A
time we would be building in any domain knowledge that we might have farther away from the chest might be just pardon me.
06:19
Speaker A
Okay. So those are my introductory comments. Now let's talk about OAK. I think it's fun just to start with the name, which may be a mystery: OAK.
06:34
Speaker A
in. So like uh I don't know a large language model everything is done as design time and when it goes out to be used in the world it doesn't do anything. My emphasis is going to be the other way around. We want I want to do
06:48
Speaker A
It comes from the idea of options and knowledge. Now, as many of you are very familiar with, an option is a pair.
07:04
Speaker A
things in it. You're not going to be have your agent know everything about the world in the factory because the world is huge and you can't know everything. Your a your robot is small compared to the world. And uh so if you
07:17
Speaker A
Well, actually, you may think it's a triple. I sort of have dropped the initiation set in my work for many last couple decades. So, for me, it's a pair of a policy, a way of behaving, and
07:32
Speaker A
I'm going to talk about this a lot. runtime. Everything has to be done at runtime and why that's actually a good way to think about it. Okay, so let's just ask the question. Yeah, should an agent's design this is a question for you guys, right?
07:50
Speaker A
a way of deciding to stop behaving in the OAK. So just those two, just a pair. And in OAK, the agent has lots of options, and it's going to learn its knowledge about what happens when you follow the option. So in this
08:09
Speaker A
I mean both answers are right. So you can um yeah if you want something to uh perform well and you want to put it out there and have it do something good right away, you know, you want to you
08:21
Speaker A
way, the agent is meant to learn a high-level transition model of the world that enables planning with larger jumps and hopefully carves the world at its joints.
08:26
Speaker A
But if you want a good design, so I'm gonna say no. My quest, my quest is that the design should not depend on the world at all. Okay. Uh it should be domain general. Now this is really just you know there there
08:44
Speaker A
Um, yeah, so that's where the name is from. I think it's a grand challenge and a grand quest. And so I show it like this: that we are seeking the holy grail, the holy grail of AI.
08:58
Speaker A
but also important is let's understand the mind in a simple way let's let's what we what we would want is a a conceptually simple understanding of what's going on inside a mind. That's that is in some sense the grand quest of
09:18
Speaker A
Let me put that up in a way that's easier to read. We want this AI design that is domain general, contains, I'm going to say, nothing specific to the world. We want a general idea. So I'll talk about that
09:33
Speaker A
so your goal your job your agent's job is to go out there and learn all the fiddly wonky little details and highle structures of the world that it's going to encounter. But you really don't want to understand what it's doing at that in
09:48
Speaker A
more in a minute. But let me just put down my three main design goals. It should be domain general. It should be experiential. That is, the mind should grow from runtime experience, not from a special training phase.
10:02
Speaker A
So it's sort of different purposes. Someone who wants to go have something that will perform really well right away or someone who wants to have a conceptual understanding of what a mind is and how intelligence works. So my
10:13
Speaker A
And third, it should be open-ended in its sophistication, in its abstractions. So that it can form any concepts in its mind that are needed to deal with whatever world it's connected to, limited only by its
10:29
Speaker A
The actual contents of minds are part of the arbitrary intrinsically complex outside world. They are not what should be built in as their complexity is endless genuinely endless. our our agent has a big computer and we don't want to
10:48
Speaker A
computational resources. So those are the three main desiderata, and we'll talk about them. I guess first I want to establish some words. I want to talk about, I'm going to talk about design time and runtime. Design time and
10:59
Speaker A
In short, we want agents that can discover like we can, not which contain what we have already discovered.
11:08
Speaker A
runtime. When you're in the factory being designed, then your robot goes out and lives its world. That's its runtime, sort of a lot the way Leslie spoke. So the agent is designed at design
11:25
Speaker A
Okay. Now second question, should the agent learn from special training data or should it learn only from runtime experience?
11:42
Speaker A
time. We would be building in any domain knowledge that we might have farther away from the chest might be just, pardon me.
11:53
Speaker A
Well, for me the agent should learn only from runtime experience should be entirely experential.
12:01
Speaker A
Okay, cool. At design time is when you're building in any of your domain knowledge, and then at runtime is when you're actually interacting with the world, learning from experience, making plans that are specific to the part of the world you're
12:21
Speaker A
the argument is going to be there's certain things that that have to be done at runtime and and but and could be done at design time and uh so basically everything has to be done at runtime maybe also at design time but it ha at
12:37
Speaker A
in. So like, I don't know, a large language model, everything is done at design time, and when it goes out to be used in the world, it doesn't do anything. My emphasis is going to be the other way around. We want, I want to do
12:48
Speaker A
those things have to be done at runtime because you're going to encounter the world. The world might not be like what you expected and certainly you won't know all the intricate details and all the abstractions that are needed for for
13:00
Speaker A
all the important things at runtime, online, on the job. And in a minute, I'll be talking about a big world, but the idea of a big world, a big complex world, is that you're not going to be able to build
13:10
Speaker A
speed you up, but since they all have to you have to be capable of being done at at runtime, why not make a conceptually simple design and just doesn't worry about trying to get a head start on the
13:22
Speaker A
things in it. You're not going to have your agent know everything about the world in the factory because the world is huge and you can't know everything. Your robot is small compared to the world. And so if you
13:38
Speaker A
talked about the big world perspective. So let's let's do that. Let's talk let's make let's gain a common knowledge amongst us of what that means. The big world perspective or the big world hypothesis something that's been floating around at
13:54
Speaker A
want to be able to learn arbitrary open-ended abstractions, you need the world to find out whether the right abstractions for the part of the world you're running into. So you got to do it at runtime. So I don't know. So
14:10
Speaker A
bigger really. It's bigger than this. It's it's big, you know, it's really big. And it's got to be much much bigger than the agent because the world contains, you know, billions of other agents and um and all of course all the atoms and
14:25
Speaker A
I'm going to talk about this a lot. Runtime. Everything has to be done at runtime, and why that's actually a good way to think about it. Okay, so let's just ask the question. Yeah, should an agent's design, this is a question for you guys, right?
14:39
Speaker A
friends and your loved ones and your enemies. Um all all those things are important to you and they have to be taken into consideration. And so you know the upshot is that nothing you the agent is going to be doing is going to be exact
14:55
Speaker A
Should the agent's design reflect the world in which it's expected to be used? Good thing about this question is both answers are wrong.
15:07
Speaker A
transition models will be much enormous enormous reductions. They're sitting inside your head. You know, your model of the world and the world is out there much much bigger, right? Um even a single state of the world, you can never hold it in your head. You can
15:26
Speaker A
I mean, both answers are right. So you can, yeah, if you want something to perform well and you want to put it out there and have it do something good right away, you know, you want to
15:41
Speaker A
Uh, and I'm citing a little paper on this that Dave Silver and and a coupe and I wrote where we just made this point. The world, you can probably see it. the if you don't have a model, you
15:57
Speaker A
put design domain knowledge in there. You want it to reflect the world.
16:12
Speaker A
you have to learn at runtime because you're going to you can't you can't have built in everything about the the whole big world at design time. You have to learn at runtime. you're going to encounter some particular part of the of
16:24
Speaker A
But if you want a good design, so I'm going to say no. My quest, my quest is that the design should not depend on the world at all. Okay. It should be domain general. Now this is really just, you know, there are multiple questions and there are multiple uses, and like, you know, I have to respect totally, you know, someone who wants to do an application, make something that's useful, make something that performs well. Sure, that's important,
16:38
Speaker A
it meets its co-workers. It has to remember the name of the guy it's working with. You know, that was not in the domain knowledge. It has to remember uh the work that they've done on that project. What's working well? What's not
16:47
Speaker A
but also important is let's understand the mind in a simple way. Let's, let's what we would want is a conceptually simple understanding of what's going on inside a mind. That's, that is in some sense the grand quest of
17:05
Speaker A
to learn during your life. You have to plan during your life. I'm Leslie Caling made this point that planning is required because the world is big.
17:17
Speaker A
AI: to understand what it means to achieve goals, what it means to understand an arbitrary world. It should be simple, and if you, the world, the actual domain that you're going to interact with is arbitrarily complex, and
17:31
Speaker A
know, what's the right joints and ideas that are involved in the world that we're encountering.
17:39
Speaker A
so your goal, your job, your agent's job is to go out there and learn all the fiddly wonky little details and high-level structures of the world that it's going to encounter. But you really don't want to understand what it's doing at that in
17:56
Speaker A
abstractions at runtime. And so, if you're going to have to create them at runtime, you know, why don't we just do it in one place and and and uh that'll be a very good start. You know, I think of a design for an AI, the
18:12
Speaker A
terms of all those domain-specific details. You want a level of understanding that is at a higher level and is based on principles and not the intricate complex details that are in the world.
18:28
Speaker A
pseudo code, you know, maybe it's three slides. Okay. I thought I think something of that that order.
18:35
Speaker A
So it's so
18:51
Speaker A
Okay. Um yeah, so this talk like hey um I was up till like the like four o'clock last night making making these slides.
19:07
Speaker A
Um this is this is um this is sort of a new talk for me. This is the first time you're the you're the first guys hearing it. I've been I've been traveling around the world and giving all these uh philosophical talks
19:20
Speaker A
and political talks and point of view talks and that's great and I've enjoyed that but but uh you know I really thought here I'm I'm at the RLC conference reinforcement learning I should do something uh substantive maybe even technical
19:37
Speaker A
um so so I um and then all during the week you know I went to every I went to all these talks I talked to all these people And uh I I keep just changing my mind about what I
19:48
Speaker A
want to say. And uh so anyway, this is all my way of of excusing myself or explaining to you that that this talk is new and it's not quite polished. In in particular, I might say some things more
20:02
Speaker A
than once. Um so and maybe it's okay to have some repetition if they're important things. Um okay, but just be aware of that. and maybe maybe kick me if I if I uh need to move on to the next
20:16
Speaker A
thing. Okay, I'll try to be quick. Runtime learning I think always wins over design time because the the world is much bigger than the agent, the big world perspective.
20:27
Speaker A
Um design time can't cover every case. Runtime learning can customize to the part of the world actually encountered.
20:34
Speaker A
Runtime learning scales with available compute whereas design time learning or anything done at design time scales with the available human expertise at design time. It was the only thing available and historically scaling with compute wins in the long run. That's what the
20:52
Speaker A
bitter lesson is is explicitly about. However, today's deep learning methods, runtime deep learning methods, continual learning, they don't work very well.
21:04
Speaker A
Okay, this is this is a a a big bitter thing for me. I wish they could worked well because I'm talking all about runtime learning and I want to use it.
21:16
Speaker A
Um yeah, if we if we one last thing about runtime learning, it does enable metalarning.
21:26
Speaker A
Metal metal learning is where you like try learning one way and then you try learning another way and you notice that oh this way works better in the future I will do this. If you were if you were
21:36
Speaker A
doing everything in one shot you couldn't do that. This idea of becoming better at learning requires uh one time you're doing learning another time you're trying a different way of learning and you pick the better one. So metalarning really requires this to be
21:53
Speaker A
done at the at the runtime. Okay. Now let's think about the problem just a little bit more. You know I like to separate things into the problem and the solution. Really almost everything I've been talking to you about is the the
22:08
Speaker A
problem the quest. What are our goals? What are our desitter? So um one more slide on that. The AI problem is to design an effective purposeful agent that acts in the world.
22:24
Speaker A
And the classic reinforcement problem is the same thing except we add the purpose is specified by scalar reward signal.
22:32
Speaker A
The reward and the world is general and incompletely known. But the world can be anything could be grid world or the human world can be stochastic, complex, nonlinear non-marov.
22:45
Speaker A
The state space of the big world is effectively infinite and its dynamics are effectively non-stationary.
22:54
Speaker A
Um let me go ahead to just talk about this one a little bit further. The purpose is specified by a scalar signal.
23:04
Speaker A
So that's we have a name for this idea. It's called the reward hypothesis. And I I wanted to bring this up because we have thought about it and uh it's not like a quick choice without intention.
23:18
Speaker A
Um so the reward hypothesis is this that all of what we mean by goals and purposes can be well thought of as the maximization of the expected value of the cumulative sum of a received scalar signal called reward.
23:36
Speaker A
Um, so there's lots of specific things there like the expectation, like the cumulative sum. Um, yeah, and that's been thought through. And the SC the idea of a scaler reward, I just want to say it's not just a uh something we
23:51
Speaker A
haven't thought about. In fact, it's a it's a great thing. It's a really clear way to specify the goal. It's become popular in many different disciplines, not just AI, but also economics and psychology, control theory, and forever people have been trying to modify.
24:07
Speaker A
They've been trying to add things like constraints, multiple objectives, risk sensitivity, and and and I I I don't know. I I hate that.
24:23
Speaker A
I mean, even if it was good to do, I I I I don't want it to be done. I don't want it to be true. I I you I think you've already gotten the sense. I like things to be simple. You know, that's like a
24:35
Speaker A
really high high uh uh deserata desire. I want things to be simple and I might even simplify them a little bit too far in order to uh to uh be clear and uh so anyway, I want things to be
24:50
Speaker A
simple. Do we need to do we need all these things to get generality? That's the real question. and Michael Bowling and others have have written this really nice paper called settling the reward hypothesis where they go through all
25:03
Speaker A
these cases and and I don't know if I want to uh they they they establish that in a certain sense the reward hypothesis is correct that that you don't um add generality um uh by adding multiple objectives or risk sensitivity or any of
25:24
Speaker A
these constraints. So they it's one way of validating that choice and you might also probably know the reward is enough paper where we argue that even a simple reward can lead to all the attributes of intelligence in a sufficiently complex world.
25:40
Speaker A
Okay. So um now I want to talk about the solution methods the architectures obvious starting place is model free reinforcement learning basic reinforcement learning where the agent constructs a policy and value function at runtime both these are functions all all these
26:02
Speaker A
are runtime RL architectures okay the model free and then you can handle the non-markoff case if you uh construct your feature uh state representation from u uh from from your data and I'll show you the the picture of that in a
26:20
Speaker A
second. Uh but still better uh would be to make a model of the world and use that model to plan with potentially better. Now the oak architecture is along the same line um of improvement of of extension and the the thing about the
26:38
Speaker A
oak architecture is it it adds to those things u auxiliary problems subpros and those sub problems are in the form of attaining individual features individual state features and in this way uh we enable the discovery of higher and higher uh levels of
26:56
Speaker A
abstraction and we we achieve this open-ended uh goal. Okay. So, as a in a picture, this this picture on the on the left is from uh the textbook, the reinforcement learning textbook. Um and this is the very last figure in the book. And so
27:15
Speaker A
maybe it's familiar to some of you. Uh we have uh the world and then the agent is everything above the world. It's the policy and value functions. It's a model of the world. It's a planner. And it's the Ubox. The viewbox. I want you to
27:29
Speaker A
notice because this is a is a um kind of reinforcement learning where we don't assume that the state is available to the agent. We only observations are available to the agent. This we send actions to the world. The world sends
27:43
Speaker A
back observations and and a reward. Okay. And then there's a a process here that that's a part of the agent uh that that computes uh something that we'll use as a state representation uh by the by the policy and value function. Okay.
28:00
Speaker A
So that's the construction process and um nowadays I draw um things uh a bit differently. Um and this the right figure is the full oak architecture and you see the many of the same components.
28:17
Speaker A
This U box is now called perception. Perception is a better name for it because what does it do? This perception process it takes in the uh the data that's happening the actions and the observations and it forms a a sense of
28:32
Speaker A
where the agent is now. That's really what perception is about. take in your sensory input, get a sense of where you are now. Use that sense to make your decisions to to as input to your policies, your value functions and your
28:44
Speaker A
models. Okay, so the oak architecture has all those things, but it also adds auxiliary subpros and those each subpro will have its own value function and its own policy. Uh so that's what's suggested by having uh these u shadow policies behind
29:03
Speaker A
the the main policy and these secondary auxiliary value functions behind the main value function. Um also each one of these subpros is going to be based on a different component of the state feature representation. So this is the thing that acts like state
29:22
Speaker A
and I want you to think of it as a feature vector and each one of these sub problems is going to be based on a different component of that feature vector. So that's what it looks like as a picture. Okay. Uh now we're going to
29:38
Speaker A
get into the a bit of the nitty-gritty. We're going to look at it. I've given all these these quick introductions to the idea tell you I've told you various properties of the oak architecture. Now I want to tell you exactly what it is or
29:51
Speaker A
it is these these what eight steps done in parallel at runtime. Okay, it's a lot of steps. Um we're I'm going to come back to this slide a bunch of times uh and develop each part of it.
30:06
Speaker A
So just relax and let's get started. Let's just learn what some of the the the characters are here and how they interrelate. So we might start with the first line. Um and this line uh is learning the policy and the value
30:23
Speaker A
function for maximizing rewards. That's like normal reinforcement learning. And I think that is is is almost done. If it was done, there would be a green check mark, but it's blue. And blue means it would be done if we could do this
30:38
Speaker A
continual deep reinforcement deep learning thing with metalarning. You know if we really could do continual learning that would be uh that would be done. And so we can do this with in very simple simplified cases. We can deal it
30:52
Speaker A
with uh some of the algorithms like continual backdrop. We can deal with um the linear case. So this is gets a blue check mark conceptually done but um really it's waiting to be done well. it's waiting to uh solve this
31:11
Speaker A
problem of continual learning for deep learning. Um now the next one uh is red because we don't really have a solution. We have lots of ideas but we don't have a specific proposal and so I'm going to come back to this later. The second one
31:28
Speaker A
is generating new state features from the existing features. And let me go over this next bit uh quicker. Let me just run all the way down verbally through all eight steps. We we we're going to have some features. We're going
31:42
Speaker A
to order the features. We're going to take the highest ranked features, the most important features according to our estimation, and we're going to create subpros of of achieving them. So if I decide that being in this lecture room
31:55
Speaker A
is a is an important sub goal, I would make a a sub problem for that would be rewarded when I when I when I not would would be would would succeed when I am here. If I think uh holding this
32:07
Speaker A
microphone up sufficiently close to my mouth is a good sub goal. Um, I would I would uh I would I would make that feature into a a good subpro and so on for finding the restroom and finding the
32:21
Speaker A
the the uh the coffee. Coffee is a really good you know feature for that flowing into your mouth and getting all the sensations involved in that. So that feature becomes a subpro for attaining it and then you learn solutions. So this
32:40
Speaker A
is the heart of of um the heart of the oak architecture is to have sub problems. You learn the solutions of the sub problems. The sub problems are the options from the from from the O and oak. And uh we also have to learn the
32:55
Speaker A
value functions that are associated with the sub problem. Okay. And then we're going to have these options. And the next step is we're going to have we're going to learn models of the option. We want to know uh what will happen if you
33:10
Speaker A
if you were to execute any any one of them. This will be part of your model of the world but it will be a high level model of the world because it'll be about uh an extended way of behaving
33:20
Speaker A
rather than about a single action. These models will enable you to plan and then and then you so those are all the the full main steps. Okay. So you've seen it once. uh you're going to have to maintain metadata on on the utility of
33:38
Speaker A
everything and and curate uh throw some things out um and propose new ones. Okay. So now we're going to go through these steps and for a while I'm going to spend a lot of time um actually on the
33:52
Speaker A
fourth one but yeah the the ordering the features it seems kind of easy but but we can't do it until we have all the rest all the other pieces done that a couple cases that will h that will
34:04
Speaker A
happen. Okay let's spend some time about the creation of the subpros one for each highly ranked feature.
34:13
Speaker A
Okay. So on acknowledge there's a long history of looking at subpros that are distinct from the main problem. People talk about curiosity, intrinsic motivation, auxiliary tasks. Some things are settled, some things are unsettled.
34:29
Speaker A
You don't have to read all this. I want to direct your question to the red part.
34:32
Speaker A
The key open questions about subpros, which are what should the subpros be? Where do they come from? How can the agent or can the agent generate its own subpros and how do the sub problems help on the main problem? So the contribution
34:49
Speaker A
of oak is to an is to propose answers to all these questions and to really uh answer questions like uh the third one how can the agent make its own subpros in the affirmative and thus get open-ended abstraction.
35:04
Speaker A
Um so I like to think of it very basically that we have problems and solutions and these interact with each other. We propose a problem to work on.
35:15
Speaker A
We work on it. We solve it. As as a part of solving it we will make new features and those features will then be the basis for new sub problems and and and then the sub problems will have to be solved
35:29
Speaker A
new features and so on in an endless cycle. roughly. That's what I'm talking about. And I wanted to give some examples from nature. Uh here's an orang ba a young orangutang um playing swinging. And so I What is he
35:45
Speaker A
doing? Like he's not getting food. He's just interested in what it feels like when he swings.
35:55
Speaker A
Yeah. Okay. That's what I think he's doing. I think it's that sensation is interesting. and and he got it once and now he's trying to get it again and understand how to control it. Um yeah, so here's and also on the other
36:09
Speaker A
slide we have an orca uh who somebody threw this big uh I don't know what to call it right now.
36:25
Speaker A
A buoy into his into his pen and he's decided trying to figure out what he could do with it. and he's managed to get it up on his back. So, that was not not random. Um, he got this idea and now
36:37
Speaker A
he's perfecting it. Yeah. So, animals play, people's play, uh, infants play, young people play. Um, so this is a sped up video of a infant playing. And, uh, this is what we want. We want the the way the the child goes from object to
37:02
Speaker A
object, learns a little bit about it, gets bored, moves on to the the next object, and just gradually develops a better and better understanding. Maybe next time when he comes back to the to an object, he'll he'll he'll uh have an
37:15
Speaker A
increased ability, be able to do new things with it. Um, this is what we want. And so I'm trying to think about them as posing sub problems for themselves, things to learn about, things to understand, things to predict,
37:27
Speaker A
and uh things to control and and figure out where it can make progress in in learning solving the sub problems. Okay, so maybe you're ready to accept this statement. The agent must create its own sub problems. Sub problems can't be
37:44
Speaker A
given to you. Can't be given to you at design time. You've got to create your own that's far too various and world dependent to have been built in. We have to give the responsibility of the the questions the problems not the solutions
38:01
Speaker A
not the features the questions okay what is the pro what is the how can we how can we um do this I mean we have much of the machinery the machine of options general value functions off policy learning planning methods these are
38:18
Speaker A
machinery to help us in this process but we want to create them in a domain independent way and that's challenging Okay. So I want to offer this this possible way to make subpros in a totally domain independent way which is
38:34
Speaker A
um when you come across a feature when you make up a new feature or you experience a new feature you can make it to be uh the basis of a sub problem uh I call it a reward respecting subpros of
38:46
Speaker A
feature attainment. So let me show you exactly what that is. um how do we create a subpro from a feature? So a feature is like yeah feature eye. It's it's a bright light that you saw once. It's a it's an
39:03
Speaker A
interesting sound that happened. It's you you're a baby and you heard the rattle make a sound. You'd like to reproduce that sound. So you have a feature I a feature index I and you have kappa which is how intensely you want
39:16
Speaker A
that feature. you have to express that uh and you'll get different um subpros if you if you want it at all costs or if you just kind of want it a little bit.
39:28
Speaker A
So the sub problem is to drive the world to a state where the feature is high without losing too much reward because you will lose some reward if you're not doing what you normally do because what you normally do is uh maximize reward.
39:42
Speaker A
There is just one reward by the way and so I don't have to qualify that is the real reward and the sub problem is to achieve a state where the feature is high uh without losing too much reward
39:57
Speaker A
without having to go through something that's painful or having lost opportunities to get something uh pleasurable.
40:06
Speaker A
Okay, so we're trying to find an option. An option is a pair. It's a policy pi termination function gamma that maximizes the value of the i feature at termination while respecting the rewards and values.
40:18
Speaker A
So here's the equation. Maybe we can understand the equation. Uh in each state you're trying to choose you're trying to maximize and you're going to choose pi and gamma to maximize. And it's the sum the sum is conditional on
40:31
Speaker A
starting uh the world in in the indicated state because um yeah for each for each state we say if we started there have a policy pi that gets you rewards um from t+1 to t. T is is the time of
40:56
Speaker A
termination. You're going to follow the option pi and you're going to terminate when gamma says terminate. And so that will fix establish the random variable which is the time which you terminate capital t. And if you look at all the
41:10
Speaker A
rewards you receive while you are following the option summing them up. Those are the things that you want to be as big as possible or not as least negative as possible.
41:21
Speaker A
And then you want to um reward yourself for achieving feature I's feature I at time of termination S sub capital T. And you know there's a there's a waiting by kappa. So you want lots of rewards. You want to be the feature to be true but
41:44
Speaker A
you know it's got to be traded off the rewards. And you also care about the state that you're in at the time of termination.
41:52
Speaker A
You don't want to like find a really good way to to uh I don't know get some coffee but has the consequence that you have to break your leg. Okay.
42:05
Speaker A
Actually, that would be a reward. That would be a bad reward. You don't want a way to get coffee that would um uh leave you in a bad state. Like let's say you got coffee but uh you know you're gonna
42:16
Speaker A
get arrested or you're going to fall down the stairs. Okay. Um those are bad states. You know I often you know the walk along the edge of a cliff but uh don't fall off. If you fall off, it
42:33
Speaker A
actually doesn't hurt very much to fall off because, you know, if you can say just terminate while you're in the air, uh the rewards are fine, but the value the value is, you know, bad rewards are coming up. So, your value will be will
42:45
Speaker A
be poor. Okay, that's how we create a sub problem. And now let's really we're getting into the heart. Uh we have these processes. Um the f the first one is we form the uh the problem just as we just
43:03
Speaker A
talked about you know given a feature form a problem. We do that with all the high highly ranked features. So we have now we have you know dozens of pro of problems. Each one we work on it to
43:14
Speaker A
produce an option. The solution to sub problem is an option. Now you've got this options. Well that defines a correct transition model for the option.
43:25
Speaker A
Is it well defined? Um what we know what the model should be. If you give me the way of behaving and the way of terminating, we know what the model should be. And so this is something you work on. You work on
43:38
Speaker A
computing this model, approximating that model. Once you have the model and that you have a models of all the different uh options for all the different subpros for all the different features and once you have the model, you use the model of
43:52
Speaker A
course to plan and to improve your behavior. Okay. So we got these three steps and there's one fourth step which is that uh you have to come up with features right we we we started with a good set uh with the highly ranked
44:05
Speaker A
features so we have to have a way of ranking the features and um and I just want to point out that we have that because all of these these these three uh pillars the the later three all use
44:18
Speaker A
features to to do their job right if you're going to find the option the option is a function of state and so you have to look at state features when when should you do the option when when it
44:28
Speaker A
when is the value function of the option uh when should it have which values it's going to look at the state features to make those decisions when you uh learn the models of the options you're going to look at the state you start in and
44:40
Speaker A
you're going to look at the state features of that state and you're going to say oh that feature I found useful that other feature was useless to me um and so and and then when you use the models you will find some models are
44:54
Speaker A
useful And that will sort of trickle back to evaluate the choice of the of the options. And that will also trickle back to evaluate the choice of the feature attainment problems. And at least uh all these learning processes,
45:08
Speaker A
predictive learning processes will use the features and they will provide feedback to the uh to the features saying these are the ones that have proven useful to us. These ones have not. Okay. So let's draw that the same
45:21
Speaker A
idea with a different picture. I'm going to have many pictures about the same idea and I'm going to say it a few times so maybe you'll you'll get it. Um this way of talking and thinking. So the perception process is going to is
45:35
Speaker A
responsible for constructing interesting state features. Um the play process or the problem posing process problem posing and solving. That's where you do the sort of core reinforcement learning things of figuring out your value functions and your policies and you
45:51
Speaker A
produce the options. And then you have to predict the consequences of those options to form a transition model. And then of course you plan with the transition model to get improved policies and values. And the feed we close the the cycle is we have
46:06
Speaker A
feedback from the later steps back to the construction of features. And that feedback is mainly say saying I have I have found that feature useful or I have not found that feature useful.
46:19
Speaker A
Okay. So here's we're back to our eight steps. U we now understand what it means to create the subpros one for each feature and we also know what it means we've talked about how you learn the solutions um and the transition models. Um maybe
46:38
Speaker A
there's I'll say one more slide about that about this this topic uh how we learn those things. Oh, but also notice they're they're in blue because although we know how to do these things, we don't really know how to do them with
46:52
Speaker A
continual deep learning or maybe we have to use Shbanch's continual backdrop. You know that this is this is this is a topic we'll come back to. This is incompletely understood. So we kind of know how to do them but we definitely
47:05
Speaker A
think we can do better. Okay, one one short slide more about that is it to a large extent we could use just standard off-the-shelf algorithms standard offtheshelf usually off policy algorithms for learning general value functions like GTD and
47:22
Speaker A
emphatic TDD and retrace ABQ uh these are prediction learning methods for generic uh GVFs and so we can use that to learn the main problem how to get reward we can learn that for learning diagrams for the sub problems. We can
47:38
Speaker A
find the transition models of the options with these methods and the planning can also be done with standard algorithms applicable to all GDFs. And this enables us to say that anything that can be learned can also be planned.
47:52
Speaker A
That's I just wanted to get to that slogan for you because it's it's a it's a good one. It's a it's a um it's a bit advanced but almost it's almost a next step.
48:03
Speaker A
Okay, so we got those things and now the other big step and uh I'm going to have to do it a little bit uh not in full detail, but we have to talk about how the planning works. Okay, how does the planning going
48:18
Speaker A
to work? And I'm going to give that a green check mark because I think we do understand this. So planning, why do we plan? Why do we want these jumpy uh temporally extended models of the world, the option models? And we we but
48:30
Speaker A
basically why do we want to plan at all? We want to plan because the world changed and the correct values change and it's easier in many cases not in every case but in many cases it's easier to get the model of the world right than
48:40
Speaker A
to get the values right. So you get the model right and then you do the planning uh to make the values consistent with your model. Um and uh uh so in this big world setting it's it's the world
48:56
Speaker A
changes or appears to change. Uh most of the world's dynamics or in many cases the world's dynamics including the reward parts don't really change but the values nevertheless change. Like it's always true that I can walk over there
49:08
Speaker A
and find the restroom but it's not always true that I want to go to the restroom. Um it's not always true that I want to get coffee. it's not always true that I want to go to the library all the
49:18
Speaker A
things or go to Edmonton. So u to prepare for these later wants um these these different values you plan and u this also has some implications for which sub problem is useful. Okay. Now how does planning work? Um I like to
49:37
Speaker A
think that planning is is by approximations to value iteration. So, so this this equation is value iteration. You may already know it, but I think maybe you should look first um here. What is the model? Model is something that that takes a state a
49:57
Speaker A
low-level model takes a state and an action gives you um a probability distribution over next states and the expected reward along the way. And so value iteration then says I'm trying to improve the values of some states. I
50:11
Speaker A
look at the possible actions and I'm going to maximize over them. I'm going to look at the immediate reward and I'm going to discount and then take the probability or the really the expected value of of the value of the next state.
50:25
Speaker A
So this is the probability of each next state. You wait by that you take the value of the next state. So you know you probably are are familiar value iteration works like that. Um and really all all planning methods in some sense
50:39
Speaker A
work like this just apply to different states in in the search tree and you know order which states are updated in such a way uh has is is is varied.
50:55
Speaker A
Okay. And um the interesting thing about planning with option models is that it's the same. It's really it's really the same.
51:05
Speaker A
Although life is lived one step at a time, we have to plan it at a higher level. So our knowledge of the world should be about the large scale dynamics. It should not be conditional on single actions but on a sustained way
51:19
Speaker A
of acting that is on an option. So if we look at the conventional model it's we receives an action an option model you receive an option still you get a probability distribution perhaps over the next states and expected reward not not a
51:36
Speaker A
onestep reward but some reward while you're following the option and then then value duration is almost unchanged.
51:43
Speaker A
We just change the actions into options and we still talk about the reward for following that option and the probability of each next state under the option. Okay, so that's good. Um that's how that gives you a flavor of how
51:58
Speaker A
planning would be done. You you would you would you you basically you say oh here's some state. What are the things I could do there? What's the best I could do? I'll update my value and here's I could imagine another state. Maybe it's
52:10
Speaker A
the state I'm in. Maybe it's not the state I'm in. But I go through this outer loop of of considering various states and then maxing over the possibilities and uh doing this equation.
52:22
Speaker A
But I'm sure you are concerned because I've been talking about via s which is the value of an individual state and uh we can't do that of course we have to have function approximation.
52:35
Speaker A
So I don't know I don't think it's that helpful to go through these equations but yeah the value V of S will become the approximate value of a state given a parameter vector W and also your model of the world will become R hat and P hat
52:53
Speaker A
they also become parametric and then after you've done that um things are much the same there is more complications which I will skip having to do with what's computationally expensive. You can ask me about that if you want. Um, if you're interested, I'll
53:10
Speaker A
say some summary things. Um, so, so we we're maybe we're almost done, but remember I said we would come back to some things. So, I said we would come back to um the first two learning how we're going to learn these things and
53:30
Speaker A
what's we the problems we have. I want to say explicitly what what the the situation is there. So two more slides.
53:38
Speaker A
Um so the oak architecture requires reliable continual learning, continual deep learning. And so this is one of those things that I said this would be we'd be all done with step one uh if we uh could do continual learning. And but
53:56
Speaker A
we can't do we can't can we do this yet? Okay. Can we do this? Okay. Do we have reliable continual learning? Well, we do have reliable continual learning for the linear case, for the tabular case, but for the nonlinear case, for the deep
54:13
Speaker A
learning case, we can't have reliable we don't yet. Do we? No, we sort of do. We have uh we have we have these catastrophic failures like u like uh catastrophic forgetting and the catastrophic loss of plasticity uh that have
54:33
Speaker A
been figured out long ago and also very recently. Um so we have these catastrophic problems but we also looks like there's a range of solution methods. So I guess this is an area that's that's right now in flux
54:47
Speaker A
and uh people are figuring things out and you know we're not there I can't give it a green check okay but um but there are lots of ideas continual backrop is one of them the metalarning of new features I think can also help
55:04
Speaker A
and related to that it's the other the other problematic uh or the other step that I wanted to talk about where we having to do with um creating generating new state features. This is this is also a super old problems going back to the
55:20
Speaker A
1960s like Minsky and Selfridge would talk about this. They would talk about representation learning. They would talk about the new terms problem and anyway I like to talk about metalarning nowadays.
55:31
Speaker A
So backrop back in ' 86 was supposed to solve this. We were supposed to you know learning representations by gradient descent. Um but it really just doesn't.
55:41
Speaker A
And um I think we're we we accept we recognize that unless we are still in love and think that gradient descent is enough for everything. Um most of the other methods other than gradient descent are based on generate
55:58
Speaker A
and test ideas where you like generate a bunch of features and then you test them to see if they're useful and so you could generate them randomly and you could test by their utility and u continual backdrop is an instance of
56:10
Speaker A
that. It also has a real old history. Leslie Keelbing did her uh did some of this work in her PhD thesis in 1993. Uh Rupam Mahmud and I did some work a decade ago. Um there's lots of ideas but
56:28
Speaker A
not yet a specific proposal for a whole network based on grad descent or any other uh there aren't I anyway don't have a specific proposal uh for solve this problem. I think it's a really really important problem and I think it's like
56:42
Speaker A
maybe it will be worked out in the next couple of years and then it will literally um take over everything that people have done with with deep learning. If we had a deep learning method that was can do everything we're
56:54
Speaker A
doing now but can also learn continually that would just um be a really big thing and I think there's no reason why it couldn't happen. Um and so I think it will I also think that something like uh my algorithm called IDBID or IDBD and
57:15
Speaker A
it's really old will be a key part of that uh solving this problem. Okay.
57:24
Speaker A
Um I'm sort of done. This this figure was just to um remind you one more time of the cyclical nature how we have state features that produce sub problems that are solved produce options that are that are used to form models and and this is
57:47
Speaker A
not doesn't look like a cycle but remembering that um there's feedback being sent from each user of the features uh information information on which features are are useful which ones are that informs the features and so it is
58:02
Speaker A
in fact a cycle. Um so in my the quest uh have we uh succeeded we have something that's doesn't it's totally domain general there's nothing in it specific to any world it's totally experential and has the claim or the hope that it will be uh
58:20
Speaker A
you know able to find unlimited open-ended abstractions uh limited only by the computational resources and so we might ask you know what do you what what should one think about this and so I think arguably reinforce Enforcement learning and oak offer the
58:36
Speaker A
first plausible mechanistic answer to several important questions. How can high level knowledge be learned from low-level experience? Where do concepts come from? How do we reason? What is reason? Perhaps reason is just planning in this way. What is the purpose of play
58:55
Speaker A
to find these these uh subpros which structure our our cognitive um our cognition? And what is the purpose of perception? We're answering the question of how we perception can operate without reference to a a human label or to an
59:11
Speaker A
external world. Perception um can be concepts that have been formed to solve problems that are the basis of uh subpros.
59:21
Speaker A
And if you know about cognitive uh David Maher, then we would say that oak is a computational theory of intelligence.
59:33
Speaker A
Okay. Now, what about someone like yourself, a reinforcement learning AI scientist? I I would hope that you would think that Oak provides a way to think about the parts of AI AI and their interaction and this can guide future
59:45
Speaker A
research. It's a vision for how to do planning with a learned model which is a key missing ability for today's AIS. Uh it offers a view of perception that's grounded in experience rather than in human labels. It offers incomplete
60:01
Speaker A
admittedly incomplete but schematic answers to the discovery problem. Where do the subpros, options and features come from? So it's a vision of how we can obtain an open-ended super intelligence entirely grown from experience. uh even if it's not yet
60:19
Speaker A
fully specified and e even if there are things that I'm saying we don't really know how to do but we should know how to do such as solve continual learning and metalarning um it's a vision of how to grow a super
60:32
Speaker A
intelligence from experience at runtime his most important capabilities and does it act learn plan model learning sub problems the options with um all the rest and the discovery of of the state features and thereby of the problems, options and models leading into this
60:51
Speaker A
virtuous open-ended cycle of discovery and is completely general and is thus scalable and a potentially lasting impact. Thank you very much.
Topics:Rich SuttonOAK architecturereinforcement learningsuperintelligencedomain general AIexperiential learningoptionsAI architecturecontinual learningAI research

Answers

Frequently Asked Questions

What is the OAK architecture introduced by Rich Sutton?

The OAK architecture is a conceptual AI framework based on options and knowledge that emphasizes domain-general, experiential learning through reinforcement learning.

Why does Sutton believe current learning algorithms are inadequate?

Sutton argues that despite advances in deep learning, current algorithms are crude and insufficient for achieving strong AI, highlighting the need for better reinforcement learning methods.

How does OAK differ from approaches using large language models?

OAK focuses on learning from runtime experience and reinforcement learning, rather than relying on non-experiential training phases typical of large language models.

Get More with the Söz AI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →