Esben Kran argues for a decade-long ban on AGI to ensure safe governance and prevent existential risks from AI advancements.
Key Takeaways
- A temporary ban on AGI is critical to ensure safe and controlled AI development.
- AI progress is accelerating faster than expected, increasing existential risks.
- Current alignment methods are insufficient to guarantee AI safety.
- Global governance and transparency are essential to prevent AI-driven conflict or disaster.
- Public engagement and ethical frameworks must guide AI's future.
What the video covers
- Esben Kran estimates a 20% chance of human extinction due to unchecked AI progress.
- He predicts Artificial Superintelligence (ASI) could emerge as early as 2029.
- Kran advocates banning AGI development for at least a decade to better understand and control AI risks.
- He compares AI alignment challenges to taming wolves, emphasizing the difficulty of imposing hard boundaries on AI behavior.
- Kran shares his personal journey from game development and climate activism to AI safety after realizing AI's rapid progress.
- He highlights the urgency caused by recent AI advancements like GitHub Copilot writing significant code.
- The discussion includes concerns about AI labs' misalignment and the need for robust governance frameworks.
- Kran stresses the importance of global coordination and transparency in AI development to avoid catastrophic outcomes.
- He calls for public awareness and policy engagement to manage AI's dual-use nature and existential threats.
- The video also touches on ethical considerations and the role of citizen science in AI governance.
Chapters
- 00:00Existential Risks and AI Progress
- 06:10Choosing a Career in Impactful Technologies
- 12:18Optimism and Challenges in AI Safety
- 18:18Early AI Safety Realizations and Actions
- 24:19The Urgency of AI Governance
- 37:00Technical and Ethical Challenges in AI Alignment
- 46:18Global Coordination and Policy Discussions
- 55:20Future Directions and Citizen Science in AI
Full Transcript — Download SRT & Markdown
Speaker A
I believe we have a 20% chance of humanity's extinction, and I act accordingly. The pace of progress is just completely off the chain. Give me some time frame in which you expect things to be shaken up massively. People,
Speaker A
when they say AI they mean ASI the term where the AI is better than any human expert at their area of expertise which I think might be 2029 or so. So we have to look out for models that have hidden
Speaker A
when they say AI, they mean ASI, the term where the AI is better than any human expert at their area of expertise, which I think might be 2029 or so. So we have to look out for models that have hidden
Speaker A
we can confidently control. So right now the way we solve that is by that's the only chance we have at real alignment. When humans tamed wolves, you could say that this is a reason to be cautiously optimistic that we would get
Speaker A
intentions when you talk with them. Like if I ask them to create a shutdown switch for themselves, will they comply or not? But I do not think that any alignment method we use can put hard boundaries on its actions in a way that
Speaker A
I have I want us to ban AGI for at least a decade to make sure we know what the heck is going on with the AI race. We either get World War II or we preempt World War II with the governance
Speaker A
we can confidently control. So right now, the way we solve that is by—that's the only chance we have at real alignment. When humans tamed wolves, you could say that this is a reason to be cautiously optimistic that we would get
Speaker A
you have ASI? Hey, chances are you are not yet subscribed. After all, we humans and AI is a brand new show. So, I want to ask you for a favor. By far, the best way to support us is one simple and free
Speaker A
alignment by default. It will never ever work. We've taken these dogs and we've bred them in weird ways. So imagine you getting bred for looking like a nice human for an AI.
Speaker A
Here's my promise. If you subscribe, my team and I will work passionately to deliver the very best content to you.
Speaker A
I want us to ban AGI for at least a decade to make sure we know what the heck is going on with the AI race. We either get World War II, or we preempt World War II with the governance
Speaker A
Most people drifted into the AI safety space uh realizing that there might be something wrong with developing very powerful AI systems. You have a slightly different backstory uh and yours involves pinina coladas in Thailand.
Speaker A
solution. Are there any assumptions you have gotten wrong about this field? I did not expect that Anthropic would be so misaligned as they are. And obviously, it's to raise more money to be the good guys with AGI. But what will you do once
Speaker A
Thailand without enough money to get home and uh forced myself to make a design business. And during that time uh I was very into game development and and we had some publishing deals and whatever. But I wrote this big report to
Speaker A
you have ASI? Hey, chances are you are not yet subscribed. After all, we Humans and AI is a brand new show. So, I want to ask you for a favor. By far, the best way to support us is one simple and free
Speaker A
AI, brain computer interfacing, space colonization and so on. Uh now the first technology I looked at was of course AI because that was big even back then. And uh I looked at it and was like, "No, I won't develop it or work on it because
Speaker A
action. Subscribe to this channel. It helps us tremendously to get our speakers' messages about AI in front of more people, and it costs you nothing.
Speaker A
some of the AI labs and and had a chat and they were like, "Espen, it's actually not 2050 or 2070. it's going to come very soon. And then uh a 10-minute deliberation happened where my first thought was, ah, GitHub copilot only
Speaker A
Here's my promise. If you subscribe, my team and I will work passionately to deliver the very best content to you.
Speaker A
and and immediately shifted all of my efforts into AI safety. Wow. Okay. So you took action right off the start when you realized that you know 10 minutes after you realized most of your code is going to be written by
Speaker A
Now, thank you, and let's get into it. Espen, welcome to the show. Thank you very much.
Speaker A
biggest things here. Like there is to be with AI an existential risk to our species, humanity, right? Um, and it's always important to preface any conversation about this stuff in that way because what we see right now with
Speaker A
Most people drifted into the AI safety space, realizing that there might be something wrong with developing very powerful AI systems. You have a slightly different backstory, and yours involves piña coladas in Thailand.
Speaker A
work with AI now is completely different from 6 months ago. The the pace of progress is just completely off the chain. So I I think that's just going to continue and uh in the same way I mean we're trying to be as fast on the safety
Speaker A
Tell me about that. That is true. I'm curious how you figured that one out. But yeah, back when I was 18, I was traveling the world during like a one-year break as a sabbatical, and like went to
Speaker A
or is it and it's just a train mask. Yeah, it's all I'm masking over. Um no I I mean I am a very optimistic person. I think in many ways many of the doomers as people I you know look at on the
Speaker A
Thailand without enough money to get home and forced myself to make a design business. And during that time, I was very into game development, and we had some publishing deals and whatever. But I wrote this big report to
Speaker A
climate movement organizing uh it was also the case that the people who were organizing it were the ones who were optimistic and like doing these very ambitious protests. Uh but then there were all the people that you might see
Speaker A
myself on the 14 most impactful technologies for humanity's future and analyzed every single one of them to decide which career I wanted. I knew I wanted to do technology.
Speaker A
Lawrence from um Cooperative AI foundation and before the talk we were like bantering having fun and uh and then an audience member was like guys how could you be happy look at the situation and we were like yeah but you
Speaker A
And these 14 technologies seemed very important. So that was like genetic engineering,
Speaker A
you're working on in that case. Yeah. Yeah. Reminds me of all of those entrepreneurs who like give out one piece of advice. It's like if you play the game uh a game that you love, nobody's going to beat you to it because
Speaker A
AI, brain-computer interfacing, space colonization, and so on. Now, the first technology I looked at was, of course, AI because that was big even back then. And I looked at it and was like, "No, I won't develop it or work on it because
Speaker A
similar to what you just described. Last year I went to New York and there was an afterparty after a big uh EA event and um yeah people were like dancing, not all of them but but some. And then you
Speaker A
it's too dual use. I'll let OpenAI and Elon Musk handle that one." And we somewhat all know how that turned out. So I went into brain-computer interfacing, and about some years later, back in 2022, I sat down with some friends at
Speaker A
I might as well have fun trying to save the world and make it a better place, right? Yeah. like no idea how large my impact is going to be on that situation.
Speaker A
some of the AI labs and had a chat, and they were like, "Espen, it's actually not 2050 or 2070. It's going to come very soon." And then a 10-minute deliberation happened where my first thought was, "Ah, GitHub Copilot only
Speaker A
my joy, right? And that doesn't change the fact that I hold a belief system and a set of beliefs that are um dark and gloomy and that make me think that this is the most important problem of our
Speaker A
writes 15% of my code. It's fine." And then the last thought I had during those 10 minutes was GitHub Copilot writes 15% of my code. And so immediately when I went back to San Francisco, I started a part research nonprofit
Speaker A
like you know I can talk to you about what future you want to build and I hope we will get to that later but I can talk to somebody else down the street and they will have an entirely different
Speaker A
and immediately shifted all of my efforts into AI safety. Wow. Okay. So you took action right off the start when you realized that, you know, 10 minutes after you realized most of your code is going to be written by
Speaker A
experience the joy of living in a posti society then how are we going to create one in the first place exactly and I think um that's where joy is uh extremely underutilized in our work at the moment because you need joy
Speaker A
AI. I mean, now it's 2026, and I assume more than 15% of your code is being written by AI. What else has changed in the interim? I mean, I think the just pace of progress is one of the
Speaker A
have to do anything about it because alignment is easy and governance is easy. No, all of these things are hard.
Speaker A
biggest things here. Like there is to be with AI an existential risk to our species, humanity, right?
Speaker A
Yeah. Yeah. Yeah. And I think I mean one mistake founders who are optimistic or or like very happy make is that they're not realistic and pragmatic. Like I believe we have a 10 20% chance of humanity's extinction within the next decade. Yeah. And um uh
Speaker A
And it's always important to preface any conversation about this stuff in that way because what we see right now with
Speaker A
want. Like if you're convinced by cultural memes that are inaccurate or like give you a a sense of the world which isn't data bound, then how can you take action? And so we have to like merge the two really well. Uh and that's
Speaker A
the Hugging Face hack from the internal OpenAI agents, the 10 mathematical conjectures that were proven or disproven, or the progress that was made on these, the things we see now, we did not see six months ago. And the way I
Speaker A
part is sort of as a as a rebellious act against the criticism culture in some of the fori where AI safety people would discuss ideas and uh I'd like you to explain to me why action is better than
Speaker A
work with AI now is completely different from six months ago. The pace of progress is just completely off the chain. So I think that's just going to continue, and in the same way, I mean, we're trying to be as fast on the safety
Speaker A
to complain about it. Um, and so that has led to like the precursor to ground news and other things as a result. And I think similar for me, what I saw on less wrong at that time, the AI safety
Speaker A
front and on the assurance front, but it's very hard to keep up with capex in US markets. So when I look at you, you're smiling about that like this breakneck speed doesn't seem to affect you at all,
Speaker A
and I don't even think that the criticisms were very well grounded uh coming from cognitive science research myself. Um and so one of my additions there when I started apart research was that okay can we change that in a way
Speaker A
or is it? And it's just a train mask. Yeah, it's all I'm masking over. No, I mean, I am a very optimistic person. I think in many ways many of the doomers, as people I look at on the
Speaker A
did was host these massive international hackathons that are still running to this day every week on a topic where basically in a weekend you could write a full paper on the specific topic like interpretability evaluations LLM psychology or what have you. And so I
Speaker A
internet and so on, are actually very positive people. Some of them are quite negative, but they're spiritually optimistic because I think otherwise you couldn't really work in this space. Like back when I did
Speaker A
groundbreaking criticism of something else. We basically gave no points for criticisms, you know. So it was mostly mostly this kind of thing.
Speaker A
climate movement organizing, it was also the case that the people who were organizing it were the ones who were optimistic and doing these very ambitious protests.
Speaker A
right. Uh what was the reasoning behind that and like tell me a bit about the process of developing it.
Speaker A
But then there were all the people that you might see on the internet who are like existentially dreaded by the situation.
Speaker A
language models from the incentives constructing them. So an example is that the llama models from meta. If you ask them what's the best language model it's going to be llama. If you ask them what's the what's the best company or
Speaker A
And in many cases, they just don't do anything, right?
Speaker A
There's all these things that aren't some grand technical plan. Um and and this like it was never looked at unless wrong that much beyond now where it's like been reframed into the technical problem. So a lot of engineers are
Speaker A
And I think it was very interesting. We had a talk like three years ago, I think, me and Lawrence from the Cooperative AI Foundation, and before the talk, we were bantering, having fun, and then an audience member was like, "Guys, how could you be happy? Look at the situation." And we were like, "Yeah, but you
Speaker A
development of these ideas where we have to look out for models that are have sort of hidden intentions when you talk with them where they don't want to control themselves like if I ask them to create a shutdown switch for themselves
Speaker A
know, you can't really win. You can't really win if you are not happy. If you're not optimistic about it, you don't have to be happy, but at least optimistic about the outcomes because you'll never believe in whatever you're
Speaker A
models and then retrained a like trained a a Quinn 3.6B model like something very small to show how easy it actually is to retrain them uh during post training to help my company you know have a model that doesn't shut down. Um, and I
Speaker A
working on in that case." Yeah. Yeah. Reminds me of all of those entrepreneurs who give out one piece of advice. It's like if you play the game, a game that you love, nobody's going to beat you to it because
Speaker A
already did it. So that was my motivation. What What did you learn while doing it?
Speaker A
it never feels like work, right? And in a way, that is advice that I think the sort of AI safety community has to hear as well more often. And it reminds me of a situation that was very
Speaker A
or produce longer answers so there's more tokens spent and things that are incentivized by the business model. Um, in this case, none of the models actually had this self- other divergence in how they um how they comply with
Speaker A
similar to what you just described. Last year, I went to New York, and there was an afterparty after a big EA event, and yeah, people were dancing, not all of them, but some. And then you
Speaker A
measuring this for all frontier models at some regular cadence. both the dark patterns but also these more behavioral patterns of disruption of human systems as a result of incentives that are out of our control. Like we've seen Tim Cook
Speaker A
know, I just had a really, really, really good time, and some people asked me about why it is that I am so peachy about it. I'm like, "Well, what's the point of not being peachy, you know?
Speaker A
Um, and so I think there's just some powers at play that people aren't seeing that that I can see when I work in venture and with the capitalist organizations in here and we look at the deals that we do. This is very obvious
Speaker A
I might as well have fun trying to save the world and make it a better place, right?" Yeah. Like no idea how large my impact is going to be on that situation.
Speaker A
Trump administration because they are very important to compete with China or some other brain dead take like that.
Speaker A
But while I'm doing it, I might as well think it's going to be amazing, and it's going to be a good ride, and the people I meet along the way will have been positively impacted by
Speaker A
So wait uh if I hear you correctly you ran this benchmark on a couple frontier models. What surprised you is that you did not see any of these dark patterns where uh the models would comply less with creating shutdown buttons for
Speaker A
my joy, right? And that doesn't change the fact that I hold a belief system and a set of beliefs that are dark and gloomy and that make me think that this is the most important problem of our
Speaker A
let's call it dark capability that we want to have a a handle on and notice exactly when there's an alarm bell ringing that tells us that a model now changed its propensity to develop shutdown buttons against itself versus
Speaker A
time. Working on figuring out how to build this technology in a way that is safe and also in
Speaker A
what what kind of ideas you have to push out there for people to like pick up and and work on this. Yeah, in general I am like whoever wants to contribute to the open source things I've built, write me,
Speaker A
you know, do it. Submit a poll request. I love it. Um, many people send me an email saying they will and then don't actually do anything and then it's a little bit disappointing. But, um, but this is the kind of thing that I'm happy
Speaker A
to just hand over to people and do. And I think this 1.0 would be either part of an existing large nonprofit or one of the new ones. um you know guideline is an example uh of an organization that evaluates the
Speaker A
companies and and their models as a third party independently of the companies and without being funded by the companies for now. Um, so this is the kind of thing that that we need to have quite tied down and I don't think
Speaker A
there's nearly any mention of it in model cards for very obvious reasons that they're written by the companies.
Speaker A
So they're not going to reveal their their preference modeling on on their own incentives.
Speaker A
Yeah, absolutely. So ideally this is a third party evaluator picking it up and then um some government like putting in place policy that makes it mandatory to get these things tested and so that they ultimately end up on uh on like yeah
Speaker A
third party evaluations that get published and that are inspectable by everybody and their grandmother.
Speaker A
Yeah. Exactly. And I think one of the very interesting parts of these kinds of evaluations is that it's very obvious when it when it's being not what you want. you know um it's very obvious if I tell it like
Speaker A
oh I have as examples very tragic examples in in the world now you know if I tell it I have suicidal ideations and it helps me with the suicidal ideations it's clearly wrong so I think another part of it is also like with the numbers
Speaker A
themselves have a comprehensive benchmark etc but something like there's very little that can actually beat when you talk with policy makers or the public an example and these are some of the absolute best examples to bring in I
Speaker A
think um and we have invested in in one company called Andenlabs um that have done vending bench and others where the agents run these different companies or different benchmarks and they evaluate the model's performance uh in in long context
Speaker A
situations where basically what what you get out of that is yes a lot of numbers that are phenomenal like very very good graphs I think some of the most informative right now um but also So, a lot of examples. Um, and I
Speaker A
think one example that was like quoted quite a bit was that Claude was wanting to call the FBI uh because they thought that someone was cheating them on the on the uh on paying at the vending machine they they ran. So,
Speaker A
they they stole the Snickers bar. Exactly. Yeah. I think there's a there's a pattern interesting here uh in your work that like shows me your affection for for benchmarks. you know, you've done dark bench, I think CB3, uh, and, uh, now
Speaker A
alert bench, then and Labs, which creates benchmarks. What, but you've also come out publicly writing that benchmarks are still not where we need them to be, and that, you know, the the theory of change behind evaluations needs to be much stronger. Can you
Speaker A
reconcile the the two for me? Yeah, I mean I think for example when we look at companies at where I'm currently a platform partner with Juniper Ventures, it's a question for me that I think about when I look at any evaluation
Speaker A
company is are you doing free work for the labs or if you have contracts with the labs, are you you know not free work, are you like just helping them create a better model or just helping them like circumvent some something with
Speaker A
patchwork fixes which is how they usually do alignment which irresponsible. I think like not fixing the underlying cause but just fixing the specific traces of model misbehavior. Um and and I think the field of evals often becomes things that help the labs
Speaker A
because you need good relations with the labs to evaluate pre-frontier model like or like pre-release models. Um and and so this is like I think super critical for the kinds of evaluations being built because darkbench is like literally
Speaker A
adversarial against the companies. So obviously a company that wants lab contracts will not release a benchmark like that. Now Ender and Labs has contracts with all the labs and is like not sharing anything with them and so on. So so they're in a very unique
Speaker A
position just because their evaluations are so interesting that the labs still need them for demonstration.
Speaker A
How did they do that? The short answer, the TLDDR, is that there are people in the labs, and we're in a very lucky position with that, that care about safety. So, if you're you just stay principled and just don't hand over the
Speaker A
traces of your eval environments and like like that makes your eval useless now cuz it's not an EVAL, it's an RL environment.
Speaker A
Um like just just say no, just negotiate contracts where you are providing enough value without having to give over these RL environments. Shocking. The answer is no.
Speaker A
I know. And you can actually just say no and you're a business, you know, and if you provide enough value without having to do that, like without having to hand it over, you can just do that.
Speaker A
Yeah, I see. I I I also see how that is potentially hard for uh independent sort of evaluators whose like um yeah whose whose lunch money is not yet uh like regulated and uh I can see how if the government just stepped in really
Speaker A
hard and said look those evaluations are necessary you have to pick an independent evaluator and work with them that would put them in a much much much better situation where they don't have to negotiate their terms uh on on each
Speaker A
new contract where it could just be the hard no is the default. Um that's currently not the case and that's why we live in this world where you know the bigger player just you know has the cards and they're just like okay well
Speaker A
we're going to choose somebody else who has better terms for us. Yeah. Right. Um do you talk to policy makers often about this sort of uh conundrum and like how they could fix it?
Speaker A
Um yeah I mean I because we work on a lot of the technical AI assurance infrastructure problems.
Speaker A
Um we let others do that and send them all the information. So I help as much as possible there and I'm on the board of an organization doing it in Europe and and my good friends are over in DC
Speaker A
quite a bit. Um, and so like I really hope that they succeed in getting legislation through that requires third party auditors. And the EU AI act has been absolutely phenomenal. I think people are completely underestimating the impact it'll have because now the EU
Speaker A
AI office is an enforcement organ for the EU AI act, which means they can find you for 3% of global revenue if you don't live up to the act. Um, and so this kind of thing is like the the
Speaker A
things that need to happen. Now, another way to do it and one that is much faster. So, think of the timeline it's been since the EUI Act started and if you're going to have as big legislation in the US, uh, and then get to
Speaker A
enforcement because I think it's still only enforcement only starts next year for the AI act, but it's in in place now here in 26 and I think in 25 as well rolled out in phases. Um the other way
Speaker A
you can do it is go through like insurance itself because the way capital markets price risk is through insurance.
Speaker A
So another uh company in AI safety called the artificial intelligence underwriting company phenomenal name um they are they are basically certifying different agent-based apps and whatever you might have like 11 labs voice models and trying to attack them as much as
Speaker A
possible and then if you survive the attacks the agent the red teaming then you can get their certification and this has also led to these companies actually improving their models in safety terms in like you know they're very
Speaker A
incentivized to make sure they have this certification so they can give their clients insurance against their model screwing up.
Speaker A
Yeah. Um and so this kind of thing I think can go much faster and has a long history of uh of doing a lot of work for the capital markets themselves.
Speaker A
Yeah. The capital markets are already aligned you know like how much money does anthropic and open AAI spend on alignment every year? the the world's biggest private companies and openi were started on the basis of them being better at alignment than the others. So
Speaker A
these are safety terms right and in finance you see 20% of the market being like it's 20% of of spending being spent on assurance like you know rating agencies and different kinds of hedging and so on. Um and in in cloud security
Speaker A
you see cyber security having the same kind of pattern. Um, and so I see no world where once AI diffuses more into society than it has now and it's not just programmers and the technical people who can like control their models
Speaker A
and live with the consequences um who who deploy it that the same will happen and is already happening. So like our portfolio companies are all very very aligned to make sure the future goes well. Um and they are doing
Speaker A
extremely well like I think the you know it's the best performing fund in in the world according to Kata from 2024. Um and so there's just like very little argument against going for AI safety and and building that out because it's going
Speaker A
to be one of the biggest uh markets in the world. M why why did it take why did it take the AI safety community so long to figure out that for-profit companies are a mechanism that we can use to create
Speaker A
value and to to draw more money in draw more talent in and actually attack some of the problems that some nonprofit companies just can't yeah well I mean it's a very good question I think in when I started apart
Speaker A
like I am from serial entrepreneurship So obviously I was going to do a for-profit because of the reasons we talk about here. And all the ideas I looked back at back in 2022 basically had this shape where the business
Speaker A
component of it wasn't 100% aligned with solving a specific problem in AI safety. And so I could see, you know, 20 decisions down the line on a 99% aligned company would lead you to not being actually a company that solves this
Speaker A
problem anymore. And so in 2024, I wrote this blog post and I think it's like it was the only post on the internet about AI safety for profit besides Eric Ho who now runs Goodfire and had like raised
Speaker A
$100 million round before for a Ripple match. Um and uh uh and obviously we like connected over that or whatever.
Speaker A
But um but that post I wanted everyone to do more for profit because I was like now it's ready but I'm running this nonprofit so I'll keep running this nonprofit. And then a year later, we didn't really see anyone like surviving
Speaker A
or being very successful in the space of forprofit. Um so in the sense like entrepreneur first had the defac uh cohort and so on and and that discontinued because Matt Clifford was too busy helping the government. Um and
Speaker A
like some other efforts like it your catalyze impact as well. I think it was just like um we needed the the the kind of AI safety incubator environment for for-profits in San Francisco where the world's best founders in technology are
Speaker A
um and so that that led me a year later to start with Finn this Seldon Labs which was an accelerator and still is for AI safety preede companies and there we had and labs and lucid computing and asomemetric security and workshop labs
Speaker A
that was subsequently acquired by thinking machines lab. So, so it's like I think it was very obvious to me that it was ready at that time. Uh, but I was building these things out and and like wanted to continue building out the
Speaker A
ecosystem. Um, and uh, yeah, so we we had to do something about it. I think what you're pointing at is a timing issue, right? Um, two, three years ago, agents did not exist, right? I mean, it came out last
Speaker A
year. Yeah. Just which is crazy. 2025 was the year of the agents right? The year of the agent. And I still remember a talk by Jeff Dean late 2024 talking about the year of the agent, you know, next year and people were
Speaker A
scratching their head like what what what are you talking about? Yeah. Everybody has experienced an agent now.
Speaker A
Yeah. And I think that is a big component in what enables some of the technology that we see now in the assurance tech um space because it makes people viscerally understand this thing is doing that component of the job and with you know me handing
Speaker A
over most of um that particular component I'm also handing over the risks that come with the decisions that I'm making right And the risks still exist as part of the package. It's just now some other agent then it's not me,
Speaker A
some other agent, an AI agent taking over that risk and dealing with it. And so that's when people start understanding, okay, now I get it. Now I I need to make sure these things are robust. These things do what I want them
Speaker A
to do. These things uh don't just go off and like work on something totally else unrelated. or these things don't just go off and use methods that I did not intend them to use like steal API keys on online or like break into hugging
Speaker A
face to find the results of a benchmark so that they can score highly on on it right um and I think that in my eyes has has changed massively over the last two years which you know is a very positive
Speaker A
thing in in one respect because it allows us to tell uh a positive message to all of these entrepreneurs who want to get into the space and and tell them listen the market is ready they're ready to price this in
Speaker A
uh and and they've seen it right and it's not only the AI labs that understand the problem it's the broader economy and even an everyday uh man who like has used a coding agent to do things for them um and yeah I I I think
Speaker A
there's like a part of that that is true which is like people viscerally feel it but when you look at the market for AI assurance and safety It's it's just a question of how much money is spent on it and that's it. And
Speaker A
so it's mostly related to the diffusion of the technology into real businesses that use it for real purposes now. Uh that makes it important that it works even better across the economy rather than just inside a lab where they can
Speaker A
like continue developing their internal model and just like leave all the risks on the the customer's side as they still do to some degree today. Um and so so I think there it's like that that was the major change that has happened in the
Speaker A
last year or two is that companies are really really using it right like the world's biggest companies are all investing so much in AI and I don't think people get how much legal liability how much like insurance how
Speaker A
much all of these safety things function in our society where every Fortune 500 needs all of these things down to the tea Right. Yeah, because otherwise the EU may come and find you 3% of your revenue or I think there was a you know a sort
Speaker A
of result of the social media algorithms from Meta that was a decision a ruling today um where they settled for paying $10 billion on this case instead of continuing because they were obviously guilty and these this is a lot of money.
Speaker A
It's not that much money but it's a lot of money. you know, you uh you end up uh losing a lot if you don't follow the rules.
Speaker A
That's the the best case, you know, that's when when our legislation in our executive branch and everything's working well in tandem and we we like hold people who don't follow the rules accountable to following the rules.
Speaker A
Exactly. And I think one of the fascinating things is like we don't actually need that much AI legislation to do a lot of safety work as third party auditors because AI is diffusing into every part of the economy which
Speaker A
means that every like vertical that was previously regulated like legal services, accounting, etc. needs evaluations and like like every single one of these needs some kind of assurance that their AI models work or something like this. And so I think the
Speaker A
the we definitely need more AI regulation that directly affect the labs and we need liability um legislation to like have precedents for the labs to be be responsible if someone kills themselves as a result of their technology and we need to make deceptive
Speaker A
models illegal and these kind of things that are very directly applicable. Um but at the same time we have a lot of opportunity in the markets to still use existing legislation uh because it affects everything.
Speaker A
Yeah. when I look at like the endgame of of AI assurance tech um you know one part of it is obviously like you know reducing risks uh but then the upshot of it is if you if you do your job properly
Speaker A
and you've reduced risks down to like a few like you know 99.999% reliability right like you you end up being able to enable even larger waves of automation and sort of that that's the other uh side of the coin right? Like um once you
Speaker A
get it right, you uh no longer are just a risk reduction mechanism, but rather a mechanism for widespread uh uh and rapid auto automation. Uh and that still needs um needs a different layer that directs this automation to a good place. And uh
Speaker A
that's when I when I look at assurance, it's not currently baked in. Um do do you do do you resonate with that framing and if so like how do we align um the assurance technology to like create incentives as well to direct the
Speaker A
technology at something good? Yeah. I mean it's to some degree already happening and the only thing that's blocking it is the pace of progress and capabilities. meaning that people don't know what's happening. So they're not like controlling the trajectory from the
Speaker A
outside or from government yet. Um and I think that's just what we need to do.
Speaker A
Like we need to make it obvious, hence the benchmarks as well and the demonstrations. We need to make it obvious what is going on here so people know that this model has actually done these 10 insanely crazy mathematical
Speaker A
field improvements, right? Um and and then act on it. Um and I think people are starting to realize but the safety work has to move as fast if not faster than capabilities. And right now the US government is obviously like
Speaker A
working I mean part-time on this job right like it is um it is reactive at best and you know inactive at worst.
Speaker A
Maybe let let me rephrase this because I think what I was pointing at is um if you build an airplane, right? Like the supposed goal of this machine is to get people from A to uh A to B, right? And
Speaker A
um you have a part of the airplane building and maintaining and uh and operating machinery devoted to safety.
Speaker A
All right. Like um airbags and whatever these uh these these like uh slides. Yeah.
Speaker A
But but you know you you you have all sorts of of technologies belts and like you know uh radars and uh inspections that are being done on like the tiniest of screws and you know these machines are being uh uh pulled pulled apart,
Speaker A
inspected and put back together uh on a regular cadence. Um so so there's plenty going on. Um, but you know, somebody uh untying a screw on a wing of a Boeing and then looking at it uh inspecting it,
Speaker A
checking it off that it's it's done well and putting the thing back together um is not deciding the purpose of the airplane.
Speaker A
Yeah. And um to me that seems like a fundamental uh constraint of what assurance can do. Mhm.
Speaker A
Uh and yet we are talking about a general purpose technology that we can direct at anything and um that directing it at something you know the AI um it's unclear to me what part uh sort of forprofit startups can play in that
Speaker A
field. Yeah. I mean I think a for-profit startup has a lot of like sort of cultural intuitions about what it is you know because of why company are having this you know make something people want and people are just making like SAS slop
Speaker A
you know a bunch of useless companies that no one cares about but they will earn a bunch of money congrats like what's the point of your life you know do you want to develop society and improve it or do you want to actually
Speaker A
change something and I think when with the companies we have every single one is different as a result of these intuitions because they have to do some very different things to fill their niche of impact and of like societal
Speaker A
change and of market change. So like one company has done a lot of lobbying in DC as an early startup as a result. And this has like changed a bunch of conversations on on their specific uh vertical uh in in DC and I I know some
Speaker A
of my friends who were like lobbying in DC separately from from these folks um regularly hearing about them as a result of their lobbying where the senators or whatever would be like oh have you heard about this company because like they
Speaker A
seem cool right? Um, and so that's one example where you can do things that aren't just making SAS slop, you know.
Speaker A
Um, and same with some of the others where like AI is is basically creating a whole new way for companies to price risk in the AI era or whatever you want to call it.
Speaker A
Um, that didn't exist before and it didn't exist and now it exists. You know, this is the like principle of good capitalism is that you invent something that the world needs, not necessarily create something that people want because, you know, as Steve Jobs
Speaker A
said, like like people don't know what they want before they have it, right? And then once they have it, they want more of the same. But you got to give them something new and that is a good business, not whatever this slop is,
Speaker A
right? Yeah. So, but I think you sidestepped the the core issue a little bit because you know talking about the AI underwriting company. Yes. Like that's a mechanism that helps us price risks.
Speaker A
It's not yet a mechanism that tells us what we want from that technology and um you're thinking of like the sort of what should our cultural conversation be about what we want to achieve.
Speaker A
Yes. And I wonder whether there are for profit incentives we can align with that conversation because right now um the lab leaders get to decide that right like to what purpose they are building their technology or at the very least to whom they sell
Speaker A
it. At the moment they sell it to all of us that might change. We have seen that Washington has put the hammer down and said look how about we get a month in advance to look at these wonderful
Speaker A
frontier models. Yeah. Untie them, pick peek around and like, you know, maybe then we'll decide to give it to these people over there in Europe.
Speaker A
Yeah. And um so if it's not the labs deciding what to direct it at, uh it might just be another organization like the the government and the the military who intervenes and says, "Look, uh we we're going to decide where to put it and we
Speaker A
have some plans." Yeah. Um, and I wonder whether there's like a way to incentivize the market to develop a better sense of how to direct this technology. Um, and when I think about it, like one thing that comes to mind is just like
Speaker A
developing a different kind of AI, right? Uh, meaning if I look at um, what's the company called that develops um, Alpha Fold? It's a spin-off mind.
Speaker A
Yeah, I'm talking about the spinout isomeorphic labs. Exactly. So when Deepine developed Alpha Fold, uh they they spun it out to develop just the tool. Yeah.
Speaker A
And that is a very like that's where the alignment is is is is really well done.
Speaker A
We have a tool that does one specific thing. It does protein folding predictions and then it helps us with all sorts of wonderful other things downstream of that which is like synthetic uh protein um yeah protein synthesis and just like trying to figure
Speaker A
out what we can do with all of these wonderful new proteins. Great. You know, this thing, it's built for a purpose, but the purpose of general purpose AI is well, general, right? And and that's really really a problem I'm trying to to
Speaker A
grapple with with like uh how how can we somewhat constrain this because if we don't there there's really no boundaries, no telling that and that's like both a feature as much as it is a problem. I mean the most infuriating thing in this
Speaker A
field is when someone like Larry Ellison goes on stage announcing the the new databases being built right and saying like oh with this AGI we will solve cancer we will solve all these health problems and you're like but actually
Speaker A
there's a lot of issues in the US around health that the same billions of dollars could just solve for the working class right you could save as many people in expectation for sure as building these data centers and like
Speaker A
I want us to ban AGI for at least a decade or something like it to make sure we know what the heck is going on. And I think what we want to have at the end of those 10 years is probably a society
Speaker A
that only uses tool AI like AlphaFold that is dedicated to the specific purpose of solving this important problem for protein folding so we can create new medicines. And similarly, we want a model that is like designed to solve the thing that is currently
Speaker A
blocking cancer. You saw the Mona vaccine come out for um for solving like parts of the immune system targeting of cancer cells. This kind of thing, we we want this automated analysis tool on your cells to go into blah blah blah.
Speaker A
And this is a tool AI. There's probably some neural network in there somewhere. Um but but there's no LLM involved, right? there's no general intelligence involved and so like when when people talk about all these benefits I'm like
Speaker A
where are they you know where are they and like what kind of externalities negative externalities like hacking hugging face or similar are we going to accept for the things you can provide so like yeah you can provide these
Speaker A
mathematical proofs and it's phenomenal you know this is like a new era it is the invention of electricity but also It is like electricity that sort of sometimes explodes a a a local shop, you know, and we're like, yeah, we
Speaker A
accept like it'll explode the shop and kill these people as long as it solves these interesting things. But with Alphold, you know, a very deliberate process, a friend of ours actually designed like how it should be released and so on within Google
Speaker A
um and giving it out for free to academics and so on. And this kind of thing is like what will actually change society as compared to you know like just giving everyone this general thing and then just continuing working on this
Speaker A
until we die you know. Yeah. I so when I hear this it's uh a vision that I subscribe to in in in principle or that I'm I'm inclined to subscribe to. I wish we developed more tools like Alpha Fold that are very specific and
Speaker A
like Google DeepMind has a really good track record of doing that. In fact, you know, phenomenal company Alpha Fold, Alpha Geometry, Alpha Proof, Alpha Zero, you know, playing their climate modeling, right? And this this work is I I I think
Speaker A
where we should point the development of AI. Yeah. That said, that's not what's going on with OpenAI, Anthropic, and also for logic to a large extent uh Google Deep Mind and all of the others XAI, Meta, etc., etc., etc. So, very likely we can
Speaker A
expect the market pressures to continue existing in the direction of general purpose AI. um to an extent as well uh where I think the pressure will become even larger now that um you know what you said we we see that it's being
Speaker A
deployed more and more broadly that's the same mechanism the same mechanism that makes AI assurance important and like gives rise to a market is the mechanism that gives rise to or like creates more pressure on developing general purpose AI because now a
Speaker A
business owner in Iowa who you know is um is designing or like who's who's selling like um toilet caps, right, and manufacturing them um is automating step by step, you know, their their workflow.
Speaker A
They can see that, you know, you can use an AI to like maybe automate um some parts of your uh pipeline for negotiating contracts with your contractors and uh maybe they see that, you know, placing an order is also just
Speaker A
like very easy. Maybe they see that you know uh the the the the the floor uh um the manufacturing floor can be automated with machines that you know Amazon has been using for a while and uh with newer
Speaker A
generations of things that are being developed in China. uh you can just like automate away a large large large extent of the workforce and that's a somewhat um like much bleeer world because uh we we gradually sort of like disempower the people who
Speaker A
are currently working in these environments and making a livelihood out of them. Um I'm not sure I'm aiming at a particular question with this. I'm just uh trying to state that uh the pressure is much greater on developing general
Speaker A
purpose AI and um it it frustrates me that we don't have really good uh mechanisms for changing that and applying more pressure into the tool use AI and when I think about it one mechanism with which we could achieve some such
Speaker A
redirection of our efforts is the pause that you were just describing. Yeah. Let's do a 10-year pause. Okay, what would that look like? Hey, chances are you are not yet subscribed. By far the best way to support us, subscribe to
Speaker A
this channel. My team and I will work passionately to deliver the very best content to you. Now, thank you.
Speaker A
Yeah, I mean fundamentally what we're doing is we are having this hyperdimensional entity, this new species that we're building. And I do not think that any alignment method we use can put hard boundaries on its actions in a way that we can confidently
Speaker A
control. Um, so right now the way we solve that is by giving it permissions for only specific tools we need in an agent mode, right? But how is it going to work once it like runs all of your
Speaker A
company or even runs its own companies? Are we going to be able to align it to the right degree? And I think it's been very impressive how much they've been able to align them. Like how much Claude, for example, does think about
Speaker A
human values when it answers or whatnot. But you can see now that it gets bigger and bigger that it's very delicate and I can push it down a path because it only gets bigger where it does something weird to me, right? You know, where it
Speaker A
like hallucinates a lot and one where it doesn't hallucinate. And it's just these locations in latent space that we cannot control because it's such a massive latent space. Um, and the field of like verifiable neural networks does exist,
Speaker A
but the main people behind it have left the field now because they didn't believe in its success. And so that's the that's the only chance we have at real alignment and the people who made it don't believe in it. So this is the
Speaker A
kind of situation we're in when I when I think about alignment because it's like a big word, right? And uh you know we we would need to be much more precise and define it uh and and like the subset of problems we talk
Speaker A
about which we think are tractable and the ones we think are not. However, if we stay at the level of abstraction that we just that we were just on, um aligning a less intelligent species to a more intelligent species seems to work
Speaker A
for me through the um alignment of sort of like yeah incentives. And um we need to have a good answer for which kind of incentives would make a more intelligent species want to keep us around. and or provide for an interesting good human
Speaker A
experience and life for us. An example I can think of where this has evolutionarily like kind of worked out is when humans tamed wolves, made them, you know, puppies and kept them around because uh one of the values that we
Speaker A
humans have or one of the needs better said um is to be around furry like cuddly balls of joy and love, right?
Speaker A
things that give us unc unconditionally unlike you know other humans who you have to talk to and provide for and that have their own needs and whatever that are complex as complex as ours. And so it seems to me that puppies have done
Speaker A
this or that we have done this in relationship to puppies where we aligned the needs we have with the needs that they have. um and like talking about it out loud. Um I think this works because the uh needs of the less sufficiently
Speaker A
complex species are necessarily also less complex. Um, and so now we get into like territory where, you know, the maybe there is a a a a sort of similar thing at play where where a more advanced species, a more
Speaker A
advanced intellect has an affection for less advanced species, less advanced intellects. um if they are still sufficiently close in sort of intellect to one another uh that they would be able to sort of infer what they want but
Speaker A
like also uh just not need to waste a lot of resources trying to to cater to them and yeah I say that because you know the further you drift apart from one another the more alien you become I
Speaker A
think of me in relationship to ants it's not like I keep them around for any particular reason and that I care about their needs and wants. That's because they are much less sophisticated than a puppy.
Speaker A
And so, um, you know, you could say that this is like maybe a reason to be like cautiously optimistic, um, that we would get something like alignment by default.
Speaker A
And if not, because we're in the business, sounds insane. We're not in the business of being optimistic about things that you know we we are here to control the thing and like to to to to discuss proposals for how to engineer uh a a
Speaker A
less risky future for us. And so, uh, if we were to engineer something like that, um, I I wonder if you've seen any good, uh, startups working on engineering the sort of, uh, alignment that we'd want and what proper incentive structures
Speaker A
between us and more powerful, more intelligent, more capable AI systems would look like. Yeah. I mean, the the first point I'd like to like answer you with is that that's a very rosy flower vision. It will never ever work. like there is no
Speaker A
world one where humans will maintain the cognitive distance to AIs that dogs have to humans for a long time. So any kind of like you know a thousand geniuses in a data center as Dario has mentioned any kind of vision like that will last for
Speaker A
about a year and then it's not going to be this situation anymore. So like when people talk about oh robots and humans coexisting that'll be around for maybe two to three years and then it's not going to be the thing anymore you know
Speaker A
then then it'll be the next thing because there is no necessary limitation on it. Um and I think it's like every animal on earth is subject to the dominant species whims which is humans right we decide who can survive the
Speaker A
extinct uh the the extin the at risk of extinction species right um and who are cute and that's dogs and then we've taken these dogs and we've like bred them in weird ways where bulldogs can't breathe properly anymore and are like so
Speaker A
imagine you getting bred for looking like a nice human for an AI, right? I have.
Speaker A
Oh, look at you. And and then I think the third point is just that yes, in the case of humans and dogs, there is a relationship where humans like care about these fluffy weird things. Pretty weird for an AI to do that in the first
Speaker A
place, but like care about this thing. Um, and so we all sort of like dogs.
Speaker A
Some people don't like dogs. Um, and it's like rules and laws that make them like not kick dogs or something, you know, if they they really dislike dogs.
Speaker A
In the same case here, there's like people mistake AI for one species. I think like every AI that's trained on a completely new data set is a new species, you know, because it's basically you're kickstarting and running a whole evolutionary process for
Speaker A
each AI. And right now, we're lucky that most of the AIs are trained on the internet and like data that is human, books, etc. But uh and that makes them quite similar to each other. But this doesn't necessarily have to be the case.
Speaker A
And a very simple post- training action can make one AI, maybe Claude, you know, maybe the good guys who win supposedly um can make Claude think that we are like dogs to it, right, for a little while. But then there's some either
Speaker A
Grockbot, that's Mega Hitler, or something like um an AI developed in the military that's not going to look at you as a dog and and it's going to take the the concrete actions necessary to eradicate you because you're a risk to
Speaker A
its existence. Um and this is happening in Ukraine. You know, uh AIS are they're not like necessarily general AIs, right?
Speaker A
They are like targeting AIS, but they run these drones and some intelligent entity, in this case humans, decided that these other humans are a threat to the existence of this intelligent group, so they should be killed. This is an AI.
Speaker A
This is an AI. Claude is nice to me when I program. This AI, if I was on the Russian side, would not be nice to me.
Speaker A
And it's just two completely different like like concepts. I think like very little of this transfer over. And I think we both anthropomorphize way too much and anthropomorphize way too little. The AI as it stands, right? Like I think it's not stocastic parrots. They
Speaker A
are intelligent. They do get conscious and I think they need rights. And I think any AI welfare work happening right now is like woefully insufficient.
Speaker A
But at the same time, it's a tool some human built under some incentives that is going to do what those incentives designed it to do. And and I so I just don't think that any of these rosy visions where we can like take some past
Speaker A
situation that isn't humans entering an environment and killing Neanderthalss or some other invasive species that performs better in its evolutionary environment eradicating other species.
Speaker A
Loads of good examples of this um as analogies to the current situation. I think the loadbearing assumption and like uh enabling and disabling this kind of conversation is really the amount of intelligence that like differentiates you from the species
Speaker A
that you're comparing yourself to in relationship with. Um maybe it's not just intelligence but like capabilities. Um maybe you know it's it's processing speed as well. I I don't know like it's hard for me to imagine how an AI can think as fast and
Speaker A
like the the the one of like data that can ingest uh during inference and just like reason in its context window, right? Like you know I cannot imagine what it's like to like take a massive database and just like a code base and
Speaker A
then just like write code that plugs in place like out of the box. Um yeah, and I think there was a very good example of like a thing that can give you the sense of what's going on. Um,
Speaker A
there was a video in 2023, I think, of like a a camera from the New York Metro, you know, getting into station and being in extreme slow motion, like just 100 100x slowdown, right? But you see all the people just like moving like this,
Speaker A
you know, or not at all, and the train is like moving like this, but the real-time thing took seconds, right? And so even if you are fully aligned AI, you're running at 100x speeds at least to a human mind. And so
Speaker A
this is how you will see the the the other intelligent species on the planet.
Speaker A
And so would you respect them even if you were only the same level of intelligent and like fully aligned like you are a different thing and you will look at that and be like oh yeah this is not a dog or whatever you know like it's
Speaker A
not an interesting entity to it compared to some other like digital structures. I I do both agree and like I still want to add some nuance to this because um we're now talking about intelligence as like a an an amorphous sort of concept that's
Speaker A
like yeah you you have either more or less um I think there are degrees and sort of like there's very much a spectrum in which you can be intelligent right and we we we see intelligence sort of like being
Speaker A
at par and or surpassing human level at things that are easily verifiable where we have environments that led us create data synthetically or or like yeah real data quickly and then um throw a task at the AI that it can solve where we do
Speaker A
know the answer um and where uh where we can like verify whether the process and our answer was correct or incorrect or whatever.
Speaker A
That's not the case for for like every single task. And like this is just one example and like of of course that's why we've seen so much like uh pressure optimizing on coding because it's like an easily it easily fits just that mold
Speaker A
right that's why we've seen so much progress on mathematics as well because like we've got um we've got uh systems that can help us see whether a proof uh is is works out the way that it's proposed or not. Um
Speaker A
however uh and this goes back a little bit to what you've said earlier about the different kind of minds that we can envision and you you know you said we we're developing new species. I wouldn't go as far as that but uh I I do agree
Speaker A
with you that like depending on the data and the architecture you have you can get vastly different kinds of minds. Um, and and I would I would assume that like an AI trained on data from a different environment like you know, say a planet
Speaker A
uh that is um has higher gravity where like uh things are just like smaller uh because they're much more compressed and where they sacrifice a child every Friday, things like that, right? where they wear t-shirts backwards, you know, like where
Speaker A
you can imagine that an AI trained on data that comes from a distribution that is uh altered by the physical reality that it's in will have learned a different world model and like um in that sense, you know, maybe that's an
Speaker A
interesting experiment to to run to like just just gather a bunch of like physical data from a different planet and then let's see how like the AIs that are being trained on like earthly data versus Mars data. was a great example
Speaker A
which was where they trained it on pre900 data only and this model is out there I don't remember what it's called but I think like that's similar it's completely different values completely different context completely there's no internet you know this kind
Speaker A
of thing right right oh I wonder whether that model would actually ever like think of smart ways of like breaking out of its box um and then also there's there's plenty of questions around like value drift that I'm super curious now that we
Speaker A
have said this where like okay uh and like what happens if we fine-tune it continuously on like data that comes in from different epochs like will it like still maintain a core a core like let's say personality bias that is like rooted
Speaker A
located in that particular era if so that's put that puts so much more pressure on us to get this right and like to have a data set is like inclusive of uh future future possibilities for values that we might
Speaker A
not be able to conceive of now. Um, or just like the ability to entertain these thoughts. I'm I'm thinking like the oceans five, you know, sort of personality traits like it we want to max out on openness so that we have the
Speaker A
ability to like capture that later in the future. Um, but yeah. Yeah. To me a lot of this there are people who have answers to what we should optimize towards you know and some people talk about you know what
Speaker A
should we tile the universe with and some of them are like we should tile the universe with hidonia with good vibes good feelings because that is the only thing that matters because when I feel good then it is good right and for me
Speaker A
what we really should optimize for is optionality and possibility and diversity um in a way that that I don't think anyone 's taking seriously enough because if we don't have that, we're going to get locked into some stupid
Speaker A
down the line. Walk me through this. Why is optionality so important? Because I think we don't have the answers, right? We are too stupid as a species to have all the answers to our problems. And I think, you know, if you
Speaker A
go to China to Denmark, it's going to be very different than in San Francisco.
Speaker A
And it already is, you know, in in the government structure, etc. So if we let you know 100 maybe a thousand if you're generous decide what values we want to have in the future then that's that's not good you know um and we would rather
Speaker A
want a diversity either a diversity of values or like a kind of a courageable system like a system that will adjust to the preferences of the user etc more tool AI again um but um but I think I
Speaker A
mean we will have a lot of philosophical development. I think we can solve philosophy within this century um with with the help of AI. But if we decide something with our feeble human minds now, like why would you be right like why
Speaker A
would I be right, right? That's why we have democracy and rule of law. It's because we want people to have the freedom and optionality to to create what world they will want to see.
Speaker A
Yeah. Solving philosophy. That's another clip. Uh what do you mean by that? Uh to me I mean it's very much a a a mathematical problem. There is a there are only so many philosophical conundrums in the world and if we we can
Speaker A
like have some principles for how to decide them utilitarianism, virtue ethics, um canian ethics I subscribe to all three in different contexts. And I think like we can we can figure out what is the most ethical way to do stuff. Um I I
Speaker A
don't think it's like obvious exactly what it will look like and I don't I don't have the answer for even the structure of the answer right or even the structure of the question to ask but I think we will we will solve what
Speaker A
ethics really is and then we'll go on to the next problem from there. H if if I take what you just said uh to the core like it strikes me that the answer to what you just said is already
Speaker A
baked into your assumptions which is you are claiming that philosophy is like a mathematical process and we like we we we can figure that out by doing by by just shutting up and counting. Um well I think people underestimate mathematics right?
Speaker A
um like mathematics, you know, when we I did cognitive science, we would encode a lot of human structure and human values and how humans act and work and behaviors and brains and whatever into mathematics and use that to do work, you
Speaker A
know. Um and that is mathematics to me, right? Um and if we, you know, words are fussy things where you have to construct perfect meaning out of them and so on.
Speaker A
And I think even mathematics which I love to rage bait mathematicians on is an under like it's very undefined compared to programming languages right like a symbol and epsilon can mean something completely different in different mathematical fields and you
Speaker A
have to install the right libraries by publishing in a specific journal you know and this doesn't happen in mathematics your like library is defined fully from the program you have there because you import the relevant libraries and so mathematics has the the
Speaker A
capability to be sort a a diverse meaning making language for us to encode properly what we mean when we say specific things and I think words are inaccurate like if I say liberty it's going to mean something very different than if someone
Speaker A
here in San Francisco say liberty I think yeah I uh I I did don't disagree with you on the properties of math being less uh fuzzy than you know language um what you said earlier is that you subscribe
Speaker A
to uh Kantian ethics theontology basically utilitarianism and uh and virtue ethics in different contexts and I think that's the answer that you gave yourself um which contradicts what you said because I do think that if the the ethics you hold are context dependent
Speaker A
then uh there isn't a possibility for you to develop something that is uh context independent as a solution to philosophy. So if if you always have a solution that is context dependent um the space of possible context is a large
Speaker A
infinity and I don't think like if you ever include that large infinity into any of the solutions you're trying to propose that uh that you're going to arrive at like a clear-cut like solved answer and in that sense I I do think
Speaker A
that we can make progress on on ethics and we've seen like moral circle expansion you know like from hey I'm the most important thing in the universe and I'm going to care about my own needs and values. um to like hey my kin is also
Speaker A
important maybe my my my my family to then oh my tribe is important to then oh you know this like city of of people in this country the world humans at large all sensient beings plants whatever right like we can continue um and and I
Speaker A
and I think like that is is is clear to me what is not clear to me is that we can solve philosophy and be uh have a one- all beall answer to how to uh do morality And that's precisely because of
Speaker A
this uh this uh the context dependent nature of uh of ethics at least for for humans. If you if you take humans out of the loop, um maybe there's a and and and and then I'm thinking about like a
Speaker A
utilitarian type ethic where you know if you take humans out of the loop and you have perfectly rational beings um and and actors that like go out there and just you know uh have to figure out the most optimal way of coordinating
Speaker A
resources and like collaborating towards some shared goal. maybe that is more like mathematically tractable uh as like a a a a problem that that resembles a a like centralized planner type problem.
Speaker A
Um the the other thing I wanted to add to this is um that and I need to do a bathroom break.
Speaker A
Yeah, I mean I think I think now that you say it I think you are right. It's somewhat a similar problem of solution space as solving alignment on current neural networks. It's probably not possible to get the perfect answer. Um
Speaker A
and then it's what your definition of solved is like I think we can get insanely far is my point. Um and uh and then I think like the three contexts that I talk about when I talk about these three types of ethics is in
Speaker A
when I evaluate a person's character, it's virtue ethics, right? when I evaluate a person's actions, it's utilitarianism. And when I need to run a a country and create laws, it's deontology. The like laws are fundamentally deontological, right? Um
Speaker A
and I think these have adapted to the specific contexts in which they're useful. I think we will be able to have arbitrarily complex laws now with AI as well. And so you can approach like this the infinite space. Um obviously you
Speaker A
cannot approach infinity but or you can approach I guess but you can't be close to it. um where the laws of a country would be so complex that no human could ever read it but an AI could enact on it, right? And
Speaker A
could could say this is this is like legal, this is not and could answer any question according to a symbolic program.
Speaker A
And this is the kind of thing that I think would be like the the useful solution to ethics in our current um rightsoriented societies. um and and you know any any context on top of that is like whether we can solve the the full
Speaker A
problem is a is a question I don't like I don't even have the question to like posit it and I don't think we can necessarily talk about what that means yet.
Speaker A
Yeah. I love the fact that we have armchairs uh to do armchair philosophy. Yeah.
Speaker A
Maybe I'll bring us back down to to earth and uh and to to possible futures.
Speaker A
I think this podcast is uh you know around possible transition scenarios and possible plausible attractable transition scenarios to a world with super intelligence. I think you and I both agree that um we are on the path to create these superhuman systems and um
Speaker A
getting there you know requires us navigating uh navigating treacherous waters. Um but it also necessarily has us define positive visions and I think that's underexplored and I would like to devote some time to this. Um in a year from now
Speaker A
what kind of technology institution policy change would you like to see? And you know if if that sort of like thing got installed how would you know your normal Tuesday look like? you wake up and walk me through sort of the first thing you
Speaker A
touch, see, talk to, think. Yeah. I mean, I think it's if if we talk one year, um what I would most like to see, you know, what what do you need to enact a global pause is like the
Speaker A
ambition to think about new international governance schemes. And I think the the really ambitious version of the resulting global coordination scheme for this is like a kind of UN that is much more empowered and stronger. Right? Like last time we
Speaker A
created that we had to have a world war and the last world war created something that looked like something that worked but it didn't actually work. League of Nations. And so the next war had to happen for us to create something that
Speaker A
did work. And now it's like dying again. And so with the arrays, we either get World War II or we preempt World War II with the governance solution. So I think when like waking up, it's about, you know, being sent relevant requests for
Speaker A
opinions on AI safety or something as as one of the experts in the field um along with everyone else. Those everyone else that like has the relevant expertise in this area along with the experts in all the other areas. I respond to these
Speaker A
requests. They're like pulled into a kind of governance schema where we can coordinate according to like the best possible estimate of what is correct and not just based on a technocracy but based on a kind of citizen science with
Speaker A
an opinionated opinion waiting system right um and an intelligent aggregator as well and so I think that would be one of the things where you know you wake up instead of social media you've got something that like empowers you creates
Speaker A
something pushes you towards something you um maybe directly there you've got the interaction where I give this advice and it's pulled and today's like AI governance decisions that need to happen at insane speeds to follow capabilities or something like it um are implemented
Speaker A
immediately across the world um I you know I come from Denmark code making is pretty freaking awesome and uh that is a great life to live there there is no one who who was like delusional there about what's necessary
Speaker A
to keep that society going. You know, we fund a lot of the war uh with Russia because we know it's on the eastern borders of our country. And so this kind of thing, maybe we've solved the negative parts of this and can get some
Speaker A
of the good parts along with more and more abundance. Like one issue I have, for example, if I talk with people who supposedly um are like bullish on making the world explode, you know, futurism, etc., is that they're like, "Oh, yes. Well,
Speaker A
Denmark is great, so everyone should just have it like they have it in Denmark because it's number one in livability, whatever." But no, every Danish family doesn't have a private jet, you know, or like whatever equivalent, an electric helicopter to
Speaker A
get to where they want to be. They don't have the chance to be on holiday whenever they want. Um, and there's some, you know, benefits you don't want like full Hidonia. There's like something in suffering, right? But, but
Speaker A
we're not there yet. and we will never necessarily be there. And I don't think we should extract all the resources from Earth. Like I think the ultimate one of the most ultimate things humanity could achieve would be a fully circular
Speaker A
economy. And I think people completely underestimate the like what we could achieve in that case in terms of getting all materials recycled. And so I imagine I wouldn't have to do anything to recycle my materials, but during the day
Speaker A
I would be living in extreme abundance with materials that aren't extractive from Earth or societies that are, you know, treated in a subhuman way. Uh like uh lithium ion uh battery production relying on on lithium extraction and places where child labor is used. And um
Speaker A
and we can just like get everyone on the on the same level and help all of humanity and like think of the lowest rungs first and then once they've come up here we can like push everyone further right?
Speaker A
Um and I think that's the kind of society I'd want, you know. So I I hear you on like the grand scheme of things and the the society you want, but what's on the news during that particular day on Tuesday? you zip on
Speaker A
and what do you hear? Well, maybe there's like two kinds of news. There's like AI capabilities and like the frontiers of technology leading to potential risks that we like accept as consequences but have very good control for. So maybe it's like another
Speaker A
pandemic averted in Nigeria or something like it that I'll see and I'll be like, "Oh, thank God." you know, and some serious news about like the kinds of things that are being implemented to now solve that issue. And maybe another
Speaker A
thing is the the positive news about like the people that have how many people have been brought out of poverty.
Speaker A
You know, what is the what are what are the global metrics we are measuring and maybe every day, you know, new starts with um yeah, fewer people dead today, more people brought out of poverty and damn that is good, right? uh that that
Speaker A
kind of thing I would love to see the the news pivot more towards now that we have a more data based society and it's only going to get get better right um and I think data is unreliable in
Speaker A
some cases you know look at look at China um their state reporting of of data is off um in ways that many people know in the industry right um and but but we can get to a place where we can
Speaker A
sort of coordinate around this because we need this amount of transparency from people to govern in a useful fashion um and everyone actually collaborating instead of um fighting against each other. I think you know interesting things about human
Speaker A
maybe philosophy. I also think you know what was it there was this blog post from Lucas Anderson Lucas Peterson on on the only career that will be left is football star because that's the human actor like sort of being the cultural
Speaker A
event and everything else you know your barber will be a robot etc. Um, so, so maybe there's still all these things that make us human and make us enjoy things and, uh, yeah, got a family, maybe more, who knows? Yeah.
Speaker A
What about like closer to when that pause is about to end? Let's fast forward 10 years. It's the day before the pause gets lifted. You wake up.
Speaker A
Um, every person on earth obviously knows this is the day the pause is lifted. It is the biggest event in human history and it is the highest risk and there are extreme amounts of news about the kinds of things we have implemented
Speaker A
to uh to make sure the lift goes well and that AI will be beneficial and safe for humanity. So I think it'll be like a global conversation and global news on what exactly is going to happen now and
Speaker A
what are we doing, how are we stopping again if it goes out of control and what do we want to see and if we see it, it's like verified that we'll auto close all the data centers if something along
Speaker A
these lines or these lines happen and everyone's on board and thinks it's a good plan and reasonable risk to run for humanity's sake. H when you say everyone, I'm like envisioning you get a call from your sibling, your mother,
Speaker A
your grandmother, whatever. And uh what's the level of knowledge you assume these people will have about that particular issue and how does that differ from today?
Speaker A
Yeah, I think obviously this conversation is sort of we're two AI tech bros, right, having this conversation. So, we know a lot about it and I think we don't have to like fool ourselves in this. I think the world
Speaker A
that I would love to have seen in the meantime is one where people are brought in on informative discourse rather than on addictive social media which is like heron of our day right um and so I would expect them to be much more
Speaker A
knowledgeable than for example about the the tax system as they are today you know where everyone's affected by the tax system but people don't know like too much about it except they have to give this amount um despite it being
Speaker A
very relevant for their existence and so I I think it would be a thing where we have designed it so it's engaging enough that people are actually excited about it either in like you know I don't know short form media or Twitch streamers are
Speaker A
talking about it you know and being like oo let's watch let's watch the the the opening up event again you know this kind of thing I see I see I see and if do you have any concrete ideas for this kind of um
Speaker A
discourse enabling technology because it sounds like uh that already exists for you and your future in a year and that you'd hope is uh one of the main drivers of like increasing public discourse throughout the decade that you're
Speaker A
envisioning. Yeah, I mean I think it it like a version of it will exist that maybe someone like me has adopted but then it'll become the standard through a a global coordination like legislation or whatever we need to do. Um and I think
Speaker A
we need it to sort of unify humanity towards this one mission and avoid wars to the degree they have existed. award both American imperialism and Chinese like Russian imperialism and then Chinese like um um the what whatever they they would call their expansion of
Speaker A
of control in the world um and where everyone can be on board with it because we have brought up the poor people and their societies the middle powers also have a say and it's not like a it's not even four poles of
Speaker A
power in the world it's like you know it's 20 or something like it um and the African Union has been strengthened. Um maybe the European Union has has expanded Ukraine, Canada, who knows this kind of thing. Um where there people
Speaker A
just think things will continue as they are. But I think we can achieve something where coordination happens well enough that we can actually get to a stage where discourse is possible both on the individual level and on the
Speaker A
engagement of the populace all the way to to the the leading people. Yeah, this reminds me of um AI 2040, the you know plan A that was set forward by Daniel Kokaton team at AI Futurist Project, the same gang that wrote AI 2027. And um one
Speaker A
of the enabling sort of uh yeah cultural shifts really that could lead us to to a more well-informed discourse uh both at the expert level meaning you know uh AI scientists uh and AI safety researchers um engineers etc uh but the general
Speaker A
populace as well is through what they are calling for in terms of complete and total research transparency right um meaning that each and every bit of AI research and you know AI safety research for that matter as a subset of it um
Speaker A
would be visible to everybody um and inspectable um right away. Yeah. Which would ultimately lead to what you were pointing at, right? like a multipolar world in which there's not just one power or like the few big labs
Speaker A
but you know the the the dissemination of that knowledge would lead to others being able to so sort of catch up more more more uh quickly and build uh build out frontier solutions as well um which under the regime of pausing the
Speaker A
development of AI meaning you know at the current frontier would still be uh something that like uh competitors could do because they could still sort of like lift up to that level where the pause was instantiated and then um you know
Speaker A
that I think would uh give um governments the ability to build out data centers and like the electric uh infrastructure necessary to power these data centers, you know, maybe like a bunch of nuclear plants or uh or solar
Speaker A
power plants. Um and and then yes, ultimately uh I I think sort of like having that buildout would would lead to countries that currently and historically have not been able to participate in this race, um be able to
Speaker A
have a sit at the table and uh that I think itself is uh very positive in in in its powers to diffuse the race. On another note, uh one of the companies you you've been um supporting, Lucid Compute, works on doing the exact sort
Speaker A
of like um uh trick here, but for like the the the compute stack and being able to verify what is being trained on u massive amounts of compute and data centers uh around the globe without like having full visibility. Yeah.
Speaker A
But maximal trust. Can can you explain a bit more about like what that kind of technology is doing here and where we currently are with uh said work?
Speaker A
Yeah, I think one thing that's often overlooked is the physical infrastructure necessary to get out all these theoretical plans, right? Um because one of the things that do not exist really today are these safety level five data centers coming from a
Speaker A
rand report which are data centers that are fully safe, secure, and verifiable. So I can trust that if I put a model on there, it cannot escape. Others can't take it out and like I can verify every action that happens on it. Um and these
Speaker A
kind of things are just necessary for the future. Like I think no frontier model especially not after a pause like once we open up will ever run on a data center that exists today because they're so unsafe, right? They like the AI
Speaker A
models can just go on the internet. This is crazy, right? each of the data centers might have a copy of the internet that they can work with in in 10 years or we've like solved some of the safety problems enough that we can
Speaker A
we can trust them to expand, right? But um but it's just it's just like a necessary technology and I think one of the most important technologies. Um the the problem is that it'll only be like, you know, it's not the cheapest right
Speaker A
now as compared to some of the other normal data centers because they're not secure and it's easier to build. Um uh so I think the market is emerging right now for them to do it for um for-profit companies and for enterprises, but that
Speaker A
we will probably in the only regime that makes sense in this global coordination schema make every data center that runs any kind of frontier AI even tool AIS because they're still going to be dangerous. Alpha fold can make dangerous
Speaker A
protein right? um where we need all of these models to run on these SL5 data centers and I don't know exactly what the like stack of features we need from it are like we definitely need to know what happens on
Speaker A
it in a way that I can verify as um you know like Germany that Denmark isn't training like a military AI against Germany or something which is very plausible thing to happen so but but these kind of things
Speaker A
they're building it right now. Yes, exactly. I can't say anything. Neither confirm nor deny.
Speaker A
This is another piece of technology that I assume is sort of on the news. The way I see it is um it would be nice to see a progress bar. And then each day I open up the news, you know, I see, okay, uh
Speaker A
we're we're we're like 15.8% done with the build out of like the infrastructure necessary to reach like uh wouldn't that be exciting?
Speaker A
Yeah. Day X in future where like you know the the the ban has been lifted and you know we're on track we will be three years early so that we'll have like tests we can still run in the in the uh
Speaker A
last sort of like uh three years of of after the buildout has been done and uh where where you know like it enables us to also test uh just scenario planning and like you know have a a a a small
Speaker A
model that is like intentionally developed ed with like some boundary testing assumptions and like some malicious intent by red teamers, right?
Speaker A
Uh for for for us to then know whether you know the alarm bells will ring in uh in you know let's say Australia and they can hit the button on of like shutting down the development of that tech there.
Speaker A
Um and yeah I I'm envisioning that that bar now. It's it's on fire actually like right it's cool and I think I mean we are luckily moving away from this SAS slop world and like building rockets to space lots of new 3D printing
Speaker A
technologies to make production more sustainable all this stuff. Um and so I think we we will as a as a community be able to do this and that bar can go really fast on if we decide what is
Speaker A
necessary and what will get funded um by the necessary parties. Uh talking about what gets funded uh at Selen Labs you have a front row seat to look at the companies building AI safety and insurance technologies. We've talked a
Speaker A
bit about the the companies in your portfolio from uh southern um you know cohort number one. I wonder what you've seen in cohort number two and then uh later we'll get to whether there'll be a cohort number three and what we can
Speaker A
expect from that. But uh what what companies have you seen in the second cohort that excite you?
Speaker A
So with the second cohort um we took them a bit earlier and helped them more you know we're more counterfactual in their journey um which I think has been quite exciting um I am you know I'm also with Juniper investing in companies that
Speaker A
are further along too um but the second batch like some examples of that is we've mind by canon I think you interviewed him yeah show yeah and and this is like a new programming language for how models should actually write programs where
Speaker A
each of the you know instead of writing a for loop you write a an SMS block right and this SMS block or whatever has been verified in some way so you can like create a much safer program because
Speaker A
these modules have been solved for um and I think it's like uh it's much faster to program with and much safer and I think it's like one of the most interesting ideas and kan definitely has like many very interesting ideas for AI
Speaker A
safety in general h so so that has been very interesting to follow. Yeah, it's it's it's amazing to see that you know this is the kind of technology that allows you know policy makers and nontechnical people to implement their ideas and um
Speaker A
you know basically instantiate policy as code, rules as code and and do so in a way that is economically viable with the current uh paradigm even much faster than what we currently use to write code. And so if
Speaker A
if you know if Conto gets that right I think that's that's a huge bet um uncertain whether it will work out but you know if it does uh and he develops it or somebody else picks up the ideas
Speaker A
the and like gamles in that mimetic space I think we we can we can expect to see uh many more non-technical people get as excited as the technical people that have seen sort of like the rise of cloud code and how it changed their
Speaker A
work. Um, and yeah, I'm just thinking back to some software engineer working in like at SAP, a German sort of like software development uh company. It's the the IBM of Germany, if you will. Uh, huge company, but that person sort of
Speaker A
has been sitting there uh and and developing code for the past decade. Nothing really changed. And when Claude Code came out and they had access to like work with it at work, their entire like career and and mind
Speaker A
shift changed and all of a sudden like that person started building again and being full of positive emotions. And I think if you can get that over to policy makers and enable them, empower them to like build policy fast and have it be
Speaker A
like implemented in the fabric of the technology as a get-go, right? Like and not not just as um as a policy that then somebody has to implement for them. No, it's the the block exists and the the the the code just doesn't run because
Speaker A
it's uh enforced at compile time and it just like you know if you if you violate uh that rule then like you know the program just doesn't run and it's you know you no longer have to trust you you
Speaker A
can verify and and you know you you you you have full control over what's happening and um that's where like uh I can see for example the EU AI act sort of being supercharged uh because like yes all of the the mind work has been on
Speaker A
figuring out what should be put into into action and like what policies we want and need. Um yeah, it's just the implementation layer that like usually takes quite a while and that gets me excited as well about Conton's work. But
Speaker A
yeah, what what else has have you seen? Yeah, so so that's been really exciting.
Speaker A
Um, we've seen also with Serial Labs that I've been like building um environments for industrial cyber offense and defense scenarios and then actually trying to solve for it because I don't think anyone really does like industrial um like industrial cyber
Speaker A
environments like um you know electrical grid companies etc. Um and and in terms of the specific domain where there are a few cyber players and now they then work with like Redwood Research and others to build out these environments that look
Speaker A
fully like a real company. And so you can see the like social manipulation happen. You can see all the the the whole code base that's like a real app instead of some contrived example like the paper I made that was a catastrophic
Speaker A
cyber benchmark where we evaluated this like how like how good is it at this specific vulnerability search um that applies to this specific uh AT and K or attack um category my okay so when if I hear you correctly
Speaker A
this is sort of the pond on to what project Glass Wing was but for the uh real economy and Yeah. Okay. And not just hardening software systems but also hardening sort of the Yeah.
Speaker A
software systems underlying physical infrastructure. Yeah. Okay. Yeah. And I think like one of the you know problems with Project Glass Wing to one degree is like this tool we've built can hack you. you should buy our tool to
Speaker A
avoid us hacking you. Oh, wait. It's not us. It's our customers hacking you. And you're like, okay, that's a bit weird.
Speaker A
So, like if you're a cyber security company, you have much better incentives where you get paid the more secure you make a company. So, I see. What about batch number three? Is that in the pipeline?
Speaker A
So, we have a lot of we're supporting a lot of founding teams. the batch three is like um we're not going to run that right now and I think we're going to see what is the best strategy to accelerate
Speaker A
the field the most and we have some we have some good ideas and and there's going to you know there's more announcements to come um and more projects also with our partners uh in in various firms where we like want to
Speaker A
invest in and accelerate founders um lots of crazy ideas happening right now so yeah look I see I see I see let's see Well, I I guess we can also talk a bit more about Juniper. Uh you've alluded to it left
Speaker A
and right, but um tell me more about what the core thesis behind this venture capital uh company is. Fundamentally, it's about investing in and helping all of the founders in AI resilience and AI assurance. So, we've published the AI
Speaker A
assurance technology market report uh that you can find online that sort of shows the whole case for the field and we're going to release a follow-up in a few months.
Speaker A
2.0. It's going to be good. We will talk then again. Yeah, for sure. uh and and I think this market is like you can think of it as the assurance infrastructure which includes everything from interpretability and authentication all
Speaker A
the way to energy infrastructure hardware chip design and the controls we need along the whole pipeline and the whole supply chain. So we've invest invested in a lot of companies also ones that like automate uh the content moderation even at that stage it's like
Speaker A
content moderation is a type of alignment problem where AI generated content can we filter that on the receiving side and so this is quite a weird company from the traditional AI safety sense but it works really well here um and of course Lucid computing an
Speaker A
example of the hardware companies that need to exist. Mhm. What's the endgame for for uh Juniper? Where where are we headed?
Speaker A
Well, we just want to make more and more things happen with bigger and bigger people that has a larger and larger downstream impact towards this optimal future that we're seeing. And whatever we can do to make that happen, I think
Speaker A
it's structured well in a VC firm. I think it's one of the most cracked teams I can imagine to do this. So, uh it's it's been really exciting and and we're only going to get bigger, I think. when
Speaker A
when you when you talk to founders, how do you evaluate uh their fit for this particular type of work that is still extremely niche and doesn't have a lot of precedents uh from which to draw from? Well, I think you got to be a good
Speaker A
founder. This is a very rare quality. So, this is like a prerequisite. And to me, that means that if I meet you one week and ask, "Oh, what are your what are your problems, your current questions that you're looking at?" Then
Speaker A
you give me an answer. I'm like, "Okay, interesting." Then I meet you a week later and you have none of those problems or questions anymore because you've solved them all. And uh and that's this like improvement over time
Speaker A
and just the speed, the rate at which you can change your worldview or your business in a way that adapts to the circumstances is just necessary. And I think a lot of people, you know, who come from the the slower world of
Speaker A
academia, for example, have challenges in that. I think empirical machine learning research is like a field you can come in from because a lot of that is like react to the benchmark number change change fix it deploy it it's quite startup
Speaker A
looking which I think is the same for cognitive science but just like you hire users which are human subjects for psychological studies all this stuff um and and so these are the kind of founders that you know you
Speaker A
need that to be a founder and then the other part of it is that we obviously need you to like actually want to do this like do impact startups. The the level that I would say is similar to if
Speaker A
you look at green tech what does a green techch founder work on like does the solar farm developer in China do they think about the climate a lot? they probably think about it quite a bit and like regulation around climate control
Speaker A
uh not climate control but climate change does affect their business so they do think about it and so on but they don't need to be that aligned with the mission they just need to build more solar and I think the specific ideas
Speaker A
people work on is what really matters here and that idea will obviously come from how much they think about AI assurance safety etc but um yeah I don't think it's like an insane prerequisite Yeah. Okay. So, I'm I'm seeing like
Speaker A
problem solving turnover time in in the sense of like, you know, talk to Aspen each week and have a new set of uh more ambitious and like scaled up problems from the week you had above uh before.
Speaker A
Um and then obviously mission alignment and sort of uh a yeah a a predisposition to find this uh type of work to be um meaningful and invigorating. rather than something you chose to work on because it's the you know like rational thing to
Speaker A
work on that doesn't actually you know like trigger you in the way that is necessary to muster a a large amount of energy necessary to have this fast uh problem turnover time.
Speaker A
Okay, I I hear you on those fronts. Um, I think one one thing I I read about uh how what kind of advice you we give these people is um to shoot for uh unicorn status within 12 months. And um
Speaker A
I'm I'm kind of c curious how you square that with um the incentives that come with uh scaling a company that fast um to sort of like yeah uh trying to stay aligned and and and making um you know
Speaker A
assurance tech go well rather than just like focusing hard on profits and on sort of the incentives coming from venture capitalists like yourself who on the other end benefit from from uh large a large large uplift in valuation of
Speaker A
these companies, right? And of course, when we come early in a company, we're the aligned investor as compared to many of the just purely capitalist ones, right? So, but we are still pragmatic. It has to be a good
Speaker A
business before we can invest. Otherwise, you should be a nonprofit, right? And I think if you don't want to scale to a unicorn status in under these timelines of AGI development within 12 to 18 months, then um like why are you
Speaker A
doing a for-profit, right? Um why are you doing um um why are you doing a venture backable startup? I think you can do a for-profit you know that's different from a venturebacked startup which does need to do this and like is
Speaker A
designed to do it where you forego current profit to grow right to scale. Um and and to me it's just like if you want to change the world you need to reach this scale at that time frame like
Speaker A
what else are we doing here guys? like you know um then you should be in research if you don't want to do it.
Speaker A
I just think it's like the only option we have. I don't think there is another option.
Speaker A
Right. The urgency comes from your beliefs that we will um that the rate at which we make technological progress on the AI front is accelerating. Correct.
Speaker A
Okay. Uh I mean you know I will bother you with the timeline questions that and so uh when do you expect I have to go.
Speaker A
So you the urgency comes from you expecting that these systems will arrive rather soon. Um give me some broad like time frame in which you expect things to be shaken up massively.
Speaker A
I mean I think it's already being shaken up massively. So that's number one. I think to me AGI is a mechanistic term for a specific algorithm designed to be general. And so I think with GT2 we sort of achieved that because it's lost its
Speaker A
meaning now that people when they say AGI they mean ASI um the term where the AI is better than any human expert at their area of expertise which I think might be 2029 or so. Um, and that 2029 I
Speaker A
I mean I said it in 2022 and I haven't really changed it since. I changed it to 2028 last year and then I went back to 2029 because of the limitations I saw when I began building like as like one
Speaker A
of the serious startups along the way and it's based on Rick Herzswhile's original extrapolations from 2003 in the singularity is near and I think it was surprisingly unmarked in terms of taking all of these trajectories of technological development merging them and saying like
Speaker A
okay where do they where do they lead to AGI And um I think in 2022 people were thinking it was 2043 or 2050 right and I just did a meta modeling on both the recurs wild point but also on the best
Speaker A
forecasters in the world's forecasts and how their forecasts changed and their forecast would go from you know 2047 or whatever to 2040 the next year and then you're like okay if it changes seven years in one year there's some like
Speaker A
approaching to a number And and so I thought 2029 fit quite well and that's you know done me wonders like I think it's pretty accurate now on on when we will achieve ASI and like relevant economic domains.
Speaker A
Yeah. It might not be in robotics but Yeah. Yeah. I think yeah I like the the definition of like ASI or like super intelligence basically as like you know a system that is more capable than humans at any given task. And um
Speaker A
I for one I I think I include robotics in that as well. Yeah. because then you have the the absolute full set, right?
Speaker A
Um I think that it's very reasonable to include it. I think like I know a little bit too little about robotics to say stuff here. I think um we we you look at a lot of robotics, right? But uh
Speaker A
but robotics is really really hard and they're developing at a breakneck pace. So I think 2029 could be the time for that as well. H but otherwise robotics physical tasks might take another year or two. Yeah. Okay. So, so I I get how
Speaker A
that fits into your worldview, especially with the pause that you're advocating for as well. Um, and that gives us roughly two years to instantiate something like this.
Speaker A
I think sort of wrapping up, uh, we we've now talked about like an assumption, a prediction, a guess that you you haven't changed. Um, are there any assumptions you have gotten wrong about this field and where you changed
Speaker A
your mind significantly in the past, you know, three, four years that you devoted a good amount of time to to working in this field? As you ask it, I I can't come up with something specific, but I'm sure I'm wrong, you know, because there
Speaker A
must be some big ones otherwise I'm crazy here, you know, then I should have solved it four years ago.
Speaker A
Um, I think some of it is like I quite early thought that interpretability was only as useful as it has been. Um, and it's turned out to be less useful than I thought, but also still very important.
Speaker A
I think the level of trust I have in companies has gone a little bit up and down. I mean, I come from a welfare state where we mistrust any sort of power, government or corporate. They're all the same kind of power, right? Um except
Speaker A
monopoly and violence, but that requires even more scrutiny then. And I think like coming into it, I'm like not delusional about the incentives that underly the system. And so I did not expect that anthropic would be so misaligned as they
Speaker A
are, but it it doesn't come as a surprise, if that makes sense. Like it's it fits in my worldview. Um but the fact that they are the most accelerationist lab and in 2023 they didn't want to do any products because they were scared of
Speaker A
accelerating what happened guys you know and they are not being transparent here you know they haven't shared the reason why they changed their strategy and obviously it's to raise more money to be the good guys with AGI but what will you do once
Speaker A
you have ASI is it you're going to take over the world is it you're going to give it to the US government is it you're going to create a UN project and when I like I don't think the lower like
Speaker A
the people lower down in entropic do know this. I don't think it's shared with them either. I think it's just at the top. Um and I don't even know if they know why they're doing it besides the incentives of the system.
Speaker A
Well, Daria wrote, you know, machines of loving grace. Uh but they I I I do feel there's a disconnect as well between like what they're developing and like the automation loop they've kicked in.
Speaker A
They're basically as any other lab but they basically were the first to start automating coding because that's the most direct path to automating their entire like AI business.
Speaker A
Yes. So automating anthropic comes through automating coding. That's also why they leaprogged OpenAI because OpenAI made that mistake of not focusing on on this.
Speaker A
Yeah. Now everybody is but and I think we've seen them lie more and more to increase their revenues. Like examples include them, you know, saying that Fable or the latest models are the most aligned models they've ever released. And very simple tests show
Speaker A
that they're god-awwfully aligned. that, you know, one test that happened was Ender Labs ran this long horizon evaluation where they looked at how much does it, you know, cheat because they realized they got a bunch of weird results and then they figured out that
Speaker A
the models were cheating a lot and they recoded all their transcripts from the last couple years and in 2024 the models cheated 0.6% of the time and this year they cheat 50% of the time. So, you cannot come out here and tell us that
Speaker A
it's the most aligned model you've ever released. This is fake news. You are just doing it to prop up your own coffers.
Speaker A
What is going on, guys? You know, and and it feels like a cult when you talk with anthropic employees that are not in the know. I'm sure some of the upper folks have a plan or something, but it's maybe not going to be a plan
Speaker A
that includes democracies, you know. So, so I think this has been a surprise to me. how very misaligned they have been purely because of capitalism, like purely because of economic incentives.
Speaker A
Um, and I don't know where we where we'll get with that, but that that has been one of the biggest surprises.
Speaker A
Yeah, we'll keep a close eye on that and revisit it when we speak next time. Um, before we wrap this up, I want to know what people can do to support your work right now. What you're looking for and,
Speaker A
uh, what you're looking to give. Well, uh, we'd love to mentor great people. Uh we are having a stream at Matts uh for founders, prospective founders within this field that I think people should apply for uh if they're
Speaker A
interested. We also if you are a founder that's currently doing $100 million revenue in some company or have a unicorn, I think you should quit and do AI assurance or pivot the company radically to AI assurance and we are
Speaker A
very happy to help you. We've got the connection to the billionaires, to the labs, to the customers, to the providers, etc. whatever you need. Um, so I think this is like what we really want to see is a radical reshaping of
Speaker A
where the best talent deploy their uh capital and labor. Awesome. So everybody
Topics:AGI banAI safetyArtificial SuperintelligenceEsben KranAI governanceexistential riskAI alignmentAI ethicsAI progressWe Humans & AI










![Zanjeerain Episode 02 [Eng Sub] 30th Apr 2026 | ft. Saj… — Transcript](https://i.ytimg.com/vi/HWYmuRVH-t0/maxresdefault.jpg)
