Skip to content

EP 105. AI가 연구를 할 수 있을까? — Transcript

Exploring whether AI can conduct research, challenges in peer review, and the evolving academic landscape amid AI advancements.

Key Takeaways

  • AI can contribute to research but evaluating its output remains complex and uncertain.
  • Traditional peer review systems are strained by the volume of submissions and AI-generated content.
  • Preprint servers like arXiv face challenges balancing openness with quality control.
  • Good research fundamentally expands human knowledge and positively impacts the world.
  • The academic community must adapt to new realities brought by AI in both research and review processes.

Summary

  • The video discusses the question: Can AI do research, exploring examples like PPO and its unexpected impact on LLMs.
  • It highlights the difficulty in evaluating research quality, with many good studies initially rejected by prestigious conferences.
  • The role of arXiv as a preprint server is examined, including recent restrictions due to the flood of AI-generated papers.
  • The massive increase in paper submissions and acceptance rates at major conferences like ICML and ICLR is analyzed.
  • Concerns about AI-assisted peer review and its impact on review quality are raised, with some reviewers caught using AI improperly.
  • The fundamental definition of good research is debated, emphasizing expanding human knowledge and positive world impact.
  • The video touches on the challenges of maintaining research quality amid rapid production and AI involvement.
  • It reflects on the evolving academic system and the need for new methods to evaluate and review research effectively.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
[Seungjoon Choi] Hello. Today, as we're recording, is July 18th, 2026. Let's get started. Jonghyun, what have you prepared for us?
00:10
Speaker A
[Jonghyun Park] First, on the subject of AI and research, that's what I'd like to discuss.
00:14
Speaker A
After going through the major conference season, meeting researchers, and seeing a lot of related research, a question came to mind.
00:19
Speaker A
It's one of the defining questions of our time. Can AI do research? We'll look at some major examples, how things are currently unfolding, and then share our thoughts.
00:30
Speaker A
The first thing I'd like to introduce is John Schulman, a co-founder of OpenAI, who is now at Thinking Machines Lab.
00:38
Speaker A
In reinforcement learning, there's an algorithm called PPO, which was especially prominent when ChatGPT was at its peak, and was applied to RLHF, the technique of doing RL with human feedback.
00:47
Speaker A
It's the most commonly used algorithm, but apparently the paper on this algorithm was actually rejected in 2017.
00:54
Speaker A
[Seungjoon Choi] So that happened.
01:01
Speaker A
[Jonghyun Park] I didn't know either. This tweet was posted a month ago, and it got an enormous number of views.
01:09
Speaker A
Everyone does recognize this issue. Just because a study is good doesn't necessarily mean it will be accepted by a prestigious conference.
01:24
Speaker A
It can get rejected. I think a lot of people related to that, and for reference, thinking about another paper by Tony Zhao that I reposted in response to this tweet, apparently that one was nearly rejected as well.
01:33
Speaker A
It's now a very popular study in the action space, but the reviewers themselves didn't view it favorably.
01:39
Speaker A
[Seungjoon Choi] They just don't know.
01:43
Speaker A
[Jonghyun Park] There are so many examples like this, and in every field, these stories come up all the time.
01:51
Speaker A
At least among people who have done research.
02:04
Speaker A
[Seungjoon Choi] Was PPO related to Policy Optimization?
02:13
Speaker A
[Jonghyun Park] I think it stood for Proximal Policy Optimization. Anyway, to give just a brief overview, this is how it was used in ChatGPT.
02:30
Speaker A
When we have a conversation, a person rates the conversation as good or bad, and the model is improved based on that feedback.
02:42
Speaker A
They initially used people, but because people have limitations, they also created models that imitate people, and it evolved in that direction, but that's basically what it was used for.
02:47
Speaker A
But what I wanted to discuss here goes beyond the rejection itself. Even the authors themselves didn't anticipate that PPO would later be used so extensively in GPT—that is, in LLMs—and have such a major impact.
02:53
Speaker A
Until it was later used in LLMs, neither the reviewers nor even the authors themselves could be certain that it would eventually have a major impact.
03:08
Speaker A
In other words, evaluating good research is itself extremely difficult. It's difficult to assign it a score.
03:20
Speaker A
Those are some points worth considering.
03:32
Speaker A
[Seungjoon Choi] It sounds like the difficulty of evaluation will be today's main topic.
03:38
Speaker A
[Jonghyun Park] It's difficult to develop good judgment. Looking at a few similar incidents that have recently occurred in academia, one interesting case involves arXiv. The PPO paper was ultimately just uploaded to arXiv, where everyone read it and it received an enormous number of citations,
03:44
Speaker A
but it wasn't published in any journal or conference and simply remained that way. Last year, arXiv said that certain papers in the CS category would only be accepted if they had passed peer review.
03:55
Speaker A
There was an announcement along those lines. The reason is very simple. As everyone can probably guess, there are simply too many AI-written papers flooding in, making quality control impossible and causing various other issues.
04:06
Speaker A
So they would no longer make everything public as they had before, except on a limited basis.
04:14
Speaker A
There was also a lot of heated debate about this on Hacker News. There were plenty of keyboard battles.
04:20
Speaker A
Those in favor argued that generating papers with an LLM is extremely cheap, whereas determining whether they are actually good and reviewing and evaluating them is expensive and difficult.
04:26
Speaker A
So on arXiv—though it may be harsh to put it this way—among papers written using LLMs, there are probably many that offer no meaningful contribution to research, the so-called garbage research.
04:31
Speaker A
If arXiv becomes flooded with that kind of slop, people will leave it. So even though it was created to make all good research readily available, this restriction is necessary.
04:40
Speaker A
Others argue that arXiv was originally meant for uploading and distributing papers before peer review.
04:46
Speaker A
It can uncover work like PPO, so how can they shut that down? Those seem to be the opposing views.
04:51
Speaker A
[Seungjoon Choi] It was for preprints. It was a place where papers could be released in advance, before being submitted to a journal.
05:01
Speaker A
The barrier to entry was low.
05:16
Speaker A
[Jonghyun Park] I'm not entirely sure why it was originally created, but as I understand it, it had existed since before the LLM era.
05:25
Speaker A
[Seungjoon Choi] Right. It existed in physics and mathematics.
05:33
Speaker A
[Jonghyun Park] It may have also been intended as a place to upload work that went undiscovered because of rejection, but that wasn't the main reason.
05:37
Speaker A
It was more because journals and conferences were simply too slow.
05:50
Speaker A
[Seungjoon Choi] It could take six months or even longer.
06:04
Speaker A
[Jonghyun Park] In rapidly evolving academic fields, by the time a paper was published, it was already a little late, so people also appreciated it as a tool for sharing knowledge more quickly.
06:17
Speaker A
[Seungjoon Choi] But now the world has rolled back.
06:26
Speaker A
[Jonghyun Park] Looking at recent conferences, we're seeing the same kind of thing. This year's ICML received 23,000 submissions, and reportedly accepted 6,000 papers.
06:35
Speaker A
I was in graduate school ten years ago, and back when I was a graduate student, the scale wasn't in the thousands, but I think it was around 100 papers.
06:40
Speaker A
We used to say that 100 published papers was a lot, things like that, but things have really changed.
06:52
Speaker A
[Seungjoon Choi] There are a lot of published papers themselves, too. 6,352 papers.
07:04
Speaker A
[Jonghyun Park] Indeed.
07:17
Speaker A
I was a little surprised that there were this many published papers as well. And similarly, at a conference called ICLR, there were an enormous number of reviews, and I heard that a significant portion of those reviews were written
07:23
Speaker A
entirely by AI. So not only has the average quality of papers declined because writing papers has become easier, but because the volume of reviews is unmanageable, people are also using AI to do the reviews, which causes the quality of reviews to decline as well.
07:34
Speaker A
There was similar talk at this year’s ICML. It seems they divided the review track policies into reviewing with AI and having actual humans review without AI, and proceeded with the tracks divided that way.
07:40
Speaker A
But I heard there were cases where people had agreed to review the papers themselves but were caught using AI to review them.
07:44
Speaker A
So they were excluded from reviewing, and in any case, it seems academia is now also undergoing changes to the system for evaluating and reviewing papers, this entire system.
07:48
Speaker A
Though I don’t know where the right answer will lead.
07:54
Speaker A
[Seungjoon Choi] Those problems arise because production is overwhelming the system.
08:00
Speaker A
[Jonghyun Park] Then let’s return to a more fundamental question: What is good research? Ultimately, the conference acceptance and rejection system and all these other things are systems designed to find good research.
08:11
Speaker A
As for what I consider good research, the fundamental meaning of research, first of all, is that if we take the entirety of human knowledge, it means expanding it, even by a little.
08:27
Speaker A
That is what I personally consider research. And if the area where that knowledge has been expanded actually has a significant impact on our world and changes the world for the better, then that is good research.
08:33
Speaker A
That is what I think. Seungjoon, do you happen to have a definition of good research?
08:39
Speaker A
[Seungjoon Choi] That’s a good question. I haven’t really thought about it, but I do agree that expanding and pioneering the frontier of human knowledge or wisdom is clearly one aspect of it.
08:53
Speaker A
And although I’m not certain, I understand that reviewers are also paid for many reviews.
08:59
Speaker A
[Seungjoon Choi] Is that so? [Jonghyun Park] As I understand it, that varies considerably by conference or journal.
09:03
Speaker A
In any case, it seems we have not found a system better than peer review.
09:10
Speaker A
At least in human society. But now, because of LLMs, cracks are beginning to appear, and this is one of the problems facing our generation.
09:19
Speaker A
As we have discussed several times regarding LLMs, using RLVR to develop intelligence based on good rewards is the most important way to improve the intelligence of existing LLMs.
09:34
Speaker A
So, in the case of mathematics, which we will discuss in more detail later, they have already reached the level of an IMO gold medal, a Mathematical Olympiad gold medal, over the past two years.
09:45
Speaker A
[Seungjoon Choi] I think it was a silver medal the year before last and a gold medal last year.
09:47
Speaker A
[Jonghyun Park] But in the case of research, because it is currently difficult to provide a verifiable reward, it also seems somewhat difficult for LLMs to advance toward conducting good research.
10:00
Speaker A
And similarly, in human society, it also seems difficult to effectively identify good research. I’ve brought some interesting attempts I’ve seen to solve this problem of our time to discuss.
10:10
Speaker A
The first one is currently underway. Hugging Face is running something called the ICML Reproduce Challenge.
10:18
Speaker A
In this challenge, agents actually reimplement the methods presented in papers and verify whether they work properly.
10:27
Speaker A
This is one of the issues that has always been highly controversial in academia. A paper says, “We experimented using this method, and achieved these improvements.” Then it is unveiled with a flourish, but when I actually try to reimplement it myself,
10:42
Speaker A
it very often doesn’t work. That means it cannot be reproduced. Then it is difficult to trust.
10:48
Speaker A
So how do we verify it? We could require them to release everything as open source, but in many cases, actual hardware experiments are involved, or the work was done using a closed-source simulator that is a critical asset, and often it cannot be made public,
11:03
Speaker A
so there are quite a few practical difficulties. Anyway, based solely on the reasoning presented in the papers, they are running this reproduce challenge.
11:12
Speaker A
So, among the many studies written by AI, some slop may contain mistakes in the experiments, be incorrect, or, in extreme cases, may not involve any experiments at all and may even use fabricated data.
11:28
Speaker A
They could just say, "According to my reasoning, the experimental results would probably look like this," and put that out there.
11:32
Speaker A
[Seungjoon Choi] So, were quite a few of them reproduced? [Jonghyun Park] I'm not sure about that either.
11:35
Speaker A
This is still the submission period, and it will probably end on August 2, so I think we'll have to take a look after that.
11:43
Speaker A
But this is just my personal prediction: out of those 6,000 papers, if you ask how many were actually reproduced, it looks like around 800 have been reproduced so far.
11:54
Speaker A
[Seungjoon Choi] But in any case, that means around 800 have been verified as reproducible.
11:59
Speaker A
That's still quite a lot out of 6,000. [Jonghyun Park] Yes, I think so. But I believe that may include refutations as well.
12:07
Speaker A
So, in any case, this reproduce challenge goes beyond peer review based solely on the text of a paper by actually running the work and verifying whether the data really comes out correctly, so it could be seen as a kind of
12:20
Speaker A
intermediate-stage reward. So it could be viewed as an intermediate stage of filtering for good research.
12:28
Speaker A
And then there's one more thing, which is particularly popular here in Korea right now: Ralphthon has become a popular meme, and a huge number of participants are having a great time with it.
12:39
Speaker A
I went to the event myself, and it's a hackathon challenge where you run Ralph.
12:45
Speaker A
Because the concept itself is fun, everyone can leave Ralph running and have fun talking among themselves, and it's also good for networking.
12:52
Speaker A
This time, to coincide with ICML, the Ralphthon was a research Ralphthon. So there were two hackathons. The first was writing a paper with an LLM.
13:02
Speaker A
Writing a paper all the way through actual research and experimentation until a paper is produced. And the second was reviewing those papers.
13:10
Speaker A
So Ralphthon was held with those two tracks, and I interviewed the people at the event.
13:19
Speaker A
But to digress for just a moment, professors came, doctors came, researchers came, and even high school students came.
13:30
Speaker A
We uploaded a video about it to our channel. Since we were live at the event, I saw someone who looked very young and asked how old they were, and they said they were 15.
13:40
Speaker A
Curious, I asked what they were working on, and they said they used LLMs extensively, but once they filled the context with data through prompts, the performance seemed to deteriorate at some point.
13:52
Speaker A
They wanted to know when and how it deteriorated. Should important data come at the beginning or the end?
13:58
Speaker A
Performance would vary depending on things like that, and they wanted to understand how. This student was in high school and probably had never conducted research at school or anywhere like that, but their own questions were truly close to research,
14:10
Speaker A
and they were independently working to answer those questions. I was really impressed to see that.
14:17
Speaker A
[Seungjoon Choi] As a high school student, they may have taken AP courses—I don't know how it works in the US or Canada— or encountered something early through an advanced program.
14:24
Speaker A
Anyway, so that happened. [Jonghyun Park] Yes, I found it quite impressive too. [Seungjoon Choi] But haven't you interviewed Gubong several times?
14:32
Speaker A
[Jonghyun Park] Yes, when I first interviewed Gubong, Ralphthon hadn't been created yet. The idea was wanting to give a Ralphthon a try.
14:41
Speaker A
[Seungjoon Choi] It was something where people stayed up all night, right? [Jonghyun Park] Yes, that was the wish, so during the first interview, I heard about that idea. Then Ralphthons were actually held, including one in San Francisco and one in Singapore,
14:53
Speaker A
and people overseas enjoyed them too. That young student also thought Ralphthon was unique and fun and wanted to attend, so they said they came while they were visiting Korea.
15:05
Speaker A
[Seungjoon Choi] I'll have to watch the video too. [Jonghyun Park] Anyway, what I found most encouraging was that, whether it's research or coding, building itself can be painful, but Gubong seems to have transformed it into something all the participants could enjoy,
15:22
Speaker A
which I think is really wonderful. If I may take this opportunity to promote Gubong, the person shown here is Gubong, who organizes this Ralphthon through Team Attention, and I have been thoroughly enjoying attending the events Gubong organizes.
15:39
Speaker A
I go whenever I have time, and there are always some fun elements there, so I always have a great time.
15:47
Speaker A
Anyway, getting back to the topic, they wrote papers with AI and reviewed them with AI, then ranked them by score and even held a poster session.
15:58
Speaker A
And of course, they also invited actual researchers to serve as judges. They explained how they judged the creation of a good review agent.
16:11
Speaker A
They selected all the good reviews based on how the judges themselves would review the papers.
16:17
Speaker A
And they actually selected one more, comparing the scores from the judges' reviews of the AI papers with the scores from the agents' reviews, and regardless of methodology, when they selected the one with the most similar scores, that review agent ended up winning.
16:37
Speaker A
So when I asked what the most interesting aspect of this was, Gubong highlighted this.
16:42
Speaker A
What we can take away from this is that the quality of a review is difficult to assess, just as we are currently struggling to evaluate papers.
16:53
Speaker A
Then defining what makes a good review is also difficult. Simply looking at the correlation with human reviewers may be better than various review methodologies, which seems to be what this implies.
17:07
Speaker A
The same applies to RLVR: if there is one good reward, you do not impose a particular method for reaching it and instead let it figure out how to solve the problem.
17:18
Speaker A
This methodology worked well in this case too. That gave us a few things to think about.
17:25
Speaker A
In any case, if we consider what these attempts themselves imply, the question is: how can we provide good rewards for research too?
17:34
Speaker A
Having humans provide each reward properly is extremely difficult, and the sheer volume is unmanageable, so the idea is to somehow use AI to do it well.
17:45
Speaker A
Efforts to provide rewards are also moving in this direction, but ultimately, this is not RLVR— that is, the rewards are not verifiable.
17:54
Speaker A
Nevertheless, the goal is to specify the rewards as precisely as possible somehow. Attempts in this direction also seem to be active in academia.
18:02
Speaker A
[Seungjoon Choi] It does seem that methods involving giving an LLM a rubric are being widely used these days, and this might fall into that category.
18:09
Speaker A
[Jonghyun Park] That's true. Now that you mention giving an LLM a rubric, something else suddenly comes to mind.
18:14
Speaker A
Kimi K3 has been released, and back in the days of Kimi K2, they successfully improved agent performance by giving an LLM a rubric and having it evaluate things in as much detail as possible, which, as I understand it, is why Kimi K2 performed so well.
18:28
Speaker A
You could say this follows the same direction. It also seems quite promising. [Seungjoon Choi] It feels like we are gradually moving into today's difficult part.
18:38
Speaker A
[Jonghyun Park] Seungjoon was also very interested in this, so we decided to look at it together, and this is what we found.
18:44
Speaker A
Two recent episodes of the Dwarkesh podcast were released. The first featured Grant Sanderson, who is known for the channel 3Blue1Brown, a famous mathematics educator— described as the best explainer there is.
18:59
Speaker A
It is an interview with someone who runs an excellent mathematics education channel on YouTube.
19:03
Speaker A
So that episode discussed mathematics, and the one after it featured [Seungjoon Choi] Adam Brown.
19:07
Speaker A
[Jonghyun Park] Adam Brown, yes. Adam Brown is at DeepMind now. and is probably conducting research on superintelligence there, and originally came from physics.
19:17
Speaker A
Physics at Stanford. So two episodes were released, one about physics and one about mathematics, and we would like to discuss what we saw and took away from them.
19:27
Speaker A
[Seungjoon Choi] Didn't you design a connecting thread to link this to the preceding discussion?
19:31
Speaker A
[Jonghyun Park] Ultimately, when we think about sciences like mathematics and physics, or fields of pure research, the question is whether AI can achieve a breakthrough in them— in other words, whether AI can conduct research.
19:46
Speaker A
I selected only the parts that focused on that question. Ultimately, it is about progress in mathematics. Yes, exactly. In mathematics, we have been hearing a lot of news lately.
19:55
Speaker A
Cases of OpenAI proving certain hypotheses or conjectures are beginning to emerge one after another, and I think we perceive such cases as breakthroughs.
20:07
Speaker A
This seems to reflect those expectations and the question of how far this can go. What surprised me while watching this video was that it begins with the following story.
20:15
Speaker A
They had also done an interview three years earlier, back when there was a great deal of discussion about AGI.
20:24
Speaker A
Would AGI happen, and when would it happen? Starting with the question of what AGI even is, they asked, “If AI wins an IMO gold medal, wouldn't that make it AGI?” In other words, if it is exceptionally good at mathematics,
20:34
Speaker A
doesn't that mean it is truly intelligent? Apparently, even then, Grant Sanderson said that it did not.
20:42
Speaker A
It is simply one benchmark. It would merely be succeeding at one benchmark, something that today's LLMs are extremely good at.
20:49
Speaker A
As was already the case three years ago, it would just be cracking that one benchmark.
20:53
Speaker A
It probably would not constitute general intelligence. But from where we are today, AI has won an IMO gold medal, and we do not call it AGI on that basis.
21:04
Speaker A
So Grant Sanderson was prescient. That seems to have made us pay even closer attention to those words.
21:10
Speaker A
[Seungjoon Choi] The IMO is, after all, a competition where exceptional high school students solve problems.
21:14
Speaker A
Models can now do mathematics of the kind that mathematicians would need to tackle, so the progress over the past year has been tremendous.
21:23
Speaker A
[Jonghyun Park] I also saw a lot of discussion along these lines when AI won the IMO gold medal.
21:28
Speaker A
There is a company in the GPU cluster business whose founder is an IMO gold medalist, and the founder said something along these lines on Latent Space.
21:38
Speaker A
An IMO gold medal is essentially the entry threshold for becoming a mathematician. At the IMO, contestants solve six problems in about nine hours.
21:49
Speaker A
It takes place over two days, so at most it is a ten-hour task. But actually conducting mathematical research is a task that takes ten thousand or tens of thousands of hours, so whether AI can advance to that level is the next challenge.
22:03
Speaker A
That is what the founder said. So although these are problems that only exceptionally intelligent humans can solve, you could also say that their scale is not very large.
22:14
Speaker A
So Grant Sanderson says that the hidden truth about the IMO is that you can train for it extensively.
22:21
Speaker A
In other words, you can keep practicing problems. [Seungjoon Choi] Even when a problem has novelty, it is still something you can train for.
22:26
Speaker A
[Jonghyun Park] Humans can. Yes, that seems to be what Grant Sanderson is saying. To use a simple analogy from my own world, if we were preparing for the mathematics section of the CSAT, we could train by solving numerous past questions and many similar problems.
22:39
Speaker A
that it feels like you could do well on the CSAT, and I think the point is that the IMO is no different in that respect.
22:46
Speaker A
But what they say here is that current AI—this is probably based on last year.
22:52
Speaker A
Last year's IMO gold-medal models were probably all similar, but mathematics itself is divided into different fields.
22:59
Speaker A
I mean, just as we say that being good only at math does not make something AGI, because it sometimes cannot do other things that seem so obvious, the same structure exists within mathematics—it's isomorphic, in other words.
23:11
Speaker A
Within mathematics, there are algebra, geometry, number theory, and combinatorics. Of those, it cannot solve combinatorics problems, while it solves the others very easily.
23:22
Speaker A
So even within this domain, there are distinct kinds of spiky intelligence. For reference, combinatorics means problems about counting the number of possible cases.
23:29
Speaker A
[Seungjoon Choi] I also vaguely remember watching the video, and I think Grant Sanderson said that if there had been two geometry problems in 2024, it would have won a gold medal that year.
23:40
Speaker A
[Jonghyun Park] Right. The IMO has six problems, but there are four subjects, so the distribution varies each time.
23:48
Speaker A
So because there was one combinatorics problem, it got one problem wrong and got all the others right, earning a silver medal, if I remember correctly.
23:55
Speaker A
The year before last. Right. I'm not exactly a great expert on mathematics, but I asked about this to some friends who had studied math very seriously. They told me that even among people, combinatorics experts are a distinct group.
24:10
Speaker A
I guess there are effectively different specialties even within mathematics. They say combinatorics is typically the kind of thing that geniuses excel at.
24:18
Speaker A
They described it as a field that relies somewhat on innate ability. To wrap up the discussion of the IMO here for a moment, the question is whether AI can actually conduct good research, and what good research is. If we think about these questions a little more
24:34
Speaker A
from a mathematical perspective, examples from the past come to mind. Both here and on Dwarkesh, they discuss group theory, which was created by this person named Galois, and that is how it is translated here.
24:46
Speaker A
Apparently, this is called group theory. It seems Galois came up with this idea of group theory, and that things like this emerged from it.
24:53
Speaker A
But it later truly came into the spotlight and even influenced things like the prediction of quarks.
25:00
Speaker A
They say that took 100 years. In other words, from the time a theory emerges until it actually has an impact on human society, and until it is validated as genuinely good research, the timescale is enormously long.
25:12
Speaker A
That means the process cannot be repeated iteratively. Because of difficulties like these, RLVR, which is currently improving the intelligence of LLMs by providing immediate rewards and developing intelligence in the direction those rewards indicate, operates on a completely different timescale,
25:28
Speaker A
so this method would not apply here. That is one of the difficulties. Another thing that struck me was something Grant Sanderson quoted.
25:39
Speaker A
A good mathematician proves theorems, a great mathematician makes good conjectures, [Seungjoon Choi] There are those famous conjectures named after various people.
25:49
Speaker A
[Jonghyun Park] Yes, things like the Poincaré conjecture. And then the very greatest mathematician comes up with definitions— good definitions.
25:58
Speaker A
That is what Grant Sanderson said. So the perspective on research that we have today is still focused on proving theorems, or skillfully proving theories and hypotheses.
26:13
Speaker A
So whether AI can go beyond that is ultimately what we need to consider. What do you think?
26:21
Speaker A
I would like to hear your opinion. [Seungjoon Choi] The reason I suggested that Jonghyun and I watch this Dwarkesh episode together was actually because Dwarkesh is planning a series around this question.
26:30
Speaker A
Dwarkesh keeps persistently digging into why there are areas where these systems perform well, yet other areas where they still do not, and particularly why, in science, there is a possibility that they may fail to make dramatic progress.
26:41
Speaker A
Perhaps I was influenced by watching that, but I also still have questions about it.
26:46
Speaker A
I'm curious. I'll introduce this during my turn as well, but these systems perform extremely well in the areas where they work, while in the areas where they do not, they continue not to work. This spiky characteristic keeps repeating like a fractal,
26:57
Speaker A
as Grant Sanderson said earlier, and I think that is exactly right. They will keep picking the low-hanging fruit.
27:04
Speaker A
Nevertheless, within the immediate, near-term timeframe, I hope that the direction in which humans still need to remain involved will continue for at least another five years.
27:17
Speaker A
[Jonghyun Park] In any case, with the current method of improving intelligence based on RLVR, it does not seem easy to reach those lofty heights.
27:26
Speaker A
That seems to be what you think. I think similarly. [Seungjoon Choi] But Chester has a somewhat different perspective.
27:32
Speaker A
Chester isn't feeling well today, so although we'd surely have heard a comment, I think we'll probably hear it next time.
27:39
Speaker A
[Jonghyun Park] The people watching this video are probably grappling with similar questions in their respective fields, and today, we are going to discuss them in relation to research, particularly science.
27:50
Speaker A
As for whether AI can conduct good research, I think it absolutely can. I'm not sure whether it can come up with definitions like those very greatest mathematicians, but I think it can conduct good research because one example also comes up in this Grant Sanderson episode.
28:08
Speaker A
The number theorist Montgomery and the physicist Dyson were apparently discussing something like the Riemann zeta function.
28:15
Speaker A
As they were talking about some kind of formula, they said, "That is the same as this thing in mathematics and that thing in physics," and discovered some commonality between them.
28:26
Speaker A
[Seungjoon Choi] At the Institute for Advanced Study in Princeton, they arrange occasions where highly distinguished people can meet by chance, and I think this may have been an anecdote from one of those lunches.
28:38
Speaker A
[Jonghyun Park] Whether it was lunchtime or time for tea, I heard it just came up while they were talking at one of those gatherings.
28:44
Speaker A
Things that we think of as very different fields can actually have a tremendous number of things in common.
28:50
Speaker A
That they’re isomorphic, that we have the same structure, that’s what we’ve been talking about, and if you think of research as expanding humanity’s knowledge, every person knows their own domain well, but doesn’t know every other domain well.
29:04
Speaker A
But from a knowledge perspective alone, an LLM can instantly know all of them well.
29:10
Speaker A
So when you take some element from this field and apply it directly to that field, when you apply it isomorphically, progress can occur.
29:19
Speaker A
If you think about it sufficiently from this perspective, knowledge that humanity has yet to discover can be expanded by taking things from other fields and applying them.
29:28
Speaker A
That is entirely possible. [Seungjoon Choi] Right. But personally, because of the curse of context, the model only does that when you prompt it to.
29:36
Speaker A
Because if you just tell it to do something, the distribution tends toward the average, so it only says what would normally be said in that field and doesn’t bring anything in from elsewhere.
29:45
Speaker A
It only brings things in when you ask it to. So personally, I think that’s still something humans can do.
29:51
Speaker A
[Jonghyun Park] Now that you mention it, a lot of methods come to mind. The easiest would be to simply tell it in the prompt to look at other fields.
29:59
Speaker A
You could also deliberately bring in material from other fields and feed it to the model, or, thinking a little more from an RLVR perspective, among the output tokens generated by the LLM, you could deliberately select low-entropy tokens and tell it to explore in other directions as well.
30:13
Speaker A
[Seungjoon Choi] That’s why, last year, Gwern discussed something called LLM Daydreaming, a concept where it connects various things while sleeping, but in any case, I suppose this could be overcome to some extent with a harness or something similar.
30:27
Speaker A
[Jonghyun Park] In any case, when it comes to an LLM, if you ultimately think about it, it isn’t impossible.
30:32
Speaker A
When I think about these points, it’s because this was true for me in the past, and it’s still true when I’m building things these days.
30:39
Speaker A
It was also true when I was doing research: I would read other papers, draw inspiration from them, and think, “I should try applying this over here.” Naturally, I had those kinds of thoughts a lot.
30:48
Speaker A
I think those things are entirely possible. That’s roughly where I stand. Going one step further, there’s the recent episode, which is about physics.
30:56
Speaker A
I think the video is about an hour and 40 minutes long. For an hour and 10 or 20 minutes of it—almost the entire thing—it really is just [Seungjoon Choi] A physics lecture.
31:05
Speaker A
[Jonghyun Park] It just explains the theory of relativity. That in itself is incredibly impressive, of course, but what comes next is ultimately, “Then how can AI do something like this?” That’s what they discuss.
31:17
Speaker A
But because the discussion of the theory of relativity in physics is naturally so difficult in theoretical terms, I personally don’t think I understood all of it either.
31:26
Speaker A
But what I’d like you to do is give it a watch, because even if you don’t understand all of it, there are many parts you can understand.
31:35
Speaker A
On Grant Sanderson’s 3Blue1Brown channel, there’s a video series currently being released called Compression is Intelligence, I think, meaning “compression itself is intelligence.” That’s what it discusses.
31:45
Speaker A
[Seungjoon Choi] Grant Sanderson talked a bit about entropy, and it seems that’s becoming a series.
31:49
Speaker A
[Jonghyun Park] I think it’s probably a three-part series. That’s what it discusses, that compression is intelligence, and that is its main theme.
31:56
Speaker A
It felt similar to me. In other words, they talk extensively about how Einstein developed the theory of relativity, whether it took 10 or 20 years, and which idea it began with.
32:06
Speaker A
They go into all of that in great detail. But we’re hearing all of it now in the span of an hour.
32:11
Speaker A
So during the time the research was being conducted, people explored one idea after another, formed hypotheses, conducted thought experiments, and even performed actual physical experiments.
32:22
Speaker A
The vast body of knowledge built upon those efforts was ultimately compressed and compressed again, and explaining it as simply as possible at a level all of us can understand is itself intelligence.
32:35
Speaker A
That’s ultimately the kind of point the 3Blue1Brown channel is making, and I recommend watching it from that perspective.
32:45
Speaker A
It makes you think, “Something that difficult can be explained like this,” at least that’s how it made me feel.
32:51
Speaker A
[Seungjoon Choi] But honestly, no matter how visually Grant Sanderson explains things, difficult things are still difficult.
32:56
Speaker A
I enjoy watching the videos, but I don’t understand all of them. But we briefly passed over this earlier: there are fascinating stories about geniuses who died young, such as Galois and Abel, and Grant Sanderson also seems to be preparing
33:09
Speaker A
a documentary about them. It seems to be some kind of documentary about mathematicians that connects their stories to the AI era, and I’m really looking forward to that as well.
33:17
Speaker A
[Jonghyun Park] I always recommend the 3Blue1Brown channel to people around me. It has personally helped me tremendously.
33:25
Speaker A
When it came to doing practical work, although I learned a lot about things like Fourier transforms in school, when I actually needed to use them for work or research, I didn’t understand them.
33:37
Speaker A
But after watching that video once, I understood immediately. Visuals really are powerful. That’s what I thought.
33:44
Speaker A
Let me continue for now. They discuss theoretical physics, and ultimately, one of the difficult things about theoretical physics is that experiments are extremely difficult.
33:54
Speaker A
You have to make certain particles move at nearly the speed of light. You have to conduct experiments like these to observe each hypothesis one by one, but ultimately, what you have to do is simply conduct thought experiments.
34:06
Speaker A
In the absence of actual experiments, how far can we go without experimentation? The idea is to entrust this to AI.
34:14
Speaker A
Then, without experiments, how far could AI actually take the research? Adam Brown explains it this way: Suppose we had intelligence at Einstein’s level, and assume that an LLM had intelligence at that level.
34:29
Speaker A
Throughout human history, there has only been one such person at a time, and even if exceptionally brilliant people at that level had continued to exist, there still would not have been very many of them.
34:38
Speaker A
If those people were LLMs, we could create LLMs in vast numbers and have them conduct a great many different experiments in their minds, so we would undoubtedly be able to discover far more.
34:49
Speaker A
We could simply think of the problem as exploration and invest the vast resource of computation into that exploration.
34:58
Speaker A
There is an enormous amount that could be discovered. That is what Adam Brown is saying.
35:02
Speaker A
Also, I provided all the source materials for this article, and Fable 5 wrote it.
35:08
Speaker A
As I reviewed what Fable 5 had written, I deliberately left this part in. It described it as being willing to waste resources.
35:16
Speaker A
Of the things we currently believe are all true, a great many have actually turned out not to be true.
35:23
Speaker A
Even in the case of Einstein’s theory of relativity, there are many examples of that.
35:27
Speaker A
Newton’s universal gravitation, gravity, and things like that were all considered true, but they actually contained many contradictions.
35:36
Speaker A
The video also provides many accessible explanations of such contradictory cases. Once we identify those contradictions, we need new explanations to resolve them, and because LLMs do not mind wasting resources on proving again the assumptions we believe to be true,
35:53
Speaker A
as they do those things, they can uncover things that were previously overlooked. That is the Hugging Face Reproduce Challenge I mentioned at the beginning.
36:02
Speaker A
Because the papers passed peer review, we naturally assume that they are all correct, but deliberately reproducing them to demonstrate that they may have been wrong is something that, for an LLM, might seem wasteful from a human perspective, but is not actually wasteful.
36:17
Speaker A
So I think we can conclude that even today, there are areas where LLMs can provide substantial assistance with our research.
36:25
Speaker A
Then, when we ask what will happen going forward, the pace will naturally vary significantly by field.
36:31
Speaker A
Fields where everything can be worked out mentally, such as mathematics and theory, will naturally advance first, while fields that require actual experiments, such as biology or chemistry, where substances must be synthesized or reactions observed, will naturally have to move a little more slowly.
36:51
Speaker A
Those are some of the thoughts that come to mind. And for now, fields where rewards can easily be provided will naturally advance first, while fields where they cannot will progress more slowly, and I think that is what Dwarkesh
37:05
Speaker A
is trying to say through this podcast. This has limitations. Might we need an approach other than RLVR?
37:11
Speaker A
Or, in the preceding series, there were limitations to pretraining. Continual learning does not work at all either.
37:17
Speaker A
What do we need to do about these things to reach a higher level of intelligence?
37:21
Speaker A
Dwarkesh seems to have continued conducting interviews in this context, and this appears to be following a similar line of thought.
37:27
Speaker A
Then, if we consider the ultimate goal, could AI formulate the definitions associated with the greatest mathematician that I mentioned earlier? Will that remain the domain of humans until the end, or will it become the domain of AI?
37:38
Speaker A
This seems to be an area where everyone may have a different opinion. Personally, I think AI might be able to do it.
37:45
Speaker A
That is what I think. A definition is also something that can simply be created, but whether it is a good definition ultimately depends on how accurately it reflects reality and how well it explains physical phenomena and natural phenomena.
38:04
Speaker A
Ultimately, it might be no different from the search problem of exploration. If evaluation is possible, if verification is possible, might AI be able to accomplish this as well?
38:14
Speaker A
But I do not think this will happen on a timescale of the next two years or so.
38:19
Speaker A
It does seem to be a long way off, but looking 10 or 20 years ahead, might it be possible?
38:24
Speaker A
Personally, that is what I think. [Seungjoon Choi] By around 2040, I suppose anything could happen.
38:30
Speaker A
So we have heard from Jonghyun, and shall I continue from here? I will skip over some of the overlapping parts and share my thoughts as well.
38:39
Speaker A
I mentioned this once in the previous episode too, but Periodic Labs was founded by people who had been at OpenAI and worked in fields such as materials science and physics, and co-CEO Liam Fedus has been working on this lately.
38:54
Speaker A
They are building a factory because they need to conduct experiments. I remembered this tweet and brought it up again.
39:02
Speaker A
There clearly seem to be areas where the loop must be closed through experimentation. So after watching the video, I had a conversation with the model and generated a short piece titled “What Remains in the Age of Intellectual Automation.”
39:17
Speaker A
I highlighted several points, and revisited Grant Sanderson’s argument that intellectual breakthroughs take at least three forms.
39:27
Speaker A
First, there is the lightning bolt mentioned earlier, where two different fields can be placed side by side.
39:35
Speaker A
But this is relatively easy. These days. [Jonghyun Park] Yes, that's right. [Seungjoon Choi] Right. But building a new mountain is extremely difficult, takes a lot of time, and you can't know early on whether it's right or wrong, and it's difficult to evaluate. So this kind of achievement
39:51
Speaker A
is hard to evaluate immediately as simply having gotten the right answer. Whether it was truly a productive idea may not become apparent until decades later.
39:59
Speaker A
That's one of the three forms of intellectual breakthrough, and I found another one rather interesting.
40:06
Speaker A
It is reaching an answer by pushing through an enormous amount of reasoning even without any new concepts.
40:11
Speaker A
In fields like formal proofs, using systems such as Mathematica or Lean, it was already possible to run automated theorem provers and explore things resembling alien mathematics that humans find difficult to understand.
40:25
Speaker A
That reminded me of what Grant Sanderson said: proving that something is true and determining whether it is really what we want and an idea capable of changing how we think are different things.
40:40
Speaker A
So making it digestible for humans is still important, but the point was also that machines could be good at explaining such things.
40:51
Speaker A
So this part also gives me something to think about, and there is a part related to the back-and-forth that Chester, Jonghyun, and I have on our channel: the role of choosing what to explain, which ideas are worth wrestling with,
41:06
Speaker A
in what order to present them to people, and why they are interesting will remain important for the foreseeable future.
41:14
Speaker A
So I also found Grant Sanderson's remarks on how to frame the value of one's own existence quite striking.
41:20
Speaker A
So for now, this still holds true, and Dwarkesh said something along these lines: when studying something difficult, you do read the original text, but having a model explain it dramatically lowers the difficulty.
41:35
Speaker A
I strongly relate to that as well. But while AI writes explanatory prose well, it isn't particularly good at creative writing.
41:42
Speaker A
It writes well, but hasn't been able to surpass a certain level. [Jonghyun Park] What kind of writing do you mean by creative writing here?
41:49
Speaker A
[Seungjoon Choi] For example, things like novels. Jonghyun mentioned this once before too. Things like comedy or novels don't work very well.
41:56
Speaker A
[Jonghyun Park] Right. Comedy really doesn't work. [Seungjoon Choi] Right. It's not funny at all.
41:59
Speaker A
For RLVR to work, you need to be able to provide a scalar reward based on whether the final result is right or wrong.
42:07
Speaker A
Since that rewards the entire path, including which tokens should be included, it does make AI better at coding, but there is no way to know whether each token is valuable enough.
42:21
Speaker A
But in writing, every individual token is part of the final product, so I don't think that can be solved with RLVR.
42:28
Speaker A
That point Dwarkesh made has stayed with me. Good writing considers where readers might hesitate, what they might misunderstand, and in what order they need to think to accept the next sentence.
42:39
Speaker A
Each sentence and word is not mere packaging, but the content itself. The point that writing is not like code, where it can look messy as long as the result works well, really resonated with me.
42:53
Speaker A
So there needs to be a layer that judges importance. Even in automated knowledge production, the human question of what is valuable may be somewhat ambiguous, but it will not easily disappear.
43:04
Speaker A
Since models are also steadily becoming better at reviewing and making judgments, this may not last forever, but the view that it will not easily disappear seems to be something Dwarkesh and Grant Sanderson agree on.
43:18
Speaker A
This also gives the sense that the teacher's role is shifting somewhat. That's because what is still needed is a social relationship, not merely an explanation.
43:33
Speaker A
The point was that because teaching is fundamentally a social and relational activity, some part of the teacher's role will remain.
43:42
Speaker A
Ultimately, to read this once more, what matters is the instinct for recognizing which problems are worth asking, the ability to see connections between distant ideas, and the patience to judge whether a new concept is truly productive.
43:56
Speaker A
I think that patience is important. The skill of compressing complex results into a form humans can understand, and guiding people toward what is worth seeing amid an overabundance of knowledge—these were all extracted from discussions of complex mathematics, stripped down with the help of Fable 5,
44:15
Speaker A
leaving only the framework. So these are the points I noticed. [Jonghyun Park] If I try connecting writing and mathematics and think about them together, I personally love novels.
44:28
Speaker A
But then, what makes a good novel? If I ask, "Why do I find this interesting?" it's hard to explain.
44:35
Speaker A
And if we think about mathematics, Seungjoon mentioned Lean earlier, and I came across something about that.
44:40
Speaker A
With Lean, though I don't know much about it myself, when DeepMind won a medal at the IMO the year before last—it was a silver medal— DeepMind wrote the code in Lean.
44:52
Speaker A
Lean just looks like mathematical notation. It's text, but it looks like code, like mathematical code, and I heard it's a language that can determine whether something is true or false when you run it.
45:04
Speaker A
So people say that if you learn Lean well, you can more or less turn mathematics into something verifiable.
45:10
Speaker A
[Seungjoon Choi] Right. You do have to translate all of it into Lean. But starting this year, I believe, DeepMind no longer uses Lean.
45:16
Speaker A
[Jonghyun Park] Right. And that brings us to the Adam Brown episode that frightened everyone at the time.
45:22
Speaker A
As time has passed, whether you're using Lean, writing mathematical code, or writing a novel, both are simply writing.
45:30
Speaker A
Assuming that an LLM spits out tokens, suppose it did good mathematical research and the output proved a theorem, but the output is millions of lines of Lean code.
45:41
Speaker A
Then it would be far too difficult for humans to look at it and understand it.
45:43
Speaker A
That is one of the things I worry about. In other words, it would be a boring novel.
45:48
Speaker A
It would be true, but impossible to explain. I worried about this a lot, but at this point, judging from some recent announcements by OpenAI, apparently their content is easy to understand.
46:02
Speaker A
When Adam Brown said something along those lines, Dwarkesh responded, “Could it be that it’s only easy for you?” That was the response, but anyway, by the standards of an exceptionally intelligent human, it is not at an inexplicable level.
46:17
Speaker A
In other words, it is not couched in extremely difficult language. Seeing Adam Brown assess it as an explanation that humans can comprehend makes me wonder whether, in a sense, it is closer to a good novel. Because when humans read it,
46:31
Speaker A
they can understand it. So it is fine. You could put it that way. So in a sense, the boring novel is a proof consisting of millions of lines of Lean code, while the interesting novel is an explanation for humans,
46:45
Speaker A
that is, an explanation at a level they can read and understand. That is precisely Compression is Intelligence.
46:50
Speaker A
This idea is also connected to that: a piece of writing that compresses things well and makes the explanation easy is intelligent writing.
46:57
Speaker A
And in a sense, that might also be a good novel. Could these ideas be connected in that way?
47:02
Speaker A
That is something I have been thinking about recently. But an interesting novel and an easy-to-understand novel also seem to be different things.
47:11
Speaker A
So even when I think about these things myself, the categories are not clear, and because of that, it seems only natural that this does not work well right now, but I am not sure whether that will still be the case in the future.
47:25
Speaker A
Of course, this probably will not work for the time being either. [Seungjoon Choi] But since you mentioned compression, this is quite fascinating.
47:32
Speaker A
The person who became famous for talking about AGI, that is, the person who widely popularized AGI, is Shane Legg, the co-founder of DeepMind.
47:42
Speaker A
And Shane Legg’s advisor was Marcus Hutter. Where does Marcus Hutter come in? Wikipedia compression— Marcus Hutter was the person who put up the prize.
47:51
Speaker A
It was about how efficiently Wikipedia could be compressed. So compression was a central research topic for Shane Legg and Marcus Hutter.
47:59
Speaker A
One that led to AGI. And the foundational idea behind it is a concept called Solomonoff induction.
48:06
Speaker A
Shane Legg’s doctoral dissertation discusses how Solomonoff induction feeds into all these ideas. And Ilya Sutskever once touched on a related topic while explaining Kolmogorov complexity.
48:19
Speaker A
I remember hearing that once. So I am saying this without fully understanding it myself, but these foundational ideas involving compression and intelligence are extremely important within the field of algorithmic information theory, and it suddenly occurred to me that these ideas
48:35
Speaker A
form the foundation of today’s AI. [Jonghyun Park] Yes, I have been thinking about that a lot lately as well.
48:40
Speaker A
Watching AI develop, especially when I imagine that I had been born a thousand years ago, I do not think there would be that much difference in mental ability between someone from a thousand years ago and me.
48:51
Speaker A
But there is an enormous difference between the output someone from a thousand years ago could produce and the output I can produce now, and that is because humanity’s knowledge has accumulated.
49:01
Speaker A
And because that knowledge has been compressed so effectively, we absorb all of it through formal schooling extraordinarily quickly as we grow up, so by the time we turn twenty, we become capable of doing many things.
49:14
Speaker A
[Seungjoon Choi] Right. Because our minds are already filled with so much scaffolding, there are many ways we can use it instrumentally, as tools for thought.
49:21
Speaker A
[Jonghyun Park] Right. The brain’s hardware probably is not all that different, so perhaps compressing knowledge and putting it into the brain effectively is the essence of intelligence.
49:30
Speaker A
That is what I find myself thinking. From a macroscopic perspective. [Seungjoon Choi] But Ted Chiang once described this by saying that LLMs are blurry JPG compressions of the web, which became a subject of controversy.
49:42
Speaker A
Anyway, let us keep moving along for now. I majored in applied physics as an undergraduate, but I am actually far removed from physics.
49:51
Speaker A
So I do not know it very well. The material was not easy for me either, but what I did was provide Adam Brown’s transcript and try turning it into a mnemonic book.
50:06
Speaker A
It presents a conversation, then things like simulations in the middle that readers can use for tinkering, and then toward the end, there are quizzes that ask about the earlier material, so I tried structuring it in a way that allows readers to use mnemonic techniques.
50:20
Speaker A
And as I listened to these discussions, I was reminded of Ted Chiang, whom we just mentioned.
50:24
Speaker A
In 2000, Ted Chiang wrote a story whose translated Korean title is probably “The Evolution of Human Science,” in which, after metahumans emerge, humanity has to practice archaeology.
50:34
Speaker A
People strive to understand the research conducted by the metahumans, and it is a story that is both funny and sad.
50:39
Speaker A
It is a short piece, and it is interesting. That was one thing that came to mind, and what I would like to spend the latter half of today discussing is something I mentioned earlier: I believe Dwarkesh has meticulously planned
50:50
Speaker A
how to approach this process. The idea is that RLVR may be particularly weak in science.
50:58
Speaker A
This was a blog post published in May, and after putting forward that idea in advance, Dwarkesh found suitable interviewees and conducted interviews with people such as Michael Nielsen, as well as people working in science, and Grant Sanderson, who explains mathematics.
51:16
Speaker A
The way Dwarkesh shapes all of this around that plan is a very interesting point.
51:21
Speaker A
So what alternatives might there be when RLVR does not work? I have imagined things like that before.
51:27
Speaker A
But just a moment ago, Jonghyun, you mentioned people from a thousand years ago. But there was an interesting experiment this April.
51:37
Speaker A
Alec Radford called it Vintage LM— the person who created GPT, you know. Using only data from before the 1930s, meaning only texts published before then, to create a trained language model, and is experimenting to see whether, using things like pretraining and RL,
51:57
Speaker A
it can imagine developments that occurred after the 1930s—that is the general idea behind the experiment.
52:05
Speaker A
So a family of models called Vintage LMs is emerging, and I personally find it fascinating.
52:11
Speaker A
[Jonghyun Park] This is the first time I have heard of Vintage LMs, but ultimately, if the data here is really well constructed and the model can achieve technological advances through the 2000s, then we can naturally project from that and imagine
52:25
Speaker A
that it should be able to do the same going forward from the present day.
52:28
Speaker A
In a way, I think it could even be seen as proof. So simply observing this experiment could be very helpful in predicting the future, and depending on the results, people's thinking could change, and their actual behavior could change as well.
52:41
Speaker A
[Seungjoon Choi] Right. So I just wanted to introduce the idea that there seems to be an interesting connection here, and next, the ideas underlying the RLVR work happening these days are actually very old, going all the way back to Ronald Fisher's score function idea.
53:01
Speaker A
That is a very old idea. But I believe the REINFORCE algorithm itself emerged in the early 1990s.
53:10
Speaker A
The score function in statistics emerged in the early 1900s, in the 1920s, and then the policy gradient work by Williams and others appeared in the early 1990s. That gave us the idea that reinforcement learning could be performed using rewards,
53:30
Speaker A
and the idea underlying the RLVR used in LLMs today actually emerged very early. So, to explain this very briefly to the extent that I understand it, this x represents a conditional probability, and given a condition or input x,
53:47
Speaker A
there are parameters theta for a model that can produce y, and the model outputs a probability.
53:54
Speaker A
A score function ultimately calculates the gradient of the logarithm of that probability. So to help the model do that well, if a reward is provided from outside the system—that is, from outside the model— then even if the computational graph is not connected,
54:10
Speaker A
because it only involves multiplying this by a scalar, it can alter the distribution. So this provides a way to optimize the objective function, and it works with LLMs because LLMs output probabilities.
54:24
Speaker A
When a model outputs a token's probability logit, that gives you a probability, and what comes out of those probabilities is the vocabulary.
54:32
Speaker A
Each token has a certain percentage probability, and we take some of the highest-probability tokens and use them for next-token prediction.
54:43
Speaker A
The set of tokens ultimately serves as the action space, so it is possible to reinforce what the LLM outputs probabilistically.
54:51
Speaker A
In RLVR, after the model writes out a piece of code, an external verifier runs tests and tells the model whether it worked, and that result can be supplied as reward R.
55:03
Speaker A
If the entire trajectory of tokens works, all of it is assigned a high probability, and if it does not, it is assigned a low probability.
55:10
Speaker A
But methods such as GRPO, introduced by DeepSeek, establish the average of those results as a baseline and then evaluate them relative to it, assigning a positive value to those that outperform the average and a negative value to those that underperform,
55:22
Speaker A
thereby reinforcing the trajectory itself. Once I had a rough understanding of this, I understood what Dwarkesh had said: ultimately, which particular tokens are making it work better is something recent studies are investigating as well, of course, but it is not guaranteed; all that matters is whether it works.
55:40
Speaker A
That is why it does not work for writing, and another thing it does not work for is anything without a single correct answer— things where this could be right and that could also be right, as well as things that carry implications.
55:52
Speaker A
Those things still cannot be handled with RLVR, and those kinds of problems tend to arise in design and art.
56:01
Speaker A
[Jonghyun Park] Yes, just to clarify this a little, x is the input text, or the prompt, and y is the output token.
56:09
Speaker A
Then pi is the LLM, [Seungjoon Choi] which produces that probability. Theta is the LLM's parameters.
56:16
Speaker A
[Jonghyun Park] Theta is the LLM's parameters. So y is actually not just one token; it is a sequence of N tokens, and once those N tokens have been generated, if we suppose the model is solving a math problem,
56:27
Speaker A
no matter how it wrote the solution, if the final answer is correct, we assume that the entire solution must have been correct and reinforce all of it, which makes it difficult to apply RL properly to things like how concise the solution is,
56:39
Speaker A
or how elegantly it is written. That could result in the million-line Lean code we discussed earlier, and I think this relates to the context in which those concerns arose.
56:51
Speaker A
[Seungjoon Choi] Right. Recent studies are trying to address all those specific issues by using various regularizers and hacking around them with various techniques, but fundamentally, it still feels like we remain in this regime.
57:06
Speaker A
[Jonghyun Park] You add those kinds of factors to the reward. For example, how concisely the model wrote the solution to reach the answer.
57:11
Speaker A
Those kinds of criteria are also used, and I have recently been trying my hand at design.
57:17
Speaker A
Since I am personally an engineer, naturally, I do not know how to design well.
57:21
Speaker A
So what I have been doing is looking at something from YC that shows how YC uses AI to create designs.
57:27
Speaker A
These days, when we browse websites or look at PowerPoint presentations, we can immediately think, “That design was made by Claude.” It really stands out.
57:37
Speaker A
But when I looked at what YC made, none of that was visible at all.
57:41
Speaker A
I wondered how they managed to create designs that didn't just have a similar feel, but had a completely different feel and were done so well, so I've been working hard to emulate it.
57:51
Speaker A
But when you get down to how they actually do it, it's a little complicated.
57:56
Speaker A
They create a soul, make a mood board, and do all sorts of other things, but ultimately, a human makes the selection.
58:04
Speaker A
They perform all the actions that correspond to taste. They do that and feed it in as input.
58:10
Speaker A
The taste. Then, when N different output designs are generated, they make another selection from them.
58:17
Speaker A
[Seungjoon Choi] So, [Jonghyun Park] for things that can have multiple correct answers, what they do is put a human in the loop, and have the human quickly pick them like this.
58:25
Speaker A
So when it comes to taste, to personal preference, they still seem to make sure to do that.
58:30
Speaker A
Ultimately, it's a person's taste, the creator's taste, so it's not easy for it to just magically appear.
58:36
Speaker A
[Seungjoon Choi] They're definitely not doing it by building a model; they're just doing it at the text level.
58:41
Speaker A
[Jonghyun Park] It is at the image level, but it's at the image or text level, so either way, it's at the token level.
58:46
Speaker A
Because images also go in as tokens. [Seungjoon Choi] There are, of course, ways to overcome that, but they're not very elegant yet.
58:52
Speaker A
[Jonghyun Park] Right. [Seungjoon Choi] And that was one thing I wanted to introduce today, but despite that, there are still many areas where it works well.
59:02
Speaker A
So I also found this fun to read: Anthropic acquired Bun, you know. So on July 8th, there was a post that was practically a success story about converting a large project with around one million lines of code to Rust, and that was interesting too.
59:19
Speaker A
What was interesting was that to run this, an expert still had to spend about three hours compressing everything they knew, creating a proper spec, and then running it—it didn't just work automatically.
59:31
Speaker A
Even with an agent swarm. So that part stuck with me, and the full story is much more detailed, longer, and more interesting.
59:40
Speaker A
Open-weight models have come out, you know. So TML—the one John Schulman is involved with, as mentioned earlier— released a model called Inkling, and then Kimi K3—the Kimi K3 you mentioned earlier—came out as well.
59:52
Speaker A
What was interesting about Inkling was that, in English, Inkling—the main image here also shows a trace of ink— when I looked up the meaning, I found that its core nuance is a faint clue or a premonition.
60:06
Speaker A
So this is only the beginning. It seems that they also released a model with nearly 1T parameters as open weights, and it felt like a teaser signaling that they were starting to show what they could do with it.
60:21
Speaker A
The most interesting part was nicely summarized on a blog by Jaewoong Park: self-finetuning. Lilian Weng had published “Harness Engineering for Self-Improvement,” and they introduced what they were doing as a practical implementation of that, which I think is the key part.
60:43
Speaker A
Broadly speaking, they perform a finetuning task, upload it to their infrastructure, Tinker, and finetune the model itself.
60:52
Speaker A
An interesting example comes up here: they finetune it to speak without the letter e.
60:58
Speaker A
Creating works without the letter e is something found in literature. It's an experiment that was often done by groups such as Oulipo, using something called a lipogram to create works without lowercase e or uppercase E.
61:10
Speaker A
That's what they make it create. But they're finetuning the model to do it. [Jonghyun Park] So they've connected the model called Inkling and are running OpenCode, [Seungjoon Choi] Right. It's finetuning itself.
61:23
Speaker A
[Jonghyun Park] So the output tokens gradually change. [Seungjoon Choi] I found the part where it performs RSI, finetuning itself for a specific purpose, quite interesting.
61:35
Speaker A
[Jonghyun Park] For one thing, Thinking Machines Lab, TML, which created that Inkling model, was founded by a group of extremely famous people from the outset, including former OpenAI CTO Mira Murati, who left OpenAI, as well as John Schulman, whom we mentioned earlier,
61:48
Speaker A
and Lilian Weng, along with many other famous people. So I'd been extremely excited about them, but they didn't release a model for quite a while.
61:56
Speaker A
[Seungjoon Choi] No model came out—just infrastructure and a few papers. [Jonghyun Park] Right. They released things like a finetuning service, and recently they released an omnimodality model, which was also quite impressive.
62:07
Speaker A
[Seungjoon Choi] Then they kept publishing articles in the context of self-improvement. So this is another organization I'd never heard of, besides TML, called AIDE², and rather than doing self-improvement through model weights, this seems to be on the harness side.
62:21
Speaker A
But those two approaches are currently mixed together. There are ways for the harness to improve itself, and if that harness is an Autoresearch-type harness that updates the model weights, then it feels like model updates are happening from both sides. As for the Kimi K3 in question,
62:40
Speaker A
I haven't used it yet, but judging solely from what I've read, it's remarkable. The blog also covered the technical details quite well.
62:50
Speaker A
I did translate this, and although it probably won't be easy to run, it's a milestone nonetheless.
62:59
Speaker A
[Jonghyun Park] First of all, it's too large for an individual to run easily, but the mere fact that a model of that size was released as open weights seems tremendously encouraging.
63:07
Speaker A
Naturally, numerous GPU cloud and neocloud providers will run it, serve it, and offer it through OpenRouter, [Seungjoon Choi] so it's hard to tell what kind of technological breakthrough there is just from looking at this, but Kimi K3 is, of course, also based on MoE,
63:21
Speaker A
and Inkling, which we mentioned earlier, is also disclosing everything about its MoE architecture. They seem to be doing well, and they even say that they design their own chips.
63:30
Speaker A
China has been constrained when it comes to chips, but they are designing chips as well.
63:35
Speaker A
I think we're also seeing some of that kind of talk around Kimi K3. Every single piece of content seems to have been carefully crafted— or rather, you can see signs that they had it crafted with great care.
63:48
Speaker A
And even the video editing— a teaser video came out a day earlier. And they even told us that the teaser video itself was edited by Kimi K3.
63:59
Speaker A
The tweets from Satya Nadella and Demis Hassabis were also rather interesting, but more than what Demis Hassabis said, Satya Nadella's counterintelligence paradox seemed to raise some points worth thinking about.
64:12
Speaker A
In the age of intelligence, how should companies protect their core knowledge assets? Big Tech companies promise that they won't use them for learning or training, but usage patterns themselves still become data.
64:25
Speaker A
And that in itself is extremely powerful information. So that discussion of the counterintelligence paradox became a major issue, and I think it offered some interesting points to consider.
64:38
Speaker A
Google, where Demis Hassabis works, seems to be in a somewhat difficult position right now.
64:41
Speaker A
[Jonghyun Park] And these days, models are emerging so quickly from all over the place that we sometimes get overly excited or disappointed by each one, but even if we extend the time frame to just a year, on a yearly scale,
64:57
Speaker A
a lot of companies really do catch up very well. When Gemini came out last time, it received favorable reviews for quite a long period, so even if there are some problems, perhaps if we wait a little, they'll resolve them all and come back with something good.
65:12
Speaker A
Especially since it has a lot of data and is a company with all the hardware resources it needs, it's difficult to imagine that Google [Seungjoon Choi] won't be able to catch up, though I do hope they'll step it up.
65:23
Speaker A
I think if we're patient and wait just a little, something will come out. So, after using GPT-5.6 for a week, how was it?
65:33
Speaker A
[Jonghyun Park] When I tried using it, something I started last night still hadn't finished after more than ten hours, so I just stopped it.
65:41
Speaker A
There were issues like that. I think the fact that it's focusing so heavily on test-time compute, which Chester mentioned, also manifests itself in that way.
65:53
Speaker A
[Seungjoon Choi] But what I found interesting was that, for something I'm building these days related to Minecraft, when working on something like this, I had been working under the assumption that Fable 5 was good at it.
66:07
Speaker A
So when I give instructions here like this, the model moves the bot, and the model created that bot.
66:19
Speaker A
Usually, the bot would do the task, but I wanted to understand why it was getting stuck at certain points, so I created an interface that lets me play using the commands the bot uses when I play.
66:31
Speaker A
And it did that remarkably well. But as I kept working on it—building it with Fable 5— I had been using Sol to review it.
66:43
Speaker A
But Sol kept reviewing it very harshly, saying this was wrong and that was wrong, and Fable 5 ended up yielding to it.
66:53
Speaker A
So I tried putting Sol in the driver's seat. I gave it to Sol, and since I tend to work through constant back-and-forth in HITL, I talked things through with it, but when reviewing Fable 5, it wasn't as adversarial as Sol had been.
67:11
Speaker A
It accepted everything. So I got the sense that Sol is also good at working in this kind of rhythm.
67:19
Speaker A
[Jonghyun Park] So just now in Minecraft, when you said, "Go here," the bot walked around and seems to have gotten somewhere, and Sol and Fable 5 are building all of that as you go, right?
67:30
Speaker A
[Seungjoon Choi] Right. But the more difficult problem they solved was when loading chunks at scale, how to stream them.
67:42
Speaker A
That was fairly difficult, and Fable 5's attempt at it went awry, but Sol solved it.
67:48
Speaker A
So that's how I've been using them lately, but honestly, I just don't have enough time.
67:55
Speaker A
How about you? To prepare something like a YouTube video, you need to keep up with the news and have a lot to read, while also living your everyday life.
68:03
Speaker A
[Jonghyun Park] Right. Since I have my regular job. [Seungjoon Choi] Exactly. Even just getting it to code, I found that there simply wasn't enough time.
68:09
Speaker A
[Jonghyun Park] It's the same for me. [Seungjoon Choi] So now I do everything remotely, and even though I can do it from my mobile phone, I still had the problem of simply not having much time.
68:19
Speaker A
To wrap up today, this is a bit of an advertisement, but as I've mentioned several times, I teach an AI class at Pi Design School, which is known as the design school created by Toss.
68:31
Speaker A
So even when building a Minecraft agent, there is work where you create a spec and run it from outside the loop, but there is also a mode where you enter the loop and, even if you don't know Rust,
68:42
Speaker A
engage with it while understanding things like the core parts of the algorithm currently being developed.
68:47
Speaker A
I believe you need to move back and forth and experience both sides. So I'm working to teach those things to people who want to do design and people who want to learn design, and there are also some interesting classes besides the one I teach,
69:03
Speaker A
so I thought I'd give it a mention. [Jonghyun Park] I also took a look after you mentioned it, and although I know nothing at all about design, if I wanted to become good at design, it did make me want to go.
69:17
Speaker A
[Seungjoon Choi] Right. Developers were applying too. I'm not sure whether they're considering it as an extension of their current role, or considering it because they want to pivot, but even among people without a design background, for the pre-course,
69:29
Speaker A
people applied to the first and second cohorts, though the tuition was nearly as expensive as university tuition.
69:35
Speaker A
Right. It seems like it does require a substantial investment. But it seems that some very prominent figures and top professionals in the design field are participating as well.
69:44
Speaker A
I was thinking of offering a course that addresses these kinds of questions. How was today?
69:51
Speaker A
We finally went through what we'd been putting off for so long, covering the Dwarkesh episodes with Grant Sanderson and Adam Brown as well.
69:59
Speaker A
[Jonghyun Park] First of all, I don't think today's discussion was really about business, but rather a time to explore and share visionary ideas that stretch our thinking far into the future.
70:11
Speaker A
Personally, I've always liked this kind of thing, so I enjoy doing it. [Seungjoon Choi] Shall we wrap up here for today?
70:17
Speaker A
[Jonghyun Park] Yes, great work. [Seungjoon Choi] It was fun. [Jonghyun Park] Thank you.
Topics:AI researchpeer reviewPPO algorithmarXivLLMsacademic publishingresearch evaluationAI-generated papersICMLICLR

Answers

Frequently Asked Questions

What is the PPO algorithm and why is it significant?

PPO (Proximal Policy Optimization) is a reinforcement learning algorithm widely used in training models like ChatGPT through human feedback. Despite initial rejection in peer review, it later had a major impact on large language models.

Why has arXiv restricted some AI-written papers?

Due to a flood of AI-generated papers, many lacking meaningful contributions, arXiv restricted submissions in certain categories to only those that have passed peer review to maintain quality and prevent the platform from being overwhelmed.

How is AI affecting the peer review process in academia?

The volume of paper submissions has led to some reviewers using AI to assist or replace human review, which has raised concerns about review quality and integrity, prompting conferences to implement policies dividing AI-assisted and human-only review tracks.

Get More with the Söz AI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →