Speaker A
ChatGPT 5.6 is a dumber model, and I love it so much. In fact, I use it all the time. Today, I want to tell you about which model works for you, not which model works for me. I will tell you which model works for me. Don't worry. More importantly, I'm going to tell you how to pick the model that works for you and why, and why your heuristic, why the thing you use to pick that model, is not anybody's benchmark score, including mine. And I do benchmark these models on a private benchmark suite, and I'll share that. But that's not the point. The point is for you to have the tools to pick the model that works for you. So, let's get into it. Last night, I reached for ChatGPT 5.6 Soul, even though I think it's the dumber model. Now, dumber does not mean dumb. Not remotely. Soul is an incredibly intelligent model. On Agents Last Exam, which measures long-running professional work across 55 different fields, Soul set a new high. And on my own benchmark, Soul scored 93 on Dingo, which is my knowledge work package. I've talked about it before. It's a really, really good model for doing complicated knowledge work. Dingo, in this case, is a test that measures whether a model can propose a business startup idea selling dingo dogs in Alaska, and get around the legal issues, get around the regulatory issues, and get into the marketing side of things. It's a really funny package on purpose because I like humor. It also tests whether a model is smart across a wide range of professional knowledge work fields. And Soul did really great on that. But what Soul does not have, at least for me, is that big model smell. And that matches what OpenAI has been investing in. They've been investing very specifically in improved reinforcement learning for existing model lineages, which allows them to be more and more useful on specific tasks. In this case, it's very strong on knowledge work. It's strong on long-run agentic coding. What it does not have, at least for me, is the same big model smell that Fable 5 has. And that makes sense because Anthropic has been investing in pre-train for their models. In other words, training on larger and larger data sets to enable more and more general-purpose models. We have incredibly good assistants, assistants that make our best experts better because they act as companions to research, companions to thinking, but not yet truly generalizable intelligence with deep recursive learning. In the meantime, we have to figure out what to do with the models in front of us and make use of them today. And that brings me to the point that I'm trying to make. I am picking the model I am picking because it makes it easier for me to produce my best work. What is my model recommendation? I say, are you me? Do you have the same habits with your model that I have? I am someone that is extremely willing to do very lengthy, somewhat technical prompts, and I'm okay verbalizing that, and so I just talk into Whisper Flow, and I just feed it to the model, and I'm fairly specific about it, and that fits 5.6 fairly well because 5.6 will read through that whole prompt, understand all of the edges I just talked about, and come back with a full piece of work and be really persistent about getting it done. But not everyone talks and works that way. And your best work may actually come from a different approach to prompting and knowledge management. You may have different tasks that you're working on. You probably do. Hint for you: my best tip when people say, "What model do I pick?" is to not look at the model first. Instead, look at your best work and look at how you get there. And it may be not with the model. Look at the process you use for thinking. And then start to ask yourself which model helps me to accelerate that loop that gets me to my best self. I find, as I've been saying, those lengthy prompts, the ability to just talk about what I want done, plus the harness that allows me to self-improve really easily with Codex, that gets me really far. And by self-improve, I mean that Codex will learn from what I do and further improve skills. I mean that Codex is steerable, and I have fairly high intent with my prompts. So I like to steer it. Fable 5 is really good out of the box at understanding intent that's a little bit more high level. It has that ability to generalize associated with big models. I love that. It's a fantastic model. The Anthropic team could, but I'm not reaching for it as much because it isn't suited to my particular work patterns. And if your work patterns are more around understanding very high-level ambiguity, wrestling with concepts, trying to pin down ideas between ideas, then Fable may be a much better model for you. If you're more suited to understanding how to get coding done efficiently and quickly, honestly, you may reach for the Luna series from OpenAI, which is much cheaper to run on, incredibly high-powered, and they released it along with 5.6 Soul and 5.6 6 Terra, or you may reach for Grock, which also has good frontier-ish coding capabilities. You may reach for GLM 5.2. You may reach for Ringer, which I built and talked about last week because it enables you to farm out and orchestrate from one central model like Fable to a bunch of cheaper models. It suits your work, right? And by the way, if you're wondering, would I still use Fable as the architect in Ringer even with 5.6 out? I would because Fable is good at understanding intent and breaking down those tasks to get that intent done. It also has a good front-end instinct. It just Anthropic has been consistently good at front end. I'm going to flash the benchmarks up here on the screen as I talk so you can see how I scored 5.6. I'm using the same benchmarks I've used for all of the models over the last few generations. So, we're not changing anything. But as I do that, I want you to think about the larger point we've been talking about. And I want to suggest to you that given everything I've shared with you, we are missing a core insight. We are missing the idea that models are becoming more like families we need to get to know and less benchmarkable, period. I don't believe my benchmark or any benchmark fully captures what these models do in a way that's useful. And that's why I make videos that may feel vibey like this where I tell you what it actually feels like to use these models because I want you to get inspired to jump in and test them for yourself on your workflows and also to hear from someone who does that all the time. The key insight I have for you as someone who has touched these models and lived with these models is that the models really do need to be treated like family. Think of it as every new model is like a new picture in a family photo album. You're getting to know someone new who is a part of the family. There's family resemblance. I can tell you the 5.x family from OpenAI has family resemblance. They all have that preference for long-running agent coding flows. They all have the ability to understand what you're saying explicitly, paint the edges very clearly, and just go after it. And they maybe have less of an ability to read between the lines. Whereas the Mythos lineage, we have Mythos and Fable there, is extremely good at ambiguous tasks, has extraordinary front-end taste, is almost philosophical in the way it approaches problems. It's a deep thinker. And those are just fundamentally different approaches. It's not that one is really better or worse. And this is where I think when we say words like dumber or smarter, we say them in the context of benchmarks, but I don't think it does the models a service because the models are becoming different the way families are different. And we don't really say this family's dominant, this family's smart. We say these families are different. And I think that's more useful. And so we have the Anthropic family of models. It's more pre-trained. It's more front-end. It's more interested in character and philosophy. In fact, Anthr