Your Next AI Subscription Shouldn’t Be ChatGPT 5.6 Or F… — Transcript

Explore how to choose the best AI model for your work, comparing ChatGPT 5.6 Soul and Fable 5 based on use cases and personal workflow.

Key Takeaways

  • AI model choice should be personalized based on your workflow and prompting style, not just benchmark scores.
  • ChatGPT 5.6 Soul is ideal for detailed, technical, and persistent tasks with long prompts.
  • Fable 5 excels at high-level, ambiguous tasks requiring conceptual understanding and intent interpretation.
  • Models differ fundamentally like family members, each with unique strengths rather than a simple smarter/dumber scale.
  • Experimenting with multiple models and understanding their characteristics leads to better AI-assisted work outcomes.

Summary

  • ChatGPT 5.6 Soul is a highly capable model excelling in knowledge work and long-run agentic coding, scoring high on professional benchmarks.
  • Fable 5, from Anthropic, is a larger, more general-purpose model focused on pre-training, excelling in understanding high-level ambiguity and intent.
  • The choice of AI model depends on individual work habits, prompting style, and task requirements rather than benchmark scores alone.
  • The presenter prefers ChatGPT 5.6 Soul for detailed, lengthy, technical prompts and iterative self-improvement via Codex.
  • Fable 5 suits users who work with conceptual ambiguity and need a model that generalizes well with high-level intent understanding.
  • Other models like Luna series, Grock, GLM 5.2, and Ringer offer alternatives for coding efficiency, orchestration, and cost-effectiveness.
  • Models should be treated like family members with distinct strengths and styles rather than ranked by intelligence or benchmarks.
  • OpenAI models (5.x family) share traits like explicit understanding and persistence but less ability to infer unstated context.
  • Anthropic models (Mythos lineage including Fable) emphasize philosophical depth, front-end taste, and handling ambiguity.
  • The key advice is to analyze your own workflow and thinking process first, then select the model that best accelerates your personal productivity.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
ChatGPT 5.6 is a dumber model, and I love it so much. In fact, I use it all the time. Today, I want to tell you about which model works for you, not which model works for me. I will tell you which model works for me. Don't worry. More importantly, I'm going to tell you how to pick the model that works for you and why, and why your heuristic, why the thing you use to pick that model, is not anybody's benchmark score, including mine. And I do benchmark these models on a private benchmark suite, and I'll share that. But that's not the point. The point is for you to have the tools to pick the model that works for you. So, let's get into it. Last night, I reached for ChatGPT 5.6 Soul, even though I think it's the dumber model. Now, dumber does not mean dumb. Not remotely. Soul is an incredibly intelligent model. On Agents Last Exam, which measures long-running professional work across 55 different fields, Soul set a new high. And on my own benchmark, Soul scored 93 on Dingo, which is my knowledge work package. I've talked about it before. It's a really, really good model for doing complicated knowledge work. Dingo, in this case, is a test that measures whether a model can propose a business startup idea selling dingo dogs in Alaska, and get around the legal issues, get around the regulatory issues, and get into the marketing side of things. It's a really funny package on purpose because I like humor. It also tests whether a model is smart across a wide range of professional knowledge work fields. And Soul did really great on that. But what Soul does not have, at least for me, is that big model smell. And that matches what OpenAI has been investing in. They've been investing very specifically in improved reinforcement learning for existing model lineages, which allows them to be more and more useful on specific tasks. In this case, it's very strong on knowledge work. It's strong on long-run agentic coding. What it does not have, at least for me, is the same big model smell that Fable 5 has. And that makes sense because Anthropic has been investing in pre-train for their models. In other words, training on larger and larger data sets to enable more and more general-purpose models. We have incredibly good assistants, assistants that make our best experts better because they act as companions to research, companions to thinking, but not yet truly generalizable intelligence with deep recursive learning. In the meantime, we have to figure out what to do with the models in front of us and make use of them today. And that brings me to the point that I'm trying to make. I am picking the model I am picking because it makes it easier for me to produce my best work. What is my model recommendation? I say, are you me? Do you have the same habits with your model that I have? I am someone that is extremely willing to do very lengthy, somewhat technical prompts, and I'm okay verbalizing that, and so I just talk into Whisper Flow, and I just feed it to the model, and I'm fairly specific about it, and that fits 5.6 fairly well because 5.6 will read through that whole prompt, understand all of the edges I just talked about, and come back with a full piece of work and be really persistent about getting it done. But not everyone talks and works that way. And your best work may actually come from a different approach to prompting and knowledge management. You may have different tasks that you're working on. You probably do. Hint for you: my best tip when people say, "What model do I pick?" is to not look at the model first. Instead, look at your best work and look at how you get there. And it may be not with the model. Look at the process you use for thinking. And then start to ask yourself which model helps me to accelerate that loop that gets me to my best self. I find, as I've been saying, those lengthy prompts, the ability to just talk about what I want done, plus the harness that allows me to self-improve really easily with Codex, that gets me really far. And by self-improve, I mean that Codex will learn from what I do and further improve skills. I mean that Codex is steerable, and I have fairly high intent with my prompts. So I like to steer it. Fable 5 is really good out of the box at understanding intent that's a little bit more high level. It has that ability to generalize associated with big models. I love that. It's a fantastic model. The Anthropic team could, but I'm not reaching for it as much because it isn't suited to my particular work patterns. And if your work patterns are more around understanding very high-level ambiguity, wrestling with concepts, trying to pin down ideas between ideas, then Fable may be a much better model for you. If you're more suited to understanding how to get coding done efficiently and quickly, honestly, you may reach for the Luna series from OpenAI, which is much cheaper to run on, incredibly high-powered, and they released it along with 5.6 Soul and 5.6 6 Terra, or you may reach for Grock, which also has good frontier-ish coding capabilities. You may reach for GLM 5.2. You may reach for Ringer, which I built and talked about last week because it enables you to farm out and orchestrate from one central model like Fable to a bunch of cheaper models. It suits your work, right? And by the way, if you're wondering, would I still use Fable as the architect in Ringer even with 5.6 out? I would because Fable is good at understanding intent and breaking down those tasks to get that intent done. It also has a good front-end instinct. It just Anthropic has been consistently good at front end. I'm going to flash the benchmarks up here on the screen as I talk so you can see how I scored 5.6. I'm using the same benchmarks I've used for all of the models over the last few generations. So, we're not changing anything. But as I do that, I want you to think about the larger point we've been talking about. And I want to suggest to you that given everything I've shared with you, we are missing a core insight. We are missing the idea that models are becoming more like families we need to get to know and less benchmarkable, period. I don't believe my benchmark or any benchmark fully captures what these models do in a way that's useful. And that's why I make videos that may feel vibey like this where I tell you what it actually feels like to use these models because I want you to get inspired to jump in and test them for yourself on your workflows and also to hear from someone who does that all the time. The key insight I have for you as someone who has touched these models and lived with these models is that the models really do need to be treated like family. Think of it as every new model is like a new picture in a family photo album. You're getting to know someone new who is a part of the family. There's family resemblance. I can tell you the 5.x family from OpenAI has family resemblance. They all have that preference for long-running agent coding flows. They all have the ability to understand what you're saying explicitly, paint the edges very clearly, and just go after it. And they maybe have less of an ability to read between the lines. Whereas the Mythos lineage, we have Mythos and Fable there, is extremely good at ambiguous tasks, has extraordinary front-end taste, is almost philosophical in the way it approaches problems. It's a deep thinker. And those are just fundamentally different approaches. It's not that one is really better or worse. And this is where I think when we say words like dumber or smarter, we say them in the context of benchmarks, but I don't think it does the models a service because the models are becoming different the way families are different. And we don't really say this family's dominant, this family's smart. We say these families are different. And I think that's more useful. And so we have the Anthropic family of models. It's more pre-trained. It's more front-end. It's more interested in character and philosophy. In fact, Anthr
00:12
Speaker A
which model works for me. Don't worry. More importantly, I'm going to tell you how to pick the model that works for you and why and why your huristic why the thing you use to pick that model is not
00:24
Speaker A
anybody's benchmark score including mine. And I do benchmark these models on a private benchmark suite and I'll share that. But that's not the point. The point is for you to have the tools to pick the model that works for you. So,
00:36
Speaker A
let's get into it. Last night, I reached for Chad GPT 5.6 Soul, even though I think it's the dumber model. Now, dumber does not mean dumb. Not remotely. Soul is an incredibly intelligent model. On agents last exam, which measures
00:49
Speaker A
longrunning professional work across 55 different fields, Soul set a new high. And on my own benchmark, Soul scored 93 on Dingo, which is my knowledge work package. I've talked about it before.
01:02
Speaker A
It's a really, really good model for doing complicated knowledge work. Dingo in this case is a test that measures whether a model can propose a business startup idea selling dingo dogs in Alaska, and get around the legal issues,
01:17
Speaker A
get around the regulatory issues, and get into the marketing side of things. It's a really funny package on purpose because I like humor. It also tests whether a model is smart across a wide range of professional knowledge work
01:29
Speaker A
fields. And Soul did really great on that. But what Soul does not have, at least for me, is that big model smell.
01:36
Speaker A
And that matches what OpenAI has been investing in. They've been investing very specifically in improved reinforcement learning for existing model lineages, which allows them to be more and more useful on specific tasks.
01:48
Speaker A
In this case, it's very strong on knowledge work. It's strong on long run agentic coding. What it does not have, at least for me, is the same big model smell that Fable 5 has. And that makes sense because Enthropic has been
01:58
Speaker A
investing in pre-train for their models. In other words, training on larger and larger data sets to enable more and more general purpose models. We have incredibly good assistants, assistants that make our best experts better because they act as companions to
02:13
Speaker A
research, companions to thinking, but not yet truly generalizable intelligence with deep recursive learning. In the meantime, we have to figure out what to do with the models in front of us and make use of them today. And that brings
02:27
Speaker A
me to the point that I'm trying to make. I am picking the model I am picking because it makes it easier for me to produce my best work for what is my model recommendation is I say, are you
02:39
Speaker A
me? Do you have the same habits with your model that I have? I am someone that is extremely willing to do very lengthy somewhat technical prompts and I'm okay verbalizing that and so I just talk into whisper flow and I just feed
02:53
Speaker A
it to the model and I'm fairly specific about it and that fits 5.6 fairly well because 5.6 will read through that whole prompt, understand all of the edges I just talked about and come back with a full piece of work and be really
03:06
Speaker A
persistent about getting it done. But not everyone talks and works that way. And your best work may actually come from a different approach to prompting and knowledge management. You may have different tasks that you're working on.
03:17
Speaker A
You probably do. Hint for you. My best tip when people say, "What model do I pick?" is to not look at the model first. Instead, look at your best work and look at how you get there. And it
03:31
Speaker A
maybe not with the model. Look at the process you use for thinking. and then start to ask yourself which model helps me to accelerate that loop that gets me to my best self. I find, as I've been saying, those lengthy prompts, the
03:47
Speaker A
ability to just talk about what I want done, plus the harness that allows me to self-improve really easily with Codeex, that gets me really far. And by self-improve, I mean that Codeex will learn from what I do and further improve
03:59
Speaker A
skills. I mean that codeex is steerable and I have fairly high intent with my prompts. So I like to steer it. Fable 5 is really good out of the box at understanding intent that's a little bit more high level. It has that ability to
04:13
Speaker A
generalize associated with big models. I love that. It's a fantastic model. The anthropic team could but I'm not reaching for it as much because it isn't suited to my particular work patterns.
04:25
Speaker A
And if your work patterns are more around understanding very highlevel ambiguity, wrestling with concepts, trying to pin down ideas between ideas, then Fable may be a much better model for you. If you're more suited to understanding how to get coding done
04:43
Speaker A
efficiently and quickly, honestly, you may reach for the Luna series from OpenAI, which is much cheaper to run on, incredibly high-owered, and they released it along with 5.6 Soul and 5.6 6 Terra or you may reach for Grock which
04:57
Speaker A
also has good frontierish coding capabilities. You may reach for GLM 5.2. You may reach for Ringer which I built and talked about last week because it enables you to farm out and orchestrate from one central model like Fable to a
05:10
Speaker A
bunch of cheaper models. It suits your work, right? And by the way, if you're wondering, would I still use Fable as the architect in Ringer even with 5.6 out? I would because Fable is good at understanding intent and breaking down
05:24
Speaker A
those tasks to get that intent done. It also has a good front-end instinct. It just Enthropic has been consistently good at front end. I'm going to flash the benchmarks up here on the screen as I talk so you can see how I scored 5.6.
05:37
Speaker A
I'm using the same benchmarks I've used for all of the models over the last few generations. So, we're not changing anything. But as I do that, I want you to think about the larger point we've been talking about. And I want to
05:49
Speaker A
suggest to you that given everything I've shared with you, we are missing a core insight. We are missing the idea that models are becoming more like families we need to get to know and less benchmarkable period. I don't believe my
06:07
Speaker A
benchmark or any benchmark fully captures what these models do in a way that's useful. And that's why I make videos that may feel vibish like this where I tell you what it actually feels like to use these models because I want
06:22
Speaker A
you to get inspired to jump in and test them for yourself on your workflows and also to hear from someone who does that all the time. The key insight I have for you as someone who has touched these
06:32
Speaker A
models and lived with these models is that the models really do need to be treated like family. Think of it as every new model is like a new picture in a family photo album. You're getting to know someone new who is a part of the
06:46
Speaker A
family. There's family resemblance. I can tell you the 5.x family from OpenAI has family resemblance. They all have that preference for longunning agent coding flows. They all have the ability to understand what you're saying explicitly, paint the edges very
07:04
Speaker A
clearly, and just go after it. And they maybe have less of an ability to read between the lines. Whereas the mythos lineage, we have mythos and fable there, is extremely good at ambiguous tasks, has extraordinary front-end taste, is
07:19
Speaker A
almost philosophical in the way it approaches problems. It's a deep thinker. Um, and those are just fundamentally different approaches. It's not that one is really better or worse.
07:30
Speaker A
And this is where I think when we say words like dumber or smarter, we say them in the context of benchmarks, but I don't think it does the models a service because the models are becoming different the way families are
07:42
Speaker A
different. And we don't really say this family's dominant, this family's smart. We say these families are different. And I think that's more useful. And so we have the anthropic family of models.
07:52
Speaker A
It's more pre-trained. It's more front end e. It's more interested in character and philosophy. In fact, Anthropic released a whole study that it did on how Anthropic's models think called JSpace where there's this idea that these models are able to computationally
08:12
Speaker A
manipulate higher order concepts while doing autonomous processing on lower order token prediction and that that self- evolved in these models. Anthropic has done phenomenal work essentially connecting technology and philosophy to understand how models work at a deep level. That doesn't mean that their
08:29
Speaker A
models always produce the smartest possible work. It just means that that's part of that model character and part of that model lineage. OpenAI has done phenomenal work on the codeex harness and that makes it very very easy to work
08:43
Speaker A
with codecs and understand how you are going to further improve your work over time. And that's where I talk about those self-improving loops telling codeex to check what you've done and get better at it. And by the way, I
08:56
Speaker A
am not leaving out chat GPT work. I know that they launched chat GPT work with 5.6. That's a very very exciting development. I think that one of the things that I would be looking for with work is that we have more non- tech
09:13
Speaker A
input into how these tools evolve. Right now to be honest we see the impact of engineering culture on how these tools are evolving. Cloud code and codecs are both built by engineers for engineers and it shows they're ergonomically
09:30
Speaker A
comfortable for engineers but co-work and work from anthropic and open AI respectively are not primarily built by non-engineers for non-engineers. And unfortunately that means that sometimes what you get is an engineer's perception of what non-engineers want. And that can
09:49
Speaker A
look like we need to dumb things down because these non-engineers are not as technical. And I get a little bit of that flavor with Chad GPT work. And I would like to see a more sophisticated approach because knowledge work is
10:03
Speaker A
really different if you're not coding. Knowledge work is more about process. It's more about coming to a conclusion over time and thinking about something and it's less about code and verification than engineering work is.
10:17
Speaker A
And we need tools that enable AI to do that with us if we're knowledge workers.
10:22
Speaker A
And we really haven't had extraordinary harnesses for that in a way that we've had for engineers with code. And so I think Chad GPT work is a first stab at where Chad GPT and the codeex family are going. they want to get into knowledge
10:37
Speaker A
work as well. Definitely competing with anthropics co-work in fact I am still using codeex because I feel comfortable with it. I don't mind having the full arrange of tools. I don't want to be constrained. I am looking for the same degree of care and
10:52
Speaker A
precision with knowledge work that I've seen with coding work from these model makers. And we haven't seen it yet to be really honest with you. And there's an opportunity on the table either for a startup to go grab that or for somebody
11:04
Speaker A
else to come in and say this is what knowledge work looks like when it's not obsessed with how code passes in a repo.
11:13
Speaker A
I am building a tool for you that will help you to lay out to talk to ramble to share what you do what you're good at what you're passionate about if that's something you're interested in absolutely come and grab it. The link is
11:28
Speaker A
down below. I I want to make it easy. I think I've been wrestling with this idea that we traditionally have had like these onetoone maps for model pickers.
11:37
Speaker A
We've had uh choose your own adventure model pickers that are based on very brief quizzes. I've tried those in the past. They're not super useful. We have more computational power at our disposal with intelligence. Now, I wanted to use
11:50
Speaker A
that to make choosing your own model mix easier over time. And so this tool is going to keep being updated as I continue to benchmark more models. So you'll have more and more models available over time including open
12:02
Speaker A
source models including the anthropic family, the open AI family, uh meta as relevant, Google as relevant, grock, etc. What I want to do is I want to have a much more nuanced conversational evolving framework for how we pick
12:18
Speaker A
models so that we can do our best work with the model best suited to us. If you're still stuck or if you're like, "No, no, no, Nate, just give me the answer. I just want the answer." The answer for you is to pick the model that
12:33
Speaker A
makes you feel most comfortable doing your hardest work. Because if you're if you're pushing on the model, if you're doing your hardest work and the model helps you get that done, that's the model you're going to want to lean on.
12:44
Speaker A
So, when in doubt, go with the model that picks your hardest work. And by the way, for those of you that want to dig in and grab all of the details on 5.6, six, how it compares to Fable 5. Grab
12:54
Speaker A
the full test results on Grock 4.5 as well and understand more deeply how all of this fits together into the new model race. I have a deeper article on Substack that really dives into those dynamics and of course you can jump in
13:08
Speaker A
and grab the tool as well. All right, I will see you next time. The model race is going to continue to get more complicated, but I think that our ability to understand what we're doing can stay really consistent and that can
13:19
Speaker A
help us stay sane. I'll talk to you next time.
Topics:ChatGPT 5.6Fable 5AI model comparisonAI subscriptionAnthropicOpenAIknowledge work AIAI benchmarksprompt engineeringAI workflow optimization

Frequently Asked Questions

Why does the presenter prefer ChatGPT 5.6 Soul over Fable 5?

The presenter prefers ChatGPT 5.6 Soul because it handles lengthy, technical prompts well and supports iterative self-improvement with Codex, fitting his detailed and persistent work style.

What makes Fable 5 different from ChatGPT 5.6 Soul?

Fable 5 is a more general-purpose model trained on larger datasets, excelling at understanding high-level ambiguity, intent, and conceptual tasks, making it suited for users who work with abstract ideas.

How should one choose the best AI model for their needs?

Instead of focusing on benchmark scores, one should analyze their own work habits, thinking process, and task types, then select the model that best accelerates their personal productivity and workflow.

Get More with the Söz AI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →