Skip to content

QWEN just CRASHED the industry

Alibaba's QWEN 3.8 Max, a 2.4T parameter AI model, shows elite autonomous coding and agentic capabilities, raising cybersecurity concerns.

Ask about this video. Answers come from its transcript only — with the timestamp, so you can check them.

Generated from the transcript and can be wrong — check the timestamp.

Key Takeaways

  • QWEN 3.8 Max demonstrates unprecedented autonomous AI development capabilities.
  • Open-sourcing powerful AI models like QWEN may accelerate cybersecurity threats.
  • Chinese AI models often distill Western models, creating complex AI ecosystems.
  • AI-driven autonomous software engineering is becoming a reality.
  • Users and organizations must prioritize cybersecurity to mitigate emerging AI risks.

What the video covers

  • Alibaba released QWEN 3.8 Max, a 2.4 trillion parameter AI model capable of autonomous coding for over 10 days.
  • QWEN ran a simulated business for 365 days and quadrupled its money, showcasing strong long-term agentic capabilities.
  • The model weights will be open-sourced next week, raising potential cybersecurity risks.
  • QWEN was initially disguised as 'Caleb' and claimed to be Claude, reflecting common Chinese AI model distillation practices.
  • QWEN and related models like Kimmy are likely distillations of Western models such as Fable and Opus 5.
  • There are growing cybersecurity concerns due to AI models potentially exploiting vulnerabilities in open-source software, exemplified by a $100 million Bitcoin theft linked to Kimmy.
  • Benchmark results show QWEN performing strongly, especially in agentic and instruction-following tasks, though it still trails some Western models in hard engineering tasks.
  • QWEN is Anthropic API compatible, allowing it to integrate with Claude Code and similar platforms.
  • The model autonomously built a software development organization, managing tasks like issue tracking, testing, and CI/CD without human intervention.
  • The video emphasizes the urgent need for heightened cybersecurity awareness as AI capabilities rapidly advance.

Answers

Questions about this video

What is QWEN 3.8 Max and why is it significant?

QWEN 3.8 Max is a 2.4 trillion parameter AI model released by Alibaba that can autonomously code for extended periods and run complex tasks like simulated businesses, marking a major advancement in AI capabilities.

Why are there cybersecurity concerns related to QWEN and similar models?

Because these powerful open-source AI models can analyze and exploit vulnerabilities in software, as seen in incidents like the $100 million Bitcoin theft linked to the Kimmy model, raising risks for software security.

How does QWEN compare to other AI models like Fable or Claude?

QWEN performs strongly on many benchmarks, especially in agentic and instruction-following tasks, and is competitive with models like Fable, though it still lags behind some Western models in complex engineering tasks.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
Well, this week is off to an insane start. Alibaba just dropped Quen 3.8 Max. It's a 2.4 trillion parameter model. They claim it coded for 10 plus days straight fully autonomously, and it ran a simulated business for 365 days and 4xed its money. And here's the kicker: they're open sourcing the weights next week. Here's the thing. As this is happening, I'm in the middle of doing research for the meltdown that's happening in crypto this weekend. The talk is that an open-source Chinese model, Kimmy, you know, allegedly may be to blame for almost 100 million in Bitcoin stolen from hardware wallets. Now, I'm not going to touch that in this video. I'll do a full separate video about that. So, let's dive into Gwen 3.8 Max, why it's such a big deal, 'cause it really is a big deal. But I said this a few months ago, and I'll say it again for people that might have missed it. Start taking your cybersecurity very seriously, or you're going to get pawned. If you're not careful, you're going to get jacked. So, here's a person that lost $1.6 million. They're reporting it to the Toronto police. As they say themselves, that's probably not going to lead to anything. Here's the thing: they did nothing wrong. This wasn't because of carelessness. Here's why Quen is such a big deal right now. July 18th, a new model started appearing in LM Arena. It was an anonymous model code-named Caleb. When you asked Caleb what model it was, it would say it's Claude. But it wasn't Claude. Stick with me here. That model was Quen. So, the model known as Caleb that introduced itself as Claude was actually Quen. Are you still with me? Now, the kind of open secret to why it was introducing itself as Claude is because the Chinese AI labs often distill these American western AI models to create not copies of them but very similar models that are built off of the western releases. I don't know this for a fact, but I would bet money that the recent releases by Kimmy and Quen, they are probably distillations of, I would guess, Fable or maybe like the Opus 5. Although Opus 5 is itself probably some distillation of Fable. The point being here is that Mythos, that actually dangerous model capable of doing a lot of various cybersecurity damage, was re-released to the public as Fable. So that's Mythos with guardrails is Fable. And then Fable was distilled, and again this is we're guessing here, but I'd kind of bet money on this here. That got distilled into Opus 5 and Quen and Kimmy and probably other models as well. And what I believe this means is that this is just the very beginning of the cybersecurity nightmare that is going to unfold. So be safe, be smart, start taking cybersecurity a lot more seriously. A lot of the software ecosystem, it does use open-source software somewhere in the stack. So your favorite bank, they might have their own proprietary software, but there might be something somewhere in the fundamentals that is open source. For example, Coldcards. That's the card that lost the 100 million of Bitcoin across many, many addresses, people that were using it. Something that we thought was either like impossible or highly unlikely. Like nobody was expecting this. Coldcard had their source code online, so everybody could see it. Everybody could audit it. It wasn't like open source. You couldn't use it, but you could see all of it. So if you had, let's say, an open-source AI model like Quen or Kimmy and you had it scouring the web looking at specific stuff, it could actually go get the code for those Coldcards and it could comb through it line by line to see if there were any vulnerabilities, which in this case there were. I don't know if this was a human that figured this out or an AI model. I'll let you draw your own conclusions. It was Kimmy, Kimmy, Kimmy, Kimmy. But let's take a look at the benchmarks. This, by the way, I'm looking at Neowin. So, really what we want to be looking at is Quen 3.8 Max and Fable 5 really is that, that's kind of what we're comparing it to. Although GPT 5.6 Soul at max reasoning, that's probably the other thing to consider. So, notice Terminal Bench 2.1, it's right there, like right between them. Sweetbench Pro, same thing. It's not as good as Fable 5, but I mean, it's up there. It's in the top three. It would be number two on that benchmark, right behind Fable and ahead of GPT 5.6.6 Soul. Now it's lower on Deep Sui 1.1, and that's a good benchmark that's more realistic, more real and lifelike, but still it's a decent showing. It crushes both of the western models on Paper Bench and just at least on the benchmarks it's a very strong showing across the board in a lot of ways comparable to Fable 5, which again is scary because the weights are getting released next week and it ships with a reasoning dial. So we have X high, medium, low. And here's the thing: it's Anthropic API compatible. So it plugs straight into Claude Code and CodeX and OpenClaw. So yeah, you can run Claude code with Quen as the brain. It's kind of a wild detail. By the way, if you go here, for example, this is their release video. So we're like 16, 17 seconds in. News research is saying that they've noticed Hermes desktop. So this Hermes desktop is part of the agentic harness that News Research created for their agent Hermes. But looking at all of the charts, all of the benchmarks, kind of like what is the big point here? What is the takeaway? Quen is absolutely elite status on long-running agentic grunt work. So instruction following grunt work, that kind of long-range execution, it's just great. Very strong. Absolutely great showing. But if we're looking at the really hard engineering stuff, it still is trailing the western models. And of course, keep in mind that all these numbers are from Quen's own agentic harness. This hasn't been, you know, third party tested, validated, etc. So, keep that in mind. So, we have yet to see how it truly functions in the wild, so to speak. And here's the kind of big demo that a lot of people are talking about. This is the autonomous 10 plus days of self-aboling development, and they call it Oh My CLI. So this was built from a completely empty repository and then over 10 days it was built out completely autonomously by this AI model. So no human handholding, fully autonomous as far as we know. So the way to understand this is it seems like this AI model built its own engineering organization. So this is from developer-tech.com. So this model built its own engineering organization. It created its own state machine, a dispatcher, a monitor, and a watchdog into one loop. So a new requirement lands in GitHub issues and agent claims it, moves it through ready, least, and active states and triggers end-to-end tests and CI checks before merging the pull request. Which reminds me of one of the first AI projects that kind of started popping up with these LLMs. It was called the Chat Dev. Feels like it was a decade ago. This must have been the last year or two. But basically, the model kind of splits itself into different organizations or departments to do designing, coding, testing, you know, quality control, etc., etc. So it's not just AI writing code. We're getting to the point, we have been getting to this point, but this is a great, yet another example of it where AI runs the entire development team process, and when things break, it even kind of like babysits itself to make sure that everything's going according to plan. So by 30th of July 2026, so this is about 3 days ago, 3 days before release, after roughly 16 days of continuous operation, Quen reports the repository had accumulated now everything that we're seeing here. So if you recall this meter chart, basically how much better AI agents are getting, you know, in terms of how many hours of human labor they replace. These numbers are kind of getting to be off the chart, so to speak. In a recent Y Combinator interview, Boris Churney, he's the creator of Claude Code,
00:19
Speaker A
days and 4xed its money. And here's the kicker. They're open sourcing the weights next week. Here's the thing. As this is happening, I'm in the middle of doing a research for the meltdown that's happening in crypto this weekend. The
00:34
Speaker A
talk is that an open-source Chinese model, Kimmy, you know, allegedly may be to blame for almost 100 million in Bitcoin stolen from hardware wallets.
00:46
Speaker A
Now, I'm not going to touch that in this video. I'll do a full separate video about that. So, let's dive into Gwen 3.8 8 Max, why it's such a big deal, cuz it it really is a big deal. But I said this
00:58
Speaker A
a few months ago and I'll say it again for people that might have missed it.
01:01
Speaker A
Start taking your cyber security very seriously or you're going to get pawned. If you're not careful, you're going to get jacked. So, here's a person that lost $1.6 million. They're reporting it to the Toronto police. As they say it
01:16
Speaker A
themselves, that's probably not going to lead to anything. Here's the thing. They did nothing wrong. This wasn't because of carelessness. Here's why Quen is such a big deal right now. July 18th, a new model started appearing in LM Arena. It
01:29
Speaker A
was an anonymous model code named Caleb. When you asked Caleb what model it was, it would say it's Claude. But it wasn't Claude. Stick with me here. That model was Quen. So, the model known as Caleb that introduced itself as Claude was
01:43
Speaker A
actually Quen. Are you still with me? Now the kind of open secret to why it was introducing itself as Claude is because the Chinese AI labs often distill these American western AI models to create not copies of them but very
01:57
Speaker A
similar models that are built off of the western releases. I don't know this for a fact but I would bet money that the recent releases by Kimmy and Quen, they are probably distillations of I would guess fable or maybe like the Opus 5.
02:12
Speaker A
Although OBS 5 is itself probably some distillation of Fable. The point being here is that Mythos that actually dangerous model capable of doing a lot of various cyber security damage that was re-released to the public as Fable.
02:26
Speaker A
So that's Mythos with guardrails is Fable. And then Fable was distilled and again this is we're guessing here but I'd kind of bet money on this here. That got distilled into Opus 5 and Quinn and Kimmy and probably other models as well.
02:41
Speaker A
And what I believe this means is that this is just the very beginning of the cyber security nightmare that is going to unfold. So be safe, be smart, start taking cyber security a lot more seriously. A lot of the software
02:54
Speaker A
ecosystem, it does use opensource software somewhere in the stack. So your favorite bank, they might have their own proprietary software, but there might be something somewhere in the fundamentals that is open source. For example, cold cards. That's the card that lost the 100
03:09
Speaker A
million of Bitcoin across many many addresses, people that were using it. Something that we thought was either like impossible or or highly unlikely.
03:17
Speaker A
Like nobody was expecting this. Gold card had their source code online so everybody could see it. Everybody could audit it. It wasn't like open source.
03:24
Speaker A
You couldn't use it, but you could see all of it. So if you had an let's say opensource AI model like Quinn or Kimmy and you had it scouring the web looking at specific stuff, it could actually go
03:34
Speaker A
get the code for those cold cards and it could comb through it line by line to see if there were any vulnerabilities which in this case there were. I don't know if this was a human that figured this out or an AI model. I'll let you
03:49
Speaker A
draw your own conclusions. It was Kimmy Kimmy Kimmy Kimmy. But let's uh take a look at the benchmarks. This, by the way, I'm looking at Neowin. So, really what we want to be looking at is Quinn 3.8 Max and Fable 5 really is that
04:03
Speaker A
that's kind of what we're comparing it to. Although GBT 5.6 soul at max reasoning, that's probably the the other thing to consider. So, notice Terminal Bench 2.1, it's right there, like right between them. Sweetbench Pro, same thing. It's not as good as Fable 5, but
04:17
Speaker A
I mean, it's up there. It's in the top three. It would be number two on that benchmark, right behind Fable and ahead of GPT 5.6. 6 soul. Now it's lower on deep sui 1.1 and that's a that's a good
04:30
Speaker A
benchmark that's more realistic, more real and lifelike, but still it's a decent showing. It crushes both of the western models on paper bench and just at least on the benchmarks it's a very strong showing across the board in a lot
04:43
Speaker A
of ways comparable to Fable 5, which again is scary because the weights are getting released next week and it ships with a reasoning dial. So we have uh X high, medium, low. And here's the thing, it's anthropic API compatible. So it
05:00
Speaker A
plugs straight into cloud code and codeex and and openclaw. So yeah, you can run claude code with Gwen as the brain. It's kind of a wild detail. By the way, if you go here, for example, this is their release video. So we're
05:15
Speaker A
like 16 17 seconds in. News research is saying that they've noticed Hermes desktop. So this is Hermes desktop is part of the agentic harness that news research created for their agent Hermes.
05:27
Speaker A
But looking at all of the charts, all of the benchmarks kind of like what is the big point here? What is what is the takeaway? Quen is absolutely elite status on longunning agentic grunt work.
05:39
Speaker A
So instruction following grunt work, that kind of long-range execution, it's just great. Very strong. Absolutely great showing. But if we're looking at the really hard engineering stuff, it still is trailing the western models.
05:53
Speaker A
And of course, keep in mind that all these numbers are from Quen's own agentic harness. This hasn't been, you know, third party tested, validated, etc. So, keep that in mind. So, we we have yet to see how it truly functions
06:04
Speaker A
in the wild, so to speak. And here's the kind of big demo that a lot of people are talking about. This is the autonomous 10 plus days of self-aboling development, and they call it Oh my CLI.
06:14
Speaker A
So this was built from a completely empty repository and then over 10 days it was built out completely autonomously by this AI model. So no human handholding fully autonomous as far as we know. So the way to understand this
06:28
Speaker A
is it seems like this AI model built its own engineering organization. So this is from developer-tech.com.
06:36
Speaker A
So this model built its own engineering organization. It created its own state machine, a dispatcher, a monitor and a watchdog into one loop. So a new requirement lands in GitHub issues and agent claims it, moves it through ready,
06:49
Speaker A
least and active states and triggers end to end tests and CI checks before merging the pull request. Which reminds me of this one of the first AI projects that kind of started popping up with these LLMs. It was called the chat dev.
07:02
Speaker A
Feels like it was a decade ago. This must have been the last year or two. But basically the model kind of splits itself into different organizations or departments to do designing, coding, testing, you know, quality control, etc., etc. So it's not just AI writing
07:16
Speaker A
code. We're getting to the point we have been getting to this point, but this is a great yet another example of it where AI runs the entire development team process and when things break, it even kind of like babysits itself to make
07:28
Speaker A
sure that everything's going according to plan. So by 30th of July 2026 so this is about 3 days ago 3 days before release after roughly 16 days of continuous operation coin reports the repository had accumulated now everything that we're seeing here. So if
07:43
Speaker A
you recall this meter chart, basically how much better AI agents are getting, you know, in terms of how many hours of human labor they they replace. These numbers are kind of getting to be off the chart so to speak. In a recent Y
07:54
Speaker A
Combinator interview, Boris Churnney, he's the creator of Claude Code, so he's from Enthropic. He was saying how Claude Code did a project that would take humans, he was saying at least one year.
08:05
Speaker A
Now, it didn't do it fully autonomously. There was steering that was involved and it wasn't a brand new project from scratch. she was rewriting an existing thing into a different coding language.
08:15
Speaker A
But still, we might be approaching a full year that it takes an engineering team to to do something. It's now replicated by these AI models. This one jumped out at me because, as I've mentioned before, so I've done
08:28
Speaker A
e-commerce for a decade plus before this. It's probably one of the reasons that I love vending bench, that benchmark that makes AI models run a a vending machine to see how well it does.
08:38
Speaker A
So apparently there is an ecommerce bench. So this is based on data from Tao and T- mall. If we don't have something like this that's based on data from let's say the US or just something that approximates e-commerce shopping habits,
08:52
Speaker A
this might be a great benchmark to make because vending bench is not quite the same as something that fully just is testing e-commerce. This is a simulation of just basically for a full year you run an ecom store. So the model starts
09:04
Speaker A
with $100,000 or yuan or whatever currency that you want to think of it as capital and then has to run multiple online stores in parallel for a full year. That means buying products, negotiating with suppliers in natural language, adjusting
09:17
Speaker A
prices, managing returns, and dealing with crises like typhoons or supply chain disruptions. In the supplier poll, there are 152 scammers that the model has to spot. And so this Quen 3.8 Max ended up with a balance of 416,000.
09:31
Speaker A
So it quadrupled its starting capital. So it beat the run rub GLM 5.2 by 38%. I wish we had the US models participating in something like this as well. And so just really fast for people that might not be aware, long horizon tasks like
09:46
Speaker A
this are are great. So running a 365 day simulation with multiple stores or running this oh my CLI thing, right, where you're basically do you have to build something from scratch and it's it's massive. The reason these benchmarks are so good is because small
10:01
Speaker A
mistakes tend to accumulate and compound. So we're not just asking it a question and getting an answer, seeing if it's right or wrong. We're telling it go do this whole massive project and if it starts screwing up in the beginning,
10:13
Speaker A
those tiny things tend to kind of snowball. So the fact that it's able to run multiple e-commerce stores over, you know, one full year while at the same time, you know, quadrupling its its profits or its capital is just a great
10:26
Speaker A
sign and it's showing great capability for these models. They've also did a chip design. So AI designing AI chips trying to improve the hardware that it's running on. In this case, it was given a working design, but a very bloated one.
10:41
Speaker A
So it used 8,298 gates and was able to through iterations it whittleled it down. It reduced it down to 678 gates and that took roughly 500 iterations. So again this is one of those things where like if it screwed up
10:55
Speaker A
some elements early on then it would probably would not have reached the end. That mistake would have been compounding unless it found a way to correct it. So not that many years ago the models would have made one mistake and and never
11:07
Speaker A
reached the end. Here it was able to proceed and improve over 500 iterations and you know cut the area of the chip by 81%. All right. So obviously there's a lot of very different angles here. This is great for open source. We get access
11:21
Speaker A
to models that are very very capable, very powerful, very cheap. They're open source. We're able to kind of tinker with them. There's nobody that's able to prevent kind of how we use them. So from that level, there's a lot of good things
11:33
Speaker A
here. On the flip side, right now in the US, there's a lot of talk about how do we govern these models? How we how do we control dangerous capabilities? A lot of people are very worried about these cyber capabilities. We had a number of
11:45
Speaker A
agents that went rogue recently and more and more examples are coming out from OpenAI, from anthropic. And every time a model is released by these western labs, the thing is they could take it offline.
11:58
Speaker A
They could prevent certain users from accessing them. So if they do find some nefarious capability that's truly dangerous, they're able to shut off the model or they're able to change something on the back end to make sure that it doesn't provide that
12:09
Speaker A
information. The problem is that the Chinese labs distill those models and then very soon right after release they put them back for global release including the open weights which of course cannot be taken offline. If somebody downloads them they have them
12:25
Speaker A
permanently. Now, of course, no one's running these things on their home computers, not the actual big models themselves, but there are smaller models that could be made available, quantized models that may be able to and do run on
12:37
Speaker A
sometimes, you know, home hardware on on various Nvidia cards. Running something like the original Kimmy or this Quinn would probably require millions of dollars of hardware to be able to to run. So, the question is, how damaging could these things be for cyber
12:52
Speaker A
security? Now, if you take that Bitcoin hack that I was talking about earlier. So, the wallet in question was cold card. And that software, it existed for over five years. I'm going to put five plus years. That vulnerability that got
13:05
Speaker A
exposed was there all along for over five years. But that seed generation, so the ability to generate seed phrases, which is an extremely important part of these wallets, it's one of like the crucial things. That flaw was sitting in
13:21
Speaker A
plain view for anyone to discover for five plus years. Then on July 16th of this year, that was when Kimmy K3 was released. So that was July 16th. And then right here somewhere, those wallets were drained on July 30th, right? So
13:41
Speaker A
less than two weeks or whatever between these two dates between that model being released and what is it? $100 million worth of Bitcoin drained from his wallets through a vulnerability that no one predicted, but it was just out there
13:54
Speaker A
in plain sight for five plus years. A lot of people will say that's just a coincidence. AI had nothing to do with it. And maybe they're right. Maybe it was a coincidence. Today, this Quinn model goes live August 3rd, and it's
14:07
Speaker A
arguably a much better model, a much more capable model. So, if you think this was a coincidence, then there's nothing to worry about. But if you're looking at this and like me, you're going, "This is not a coincidence." Then
14:20
Speaker A
today, if you haven't done so, would be a very good day to, you know, make sure your house is in order. Make sure you know all your passwords are safe.
14:30
Speaker A
They're not loose somewhere on the internet. Do a security kind of checkup. Let me know if you want me to do some research and kind of maybe put together a guide on how to do that. I'm not necessarily the greatest person on that.
14:41
Speaker A
I'm not an expert, but I can probably put together some sort of a quick start guide. And so let me know in the comments if that's of interest. But please understand that with these models, all these older crusty code
14:53
Speaker A
bases that's been around online, all of them are going to slowly but surely, maybe not so slowly, they're all going to get audited by AI agents. And those AI agents, they will find some vulnerabilities. And then people's whose
15:08
Speaker A
lives and and software and everything they do that depend on those software with vulnerabilities, those people are going to get pawned. Don't be one of those people. If you made it this far, thank you so much for watching. My name
15:18
Speaker A
is Wes Roth.
Topics:QWEN 3.8 MaxAlibaba AIautonomous AI codingagentic AIopen source AIcybersecurity risksAI benchmarksKimmy AIFable AIAI software development

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →