Is This the Biggest AI Release of 2026? (China’s New De… — Transcript

China's Moonshot AI releases Kimmy K3, a 2.8 trillion parameter open AI model rivaling top US closed models, signaling a major shift in AI leadership.

Key Takeaways

  • Kimmy K3 is a groundbreaking open AI model that challenges US dominance in AI technology.
  • Its efficiency and cost advantages could disrupt the market for closed AI services.
  • The model’s ability to handle long, complex tasks with multimodal inputs is a significant advancement.
  • China is positioning itself as a central player in global AI development and governance.
  • Open models are rapidly closing the performance gap with closed models, accelerating AI innovation.

Summary

  • Moonshot AI unveiled Kimmy K3, a 2.8 trillion parameter open-weight AI model, the largest ever announced.
  • Kimmy K3 rivals and in some tasks outperforms leading American closed AI models like Anthropic's Fable 5 and GPT 5.6 Soul.
  • The model is designed for long-running tasks, especially software development, capable of processing up to 1 million tokens in one session.
  • Kimmy K3 supports multimodal inputs including text, images, and video, and can self-edit outputs such as promotional videos.
  • Moonshot’s model uses a mixture of experts design with 896 experts, activating only 16 at a time for efficiency.
  • The model is open-weight but requires significant hardware resources, recommended to run on systems with at least 64 AI accelerators.
  • Kimmy K3’s pricing is significantly cheaper than comparable US models, posing a competitive threat to closed AI providers.
  • Demonstrations include building 3D games, GPU code optimization, chip design, and astrophysics analysis, showcasing versatility.
  • China aims to lead a new AI order by spreading its models globally and influencing AI regulations.
  • The release challenges the notion that open models lag behind closed systems by months, suggesting the gap is now days.

Full Transcript — Download SRT & Markdown

00:02
Speaker A
China may have just created its biggest AI shock since DeepSeek, but this time the story is much larger than one powerful new AI model. On the same day that Moonshot AI unveiled a giant new model called Kimmy K3, Xiinene Ping stood on a stage in Shanghai and told the world that China doesn't plan to follow America's rules for artificial intelligence. It wants to help write the rules, spread its own models across the developing world, [music] and build a new AI order with China at the center. And Kimmy K3 gives that speech real weight because this isn't another model that looks impressive only inside a company presentation. Now, K3 has 2.8 trillion parameters, making it the largest open-weight AI model ever announced. Moonshot says it's the first open model to approach the 3 trillion mark and its full weights are expected to be released on July 27th. Companies and researchers will then be able to download it, modify it, and run it on their own infrastructure instead of being permanently locked inside Moonshot's app. But the size isn't why this release is sending shock waves through the industry. And the real reason is that Kimmy K3 appears to be dangerously close to the best American closed models. And in some tasks, it's already beating them. Moonshot says K3 can compete [music] with Anthropic's Fable 5 and outperform Claude Opus 4.8, GPT 5.5, and parts of GPT 5.6 in demanding coding work. And independent testing actually tells a similar story. Artificial analysis gave K3 a score of 57 on its intelligence index. Claude Opus 4.8 scored around 56, GPT 5.6, Six, Terra scored 55, and Gemini 3.1 Pro landed at about the same level as K3. So only Claude Fable 5 and GPT 5.6 Soul remained ahead and even there the gap [music] was only two or three points. For years, the standard belief was that open models were 6 to 12 months behind the best closed systems. More recently, people started saying Chinese labs were perhaps 3 to 6 months behind. One Reddit commenter described the new reality more bluntly. Maybe they're six days behind. Now, the most dramatic result came from arena.i's front-end code arena, where developers ask models to build real websites and interfaces, then vote on which result is better. Kimmy K3 entered at number one with 1,679 points. Claude Fable 5 scored 1,631. GPT 5.6. Saul scored 1,618. K3 finished first in six of the seven areas tested, [music] including brand and marketing work and data dashboards. Moonshot's previous model, Kimmy K 2.6, had been sitting in 18th place. K3 jumped 17 positions in one generation and went straight to the top. Even Arena's CEO called it potentially the biggest AI release of the year and suggested it could mark the moment China moved ahead of the United States in at least one major part of the model race. Now, Moonshot isn't pretending K3 wins everything. The company admits it still trails Fable 5 and GPT 5.6 Saul in overall user experience [music] and some broader tasks. But K3 doesn't need to dominate every category to be disruptive. It only needs to make companies ask an uncomfortable question. Why keep [music] paying premium prices to American providers if an open alternative is almost as capable? So, Moonshot built K3 mainly for long-running work, especially software development. And this isn't a model designed only to answer one prompt and stop. It's meant to inspect a large codebase, create a plan, use tools, [music] make changes, check whether they worked, and continue for hours with limited human help. It can also look at what appears on the screen. Moonshot calls this vision in the loop. The model writes code, looks at the result, notices what's wrong, changes the code, and checks again. That makes it useful for websites, games, design tools, animation, and other projects where [music] it needs to see whether its own work actually looks right. And since we're already talking about how businesses are using AI, here's something useful. Most of you voted for seven AI agents businesses are paying $3,000 to $10,000 for right now in yesterday's poll. So, we're giving it away for free. It breaks down the exact agents [music] businesses are buying, what they cost, who needs them, and how to spot the right clients. It also covers the ROI, setup fees, retainers, [music] and what a realistic $10,000 per month client list could look like with zero coding required. You can grab it through the link in the video description. Now, back to Kimmy K3. In one demonstration, K3 built a 3D open world game inside a browser using 3.js, JS WebGPU, and GPU [music] compute. It generated the environment and used outside tools to create a rider and horse. Moonshot also showed it building a simulation of China's Long March 10 rocket [music] and a Game Boy Advance emulator. And the company says K3 spent 15 hours improving GPU code and cut the required compute time by more than half. [music] It created a small GPU compiler called Mini Triton from scratch and reportedly designed a working chip over 48 hours using open-source engineering tools. Moonshot also says K3 reproduced an astrophysics analysis in about 2 hours. It reviewed more than 20 papers and wrote over 3,000 lines of Python code. According to the company, an experienced team might normally spend between 1 and 2 weeks on similar work. Now, these are company demonstrations, so they still need wider testing, but they show the goal clearly. K3 is supposed to stay on a difficult task for an entire afternoon, an entire night, or even several days. It can process up to 1 million tokens in one session, roughly 750,000 words. That means it can take in huge code bases, long documents, research papers, or project histories without immediately losing track of what came earlier. It also handles text, [music] images, and video inside the same model. Moonshot says K3 even edited its own promotional video from 56 clips. Its Kimmy work platform is adding interactive widgets and dashboards so users can build persistent workspaces [music] rather than using the model only through a normal chat box. The model uses a mixture of experts design and the simple explanation is that it contains hundreds of specialized sections, but it doesn't activate all of them for every question. K3 has 896 experts with only 16 active at a time. That lets Moonshot build a 2.8 trillion parameter system without using the entire model for every word it generates. Moonshot also created Kimmy Delta attention and attention residuals, two systems designed to help K3 [music] keep track of information through very long tasks. Combined with new training methods, the company claims they make K3 about 2.5 times more efficient at scaling than Kimmy K2. Now, the model was trained with MXFP4 weights and MXFP8 [music] activations, lower precision formats intended to reduce the hardware burden. That doesn't mean anyone will be running it on a normal laptop. And Moonshot recommends systems with at least 64 AI accelerators. The local LLaMA community immediately started joking that all you need is 2 terabytes of VRAM, several Mac Studios, a pile of storage drives, and enormous patience. [music] So, that's the strange reality of K3. It's open, but it isn't small. Most people will never host it at home. The companies that can host it, however, [music] are exactly the ones that matter. Large firms already spend millions every month on AI from OpenAI and Anthropic. If they can run K3 on their own servers, customize it, and keep their data private, [music] this becomes a serious threat to closed AI providers, and the price makes that threat stronger. K3 costs $3 per million input tokens when the input isn't cached, 30 cents when it's cached, and $15 per million output [music] tokens, including reasoning. Those prices stay the same even with long context. Fable 5 costs around $10 per million input tokens and $50 per million output tokens. GPT 5.6 Soul costs about 50 cents per million input tokens and $30 per million output tokens. So K3 is around
00:17
Speaker A
stood on a stage in Shanghai and told the world that China doesn't plan to follow America's rules for artificial intelligence. It wants to help write the rules, spread its own models across the developing world, [music] and build a
00:29
Speaker A
new AI order with China at the center. And Kim K3 gives that speech real weight because this isn't another model that looks impressive only inside a company presentation. Now, K3 has 2.8 trillion parameters, making it the largest openweight AI model ever announced.
00:47
Speaker A
Moonshot says it's the first open model to approach the 3 trillion mark and its full weights are expected to be released on July 27th. Companies and researchers will then be able to download it, modify it, and run it on their own
01:01
Speaker A
infrastructure instead of being permanently locked inside Moonshot's app. But the size isn't why this release is sending shock waves through the industry. And the real reason is that Kimmy K3 appears to be dangerously close to the best American closed models. And
01:16
Speaker A
in some tasks, it's already beating them. Moonshot says K3 can compete [music] with Anthropics Fable 5 and outperform Claude Opus 4.8, GPT 5.5, and parts of GPT 5.6 in demanding coding work. And independent testing actually tells a similar story. Artificial
01:34
Speaker A
analysis gave K3 a score of 57 on its intelligence index. Claude Opus 4.8 scored around 56, GPT 5.6, Six, Terra scored 55 and Gemini 3.1 Pro landed at about the same level as K3. So only Claude Fable 5 and GPT 5.6 Soul remained
01:53
Speaker A
ahead and even there the gap [music] was only two or three points. For years, the standard belief was that open models were 6 to 12 months behind the best closed systems. More recently, people started saying Chinese labs were perhaps
02:06
Speaker A
3 to 6 months behind. One Reddit commenter described the new reality more bluntly. Maybe they're six days behind.
02:13
Speaker A
Now, the most dramatic result came from arena.i's front-end code arena, where developers ask models to build real websites and interfaces, then vote on which result is better. Kimmy K3 entered at number one with 1,679 points. Claude Fable 5 scored 1,631.
02:32
Speaker A
GPT 5.6. Saul scored 1,618. K3 finished first in six of the seven areas tested, [music] including brand and marketing work and data dashboards.
02:43
Speaker A
Moonshot's previous model, Kimmy K 2.6, had been sitting in 18th place. K3 jumped 17 positions in one generation and went straight to the top. Even Arena's CEO called it potentially the biggest AI release of the year and suggested it could mark the moment China
03:00
Speaker A
moved ahead of the United States in at least one major part of the model race.
03:05
Speaker A
Now, Moonshot isn't pretending K3 wins everything. The company admits it still trails Fable 5 and GPT 5.6 Saul in overall user experience [music] and some broader tasks. But K3 doesn't need to dominate every category to be disruptive. It only needs to make
03:22
Speaker A
companies ask an uncomfortable question. Why keep [music] paying premium prices to American providers if an open alternative is almost as capable? So, Moonshot built K3 mainly for longrunning work, especially software development.
03:37
Speaker A
And this isn't a model designed only to answer one prompt and stop. It's meant to inspect a large codebase, create a plan, use tools, [music] make changes, check whether they worked, and continue for hours with limited human help. It
03:50
Speaker A
can also look at what appears on the screen. Moonshot calls this vision in the loop. The model writes code, looks at the result, notices what's wrong, changes the code, and checks again. That makes it useful for websites, games,
04:05
Speaker A
design tools, animation, and other projects where [music] it needs to see whether its own work actually looks right. And since we're already talking about how businesses are using AI, here's something useful. Most of you voted for seven AI agents businesses are
04:20
Speaker A
paying $3,000 to $10,000 for right now in yesterday's poll. So, we're giving it away for free. It breaks down the exact agents [music] businesses are buying, what they cost, who needs them, and how to spot the right clients. It also
04:34
Speaker A
covers the ROI, setup fees, retainers, [music] and what a realistic $10,000 per month client list could look like with zero coding required. You can grab it through the link in the video description. Now, back to Kimmy K3. In
04:48
Speaker A
one demonstration, K3 built a 3D openw world game inside a browser using 3.js, JS WebGPU and GPU [music] compute. It generated the environment and used outside tools to create a rider and horse. Moonshot also showed it building a simulation of China's Long March 10
05:06
Speaker A
rocket [music] and a Game Boy Advance emulator. And the company says K3 spent 15 hours improving GPU code and cut the required compute time by more than half.
05:16
Speaker A
[music] It created a small GPU compiler called Mini Triton from scratch and reportedly designed a working chip over 48 hours using open-source engineering tools.
05:27
Speaker A
Moonshot also says K3 reproduced an astrophysics analysis in about 2 hours. It reviewed more than 20 papers and wrote over 3,000 lines of Python code.
05:38
Speaker A
According to the company, an experienced team might normally spend between 1 and 2 weeks on similar work. Now, these are company demonstrations, so they still need wider testing, but they show the goal clearly. K3 is supposed to stay on
05:53
Speaker A
a difficult task for an entire afternoon, an entire night, or even several days. It can process up to 1 million tokens in one session, roughly 750,000 words. That means it can take in huge code bases, long documents, research papers, or project histories
06:10
Speaker A
without immediately losing track of what came earlier. It also handles text, [music] images, and video inside the same model.
06:18
Speaker A
Moonshot says K3 even edited its own promotional video from 56 clips. Its Kimmy work platform is adding interactive widgets and dashboards so users can build persistent workspaces [music] rather than using the model only through a normal chat box. The model
06:34
Speaker A
uses a mixture of experts design and the simple explanation is that it contains hundreds of specialized sections, but it doesn't activate all of them for every question. K3 has 896 experts with only 16 active at a time. That lets Moonshot
06:49
Speaker A
build a 2.8 trillion parameter system without using the entire model for every word it generates. Moonshot also created Kimmy Delta attention and attention residuals, two systems designed to help K3 [music] keep track of information through very long tasks. Combined with
07:06
Speaker A
new training methods, the company claims they make K3 about 2.5 times more efficient at scaling than Kimmy K2. Now, the model was trained with MXFP4 weights and MXFP8 [music] activations, lower precision formats intended to reduce the hardware burden.
07:23
Speaker A
That doesn't mean anyone will be running it on a normal laptop. And Moonshot recommends systems with at least 64 AI accelerators. The local llama community immediately started joking that all you need is 2 terb of VRAM, several Mac
07:37
Speaker A
Studios, a pile of storage drives, and enormous patience. [music] So, that's the strange reality of K3. It's open, but it isn't small. Most people will never host it at home. The companies that can host it, however, [music] are exactly the ones that matter. Large
07:54
Speaker A
firms already spend millions every month on AI from OpenAI and Anthropic. If they can run K3 on their own servers, customize it, and keep their data private, [music] this becomes a serious threat to closed AI providers, and the
08:08
Speaker A
price makes that threat stronger. K3 costs $3 per million input tokens when the input isn't cached, 30 cents when it's cached, and $15 per million output [music] tokens, including reasoning.
08:21
Speaker A
Those prices stay the same even with long context. Fable 5 costs around $10 per million input tokens and $50 per million output tokens. GPT 5.6 Soul costs about 50 cents per million input tokens and $30 per million output
08:36
Speaker A
tokens. So K3 is around five times more expensive than some earlier Kimmy models, [music] but it's still aggressively priced against Frontier Western systems. Moonshot says it launches [music] at maximum thinking effort by default with cheaper modes coming later. The company also admits K3
08:54
Speaker A
has weaknesses. If an agent system fails to return its full reasoning history, performance can drop. And when instructions are vague, K3 may make decisions on its own. So users who need strict control must set very clear rules. And even with those limits, the
09:10
Speaker A
release immediately reminded people [music] of DeepSeek R1. When Deepseek showed that a Chinese lab could match far more expensive American models, roughly $1 trillion was wiped from major technology stocks during [music] the panic that followed. Washington grew more concerned and the Trump
09:26
Speaker A
administration pushed even harder on technology export restrictions. K3 hasn't triggered that kind of market collapse, at least not yet. But it hit Chinese competitors almost [music] immediately. JIEPU shares fell 21.9% in Hong Kong while Miniax dropped [music] 13.8%.
09:43
Speaker A
So investors clearly understood what had happened. A new model had arrived with enormous scale, frontier level performance, open weights on the way, and pricing that could pressure almost everyone else. And Miniax is reportedly preparing its own 2.7 trillion parameter
09:59
Speaker A
model for release as early as the third quarter of 2026 along with a Frontier multimodal model called H3. Before K3, Mtoan's Longat 2.0 and Deepseek V4 Pro were among China's largest systems at around 1.6 trillion parameters. Several Chinese labs have now crossed the 1
10:20
Speaker A
trillion mark. But K3 isn't an isolated event. Z.AI's AI's GLM 5.2 recently shocked analysts by getting close to the best American closed models. Deepseek is still advancing. Miniaax is preparing larger systems and Chinese releases are becoming faster, cheaper, and more
10:38
Speaker A
capable. And Moonshot itself is backed by Alibaba and Tencent. Bloomberg reported that it's trying to raise $2 billion [music] at a valuation of around $30 billion ahead of a possible Hong Kong listing. There's also a political fight building around how these models
10:54
Speaker A
are trained. Anthropic previously accused Moonshot, Deepseek, and Miniax of using model distillation to copy capabilities from Claude in violation of its rules. Distillation is a common technique where one model helps train another. But US officials have started describing some forms of it as an
11:11
Speaker A
adversarial tactic. [music] And critics pointed out the irony. American AI companies trained their own systems on huge parts of the public internet, then became angry when other companies learned from their models. So, expect that argument to intensify when K3's
11:27
Speaker A
weights are released. There will be accusations about copied capabilities, scraped data, [music] export controls, and national security.
11:35
Speaker A
But those arguments may matter less if Chinese models keep improving this quickly. And that's where Xiinping's speech enters the story. At the World Artificial Intelligence Conference in Shanghai, Shei presented China as the leader of a new global AI order. He
11:50
Speaker A
called open-source AI a rare and historic opportunity and warned that unequal access could create new historical injustices. He compared AI to the invention of the steam engine and electricity, then offered developing countries something very different from the American model. lowercost open
12:07
Speaker A
technology, Chinese training, Chinese expertise, and a larger role in deciding how AI is governed. She promoted the new World AI Cooperation Organization or WICO, which signed up 29 countries the [music] day before his speech. He called it a milestone in AI history. China also
12:26
Speaker A
plans to build cooperation centers and training programs with BRICS, Azion, Latin America, and the African [music] Union. So this is a direct challenge to the US-led PAX silica initiative which is trying to secure AI infrastructure, semiconductor supply chains, and
12:43
Speaker A
critical minerals among American partners. She didn't name the United States, but he didn't need to. His message was that China won't accept American control over AI standards, advanced chips, global supply chains, [music] or access to powerful models. He
12:58
Speaker A
also made his strongest comments yet on AI safety. She said AI must remain under human control and called for early warning systems, emergency plans, and protection [music] against autonomous systems escaping human oversight. China is therefore presenting itself as the
13:15
Speaker A
country offering open access to the world while also claiming it can lead on safety and global standards. The World AI conference runs from July 17th to [music] July 20th. Attendees include major Chinese technology companies, UN Secretary General Antonio Gutirez,
13:32
Speaker A
Kazakhstan's President Kasim Jomar Tokayv [music] and Thailand's Prime Minister Anutin Charvira. It also comes just before the first government level AI talks between China and the United States under President Donald Trump. At a UN AI meeting last week, American officials
13:48
Speaker A
argued that too much regulation could slow innovation. China pushed its own message. lowcost open models could reduce the global technology gap. So now China has Kim K3 to point to. It's no longer promising that its open AI ecosystem might become competitive one
14:05
Speaker A
day. It's released a model that's already beating some of America's biggest names, costs less to use, and will soon be available for companies to run themselves. All right, that's it for this one. Let me know what you think in
14:17
Speaker A
the comments. Thanks for watching, and I'll catch you in the next one.
Topics:Kimmy K3Moonshot AIopen AI modelChina AIAI competitionlarge language modelAI developmentmultimodal AIAI pricingAI industry impact

Frequently Asked Questions

What makes Kimmy K3 different from previous AI models?

Kimmy K3 is the largest open-weight AI model with 2.8 trillion parameters, designed for long-running tasks and multimodal inputs, and it rivals top US closed models in performance.

How does Kimmy K3 compare to American AI models?

Kimmy K3 competes closely with models like Anthropic's Fable 5 and GPT 5.6 Soul, outperforming some in coding tasks and scoring similarly on intelligence indexes.

Who can run Kimmy K3 and what are the hardware requirements?

Kimmy K3 requires substantial hardware, recommended to run on systems with at least 64 AI accelerators, making it accessible mainly to large companies and research institutions.

Get More with the Söz AI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →