Speaker A
China may have just created its biggest AI shock since DeepSeek, but this time the story is much larger than one powerful new AI model. On the same day that Moonshot AI unveiled a giant new model called Kimmy K3, Xiinene Ping stood on a stage in Shanghai and told the world that China doesn't plan to follow America's rules for artificial intelligence. It wants to help write the rules, spread its own models across the developing world, [music] and build a new AI order with China at the center. And Kimmy K3 gives that speech real weight because this isn't another model that looks impressive only inside a company presentation. Now, K3 has 2.8 trillion parameters, making it the largest open-weight AI model ever announced. Moonshot says it's the first open model to approach the 3 trillion mark and its full weights are expected to be released on July 27th. Companies and researchers will then be able to download it, modify it, and run it on their own infrastructure instead of being permanently locked inside Moonshot's app. But the size isn't why this release is sending shock waves through the industry. And the real reason is that Kimmy K3 appears to be dangerously close to the best American closed models. And in some tasks, it's already beating them. Moonshot says K3 can compete [music] with Anthropic's Fable 5 and outperform Claude Opus 4.8, GPT 5.5, and parts of GPT 5.6 in demanding coding work. And independent testing actually tells a similar story. Artificial analysis gave K3 a score of 57 on its intelligence index. Claude Opus 4.8 scored around 56, GPT 5.6, Six, Terra scored 55, and Gemini 3.1 Pro landed at about the same level as K3. So only Claude Fable 5 and GPT 5.6 Soul remained ahead and even there the gap [music] was only two or three points. For years, the standard belief was that open models were 6 to 12 months behind the best closed systems. More recently, people started saying Chinese labs were perhaps 3 to 6 months behind. One Reddit commenter described the new reality more bluntly. Maybe they're six days behind. Now, the most dramatic result came from arena.i's front-end code arena, where developers ask models to build real websites and interfaces, then vote on which result is better. Kimmy K3 entered at number one with 1,679 points. Claude Fable 5 scored 1,631. GPT 5.6. Saul scored 1,618. K3 finished first in six of the seven areas tested, [music] including brand and marketing work and data dashboards. Moonshot's previous model, Kimmy K 2.6, had been sitting in 18th place. K3 jumped 17 positions in one generation and went straight to the top. Even Arena's CEO called it potentially the biggest AI release of the year and suggested it could mark the moment China moved ahead of the United States in at least one major part of the model race. Now, Moonshot isn't pretending K3 wins everything. The company admits it still trails Fable 5 and GPT 5.6 Saul in overall user experience [music] and some broader tasks. But K3 doesn't need to dominate every category to be disruptive. It only needs to make companies ask an uncomfortable question. Why keep [music] paying premium prices to American providers if an open alternative is almost as capable? So, Moonshot built K3 mainly for long-running work, especially software development. And this isn't a model designed only to answer one prompt and stop. It's meant to inspect a large codebase, create a plan, use tools, [music] make changes, check whether they worked, and continue for hours with limited human help. It can also look at what appears on the screen. Moonshot calls this vision in the loop. The model writes code, looks at the result, notices what's wrong, changes the code, and checks again. That makes it useful for websites, games, design tools, animation, and other projects where [music] it needs to see whether its own work actually looks right. And since we're already talking about how businesses are using AI, here's something useful. Most of you voted for seven AI agents businesses are paying $3,000 to $10,000 for right now in yesterday's poll. So, we're giving it away for free. It breaks down the exact agents [music] businesses are buying, what they cost, who needs them, and how to spot the right clients. It also covers the ROI, setup fees, retainers, [music] and what a realistic $10,000 per month client list could look like with zero coding required. You can grab it through the link in the video description. Now, back to Kimmy K3. In one demonstration, K3 built a 3D open world game inside a browser using 3.js, JS WebGPU, and GPU [music] compute. It generated the environment and used outside tools to create a rider and horse. Moonshot also showed it building a simulation of China's Long March 10 rocket [music] and a Game Boy Advance emulator. And the company says K3 spent 15 hours improving GPU code and cut the required compute time by more than half. [music] It created a small GPU compiler called Mini Triton from scratch and reportedly designed a working chip over 48 hours using open-source engineering tools. Moonshot also says K3 reproduced an astrophysics analysis in about 2 hours. It reviewed more than 20 papers and wrote over 3,000 lines of Python code. According to the company, an experienced team might normally spend between 1 and 2 weeks on similar work. Now, these are company demonstrations, so they still need wider testing, but they show the goal clearly. K3 is supposed to stay on a difficult task for an entire afternoon, an entire night, or even several days. It can process up to 1 million tokens in one session, roughly 750,000 words. That means it can take in huge code bases, long documents, research papers, or project histories without immediately losing track of what came earlier. It also handles text, [music] images, and video inside the same model. Moonshot says K3 even edited its own promotional video from 56 clips. Its Kimmy work platform is adding interactive widgets and dashboards so users can build persistent workspaces [music] rather than using the model only through a normal chat box. The model uses a mixture of experts design and the simple explanation is that it contains hundreds of specialized sections, but it doesn't activate all of them for every question. K3 has 896 experts with only 16 active at a time. That lets Moonshot build a 2.8 trillion parameter system without using the entire model for every word it generates. Moonshot also created Kimmy Delta attention and attention residuals, two systems designed to help K3 [music] keep track of information through very long tasks. Combined with new training methods, the company claims they make K3 about 2.5 times more efficient at scaling than Kimmy K2. Now, the model was trained with MXFP4 weights and MXFP8 [music] activations, lower precision formats intended to reduce the hardware burden. That doesn't mean anyone will be running it on a normal laptop. And Moonshot recommends systems with at least 64 AI accelerators. The local LLaMA community immediately started joking that all you need is 2 terabytes of VRAM, several Mac Studios, a pile of storage drives, and enormous patience. [music] So, that's the strange reality of K3. It's open, but it isn't small. Most people will never host it at home. The companies that can host it, however, [music] are exactly the ones that matter. Large firms already spend millions every month on AI from OpenAI and Anthropic. If they can run K3 on their own servers, customize it, and keep their data private, [music] this becomes a serious threat to closed AI providers, and the price makes that threat stronger. K3 costs $3 per million input tokens when the input isn't cached, 30 cents when it's cached, and $15 per million output [music] tokens, including reasoning. Those prices stay the same even with long context. Fable 5 costs around $10 per million input tokens and $50 per million output tokens. GPT 5.6 Soul costs about 50 cents per million input tokens and $30 per million output tokens. So K3 is around