A detailed comparison of Opus 5.5, Sonnet 5.5, Astra 6, and Sol 6.1 AI models in game development tasks including 3D modeling, sound generation, and Unity game assembly.
Key Takeaways
- Opus 5.5 leads in quality and completeness of game development tasks but at a higher cost.
- Sonnet 5.5 offers a balance between cost and quality but lacks some advanced features like Unity scene setup.
- GPT-6 Astra and GPT-6.1 Sol are cheaper alternatives but produce less polished 3D models and have some functional limitations.
- Procedural generation via AI code is viable for creating complete game assets without external resources.
- Detailed cost and resource consumption analysis is crucial when choosing AI models for game development.
What the video covers
- The video tests four AI models—Opus 5.5, Sonnet 5.5, GPT-6 Astra, and GPT-6.1 Sol—across five game development tasks.
- Tasks include creating 3D robot models in Blender, generating procedural game sounds and music, animating characters, building levels, and assembling a full game in Unity.
- Opus 5.5 produced the best 3D models with full destructibility and even prepared a Unity scene for detailed examination and interaction.
- Sonnet 5.5 delivered slightly lower quality models and did not prepare Unity scenes but was cheaper than Opus 5.5.
- GPT-6 Astra and GPT-6.1 Sol showed a plasticky style in modeling, with Sol being significantly cheaper but with some functionality issues.
- All assets, sounds, and models were generated procedurally via code without using any downloaded or pre-made resources.
- Opus 5.5 uniquely provided detailed sound generation references and comparisons between source sounds and AI-generated results.
- The video includes cost, token usage, and time statistics for each model’s performance on the tasks.
- The game concept is a Battle Royale with destructible combat robots using lasers and rockets in an arena environment.
- The presenter emphasizes the maximum reasoning level was used for all models to ensure a fair comparison.
Full Transcript — Download SRT & Markdown
Speaker A
Today, we will test the newest and top-tier models currently available. We will perform five very different game development tests and finish by building a game in Unity. It’s going to be interesting. Hello everyone, and welcome to the Game Studio channel. My name is Igor, and today we are going to make a game. Two new models were released just this week. They are Sonnet 3.5 from Anthropic and its direct price-point competitor from OpenAI, GPT-4o. And today, of course, we are going to test them. To make things interesting, we will also test Claude 3 Opus. GPT-4o. In total, we will be doing five very different tests today. The neural networks will be making 3D models in Blender, generating music and sounds for the game, animating our characters, building a level based on a reference, and finally, each network will assemble a full game in Unity. Now, a bit about the game that each neural network will be creating today. I have prepared this concept. In the image, you can see that we control a robot and fight other robots on a level that looks like an arena. So, it will be a Battle Royale featuring combat robots. To make it even more interesting, we will make the environment fully destructible, along with the robots themselves. We will destroy them using several types of weapons: lasers and rockets. I have prepared detailed descriptions of all the tests that each neural network must complete, and I’ve created five files which I placed in the project folder. You will be able to see the text for each test on the screen. You can pause the video if you want to read them in more detail. The first test is to create 3D models for all three robot types: large and heavy, medium, and small and fast. In addition to this, it is necessary to immediately account for full robot destructibility—well, as much as possible—and prepare the models for import into Unity. We will work with Anthropic models in Claude, and with OpenAI models in ChatGPT. I will be using the maximum reasoning level for all models participating in our test today. Let’s move on to the results of the first test. And we will start with the result from Opus 3.5. Right now on the screen, you can see the image of the robots. On the left is the reference, and on the right is the result that Opus 3.5 created in Blender. Getting ahead of myself, I will say that this is the best result in this test among all the models we will be testing today. It is also worth noting that Opus didn't stop there; it additionally created a Unity project and set up a scene where all three robots stand in a row so they can be examined in more detail. This is how the large robot looks, the one that is the biggest and heaviest. This is what the medium robot looks like, and finally, the smaller robot. Furthermore, a separate mode is available in this scene where you can shoot at these robots and watch them get destroyed. This is also done quite nicely. You can immediately see that Opus really tried hard to complete this task. Perhaps some might think it did too much or things it wasn't really asked to do. But if you read the task it was given in detail, there is a line like this. You can read other tasks in the folder for context to understand what we will do next. So, it turns out that Opus 5.5 is simply playing ahead of the curve. Next, let’s look at the result from Sonnet 5.5. Now on the screen, you can also see the image of the robots, references, and the results that Sonnet modeled in Blender. It is worth noting that Sonnet did not prepare the scene in Unity. Nevertheless, there are renders like these where all three robot models stand in a row. And there are separate renders showing the destruction of these robots. If you look closely, the result seems a bit worse than what Opus did. This is especially noticeable on the middle robot model. Its legs look like they are slightly twisted at the knees. Although I can say that, overall, the result is quite similar to what Opus 5.5 did. If we compare prices, Sonnet 5.5 is cheaper than Opus, but the quality of work is lower, and the volume of work performed is also much less. As soon as we look at all four models, I will show how much it cost in terms of AI expenses. And at the very end of the video, we will see how many limits were consumed for each model to complete all these five tests and build the game. And now let’s take a look at the result from GPT-6 Astra. Now you can also see the references on the left and the modeling results from GPT in Blender on the right. I can’t say it’s very bad, but I wouldn't call this result very good either. Overall, it seems that other models not only handle Blender modeling better but also maintain a consistent style better. Whatever I model with GPT-6 Astra in Blender, there’s a certain plasticky feel. I don’t know how to put it into words. So, you can see the modeling results for yourselves. Astra did not prepare the scene in Unity, nor did it prepare any renders or scenes where one could see the destructibility of the robots themselves. In terms of money, it turned out a little cheaper than what Sonnet did. And now let’s see what GPT-6.1 Soul prepared for us. According to OpenAI's claims, GPT-6.1 Soul is five times cheaper than GPT-6 Astra, but in quality, it's almost at the Astra level. On the whole, one could probably say that’s true. You can see the results for yourselves on the screen right now. And indeed, it all looks similar to what GPT-6 Astra made for us. It also has that certain plasticky feel. The general GPT style is recognizable, but additionally, GPT 6.1 Sol prepared a web page for us with a preview of all three robots, where you can look around and rotate the result, meaning you can view each robot from the front, back, and take them apart piece by piece. Besides this, it was intended that one could also bombard these robots here to see their destructibility. But, unfortunately, this function did not work properly. And as for the money, it turned out not five times cheaper than GPT-6 Astra, but about 15 times cheaper. And right now on the screen, you can see the statistics for the first test: how much time and how many tokens were spent on this task for each model, and how much it costs based on AI pricing. And I will show how many weekly subscription limits I spent at the end of this video. Another important point worth mentioning: everything you will see today was generated procedurally, without using any assets, music, sounds, textures, ready-made models, or anything else downloaded from the internet. That is, literally everything the models did, they did using code. But we continue and move on to the next test. And the next test consists of creating sounds and music for the game. In the task description, I provided a hint. It is necessary to find references for game sounds and music, and specifically what can be used to generate these sounds. For example, noise reduction for servos, wave dispersion for lasers, modal synthesis for metal parts, and so on. Literally with a description of mathematical formulas. The number of sounds is left to the discretion of the model itself. But for the music, we need to generate two tracks. One for the main menu of the game and one track for combat. All of this the models will do procedurally via code. And now let's listen to the result. And first, let's see what Opus 5.5 made for us. Like the other models, it prepared a preview website like this for us. And here you can listen to the sounds of footsteps, servos, lasers, rockets, explosions, and so on, as well as the music. The most interesting thing that I liked, which Opus has and other models don't, is that for each of these sound categories, it offers a reference to listen to on the left, and on the right, it lets you listen to the result it generated with code. Besides this, it shows exactly which references from the internet were used, and exactly how it reproduced this sound wave. I want to note once again that no samples were used here.
Speaker A
Igor, and today we are going to make a game. Two new models were released just this week. They are Sonnet 3.5 from Anthropic and its direct price-point competitor from OpenAI, GPT-4o. And today, of course, we are going to test
Speaker A
them. To make things interesting, we will also test Claude 3 Opus. GPT-4o. In total, we will be doing five very different tests today. The neural networks will be making 3D models in Blender, generating music and sounds for the game, animating our characters,
Speaker A
building a level based on a reference, and finally, each network will assemble a full game in Unity. Now, a bit about the game that each neural network will be creating today. I have prepared this concept. In the image, you can see that
Speaker A
we control a robot and fight other robots on a level that looks like an arena. So, it will be a Battle Royale featuring combat robots. To make it even more interesting, we will make the environment fully destructible, along
Speaker A
with the robots themselves. We will destroy them using several types of weapons: lasers and rockets. I have prepared detailed descriptions of all the tests that each neural network must complete, and I’ve created five files which I placed in the project folder.
Speaker A
You will be able to see the text for each test on the screen. You can pause the video if you want to read them in more detail. The first test is to create 3D models for all three robot
Speaker A
types: large and heavy, medium, and small and fast. In addition to this, it is necessary to immediately account for full robot destructibility—well, as much as possible—and prepare the models for import into Unity. We will work with Anthropic models in Claude,
Speaker A
and with OpenAI models in ChatGPT. I will be using the maximum reasoning level for all models participating in our test today. Let’s move on to the results of the first test. And we will start with the result from Opus 3.5.
Speaker A
Right now on the screen, you can see the image of the robots. On the left is the reference, and on the right is the result that Opus 3.5 created in Blender . Getting ahead of myself, I will say
Speaker A
that this is the best result in this test among all the models we will be testing today. It is also worth noting that Opus didn't stop there; it additionally created a Unity project and set up a scene where all three
Speaker A
robots stand in a row so they can be examined in more detail. This is how the large robot looks, the one that is the biggest and heaviest. This is what the medium robot looks like, and finally, the smaller robot. Furthermore
Speaker A
, a separate mode is available in this scene where you can shoot at these robots and watch them get destroyed.
Speaker A
This is also done quite nicely. You can immediately see that Opus really tried hard to complete this task. Perhaps some might think it did too much or things it wasn't really asked to do.
Speaker A
But if you read the task it was given in detail, there is a line like this.
Speaker A
You can read other tasks in the folder for context to understand what we will do next. So, it turns out that Opus 5.5 is simply playing ahead of the curve.
Speaker A
Next, let’s look at the result from Sonnet 5.5. Now on the screen, you can also see the image of the robots, references, and the results that Sonnet modeled in Blender. It is worth noting that Sonnet did not prepare the scene
Speaker A
in Unity. Nevertheless, there are renders like these where all three robot models stand in a row. And there are separate renders showing the destruction of these robots. If you look closely, the result seems a bit worse than what Opus did. This is
Speaker A
especially noticeable on the middle robot model. Its legs look like they are slightly twisted at the knees.
Speaker A
Although I can say that, overall, the result is quite similar to what Opus 5.5 did. If we compare prices, Sonnet 5.5 is cheaper than. Opus, but the quality of work is lower, and the volume of work performed is also much
Speaker A
less. As soon as we look at all four models, I will show how much it cost in terms of AI expenses. And at the very end of the video, we will see how many limits were consumed for each model to
Speaker A
complete all these five tests and build the game. And now let’s take a look at the result from GPT-6 Astra. Now you can also see the references on the left and the modeling results from GPT in Blender on the right. I can’t say
Speaker A
it’s very bad, but I wouldn't call this result very good either. Overall, it seems that other models not only handle Blender modeling better but also maintain a consistent style better.
Speaker A
Whatever I model with GPT-6 Astra in Blender, there’s a certain plasticky feel. I don’t know how to put it into words. So, you can see the modeling results for yourselves. Astra did not prepare the scene in Unity, nor did it
Speaker A
prepare any renders or scenes where one could see the destructibility of the robots themselves. In terms of money, it turned out a little cheaper than what Sonnet did. And now let’s see what GPT-6.1 Soul prepared for us.
Speaker A
According to OpenAI's claims, GPT-6.1 Soul is five times cheaper than GPT-6 Astra, but in quality, it's almost at the Astra level. On the whole, one could probably say that’s true. You can see the results for yourselves on the screen right now. And indeed, it
Speaker A
all looks similar to what GPT-6 Astra made for us. It also has that certain plasticky feel. The general GPT style is recognizable, but additionally, GPT 6.1 Sol prepared a web page for us with a preview of all three robots, where
Speaker A
you can look around and rotate the result, meaning you can view each robot from the front, back, and take them apart piece by piece. Besides this, it was intended that one could also bombard these robots here to see their
Speaker A
destructibility. But, unfortunately, this function did not work properly. And as for the money, it turned out not five times cheaper than GPT6 Astra, but about 15 times cheaper. And right now on the screen, you can see the statistics for the first test: how much
Speaker A
time and how many tokens were spent on this task for each model, and how much it costs based on AI pricing. And I will show how many weekly subscription limits I spent at the end of this video . Another important point worth
Speaker A
mentioning. Everything you will see today was generated procedurally, without using any assets, music, sounds , textures, ready-made models, or anything else downloaded from the internet. That is, literally everything the models did, they did using code.
Speaker A
But we continue and move on to the next test. And the next test consists of creating sounds and music for the game.
Speaker A
In the task description, I provided a hint. It is necessary to find references for game sounds and music, and specifically what can be used to generate these sounds. For example, noise reduction for servos, wave dispersion for lasers, modal synthesis
Speaker A
for metal parts, and so on. Literally with a description of mathematical formulas. The number of sounds is left to the discretion of the model itself.
Speaker A
But for the music, we need to generate two tracks. One for the main menu of the game and one track for combat. All of this the models will do procedurally via code. And now let's listen to the result. And first, let's see what OPUS
Speaker A
5.5 made for us. Like the other models, it prepared a preview website like this for us. And here you can listen to the sounds of footsteps, servos, lasers, rockets, explosions, and so on, as well as the music. The most interesting
Speaker A
thing that I liked, which Opus has and other models don't, is that for each of these sound categories, it offers a reference to listen to on the left, and on the right, it lets you listen to the result it generated with code. Besides
Speaker A
this, it shows exactly which references from the internet were used, and exactly how it reproduced this sound wave. I want to note once again that no samples were used here. All the models made the sounds procedurally via code.
Speaker A
I can't say whether I like or dislike this result. Some sounds turned out to be really quite good. For example, the sounds of explosions and rockets. Well, probably all the models produced decent ones. But for Opus, for example, I also
Speaker A
like the hydraulics sounds. That turned out pretty cool too. But unfortunately, some sounds are unsuccessful. They appear in Opus as well as in other models. But what we can definitely evaluate is the quality of the composed music. And now, let's listen to the
Speaker A
music for the main menu. And this is what the battle music sounds like. Write in the comments if you liked it or not, and which model's music you liked the most. Now let's look at the result from Sonnet 5.5, or rather,
Speaker A
listen to it. Visually, the webpage is designed in the same style as the one made by OPUS 5.5. Don't think the models were peeking at each other. I performed these tests sequentially, not in parallel. And before running each
Speaker A
new test for each new model, I completely deleted the project folder from the previous model and cleared the caches. So, now we can also listen to the sounds that Sonnet 5.5 prepared for us. Overall, it looks and sounds very
Speaker A
similar. And probably the only thing we can really evaluate is the quality of the composed music. And let's listen to it. This is how the main menu music sounds. And I would like to specifically note the music composed
Speaker A
for the battle. I would probably set that as my ringtone or alarm clock. Now let's listen to the result from GPT-6 Astra. The sound of a heavy machine.
Speaker A
The model also prepared a preview webpage for us in its own style. You can listen to different sounds. There is a description of how these sounds were generated, as well as formulas and references. And now let's listen to the
Speaker A
music. And this is what the main menu music sounds like. That reminds me of something very strongly. And this is what the battle music sounds like. And let's see what GPT-6.1 Soul did for us.
Speaker A
And Soul also prepared a worthy result for us that doesn't look like the results of other models. Hear the weight of the machine. What I liked here was this "battle scene" button.
Speaker A
When you click it, you can hear a composition of sounds and music, showing roughly how the game will sound . In my opinion, that's very cool. And further, similar to how GPT-6 Astra did it, you can listen to the sounds. There
Speaker A
is a description of how these sounds were created, which formulas were used, and which synthesis method. And now let's listen to the music from GPT-6.1 Soul. And this is how the main menu theme sounds. And this is how the
Speaker A
battle theme sounds. I cannot single out any one winner here. Overall, everyone did a pretty good job. As I said, everyone managed to make the sounds of explosions and rocket fire quite well. But everything else is a matter of taste. Write in the comments
Speaker A
which result you liked the most. Right now on the screen, you can see the token count, time, and money spent on this task for each model, calculated based on the AI cost. We are continuing , so let's move on to the next test. If
Speaker A
you like what I do, you can support the channel on Boosty. You can subscribe or make a one-time donation. You can see the QR codes on the screen. This really helps the channel grow. The next test is quite complex and extensive. The
Speaker A
models will need to create a level for a game in Unity, as close as possible to the provided reference. All models for the level will also need to be created in Blender. To do this, we will need to bring in sub-agents, as with
Speaker A
all other tasks. You can also see the full text of the test on the screen right now. As a result, the neural networks need to produce a video demonstrating the level in Unity, which they will record directly in the engine
Speaker A
. And we will watch the first level video from Opus 3.5. I want to note that Opus shows not just the level, but also immediately places mechs in it that shoot at houses and various objects, showing some interactivity. So
Speaker A
you can immediately see the level's destructibility and how it might look in the game later. Overall, it's done quite well. I can't say that Opus 3.5 made it look exactly like the reference , but the result was truly decent for a
Speaker A
single run. Let's watch a bit of this video. And now, let's look at the result from Sonnet 3.5. This is probably one of those tests where Sonnet loses to Opus both in the quality of the final result and in cost
Speaker A
. Sonnet tried very hard and did everything in its own unique style, although, in my opinion, it turned out a bit too acidic. The level also has mechs shooting, and Sonnet invested heavily in special effects. When the robot fires missiles, you can barely
Speaker A
see anything through the clouds of smoke, although the effects themselves don't look all that impressive. Sonnet spent 270 dollars on this task in terms of API costs, compared to 190 for Opus.
Speaker A
And now, let's watch a little bit of the video result from Sonnet. And now let's take a look at what GPT-4o has prepared for us. Here is the video of our level result, how it will look in Unity according to GPT-4o. And the
Speaker A
result turned out so bad, in my opinion , that I asked it to improve it. I literally asked the model to work for a while longer and do better, but the second iteration wasn't much better.
Speaker A
And you can see the final result of the level on the screen right now, how it will look in the game. The model spent about 110 dollars on all of this. In my opinion, something like Claude or Qwen
Speaker A
would have handled it better. But we aren't giving up, so let's take a look at the results from GPT 6.1 Sol. And I have to say, this result is truly on par with GPT6 Astra and maybe even slightly better than what GPT6 Astra
Speaker A
produced. And, by the way, it's not five times as expensive. Again, GPT6 Astra did this level for 110 dollars, while GPT 6.1 Sol completed the same task for 2.5 dollars. I don't even know how. I just can't fathom how GPT 6.1
Speaker A
Sol manages to operate so cheaply. You can see the results on the screen, but it definitely didn't turn out much worse than GPT6 Astra. By the way, on this task, I stopped tracking how many of my weekly limits I was spending when
Speaker A
working with GPT 6.1 Sol. And I think for completing all the tasks and generating the entire game from scratch , I used, well, maybe one, at most 2%of my weekly limits. Most likely one, because after this task I turned on the
Speaker A
high-speed mode, that "fast mode" in the IDE, which is one and a half times faster and uses significantly more limits, but I still didn't see any noticeable consumption. And now on the screen, you can see detailed statistics for the third test. How much time was
Speaker A
spent, how much money was spent calculated by API costs, and how many tokens all the neural networks used to complete this task. And we are moving on to the next test. And now for the next test. Now our neural networks will
Speaker A
create animations for our robots. And we will use an approach called inverse kinematics. The essence of this approach is that, first, the animation is calculated procedurally in the code, and second, it is calculated from the foot, not from the hip. That is, the
Speaker A
robot first places its foot on something, on the ground or maybe on an obstacle, and then, from that point and position, the angle of all other joints is calculated. And this way, the foot lands precisely on a piece of debris or
Speaker A
perhaps a slope, and the knee bends exactly as it should. We will also make this a small part of the gameplay, where you can destroy objects by stepping on them or crushing cars or something else. You can see the task on
Speaker A
the screen, as usual. The essence is simple. Create inverse kinematics for me in Unity. Also, implement WASD movement and have the body rotate to follow the mouse cursor. And at the end , I want to be able to walk around as
Speaker A
the robot in the Unity scene. First, of course, we look at the result from Claude Opus 5. I launch the project in Unity and start our game. And I immediately see that everything is done more than well. So, the robot walks,
Speaker A
moving its legs properly depending on whether it’s moving fast or slow, starting to walk, or perhaps stopping; it has some kind of acceleration or deceleration as it moves. That also looks cool. I also want to note that OPC is the only neural network that did
Speaker A
the controls properly. I mean, the camera and the robot's body really turn independently of how I’m moving, using the mouse cursor. And that is very convenient. There will be various other options. Today, of course, we’ll take a look at all of them.
Speaker A
I’ll talk about that later. Besides movement, in principle, one could probably say that OPUS 5.5, as usual, jumped ahead and did almost everything.
Speaker A
You can walk, you can shoot, you can use a laser, you can fire rockets at houses. The houses are destroyed.
Speaker A
Overall, it all looks almost like a game already. I even got a bit hooked on this whole thing for a while. But let's not get distracted; let's look at how the inverse kinematics were implemented. And yes, indeed, the
Speaker A
robot's leg steps onto the cars. The cars get crushed, and the legs also handle various types of obstacles and slopes properly. And in general, it all looks very, very good. And today, this is probably the best result we will see
Speaker A
. And next in line is Sonet 5. I’m launching the scene in Unity. It all looks something like this. There is already a menu where you can choose one of the scenes. We choose the city to walk around with the mech. And I want
Speaker A
to note the controls right away. Unlike Opus, all the other neural networks added camera rotation on the Q and E keys, even though it wasn't in the original prompt regarding how to handle controls. Literally, in the task, I
Speaker A
wrote: "I want to move with WASD and rotate the camera with the mouse." That's it. Nevertheless, Sonet, as well as Astra and Sol, for some reason, added camera rotation to Q and E, with aiming on the mouse, and that is not
Speaker A
super convenient. But let's look at the result. There are both pros and cons here. Among the cons, I want to note something that has been trailing along since the very first task. These are the crooked legs on the robots. They
Speaker A
look as if they are pointing in different directions. And the leg movement, the animation itself, wasn’t done, well, not exactly very well. They move somehow, well, not as naturally, let’s say, as Opus did, but as if they are sticking a bit to
Speaker A
the ground. Sonet also allows you to shoot. Use a laser and rockets to destroy houses and trample surrounding obstacles. And it all looks like some kind of bacchanalia of explosions, smoke, splashes, fountains, and dust.
Speaker A
And in general, it is very difficult to make out what is happening on the screen. But we are interested in inverse kinematics. Of course, Sanet can also be said to have coped with this task. The legs step onto
Speaker A
elevations and various obstacles. It's okay. This result can be counted as successful, but I can't say that I really like how it was all implemented.
Speaker A
And now let's see what GPT6 Astra has prepared for us. Well, Astra prepared a scene in Unity for us where you can walk with a mech. I asked to fix these awkward controls, remove the camera rotation on the Q and E keys, and use
Speaker A
the mouse instead. And it probably only got worse, because the crosshair was moved somewhere behind the mech's back.
Speaker A
But overall, it's fine; I didn't ask to change anything else in the end and left it as is. We can evaluate how the walking works, how the animation works, and how the kinematics work. You can step on cars, and you can also step on
Speaker A
buses like this. Nothing happens to the buses, of course. And similarly, at the end, I stepped on a fountain and ended up getting stuck. But overall, the kinematics were implemented. We consider that Astra successfully handled the task. Nevertheless, not
Speaker A
without bugs. That is, there are indeed moments that literally ruin the entire gameplay. You can simply get stuck in a fountain. And let's look now at what GPT 6.1 Soul did for us. GPT 6.1 Soul gave us a result five times cheaper
Speaker A
than GPT6 Astra. And, in my opinion, even better. You can see how the mech moves its legs. It can actually step on cars, although, unfortunately, not all leg positions are processed correctly.
Speaker A
That is, sometimes when stepping on a car, the mech flies up into the air.
Speaker A
This, of course, should not be happening. That is, one leg should remain on the ground, and the other leg should remain on the car. Nevertheless, GPT 6.1 Soul had fewer bugs, and I never managed to get stuck in the
Speaker A
fountain. And now on the screen, you can see detailed statistics for the fourth test. How many tokens were spent , how much money, and how much time.
Speaker A
But we continue. My channel recently started a Discord community. You can see the QR code on the screen. The link will also be in the description. There, you can discuss neural networks, game coding, and various tools and approaches. Join us, it's interesting
Speaker A
there. And we are moving on to the next test. And the next test is the final one. Now each neural network will assemble a full-fledged game in Unity using everything it did in the previous tests. This will be a mech arena in
Speaker A
Battle Royale mode. If you want to read the full text of this test, it is available on the screen right now. And we will begin testing in reverse order and see what GPT 6.1 S made for us. I
Speaker A
am launching the game in Unity. And we are greeted by this starting menu. You can choose a mech here. And I will start playing with the heaviest one, which is called Bastion. Next, we have a 3-second countdown. And after that,
Speaker A
the battle begins. As I already said, aiming is quite difficult. I lose almost immediately. I tried all sorts of different tactics. The most effective one is simply to stand aside and wait for all the robots to shoot each other. And that way, you can take
Speaker A
second place. I also tried rushing into battle, taking third or fourth place, or looking for obstacles to hide behind buildings. And, perhaps, the best result can be seen in my last run, when I switched to a medium mech, which is a
Speaker A
bit faster, although it has less health and less armor, but I managed to fully participate in this battle and even win . So that is the kind of game GPT 6.1 Soul produced. It is very far from what
Speaker A
was in the reference, but it is very cheap. I mean, even including the cost of the AI, the price of this, let's say , prototype is very small. And now let's see what GPT6 Astra made for us.
Speaker A
Next, we test the result from GPT6 Astra. Similarly, I start the game in Unity. And in the starting menu, you can also choose one of three mechs. I choose the heaviest one, and the battle begins. What can I say about the game
Speaker A
from GPT6 Astra? No matter how many times I tried to play this game with various mechs and tactics, I always died first. Well, perhaps with rare exceptions. It seems to me that the Battle Royale wasn't really implemented , because as soon as I appeared on the
Speaker A
level, it didn't matter if I was in the field of view of other mechs or not, they all started attacking me. In the end, only at the very end, I think with the medium mech, I managed to take
Speaker A
second place. And that is probably only because this tactic, where you hide to the side while they fight each other, is the most effective one. But as soon as you enter their line of sight, the enemies always switch to you. And let's
Speaker A
not forget that, on top of everything else, you can get stuck in a fountain.
Speaker A
So that is the result, that is the game that GPT6 Astra produced. And next up is Sonet 5.5. I start the game in Unity , and we are greeted by this neon-colored start screen. Here you can also choose a robot. I choose the
Speaker A
heaviest one. And we start the battle from the roof of a building. Similarly, there is a countdown timer before the fight begins, and then some kind of chaos starts. Rocket fire, smoke, lasers, splashes, dirt, dust—you can't see anything at all. I tried
Speaker A
playing this for a while. You can see it for yourselves on the screen. It’s unclear what is happening. Plus, the controls are really clunky. In general, in my opinion, SNEET went way too far.
Speaker A
It also spent way more on limits than OPUS did. But more on that a little later. For now, I suggest you just take a look at the result for a bit. And then we will move on to what OPUS 5.5
Speaker A
has done. To stay updated with the latest news and announcements, subscribe to my Telegram channel. You can see the QR code on the screen, and the link will be in the description.
Speaker A
Well, let’s continue. So, let’s look at the result from OPUS 5.5. And here, in my opinion, it turned out, well, really quite good. Starting from the main menu, we can change a large number of settings and individually select a robot in the hangar to play
Speaker A
with. By the way, the previews during the robot selection are also made nicely. The camera shifts to the selected robot, and they fire into the air. Then we start the battle. And the first thing we need to do is choose a
Speaker A
landing point. We tap a random spot on the map, and our robot arrives in the city from above. And that already looks quite spectacular. The graphics are implemented at a very decent level.
Speaker A
This result reminds me more of the original concept than any of the others . And I remind you that we did this, well, literally in one shot. If we give this process a little more control, we can get an even better result. Moving
Speaker A
on to the effects and controls, everything is done much better than with any other participant. There is a moderate amount of effects, and the controls are convenient. It was actually fun to play. Just like in the previous test, I got a bit hooked while
Speaker A
testing the movement. And here, too, I spent several gaming sessions. And in the end, I even managed to win. And it was a fair fight, not just standing aside behind an obstacle and shooting back. Every time, I had to think
Speaker A
through the landing spot to roughly plan where the other robots would land. And then, before the zone closes in, you need to manage to defeat all the other enemies on the level. Besides that, the enemies fight each other, and
Speaker A
it's not like what happened, for example, with GPT6 Astra. As soon as I get into the enemies 'line of sight, they don't just start attacking me. No, there really is a battle going on here, and there's a real sense of, well, some
Speaker A
kind of living process to a certain extent, as if you're playing with real players. Once again, I want to note that all of this looks very cool. I really like the effects. The graphics turned out fun and are quite
Speaker A
well-developed. And I suggest we just take a look at the result that Opus 3.5 created. And now on the screen, you can see detailed statistics on how much time, money, and tokens were spent on this final fifth task. And now you can
Speaker A
see the statistics on the total time, money, and tokens spent on the entire game by all participants. You can pause the video to take a closer look. But I didn't spend all this money via API on creating these four games. I have two
Speaker A
subscriptions for 200 dollars each. One for Claude, the other for GPT. And how many of my weekly limits did I use up?
Speaker A
Claude Opus 3. Sonnet 3.5 used 19%of the weekly limit. GPT-4o Astra used 13% of the weekly limit, and GPT-4o mini used 1%. Well, maybe a maximum of two.
Speaker A
At some point, I stopped counting because I switched to Fast Mode. And all this on a 200-dollar subscription.
Speaker A
If you want to see the numbers converted for other types of subscriptions, they will also be shown on the screen now. You can pause and see what this looks like in limits for the 100-dollar and 20-dollar plans.
Speaker A
There is one important nuance. Anthropic does not disclose how many weekly limits they actually provide for 200-dollar subscriptions relative to other plans. But there is official documentation they published a year ago , where they wrote that the limits on
Speaker A
the 200-dollar subscription are 1.7 times higher than on the 100-dollar subscription. Then they stopped publishing that documentation.
Speaker A
Independent measurements today confirm this information, that the limits have remained at approximately the same level. What conclusions can be drawn from this experiment? If you just want to one-shot a game and make it look as close as possible to the references,
Speaker A
documentation, and description you provided to the AI, use Opus 3.5. If you consider yourself an AI tuning expert and can find a Sonnet configuration where, for example, it works faster and more efficiently on " high" than Opus on "medium," you can do
Speaker A
that. But the result will not always be predictable. If you are using GPT-4o Astra, keep using it. It is a good model. It handles tasks quite decently.
Speaker A
Some of you might notice, well, like, it worked faster or it spent less money via API. If it had spent the same amount of money via API—well, in terms of API costs, I mean—then the result would 100%be on par with Opus or
Speaker A
Sonnet. I have to disagree with you there, because you saw the weekly limit usage yourself. If we make Astra work four times as much, or two or three times as much, the amount of the weekly limit it uses up on the $ 200
Speaker A
subscription will also be higher. Than the limits consumed by the ORC. Therefore, in my view, it is not very cost-effective. And now, GPT 6.1 Sol.
Speaker A
What can be said about it? Overall, it can be used for everything. For everything when you need to get results quickly and cheaply. It is an excellent model, and as OpenAI calls it, a workhorse. The results were indeed
Speaker A
comparable to the results shown by GPT 6 Astra. And perhaps Sol has gotten smarter, or maybe Astra has gotten dumber. But today, I would rather use Sol on a regular basis. And that is all for today. Leave comments on which game
Speaker A
you liked most and whether you agree with my conclusions. Hit the like button. That way, I know you enjoy this type of content. Subscribe to the channel; it really helps it grow. And goodbye, everyone.
Topics:AI game developmentOpus 5.5Sonnet 5.5GPT-6 AstraGPT-6.1 SolBlender 3D modelingprocedural sound generationUnity game assemblyAI comparisonrobot battle royale game











