Tim Gabe explains the four levels of app design using AI, highlighting the importance of creative direction over just better AI models.
Key Takeaways
- AI can generate decent app designs but lacks true creative direction.
- More context and data can confuse AI and degrade design quality without human oversight.
- Human vision and art direction are essential to create standout, competitive apps.
- The best AI results come from clear, deliberate creative guidance, not just more data.
- The app design process benefits from balancing AI capabilities with human judgment.
What the video covers
- The app design space is compressed into four levels, from basic AI-generated apps to highly curated, vision-driven apps.
- Level one apps are easy to create with AI but lack differentiation and creative decisions, scoring low on point of view and execution.
- Level two involves giving AI better design vocabulary and direction, improving consistency but only slightly enhancing creativity.
- Level three adds excessive context and references, which often leads to inconsistent and disjointed designs, scoring worse than level one.
- Level four introduces human art direction to decide what to keep and cut, creating apps with character, personality, and intentionality.
- AI alone cannot provide a point of view or true art direction; it produces average results without human vision.
- Improving AI models raises the baseline quality but does not raise the creative ceiling without strong direction.
- The real competitive advantage lies in providing AI with better creative direction, not just waiting for better AI models.
- A scoreboard with three dimensions—point of view, consistency, and execution—is used to evaluate apps at each level.
- The video uses the example of building a wake-up alarm app four different ways to illustrate these concepts.
Full Transcript — Download SRT & Markdown
Speaker A
I spent the last decade designing apps and products. And in the last year, everything changed. Apps all of a sudden became easier than ever to create, but harder than ever to differentiate. And what I realized is that the app space essentially got compressed to four distinct levels. Level one, where most founders start and get stuck. Level two, a level surprisingly few founders graduate to. Level three, where lots of founders end up and where everything kind of goes sideways. And level four, a place most founders never reach, but also a place where if you do reach it, your app is far more likely to be competitive. And in this video, I'm going to walk you through all four of these levels in detail. And I'm going to explain what it takes to reach them by building a wake-up alarm app four different ways with AI. To make sure we're comparing apples to apples when evaluating the apps, we're going to use a scoreboard with three different dimensions, including one, point of view. Does the app actually stand for something or is it just kind of there? Two, consistency. Does the app hold together like one system instead of five different apps stitched into one? And three, execution. Is there any interesting product craft or thoughtful execution in this app at all? Each app and level will get scored on these dimensions from one to five. Starting with level one. Again, this is where most founders begin. You type a prompt, you hit go, and the first draft looks something like this. And honestly, it's usually pretty decent. This one is clean. If you had shown me this just a couple years ago before the AI era, I would have told you a real designer made this. But since we're in the age of AI, it now looks like a dozen other apps you've come across in the last week. And the reason is this. There's no decision in here. Nobody sat down and deliberately chose what this should be. So when a founder tries to push it, saying make it premium, make it more modern or make it feel like Apple, it moves a little, but it never really blows you away. And this is clear if we bring out the scorecard point of view. Not much to talk about really. So a one. Consistency actually pretty solid. It's a clean, tidy app. This is what AI is decent at today. So a three. But product execution, it's still a very basic app pattern. Nothing revolutionary going on here. A one, which nets out to five out of 15. Pretty clean but completely empty. Nothing that really makes you stop. And again, this is the level of apps we see everywhere today. What we call AI slop. Right now, here's the interesting thing. Every time a new model drops, this level one gets a little better. The base looks a little nicer. The slop is a bit less sloppy. So, it's easy to think, give it a year or two, and this gets so good that nobody needs an eye for design at all to achieve great-looking apps, and I get why you'd think that, but there's a study going around right now made by Contra that shows something important. They put the best AI models head-to-head on design, and the ones with the best taste won easily, but on vague briefs. The second someone handed them a real specific creative direction, those same models lost to a model that just followed the direction really well. This tells us that the real ceiling is the direction. And it's something I see personally every day working with early stage startups. The better the models get, the more that direction is worth. While a floor raises for everyone as the models improve, the ceiling flies away at the same time. Because with new models, we can continue pushing what's creatively possible. So this tells us a clear story. The real move and the real moat isn't waiting for a better model. It's giving the AI better creative direction. And that is exactly where level two comes in. At level two, you stop leaning it to chance. You start handing the AI real design vocabulary. A few skill files, maybe a basic design system, a way for it to understand what you actually want and maybe even why. And that works. Watch this screen tighten up. The spacing gets more deliberate. The typography use is cleaner. The app just looks a bit more intentional overall. It's still a similar basic prompt, but now the AI has more honed-in direction to work with. Here's the catch, though. Unless you've got a great eye for design, it only really gets better by a couple of percentage points. And it's a bit of a gamble, too, because AI is unpredictable. Sometimes you nudge it and it really lands. Many times you just want to punch your freaking computer screen, right? But let's score it and see where we land. Point of view. It's leaning somewhere, but still not really differentiating. A two. Consistency. This is where it climbs. Even cleaner and more deliberate. A four. Surprisingly good actually. Product execution. Slightly better, but still no out-of-the-box thinking. Another two. Call it eight out of 15. A real step up from level one. But you can still feel it. Nobody art directed this. There is no taste. So let's not stop here. Let's give it even more. More reference images, world-class apps to pull from, a whole design system to build on. Let's pump it full of everything we've got and see how far that takes us. That's level three. Every scrap of context you can find all at once and you just let it cook. And at first glance, it looks like more is happening. It's flashier. To an untrained eye, you might genuinely think it got cooler. But let's break it down together. Now, all of a sudden, components don't really match each other anymore. The spacing goes haywire. Same with font sizes. The whole system starts coming apart at the seams because we're forcing the model in too many directions. And nobody is in charge of a single one of them. It's what I call a Frankenstein, which shows in the scoring point of view. Level two had it leaning somewhere and got two points, but crammed in this much at once and every direction cancels out. So, it washes right back to a one. Consistency is broken. This is the one thing previous levels actually did a very good job at and now it's gone back down to a one. And for execution, not really worse than level two, but not better either. Still no real product decision anywhere. So a two. That adds up to four out of 15, telling us that level three doesn't just stall. It drops below level one. And this is not a small point because it shows us that even if we give more direction to AI, that doesn't guarantee us a great result. Past a certain point, more context and direction can make it objectively worse. And it's funny because this is the exact opposite of how it works with a human. When you have a person with real judgment, more context, like more info about the business, the customer, the goal, then the result typically gets sharper, not messier. But as we see here, when we do the same with AI, everything can very easily break. So what do we do to overcome this mess? Well, meet level four. Now, for the first time, you put a person with a vision behind the wheel and not to add more. No, no, no, no. To decide what to cut, what this should feel like. This is where it stops being about just looking good and starts being about character, a personality, a voice, something that looks like it was actually handcrafted with intention. That is what true art direction really is. And it's the first thing AI genuinely cannot bring because it doesn't have a point of view. It just has an average. And if you don't steer it, it will always stay average. It will continue looking like all the other apps out there that are built without direction based on the same models with the same averages. And like we touched on before, even though the models raise the average all the time, it still stays the average, the ordinary, what everyone else has access to, no matter how good it gets. So this gives you as a founder a couple of choices if you want to level up. Option one, learn it you
Speaker A
that the app space essentially got compressed to four distinct levels. Level one, where most founders start and get stuck. Level two, a level surprisingly few founders graduate to.
Speaker A
Level three, where lots of founders end up and where everything kind of goes sideways. [music] And level four, a place most founders never reach, but also a place where if you do reach it, your app is far more
Speaker A
likely to be competitive. And in this video, I'm going to walk you through all four of these levels in detail. And I'm going to explain what it takes to reach them by building a wake up alarm app four different ways with AI. To make
Speaker A
sure we're comparing apples to apples when evaluating the apps, we're going to use a scoreboard with three different dimensions, including one, point of view. Does the app actually stand for something or is it just kind of there?
Speaker A
Two, consistency. Does the app hold together like one system instead of five different apps stitched into one? And three, execution. Is there any interesting product craft or thoughtful execution in this app at all? Each app and level will get scored on these
Speaker A
dimensions from one to five. Starting with level one. Again, this is where most founders begin. You type a prompt, you hit go, and the first draft looks something like this. And honestly, it's usually pretty decent. This one is
Speaker A
clean. If you had shown me this just a couple years ago before the AI era, I would have told you a real designer made this. But since we're in the age of AI, it now looks like a dozen other apps
Speaker A
you've come across in the last week. And the reason is this. There's no decision in here. Nobody sat down and deliberately chose what this should be.
Speaker A
So when a founder tries to push it, saying make it premium, make it more modern or make it feel like Apple, it moves a little, but it never really blows you away. And this is clear if we bring out the scorecard point of view.
Speaker A
Not much to talk about really. So a one consistency actually pretty solid. It's a clean, tidy app. This is what AI is decent at today. So a three, but product execution, it's still a very basic app pattern. Nothing revolutionary going on
Speaker A
here. a one which nets out to five out of 15. Pretty clean but completely empty. Nothing that really makes you stop. And again, this is the level of apps we see everywhere today. What we call AI slop. Right now, here's the
Speaker A
interesting thing. Every time a new model drops, this level one gets a little better. The base looks a little nicer. The slop is a bit less sloppy.
Speaker A
So, it's easy to think, give it a year or two, and this gets so good that nobody needs an eye for design at all to achieve greatlooking apps, and I get why you'd think that, but there's a study
Speaker A
going around right now made by Contra that shows something important. They put the best AI models head-to-head on design, and the ones with the best taste won easily, but on vague briefs. The second someone handed them a real
Speaker A
specific creative direction, those same models lost to a model that just followed the direction really well. This tells us that the real ceiling is the direction. And it's something I see personally every day working with early stage startups. The better the models
Speaker A
get, the more that direction is worth. While a floor raises for everyone as the models improve, the ceiling flies away at the same time. Because with new models, we can continue pushing what's creatively possible. So this tells us a
Speaker A
clear story. The real move and the real moat isn't waiting for a better model.
Speaker A
It's giving the AI better creative direction. And that is exactly where level two comes in. At level two, you stop leaning it to chance. You start handing the AI real design vocabulary. a few skill files, maybe a basic design
Speaker A
system, a way for it to understand what you actually want and maybe even why.
Speaker A
And that works. Watch this screen tighten up. The spacing gets more deliberate. The typography use is cleaner. The app just looks a bit more intentional overall. It's still a similar basic prompt, but now the AI has more honed in direction to work with.
Speaker A
Here's the catch, though. Unless you've got a great eye for design, it only really gets better by a couple of percentage points. And it's a bit of a gamble, too, because AI is unpredictable. Sometimes you nudge it and it really lands. Many times you just
Speaker A
want to punch your freaking computer screen, right? But let's score it and see where we land. Point of view. It's leaning somewhere, but still not really differentiating. A two. Consistency.
Speaker A
This is where it climbs. Even cleaner and more deliberate. A four. Surprisingly good actually. Product execution. Slightly better, but still no out ofthe-box thinking. Another two.
Speaker A
Call it eight out of 15. A real step up from level one. But you can still feel it. Nobody art directed this. There is no taste. So let's not stop here. Let's give it even more. More reference images, world class apps to pull from, a
Speaker A
whole design system to build on. Let's pump it full of everything we've got and see how far that takes us. That's level three. Every scrap of context you can find all at once and you just let it cook. And at first glance, it looks like
Speaker A
more is happening. It's flashier. To an untrained eye, you might genuinely think it got cooler. But let's break it down together. Now, all of a sudden, components don't really match each other anymore. The spacing goes haywire. Same with font sizes. The whole system starts
Speaker A
coming apart at the seams because we're forcing the model in too many directions. and nobody is in charge of a single one of them. It's what I call a Frankenstein, which shows in the scoring point of view. Level two had it leaning
Speaker A
somewhere and got two points, but crammed in this much at once and every direction cancels out. So, it washes right back to a one. Consistency is broken. This is the one thing previous levels actually did a very good job at
Speaker A
and now it's gone back down to a one. And for execution, not really worse than level two, but not better either. Still no real product decision anywhere. So a two. That adds up to four out of 15, telling us that level three doesn't just
Speaker A
stall. It drops below level one. And this is not a small point because it shows us that even if we give more direction to AI, that doesn't guarantee us a great result. Past a certain point, more context and direction can make it
Speaker A
objectively worse. And it's funny because this is the exact opposite of how it works with a human. When you have a person with real judgment, more context, like more info about the business, the customer, the goal, then the result typically gets sharper, not
Speaker A
messier. But as we see here, when we do the same with AI, everything can very easily break. So what do we do to overcome this mess? Well, meet level four. Now, for the first time, you put a person with a vision behind the wheel
Speaker A
and not to add more. No, no, no, no. To decide what to cut, what this should feel like. This is where it stops being about just looking good and starts being about character, a personality, a voice, something that looks like it was
Speaker A
actually handcrafted with intention. That is what true art direction really is. And it's the first thing AI genuinely cannot bring because it doesn't have a point of view. It just has an average. And if you don't steer it, it will always stay average. It will
Speaker A
continue looking like all the other apps out there that are built without direction based on the same models with the same averages. And like we touched on before, even though the models raise the average all the time, it still stays
Speaker A
the average, the ordinary, what everyone else has access to, no matter how good it gets. So this gives you as a founder a couple of choices if you want to level up. Option one, learn it yourself. Read up on the fundamentals. Lean on the free
Speaker A
skill stacks people like Emil Kowalsski and a bunch of design engineers put out on Twitter. It genuinely works. The catch is time because taste comes from reps, not from one weekend of reading and installing stuff in your terminal.
Speaker A
So this is the slow road. Option two, you bring in a real one, a senior products savvy founding designer or an in-house hire who has actually shipped things people use. This is a great long-term option, but it's usually also
Speaker A
the most expensive because someone that good costs a Silicon Valley salary and oftentimes b equity. Last but not least, option three, you borrow it. You get outside help, a freelancer or an agency that specializes in this kind of stuff.
Speaker A
This is the part where I'm supposed to plug my agency. So consider it plugged and take it with the grain of salt it deserves because of course I'm a little bit biased. In the end though, most founders end up doing some mix of all
Speaker A
three. Now regardless of what you do, remember that this person or this team isn't in the room to outdraw the AI.
Speaker A
They're there to steer the product design. Every single line on that scorecard is essentially their job description. For consistency, they need to consider what is our visual system here and how do we build something that still holds together on the hundth
Speaker A
screen, not just the first five. When it comes to point of view, they need to help you figure out what you actually believe this thing should be. And how do we make it feel like the top products of
Speaker A
the world while still feeling uniquely like us? A designer is there to help you make those calls. And then AI is 100% still going to be a big part of the process because once you have these puzzle pieces, AI will amplify them at a
Speaker A
speed and volume no person or team could ever hit on their own. And that's the whole flip. AI is now the ship while a human is the captain. So let's look at what this gets you. Same alarm app, but
Speaker A
now the entire thing runs on a single idea, waking up gently. This super smooth sunrise gradient shows up on every screen, so the whole app feels like one place. The copy actually talks to you in a fun human way. GM, Alex,
Speaker A
don't fight it. Time to get up, bro. And it's cleaner in general. Numbers use a considered size hierarchy, making it easier to scan. Even the mood check is just one soft emoji on a single glowing card. Somehow, it just feels obvious
Speaker A
that none of this came out of a prompt. You can just tell that somebody decided this app should feel like a calm sunrise. And then every screen aims to honor that one decision. Now, I want to be honest about something though. When
Speaker A
it comes to purely the visual layer, AI is making great progress. We saw Porto Rosha deliver an insane brand on the back of Nano Banana in early 2026. Of course, they're one of the best brand agencies in the world, but the point is
Speaker A
AI power is there right now to achieve great results. if you nudge it in the right way. And I believe that nudging will become easier and easier even for people who don't have taste in the coming years. Like I wouldn't be
Speaker A
surprised if we see people with no design experience doing objectively great art directed visual design work on the back of top tier AI models in a few years. However, what I think will still remain a very hard nut to crack for AI
Speaker A
is the deeper product experience because for that we still need someone actually asking very human native questions like what are we building here and who are we building it for? This is a crucial point because if we look at the scorecard now
Speaker A
point of view finally yes for the first time this app actually feels different like crafted and for consistency yes it all holds together. Now, when it comes to execution though, this is where it truly pulls away. And what I'm referring
Speaker A
to when I'm talking about human native aspects of design, let me just show you.
Speaker A
Imagine this. The alarm goes off. You, as a human user, are in bed. Your mind wakes up while your body is still basically asleep. So, your mind instructs your half asleep arm to reach over and tap a screen to kill it. But
Speaker A
tapping that screen doesn't really wake your body up. So you tap it and then you fall right back asleep. The alarm technically did its job, right? But if we think about it, it didn't do its job at all. It didn't wake you up. So here's
Speaker A
a decision a human makes that an AI won't do anytime soon. To turn the alarm off, you have to shake the phone for 15 seconds with a bar that fills up dynamically as you shake it long enough that your body has to move and actually
Speaker A
wake up. On top of that, there is no snooze button. This app is about waking up when your alarm rings. And just like that, the function finally matches the goal. The one score that stay low the entire way up finally earns full marks.
Speaker A
Five out of five and a total of 15. In previous levels, the AI didn't really take the actual experience or goal into account because it never considered the bed or the body or the moment of waking.
Speaker A
It was just trying to give you a cool looking alarm app. And again, outside of our direction, this is AI's biggest issue. It cannot read between the lines.
Speaker A
It doesn't intuitively understand what nobody actually said out loud. That is the key part only you or another human on your team can bring to the table. So, let's put it all together. Four levels, one app. Level one, the average everyone
Speaker A
begins with. Level two, better once you give it real direction. Level three, the Frankenstein, where piling on more of everything actually makes it worse. And level four, where a human steers the ship to create magic together with AI.
Speaker A
Now, if you want to talk about how any of these lessons could apply to your own products, we offer free design strategy calls monthly at my agency, Zip. You can check it out in the link down below. And
Speaker A
if you enjoyed this breakdown, I'm actually sure you'll love this playlist here somewhere with my best videos on how to grow your business with design.
Speaker A
Now until the next one, have a great
Topics:app designAI designcreative directionproduct designuser experiencestartup appsdesign levelsTim GabeAI creativityapp differentiation











