Learn how to maximize your Claude subscription with agent threads, token usage, and practical strategies for efficient AI workflows.
Key Takeaways
- Running multiple agent threads can significantly enhance productivity and token utilization.
- Subscription-based token access can be more cost-effective than API usage for high-volume needs.
- Fast and accurate search integration is crucial for agent efficiency and practical use.
- Many large companies use personal tier subscriptions strategically to save costs and scale.
- Understanding and navigating terms of service and potential bans is important for sustainable usage.
What the video covers
- The video discusses how to effectively run multiple agent threads simultaneously to leverage Claude subscriptions.
- It explains the token economics of Claude API subscriptions versus direct API usage.
- The presenter shares practical tips on accessing and using tokens efficiently for both small and large scale projects.
- There is a focus on real-world applications, including scaling strategies used by large companies like AWS and Microsoft.
- The video covers both best practices and common pitfalls in agent management and token usage.
- It includes a demonstration of fast and accurate search capabilities using the sponsor's product, Parallel.
- The presenter addresses anticipated criticism and misconceptions about subscription usage and API costs.
- The video warns about potential bans and terms of service issues but shares experience-based mitigation strategies.
- It highlights the importance of automation and running agents even when the user’s laptop is asleep.
- The content aims to inspire viewers to innovate with agents and share new solutions with the community.
Chapters
- 00:00Introduction and Agent Thread Question
- 03:15How to Think About Agent Usage and Productivity
- 06:30Token Economics and Subscription Tiers Explained
- 10:32Addressing Common Criticisms and Misconceptions
- 14:40Practical Setup and Usage Tips for Agents
- 21:59Dashboard Overview and Monitoring Agents
- 32:54Remote Development and Infrastructure Recommendations
- 37:12Optimizing Agent Performance and Avoiding Bottlenecks
Full Transcript — Download SRT & Markdown
Speaker A
Before we get started with this one, I have a question for you. How many agent threads do you have running right now? I don't mean have you run today, I don't mean are you going to run later, I mean
Speaker A
in the exact moment where you clicked this video, did you already have things running in the background? If the answer is less than five, you really, really need to watch this video because you are not taking advantage of the genuinely
Speaker A
awesome and unique opportunity we live in today. And if the answer is more than five, you should also probably watch this because you'll get a lot of nice edges and solutions to problems you might not have even known about and
Speaker A
you'll have a great resource to share with your friends who are not taking advantage of where we're at today.
Speaker A
I'll be real with you guys though, this video is not my norm and as such, this doesn't feel quite right. I don't think we need caffeine for this one.
Speaker A
We're going full rodeo. Much better. This is the video that Anthropic does not want you to watch and I'm going to make it very hard for them to do anything about it. So, before you get too worried about bans, I promise I've
Speaker A
done everything in my power to prevent it and of the bans that have occurred across me and my friends who do all of these things, of the 30 plus Claude accounts we're hopping between, I've had one friend get banned twice because he
Speaker A
was on a VPN when he registered an account with a credit card for a business that was not his, it was mine, and those two accounts got temporarily banned and reinstated almost immediately after. So, while this is probably not
Speaker A
something the companies want you to do, it is a thing that they are not stopping us from right now. In fact, the only thing that can really stop us is a quick break for today's sponsor. I have two
Speaker A
quick questions. First, have you ever used an agent without search? If you have, you know how painful and miserable it is. It basically can't do anything.
Speaker A
My second question is the opposite. Have you ever used an agent with insanely fast and accurate search? I personally hadn't until I started using today's sponsor, Parallel, because they have the best and fastest search results for pretty much every single thing you can
Speaker A
measure. Historically, I've really liked the search built into OpenAI, but when you compare it to Parallel, it's just night and day. Watch and see just how fast Parallel can get you good results.
Speaker A
Going now. All real-time, of course. It took just over a second for Parallel. And OpenAI still going, still going.
Speaker A
Meanwhile, OpenAI's endpoint took over 6 seconds long. If search was all they did, they'd be one of the best options available, but they also have everything else you would need on the web from a proper monitor that will send your
Speaker A
agents info when things change on different pages, to a traditional response API for when you want to actually get synthesized results from your search queries, to their extract endpoints to let you send a URL and get back the data that your agents actually
Speaker A
want and need. Even parsing JS heavy pages, by the way. There's even an MCP so you can expose Parallel to your existing agents in whatever tool you're using now. Every month you'll get 5,000 requests for free. And on top of that,
Speaker A
if you sign up today, you'll get $80 in credit. What are you waiting for? Join now at selective.link/parallel.
Speaker A
This is a real, real rough order of events that I have planned here. We're going to start with how do you actually get enough tokens to do the terrible things I'm about to show you guys. Right after that, we're going to talk about
Speaker A
actually accessing them once you've taken advantage of the ways to get more. Then I'm going to talk about how you actually use them practically well. And of course, right after, poorly, which I'd argue is more fun and much more
Speaker A
important. And then at the end, the thing that we all want, how do you use your tokens when both you and your laptop are asleep? While I will do my best to make this a top-level guide on like what things you should do, the
Speaker A
thing that will matter more is the little pieces that I show. Not because you should copy all of them, but they should help you understand how I think and operate with agents and how I've been able to do like 100 plus PRs a day
Speaker A
for a bit now, part-time. The point here is not to give you the exact formula to copy. It's to get these ideas in your head and these examples floating around in your brain, so that you'll apply some of these lessons in
Speaker A
hopefully your own ways. Like, the best possible outcome of this video would be you guys going off, building more cool with agents, and then coming to me with solutions for problems I hadn't even thought about yet that make my life
Speaker A
and my agent maxing better. I can already tell this video is going to cause huge blowback on Twitter. So, I'm going to just list the questions they're all going to bring up to try and dunk on me for making this ahead of
Speaker A
time. So, let's go through the common questions and comments that I know are going to happen the moment this video goes live.
Speaker A
This only works for side projects, not real apps. I have a good feeling if you're saying that, my apps are more real than yours. I also happen to know that a lot of the strategies I'm showing here have worked at companies scaling as
Speaker A
big as, I don't know, AWS itself, where I have friends who are learning these lessons from me, as well as at Microsoft, where I have friends who are learning these lessons from me, and even some Fortune 500 companies that are
Speaker A
doing some of the subscription abuse stuff that I'm going to show y'all. We'll talk about that in a bit. Not all of us are rich. Cool. I'm sorry.
Speaker A
This is a great way to save a load of money though. These strategies work at small scale and large scale, and the crazy numbers you've seen me spending in the 300k plus range are not spent at all. I'm putting in like $1,000 for
Speaker A
every 50,000 of tokens I'm taking out, like worst case. And if you're doing worse than that, you're paying API prices.
Speaker A
We'll get there in a bit. Of course a JS dev thinks this. Oh, yeah, the devs who understand the importance of having software people can use instead of software that people can gawk at and not use. Yeah, you're right.
Speaker A
I do know what people want, which means I am pre-qualified to talk about what people want in building real software.
Speaker A
Wow, OpenAI has paid you off. I have paid OpenAI more money than you're worth. Wow, Anthropic has paid you off.
Speaker A
I've paid Anthropic more money than you're worth. Now that we've got this out of the way, let's actually be useful.
Speaker A
First and foremost, getting more tokens. Let's do some math here. Who here knows how much money and tokens you get when you put in $200 on the Claude API?
Speaker A
Give you a hint. It's $200. What about the Codex API? I'll give you a hint. The answer is the same.
Speaker A
Which is why what we will be doing today does not involve the API. You might think this is silly or dumb or not viable for real businesses. If I was to list the big companies I know that have
Speaker A
their employees subscribing to personal tier accounts and just turning off the setting for data sharing, they'd all be really mad at me because these companies are very private and they don't want these things known, but they're all doing it. You have no idea how
Speaker A
many companies are just paying for the subscriptions. And when you do the very basic math, you'll understand why.
Speaker A
Because a $200 Claude sub gets you around $8,000 in tokens. That's 200 bucks a month for 8K in tokens. There is a catch for this one, though. Only 4,000 of that is Fable.
Speaker A
Because they have the 50% limit on Fable separately. That still means you're putting in $200 and taking out 4,000 of Fable. That's a very good deal. The $100 tier is weird because you do actually get half of this total amount, but you
Speaker A
get a fourth of the hourly limits. So, the 5-hour limit is much more strict, which means it's harder to burn that all on the $100 account. So, personally, I don't really recomme
Speaker A
I think you should save a little more, do the 200, burn it, find a way to make money off how you burnt it, and then reinvest and keep going. And the $200 Codex sub is even more interesting cuz
Speaker A
it's roughly $12,000 of inference. And there is no distinction between what you're doing with Astra and with other cheaper models like Soul.
Speaker A
They don't split it that way. You just get the tokens. And if you guys think these subsidies are brutal, I got an even crazier one.
Speaker A
Y'all know how much Fable costs? It's $10 per mil in and $50 per mil out.
Speaker A
That sounds expensive, especially when you compare it to other cheaper models, especially in the open weight world where there are things that are doing real work for a tenth or less that price. What if I told you this number
Speaker A
was also more than 50% off. Fun fact, when Anthropic first announced Claude Mythos preview in April, they actually quoted a price. Even though the model wasn't out, this was just for internal use at companies for project Glass Wing for securing their stuff.
Speaker A
Those companies were offered to continue using the model after at $25 per million in and $125 per million out. That means the price we're paying is like almost 70% cheaper than what they intended to charge. Why would they ever make it so much cheaper?
Speaker A
It's cuz their margins are insane. The prior rumored margins for how much money they made selling tokens at API price compared to like the energy costs and the eventual usage of the compute causing it to fail over time roughly
Speaker A
calculated from various napkin math that fishy and autos I know to around a 95% profit margin. So, if they charge you 100 bucks, they made 95 and they spent five on replacing dead compute and paying for electricity. Or in the case
Speaker A
of Anthropic, renting the compute from SpaceX. So, why would they knock this price down so much? Well, Anthropic knocked the price down because they know everyone has the margins. They were already in the lead.
Speaker A
And they wanted to maintain their lead. So, they decided to make the price lower in order to make it harder and harder to justify using competitors. I would argue that at 25 in 125 out, this model was, especially at the time,
Speaker A
worth it. But at $50 out, it's a bargain. And at $50 per mill out, combined with the 200 to 8,000, so the 1 to 40 ratio of subsidization means you're getting some really good deals. So, you can kind
Speaker A
of mentally double the numbers for Claude as well. And also, Codex, because there is no world in which OpenAI actually plan to charge the literal exact same rate as Mythos. OpenAI does numbers very simply. They keep them the
Speaker A
same, they lower them 20 to 80% or they double them. Those are the only things Open AI knows how to do. Except for after Fable destroys them, and after that, all of a sudden, they're going to nudge the price a
Speaker A
little bit. So, the reason that Astra is as cheap as it is is because of Fable.
Speaker A
And the reason that Fable's so cheap is they decided to eat their margins a little bit in order to guarantee a win at least at the time. And this is all without even talking about the resets.
Speaker A
Last I saw someone do the numbers, Tibo gave out 14 resets in 30 days. That means you're averaging a reset every two to three days with your Codex subs. That means that the weekly number, which is four grand, is more like a three-day
Speaker A
number, which means you can double the Codex sub here with 24 grand if you're really maxing it out. It's so bad right now that they've temporarily paused new subscriptions on the $200 tier for Codex, which sucks due to the nature of
Speaker A
this video, and I have a bad feeling this video is going to make the problem worse. Quickly. Sorry in advance. Feel a little bad exposing these things to the world, but my job in my loyalty is as a
Speaker A
journalist and informant, not as somebody who's trying to protect and shelter the few hundred people who know how abusable this is. I think everyone should know, and I'm going to share this info with the world, and how things proceed after is how things
Speaker A
proceed after. Don't be mad that I'm sharing, be mad that I'm right. Okay. When you set up your account, the first thing you should do in Codex is hop over to settings, go to data controls, go to improve the model for everyone, and turn
Speaker A
it off. As soon as you do that, the terms of service for your personal sub are effectively identical to the terms of service for a team account. There is no real difference. So, if you are paying for API prices because you think
Speaker A
you're not going to have your data trained on, you're not saving anything. You're just wasting money. Every dollar you spend on API is 80 to 90 cents you've thrown away. And if you're not thinking of it that way,
Speaker A
that's on you. There's a similar setting in the Claude code settings. I don't feel like going to find it. You get what I'm talking about. So now with all of this established, the best way to get more tokens is obviously to get more
Speaker A
accounts. And since we've established the accounts aren't meaningfully different from API pricing, you get a good deal. You should still use the API for a couple of things though. One is very important. Anything facing users.
Speaker A
If you are writing your prompts or your agents are writing your prompts, or the prompts are running on your system on your code base from things that are in your GitHub issues or whatever, all that is fine.
Speaker A
That is awesome. But if you are trying to serve public traffic, where users can submit a request and it runs on your inference, that is entirely against the rules with the accounts. If you're using this as an alternative to the API to serve user
Speaker A
traffic, you deserve your ban. That's not what these accounts are for at all. They are being very generous letting us experiment with coding with these huge subsidies.
Speaker A
You should abuse the generosity by building cool not by reselling things to squeeze margin that they aren't. I'm giving this advice up front because the things I'm about to show would make it hypothetically easier to do that. I think it's time to get there.
Speaker A
It's time to talk about how we actually access the tokens. In my first token maxing video, I gave the example of using two Claude code subs, and I showed that if you have Claude code running in the terminal and it's completing a task,
Speaker A
and you use another tab to change the auth for your Claude code instance, it will correctly recover and keep going.
Speaker A
That is great when you have one or two computers and one or two threads with two accounts. That does not work when you are doing what we are going to be doing here. You need systems that manage this, not just for your local Claude
Speaker A
code with your one or two accounts. You need a solution that will handle the insanity that is how often the Claude code auth breaks in a single shared bucket on a residential IP address that you can route all your other traffic
Speaker A
through. So, let's go through some pro tips here. So, tip number one, I love VPSs, we'll talk a lot about VPSs. I highly advise you do not sign in to Claude Coder Codex, but especially Claude Code, on a
Speaker A
VPS. Anthropic in particular is very, very nervous around people reselling their Claude Code subs in VPSs serving traffic to other people with Claude Code because it lets them sell it for cheap and also collect a bunch of data they
Speaker A
can use for distillation. So, they're being extra aggressive with sign-ins and often requests coming from IP addresses that are known to be servers, like anything on AWS or Hetzner or any of these other providers. So, I highly recommend you do not do this over
Speaker A
a VPN or on a cloud that you're serving the traffic from. I would also advise against having that traffic come from multiple different computers at the same time because that's how they know you're doing some sketchy stuff. So, how do we
Speaker A
avoid that? How can we do our best to make sure most traffic comes from one place, ideally one residential IP address, regardless of what you're doing, what account you're using, and where you're using it from, especially with something like Claude Code where
Speaker A
the auth breaks every two to three days. Dear from the future cuz I forgot an important disclosure. Before we go any further, this might get you banned, this might be against the terms of service. I am just speaking from my experience, do not hold
Speaker A
me accountable. I am literally drinking and filming as I talk about all these nerdy things.
Speaker A
If you get banned and I didn't, that's unfortunate and I'm sorry. I can't be held liable, you'll probably be fine though. Just don't do this with your Google accounts because when your Google account gets banned, you're Thankfully, there's no use for Gemini
Speaker A
models anyways, so there's no reason to deal with that. Cheers. As we were saying, it's time to talk about the proxy.
Speaker A
CLI proxy API is a thing that's come up in a lot of my videos, but not really been showcased in most of them.
Speaker A
The way that this works is you sign into it using your off for Codex and Claude.
Speaker A
And once you have signed in, it will maintain the off for you and give you an API that you can use in things like Claude code and Codex, where you point it to this instead of pointing it to the
Speaker A
traditional off servers. And now everything is routed accordingly. It is very very convenient. We'll get back to this in a sec cuz I realized I forgot one other thing we need to talk about in the getting more token section. So one
Speaker A
last thing here so I can actually wrap it up. You might have noticed the subscriptions I was talking about. I was talking about Claude subs and Codex subs. I didn't include other subscriptions here. Why did I not include
Speaker A
other things like the open code sub or the subscription to T3 chat or the subscription to cursor even?
Speaker A
The subsidies are comically comically less. Cursor does probably get a deal with Anthropic. The deal is probably 30 to 50 at best percent off.
Speaker A
That is very different from 95 to 98% off. You get a better deal than Cursor does. You get a better deal than open code does. You get a better deal than pretty much everyone does. Did a reset just get announced for Codex? Really?
Speaker A
Kind of crazy that even in a compute crunch like they have right now at Open AI where they don't have enough compute for all the answers they had to cancel the subs, they're still giving us resets. I will point out it is currently
Speaker A
Friday evening which makes it a lot easier for them to do this because they are hoping the power users like us, yes, I'm including you now viewer, you're joining us on this journey. The power users like us will burn it all to the
Speaker A
ground before Monday when the enterprises wants to pay API prices again. So, yeah. The only subscriptions that you can really massively benefit from in terms of you put in money, you get more tokens out are Claude and Codex. If Grok 47 comes out good, it
Speaker A
might make sense, but I just It is hard to token max with Grok because the Grok models just cannot stay coherent for as long yet. So, back to where we were with accessing the tokens. I gave you the
Speaker A
pre-reqs. You want it on your home network. You want it in one place ideally, and you want everything to route through. CLI proxy is a very cool place to start.
Speaker A
But, CLI proxy, not necessarily the best UI. The guy who got me into CLI proxy didn't know it had a dashboard because he has avoided it to the best of his ability. Technically, I think the version he uses and that I started from
Speaker A
was a fork called Vibe proxy, but I have extensively forked this project. I have made a lot of changes, and when you set it up and look at the dashboard, it won't look anything like this. And if you want to change that, then congrats,
Speaker A
you have your first token maxing task. Give it a screenshot of this and tell it you want it to look like this, and it'll probably figure it out fine, especially if you use Fable instead of Astra. There is one other change you
Speaker A
almost have to make, and I would argue these changes combined would be worthwhile if somebody creating a like real fork with an easier setup like flow. Because the other change you need to make is account prioritization. By default, the way CLI proxy works is it
Speaker A
routes traffic to all of your accounts evenly, and it distributes across them, which sounds great until you look at my usage here, where you see I have one account that expires at 1:00 p.m.
Speaker A
tomorrow, and I have another that expires in 4 days. Obviously, I want to burn as much as possible of the one that expires tomorrow and not touch the ones that have another few days. The only reason you wouldn't want to do this is
Speaker A
cuz you're constantly hitting 5-hour limits even with multiple accounts. I can't believe I'm saying this in the token maxing video. What I would say there is slow down a little because a 5-hour limit gets you through 40% of
Speaker A
your 7-day Fable. I've done the math. I know this one for a fact, so please do not give me I know for sure that 100% of the 5-hour on the 20x plan is exactly 40% on the 7-day. It's
Speaker A
only 20% of your full 7-day limit, but there's a reason I have this column grayed out. This is the Opus bucket from hell. Ignore that right column, it doesn't matter. We only care about the 5-hour and the 7-day, and we really only
Speaker A
care about the Fable 7-day. I often accidentally call my Claude accounts my Fable accounts because Opus is a trash model and you should avoid it to the best of your ability. Sup nerds, less drunk future Theo here. Wanted to call
Speaker A
out something that didn't exist when I filmed this a few days ago. Opus 5.5.
Speaker A
I'll be clear, all the advice in this video still applies in Opus 5.5 behaves very similarly to Fable 5.1, so anything I say about Fable for the most part applies here, too.
Speaker A
The big difference is you might not need five to 10 Claude subs anymore because Opus 5.5 on just one sub, I have found pretty hard to like hit your limits with. So, take that as you will. Maybe you don't need as many subs to take
Speaker A
advantage of these patterns now. You might not even need to set up something like the proxy I'm about to teach you about, but yeah. Thought it was worth calling this out cuz 5.5 is really solid overall and has changed my way of using
Speaker A
things. Opus 5 was a useless dumpster fire that I had to work around a lot.
Speaker A
5.5, pretty good model. Back to whatever drunk Theo was rambling about. Once you've got all of these signed in and you've had one of your models go through and make these changes, changing the routing to optimize for which accounts have the
Speaker A
next reset up soonest. And the other one I think is important is by default, your Codex subs won't go through web socket.
Speaker A
So, there's a few changes you can make to fix that, which hugely increases performance just like the time passing messages back and forth goes down a ton.
Speaker A
So, make sure you have the web socket stuff working. It's a little annoying, but you can. One last piece that's important, and thankfully my agents are a bit smart enough to do it right, is called session affinity and account
Speaker A
affinity. What this means is if I have a thread going, ideally the next prompt in that thread goes to the same account because caches are account specific. So, if you are 800k tokens into a thread and you switch to a different account
Speaker A
quietly behind the scenes of CLI proxy, Claude doesn't give a It doesn't know the difference. But, the API then has to go recreate those cache entities.
Speaker A
You only have to do it once per switch. But, if your stuff is switching per prompt or worse per tool call, then you're going to eat a lot of cache rates that you probably don't need to.
Speaker A
Thankfully, the affinity stuff works very well by default in CLI proxy and the caches only last 5 minutes anyways.
Speaker A
So, it's not too too bad. Just thought you should know because if you're using a dumber model to set this up and it gets it wrong like if you use Opus, it might get it wrong. So, keep an eye on
Speaker A
that. Make sure that you are looking at how much cache reading and writing you're doing and that those numbers aren't getting crazy. You'll get a good feel for it soon. You can always just ask the agent, "Hey, how are we handling
Speaker A
session affinity? Are we writing cache more than we should be?" And it can probably get an answer. For this next piece, I will request that you look up.
Speaker A
At the top of the screen, you'll see the URL I'm BB1 which is my framework, my desktop in the other room. That port is micro.ts.net. This is my tailscale.
Speaker A
This is my tailnet. 4318 is just a default and then I'm on the management HTML page. This URL is being accessed this way because I am using it over tailscale. Tailscale is very very good to have for a setup like
Speaker A
this because it makes it easy for other machines to access this proxy endpoint. And it means you don't even need off. I was experimenting with some things yesterday and I set up an API key for a different thing and it broke my
Speaker A
inference on everything because by default, it just works with no API key. This sounds horrible until you realize it's only exposed to other things on my tailnet. So, if you don't have my Google account that I use to sign into
Speaker A
tailscale, you cannot hit this endpoint. So, it doesn't matter. So, you're fine. And this also makes it really, really, really convenient to set up other machines. For this next part, I'm going to have to do a thing you
Speaker A
probably didn't expect from me in 2026. We're going to look at a real repo.
Speaker A
This is a project that I have on GitHub and on my machine. And despite the fact that most of my projects are on most of my machines, this one's only on this machine. That's because this is my fleet
Speaker A
management repo. This is the repo that explains how all of the other computers I do my work on function, including the one that hosts my proxy and manages the Tailscale for it so that everything else can connect. One of the files in here is
Speaker A
how to set up a box. This describes all the things I want set up in my boxes and all the random tools and things like I want RGFDJQ tmux btop node python and build tools all of this yet. Codex and
Speaker A
Claude code installed. I want them to use the proxy and all these other things. All the deals all the machines are in here gets updated whenever I add a new one. So, when I get a new machine, step one, I set up as a SSH. Step two, I
Speaker A
go to this repo. I open up Claude or Codex or in my case usually T3 code and I say, "Look, here's a new box. Here's the SSH key. Go set it up the way I like." And in not very much time, I'll
Speaker A
all of a sudden have a page open in helium for a Tailscale approval. And I go, "That was probably my agent running." I blindly hit accept and now I have the machine set up and it can do inference on these other boxes easily.
Speaker A
So, now I have what I needed. I have one machine on my home network going through a residential IP address that all five of my Claude accounts and all four of my Codex accounts go through and all my
Speaker A
machines all around the world connect to that over Tailscale and it is impossible for Anthropic or OpenAI to know that I'm using them for servers.
Speaker A
And setting this up in Claude code and Codex is trivial. It's so trivial that I'm not going to tell you how. I'm going to tell you to tell your agent to do it.
Speaker A
It will swap two variables in the configs and you're done. Thankfully, both Claude Code and Codex have to work with API endpoints because they're being used so heavily with companies that are doing all their inference through Bedrock on AWS, which means they need to
Speaker A
put a URL in. So, this will work indefinitely. It would be bad for them if it didn't. In fact, the Codex desktop app handled this poorly and enough people on AWS complained when Bedrock happened that they fixed it. And now, if
Speaker A
you're using Codex with this setup, not only does it work great, it'll even show you in the corner of your proxy. It's pretty legit. And if you're skeptical or you really, really want to have your normal install with a normal auth, you
Speaker A
can still do that and you can make another instance of Codex or Claude Code with a different home directory that has it configured. So, you have a different command for each. The other real cool benefit of this is if you do set it up,
Speaker A
you get access to all the models that are in that setup inside of your Claude Code. Notice what I said there, inside of Claude Code.
Speaker A
Do not, under any circumstance, use this to use your Claude Code sub in something other than Claude Code unless you really, really like dealing with Anthropic support and getting accounts banned all the time. You will be banned, you will be upset. I highly recommend
Speaker A
you only ever use Claude models through your Claude subs in Claude Code over the proxy. One last thing I think is worth knowing about the setup because it catches me off guard every time.
Speaker A
Anthropic resets work different than Codex ones. The difference is the dates. When Codex does a reset, it resets your weekly entirely. So, if your weekly was going to renew in a day or 10 minutes or whatever else anyways, your reset is now
Speaker A
7 days off. So, when the reset T was just promised hits at midnight, all my accounts that are currently 3 days away from a reset will now be 7 days away from a reset. And this means you end up
Speaker A
with a weird pacing issue where you can quickly get to a point where you're 6 days away from a reset across all your accounts, which really sucks. It gets kind of balanced out with the banked resets, which across my accounts, I have
Speaker A
3 6 7 8 I have two more in my other, so I have 10 resets across my Codex subs.
Speaker A
And the two types of resets are banked and immediate. Banked resets are the ones you can click whenever, and those cost them more money because the reason they can do these resets so freely is they do them in hours where they're not
Speaker A
competing for the inference with their enterprise customers. So, they can freely give them on Fridays and Saturdays, especially Friday nights.
Speaker A
They can't give them so freely on a Tuesday morning when they were to have all their customers trying to use Astra at work that are paying API prices. So, the banked resets cost them a lot more, literally, because of opportunity cost.
Speaker A
I do see a future where they start pushing us to only use these subsidized subscriptions at off hours. I'm legitimately at the point where I would shift my sleep schedule in order to maintain this level of subsidization because it's so good and I'm
Speaker A
addicted, let's be real. And we all will be soon. So, now that this is all established, the difference with Claude is there's only one type of reset. It's a limit reset, and not the timer for the limit, just
Speaker A
the percentage for the limit. So, if Claude was to do a reset right now, the only thing that would change is all of these numbers become 100% again. The dates and times all stay the same. So, if I was to get a reset from Claude
Speaker A
right now, and my next actual one was in a day, that means I need to turn on the furnace immediately. I need to burn all the tokens I can because they'll vanish at that point. And this is the mindset
Speaker A
shift I need y'all to get in. When your normal reset hits, so at 1:00 p.m.
Speaker A
tomorrow for me with this account, whatever I have left here is money I lost. Don't think of this as you spent $200 and you got more, think of this as you've been generously gifted $4,000.
Speaker A
But every week where you don't spend a thousand of it, it disappears forever. This is the opposite of the mindset that in when you're paying API prices, where you're trying to get as much as possible for as few tokens, you need to have the
Speaker A
depressed mindset where you feel yourself losing money because you are and we don't want to be stuck in the permanent underclass. Burn your tokens.
Speaker A
Cool. That's most of what I need to show in this dashboard. I'll probably be back here. I know I will be back here. Let's be real.
Speaker A
I guess this is most of the accessing your tokens bit, which means next we need to talk about using the tokens well. This one is going to be full of all sorts of different layers. What I'll say for the core of this one is I have
Speaker A
another video that's probably already out by now that's about how I code without one of my hands. It's a video all about my life after accepting that I can't really type even after my surgery and I get my hand back. I might not be
Speaker A
able to type very well, which means I have to use my computer different. Not just like voice to text, but I don't want to switch between apps as much cuz I can't command tab. I don't want to be
Speaker A
staring at the thread as it generates because I have other things to do. I don't want to be at my computer that much. I'm sitting at my computer less and I'm coding more and these are the strategies that will get you there. That
Speaker A
video has a lot of the good examples, but I do want to give a couple important points I don't necessarily think were in that and then also show you some dumb examples that weren't there that could be useful. The question I want you to
Speaker A
start getting into your head is how can I use more tokens to care less about this? When you have a problem that you're trying to solve, how early can you pull in the agent to start solving the problem and how long can you have it
Speaker A
go until you need to take another look? This is a twofold thing. You could even think of this in terms of a specific important button.
Speaker A
Merge. You want to derisk both sides of the merge button. You want it to be more likely that by the time you go to GitHub and you're looking at the merge button that it is safe to hit it because the
Speaker A
agent has addressed as many of the potential problems as possible. It's kind of crazy to think of it this way, but I would honestly guess that of my token use, maybe 10 or 15% is actually coding. And the other 80 to 85% and the
Speaker A
other 85 to 90% is being spent verifying the code. So, by the time I am actually looking, the work's done. And not like it's kind of done, but it has these bugs. If you can knock down the likelihood that the PR sucks or has some
Speaker A
small issue from 5% to 1%, you'll be able to do way, way more because that means you can focus less on each PR. So, you want to de-risk merge on that side, and you also want to after merge. It
Speaker A
needs to be easier to revert the things that are failing. You need systems that catch things before they hit your users or systems that will revert things after they hit their users and you've noticed the problem. If undoing a bad change
Speaker A
takes more than 15 seconds, you probably shouldn't be vibe coding at all yet. You probably need to fix your systems so undoing something broken is way cheaper and faster. Then you can go a little harder. Holy I think I just saw
Speaker A
the actual worst take of all time in my chat. The argument that I'm making is essentially you need to watch as much Netflix as you can because you need to maximize your subscription.
Speaker A
I I'm sorry. I We may be using a different Netflix. I I just I personally don't know how I would use my unlimited movie watching to build real businesses.
Speaker A
I don't see how I would use that to save 40x plus. I don't see how Netflix can help you escape the permanent underclass.
Speaker A
I've actually never I I hope you're rage baiting cuz this is legitimately the worst take I think I've ever seen, and it's my job to read shitty takes.
Speaker A
Which means congrats. You can probably provide a lot of value to the world through Netflix because you've successfully baited me in the middle of what will be a very good video by sending the actual stupidest possible message. So, seriously,
Speaker A
congrats. Salute. Hats off. Fantastic work. Back to real world. So, we've talked about de-risking before merge by making the code more likely to be good. We talked about de-risking after as well. So, if the code is bad, it's easy to fix.
Speaker A
Whether or not you are using agents, these are things worth doing. Make it easier to verify code is good before you look at the PR and make it easier to revert if the code is bad. But, there is
Speaker A
one other phase here, which is not really a button, which means this isn't right UI, but whatever, you get the idea. Knowing what you want. There's a very good chance if you're watching this video and you're using agents for
Speaker A
coding, that by the time you have sent the prompt to the thread, you probably know what you want. I'm saying this because this is the case for me even just a few weeks ago. That is not the case anymore. I have
Speaker A
finally rewired my brain where I don't bring in the agent when I'm done thinking. I bring in the agent when I start thinking. So, we can talk it out.
Speaker A
Maybe we find a shortcut that I hadn't thought about before. I've had times where I thought about a problem not like fully, but like pretty actively for 3 weeks. And then I went to an agent to build it. And I just asked, is there a
Speaker A
stupid simple solution I'm not thinking of here? It gave me one. And I realized I had just wasted 3 weeks of my time. I want to be really clear here because after the video about knowing your code base, I realized you guys don't listen
Speaker A
very well. So, I'm putting this in largely to have a quote that I can grab when people misquote me from this video.
Speaker A
I am not saying you should stop thinking. I am saying you should figure out if something's worth thinking about before you think about it. Because you don't know until you put the time in thinking about it. But, if the agent can solve it
Speaker A
before you have to think about it, you both just saved a bunch of time. And tokens too.
Speaker A
Because if I ask the agent about a problem and the agent has a good solution, then I don't have to trick it with my bad solution and run in circles a whole bunch with it. You save your tokens and your brain if you give the
Speaker A
agent the problem instead of the solution. And this was a hard habit for me to break. When people would DM me a bug in T3 code, for example, I would think through the bug, I would think through where the problem was, and I would go to
Speaker A
my agent with a solution. I would tell it, "I want you to change these things in this way." And then it would change it, and then I would look at it, and I would realize it doesn't actually quite
Speaker A
solve the problem I wanted to. So, now I just hand the screenshot to the agent and say, "Fix it." And it often does.
Speaker A
And if it fails to, I now know this problem is too hard for the agent to do itself. At which point, I know it's time to think more about it.
Speaker A
I think it's been really cool to learn that a lot of the problems I would have used a bunch of my mental energy on didn't need it. And also, and this is even cooler, some of the problems I
Speaker A
thought were really simple weren't. So, the agent outright failed. You ready for the spicy take here? This is similar to how it felt to be a manager. When I realized that my team doesn't need to be handed solutions, it needs to be handed
Speaker A
interesting problems, and they would usually come back with solutions. And if the first few times they come back with a solution, it's not quite right, then I know this problem is novel and difficult in some weird way, and I have to dive in
Speaker A
and help steer it a bit. So, instead of diving in once you know the solution, dive in when you discover the problem.
Speaker A
See if it can solve it autonomously, and if it can't, whatever, buy another Claude account. Or get your boss to.
Speaker A
Usually, though, if you're not employed as a dev, don't stack subs until you're making money. Like, I know this shit's expensive, but devs make a lot of money, and the companies hiring devs also make a lot of money. Burn the money when you
Speaker A
have it, burn somebody else's when you don't. So, the core thing I'm trying to say here is that at each of these stages, from problem to knowing what you want to do to the code being written and merged,
Speaker A
too many of y'all live not even this whole range, but like between a small set here, where you are just letting the agent operate between knowing what you want in PR filed. I want you to do this.
Speaker A
Go all the way from where you first hear about the problem to when you hit merge, and maybe even let the agent merge itself once you build more confidence.
Speaker A
And now your job is talking to users, and of course SRE for when things do inevitably fail.
Speaker A
The time thinking about this changes you in important ways. So again, Dan, I respect you heavily.
Speaker A
You were fantastic to work with, but I will push back on this specifically. Because I again, I even gave myself this quote earlier to make sure I have my get out of jail free card. I'm not saying think less.
Speaker A
And I personally believe it is a better use of our time to think about problems agents can't solve than the ones agents can. You will learn more and better yourself more thinking about the problems that your agents can't solve or reading through
Speaker A
the solutions that they can solve than you would get thinking about a problem for days that an agent can solve in minutes.
Speaker A
I'm not saying that you should replace your thinking with the agent. I'm saying that you should optimize for thinking about things that matter more. Don't outsource thinking, scale it. Yes, exactly.
Speaker A
Pull in the model to figure it if you need to think about the thing, and if you do, that's where your brain energy goes.
Speaker A
And we're going to get to this tip in a bit, but part of the skill here isn't to give it the task, let it go do the thing while you go and play a video game.
Speaker A
Since the agent's doing the task, it's time for you to do the next task, and then the next one, and then the next one, and then you see the first one done, so you go back and check it. And
Speaker A
that's where this starts to get really cool. If you let the agent run for this whole path, this can take hours. It is possible that from when you show it the problem to when the PR is ready to go, the agent
Speaker A
has to run for a couple hours. Maybe it runs for 30 minutes and it tests the thing quick, it gets up a pull request, it waits 15 minutes for getting, I don't know, like a review from any of our
Speaker A
awesome AI code review sponsors, and it spends 10 minutes fixing the changes, puts it up again, waits another 15 for a follow-up review, addresses those, and now it's good to go. That's an hour plus that you could be spending
Speaker A
sending more prompts to other threads to do more things. Here's where that ends because I'm going to be so real. I have been very kind and polite to the alternatives to T3 Code, but they suck when you have more
Speaker A
than three threads going. And here is where we need to talk about a very good question we just got from Strawman Twitch.
Speaker A
How do you keep them all straight in your head? I'm going to be so real with you. If you're not using T3 Code, I do not know how you do it. I I hate that this is the case because the competitors
Speaker A
have been trying to copy us, and they are not succeeding. And it is really genuinely frustrating.
Speaker A
I want the Codex activity thing to be good. I've offered to go there for free and write the code for them because I want to use these things well more than I want us to win with our open-source
Speaker A
that makes no money. So, what the hell am I talking about this arrogantly? I'm not actually arrogant about this.
Speaker A
I'm pissed off about this because the solution was really simple. It was admittedly one that I thought about for far too long cuz I was trying to figure out how do I deal with the fact that I have, let's be real, far too
Speaker A
many threads at any given time. I realized that I'm not treating threads properly. Threads are not histories that you actively go back to all the time.
Speaker A
Threads are tasks. They're to-dos. And when they are not working, they should not matter.
Speaker A
I have a bunch of stuff here cuz I've been streaming, so all of these things are done. Also, I was at Demo Day yesterday, but normally at any given time, I got 10 plus things running in five plus in the done state. And this is
Speaker A
where things get really, really cool with T3 Code specifically. When you are done with a thread, you check settle, and now it is gone.
Speaker A
I know, I know, Theo thinks he's so cool for giving a new word to archive.
Speaker A
This is not that. Settle is different. The goal here is to turn your sidebar into inbox zero. You should try to end your day with all your threads gone or running, so when you go to bed and wake
Speaker A
up the next day, you have some cool things to look at. But, that doesn't mean we have cured the context switching costs. We haven't. We have just made it easy to see which context needs your attention.
Speaker A
And it's really to start rewiring a bit. This is where if you have ADHD, you have a solid advantage. Because once you get in the mindset of thread is up, not my problem now, you're not going to get out
Speaker A
of it. What's something I want to add to T3 code? I'm going to be so real, T3 code has progressed so much and added so many of the things I wanted, I don't really think about it that much anymore.
Speaker A
Like I I don't I'm running out of things to add, which is great. It means we're doing very well. Once orchestrator V2 is in, that will change. But, let's say I want some things, and I do. I have a
Speaker A
couple in mind. Here's the first one. I would really like a simple and minimal queuing system where I can send a message and it will show as pending, and it will go up when the next tool call is
Speaker A
completed, or I can click steer and immediately send it. The attached screenshot shows an example of what I'm thinking of here from another app. Since I don't have a good way to get the screenshot cuz my codex don't reset for
Speaker A
another 3 hours or so, you should be able to find a good example of this in codex. Look for screenshots online if necessary.
Speaker A
I have two tips I'm about to give you that are important, they're going to be rapid fire, so pay attention right now.
Speaker A
Tip one. Look at the computer I have selected on the bottom left. This is one of the things I think T3 code does exceptionally. Right now, it's Theo's MacBook Pro.
Speaker A
That is the computer I'm currently using. This computer is going to have to be closed when I go downstairs later. I don't know if this thread will be done by then. So, I am not going to to this
Speaker A
on this computer because I don't run anything on this computer other than my fleet management, cuz I'm doing that in the loop. So, we're going to click here.
Speaker A
This is going to be any repo, and as long as the origin for Git is the same on the different machines, they will all be bunched under here. So, I can pick any of my servers and pick where I want
Speaker A
this to run. Usually, I pick one of these three cuz they're my Linux boxes, but if I really want this to work on mobile or do computer use, which right now is better in macOS, I can pick my
Speaker A
two Macs at the bottom here to do it. I don't care for this one, so I'm just going to throw it on a random cloud server.
Speaker A
Cool. Now, it's going. And I just realized I missed the second tip. So, I will give that after showing another thing that I want to fix. If I hop in here to connections and scroll down, you'll see all the
Speaker A
devices that currently have connected. I hate the UI here because you have update or disconnect and no remove for my T3 Connect machines, and disconnect doesn't seem like it is or isn't permanent. So, it's not very clear.
Speaker A
So, here's what I'm going to do. I'm going to grab a screenshot of the whole UI here.
Speaker A
Screenshot grabbed. Going to go back. Command shift O, enter, new thread, cool. Paste. I really want to rethink the UI for the devices connected in T3 Connect in settings on the desktop app.
Speaker A
There's a couple issues I have. First, I feel like disconnect doesn't seem as temporary as it is. It would be nice if that was a toggle instead. I also don't like that when update is an option, remove isn't. I should always be able to
Speaker A
remove, and it should be very clear that it's a permanent destructive action. If you think this is simple enough to do directly, go do it and send me a screenshot. If you feel like you aren't quite sure, make me a few mocks using my
Speaker A
HTML scale, and we can decide between them. Cool. There's a couple things I did here that might be useful. First, I gave it the exact issue I have with not a prescribed solution, but the problem, and then some ideas of solutions. I then
Speaker A
told it that it can go do it directly if it has a solution it's happy with. And I also told it that if it doesn't, to give me mocks using a skill I built so that we can make a better decision together.
Speaker A
And now if it does make one solution I don't like it, I can tell it again like go make the mocks and it knows what I mean. And here is where one of my favorite T3 code pro tips comes in. I'm
Speaker A
not going to press enter, I'm going to press command enter, which sends off the thread in the background and leaves me exactly where I am. So I can now send off another prompt without having to do anything. It's very nice. And this helps
Speaker A
you get into the mindset of oh, that's a problem. I will go fire it off and I will look when I am done. I do not check my threads until they say done or input in the corner. And if you're too lazy to
Speaker A
even pick which server to run, Maria has built an awesome feature for you in settings here.
Speaker A
Load balancing. You can now set up auto load balancing across all of the servers that you have connected in T3 code, so it will fire off the threads across your different machines. And once you work this way, your job becomes different.
Speaker A
You're no longer sitting there carefully babysitting the exact change. You are firing off different things you want to have done, and then you go through them when they are ready for your attention.
Speaker A
And here is the harshest reality I need you guys to get through your thick skulls, and this was so hard for me, that's why I'm being mean. It took me forever to accept this.
Speaker A
Agents can be multi-threaded. Humans are single-threaded. You cannot focus on two things at once. You cannot do it. This means our job's a bit different now. We are the bottleneck. So how can you get the agents to unblock themselves as much
Speaker A
as possible so you only have to come in when they need you? And how do you make it so they need you less and that you come in later? So look at this one that I filed cuz people were mad about how I
Speaker A
think about streaming. I wanted to move to a chunked by paragraph solution potentially. So I sent a prompt asking specifically how hard would it be to do this. If I really was committed and wanted this, I would have just told it
Speaker A
to go do it. But I wanted to see if there's any difficulty here before doing it cuz if it's any friction I just don't want this feature. So I asked and it said not hard and then it hallucinates
Speaker A
it about half a day all in one No, I don't care about what you think half a day of work is. It's kind of funny that these models still don't know how long work takes. Yeah, eventually it'll be fixed. Regardless, it said it
Speaker A
would be pretty easy. It said a bunch of variable names it seemed fine. I then said, "Can you build it for me? Once you get it working, file a PR and spin up an environment with Tailscale for me to
Speaker A
try." Now, and this is very important, imagine I didn't do the second sentence here. Imagine I just said, "Can you build it for me?" And then in 5 minutes it comes back, "Okay, I built it." Then I'm like, "Okay, can you spin it up on
Speaker A
Tailscale so I can try it quick?" I leave and then I get the ding. I go back and then I click the link and I try and I'm like, "Okay, this is great. Can you file the PR?" And then it does. And I
Speaker A
look and I'm like, "Oh, cool. You should babysit this, too." I probably should have included babysit here as well. So I'll do that now. Can you babysit this PR and make sure everything is good?
Speaker A
Cool. Now it will continuously monitor that PR and it's not my problem again. So your goal here is to make sure everything you need is there the next time you click. Ideally, you want to maximize the chance that the next time
Speaker A
you check that thread, you're ready to merge. And really think about that. What can you add? What context can you give?
Speaker A
What tools can you let your agent use to make it more likely that by the time you go back to that thread, you can merge the PR. And this is where one of my favorite T3 Code features comes in.
Speaker A
When you merge the PR, the thread disappears. Which means if you tell the agent, "You can merge the PR if it passes these conditions or meets these requirements," then you can send the prompt and never see the thread again.
Speaker A
I would honestly guess that around half my threads are archived without me ever seeing their final message because it doesn't matter. I see people confused about the agents stepping on each other's toes while they're working.
Speaker A
I thought we were developers, guys. We're not five coders. We made the right default in T3 Code for this for a reason. By default, [clears throat] when you start a new thread, it starts in a new work tree. Not going to pretend work
Speaker A
trees are perfect, but they are good enough. And now that the models are smart enough to deal with the weird that is get, they will fix your work tree issues for you. For example, one I run into a lot is I have a work
Speaker A
tree with a branch that makes a pull request, and then I want another agent, usually another model, to review it and give thoughts. And if those both are on the same machine, and they're both on work trees, they can't both have the
Speaker A
same branch, and then get freaks out. Previously, models were dumb enough they would get stuck there. Now they can figure it out. They have workarounds.
Speaker A
They make a new branch. They'll like pull it down in some other way. They'll make a clone. They'll do whatever they have to. It doesn't matter. I don't look. I don't care. Models are smart enough that if I give it a pull request
Speaker A
with a branch that it already has in a work tree, it'll figure out how to read it. Don't spend your time thinking about those things anymore. They don't matter anymore. This is another habit I've gotten into. By default, T3 code auto
Speaker A
settles threads that you have not touched for at least 3 days, which I think's the right call. I bumped it to seven cuz I usually have better discipline about clearing out my inbox, but I'm not good enough at it here,
Speaker A
which is why my sidebar is a little chaotic. I was exploring thread pop-outs. This one I actually do want to play with later, but not now, so I'm going to use this news feature to make this my problem later. You'll start to
Speaker A
see my mindset as I go through this. I want to make it so I don't have as much stuff trying to take my focus. This is why I've done some strategic things with the UI in the sidebar.
Speaker A
Notice that the ones that are working are semi-transparent. I wanted to hide working threads, but nobody would let me. So, I made them more transparent so you don't look as closely at them. I like this a lot. It
Speaker A
has made it much easier for me to only prioritize things that need me, things that are done or things that are waiting for input. I also really, really, really want to emphasize there are few worse uses of your time than watching an agent
Speaker A
as it works. If it works and it succeeds, awesome. Merge the code. If it works and it fails, look, maybe read the reasoning trace or even better, ask the agent why. And this is a good opportunity to pivot into the next
Speaker A
section here, which is using your tokens poorly. I know this sounds silly, but I'm actually going to be framing these things in ways that could genuinely be helpful. God, I One last thing I want to crash out on a
Speaker A
little bit. You can censor the chatter if you prefer Jeff in the edit. I see a lot of sentiment like this still. Like, if it works, that's fine, but what if it installs some malware on the side?
Speaker A
I need to be so real with y'all. If you think your agent is more likely to install malware than you are, then you are dumber than your agent. If you are smart enough to realize that won't happen, then you're both smart enough to
Speaker A
not install malware. If you think the agent is more likely than you are, you have installed malware recently and you probably have someone on your machine.
Speaker A
You should consider a new Windows install because I know you're also using Windows. Let's be real here.
Speaker A
Anyways, we're going to go back to talking to real engineers cuz that's what we're here for. Let's talk more about using tokens poorly.
Speaker A
One of the things I want you to think about is when you face a problem or have a question or uncertainty about something, how can you use your tokens instead of your brain? When I have a pull request that is hard to parse that
Speaker A
a teammate filed, how can I use tokens to figure out what it is? If I'm curious what's changed in the orchestrator V2 rewrite that Julius is working on, I could ask him, but he's busy. He's prompting. I could rather just ask my
Speaker A
agent to get me the info on what he has been working on. If I've lost track of all my pull requests and I don't know what I should focus on today, I ask my agent to go through them and find
Speaker A
something that I should be more focused on. If I can't find an email for one of my five email inboxes, I open chat GPT and tell it to go find it and it does.
Speaker A
If I don't want to sit there and download all my medical records for my surgeries, I don't ask my assistant to do it like I used to, I ask Codex to go through my inbox and go through all of
Speaker A
my dashboards for my medical on my browser and get it all for me. And what's even more fun is I would have my computer going through and spending 30 plus minutes collecting all my medical records, and while it does that, I go
Speaker A
back to T3 Code and I fire off four more prompts for things that I notice that are annoying me or for a feature I want to iterate on. And every couple minutes, I go back, I take a quick look and I
Speaker A
see, "Oh, these things are done. These things are still working. Cool." I really need to finish the durable objects if it's going to save us a load of money if I can get it right.
Speaker A
This is how I think about it. The issue with the other tools that you can use for this is that the sidebar doesn't help you prioritize the work that you're doing and the work that is done. It doesn't make completed work go away. The
Speaker A
search isn't trustworthy enough to get old work back, and you can't use the same sidebar across multiple machines.
Speaker A
Like, of the threads here, this first one is on one of my MacBooks. This one is on my server Alvin. This one's also on Alvin cuz I just spawned it. This one was too cuz I just spawned it. This
Speaker A
one's on BB1. This one's on my MacBook. This one's on this MacBook. This one's on the other MacBook. This one's on my main cloud server. They're all on different boxes and it doesn't matter.
Speaker A
They're all on the same UI and they're all on my phone app, too. I will say that right now, sadly, if you're not using Tailscale, you'll only be able to connect three devices at once with T3 Code for free by default. We're working
Speaker A
on it. I want to bump the number a bunch, but since you're already going to be using CLI proxy, that means you're also going to be using Tailscale. That means you don't even need to use T3 Connect. You can just connect directly.
Speaker A
This all works without using our servers at all. T3 Code is fully free and open source. It fully supports your subscriptions. It works incredibly with CLI proxy. So, if you have a couple boxes that have Claude Code and Codex,
Speaker A
and you have them CLI proxied, you go to one of those boxes. Here, I'll show you just how hard it is to set up T3 Code on a new box. I don't have a new box to set it up on. All my boxes are configured.
Speaker A
Let's say they weren't. You go to the box, you SSH in, you're on NPX T3 connect.
Speaker A
And then you click the link. Oh, email leak, great. I should work on that, but yeah. You click the link that comes up, you sign in, and now as long as you're signed in on the client, on the website, the
Speaker A
mobile app, and then once you're signed in on the mobile app, the website, or the desktop app, you can now control that machine. You don't have to install anything, you just need Claude or Codex, ideally with a sub or a proxy. You're on
Speaker A
the T3 connect command, and now you can connect through our layer. Or, you can do NPX T3 serve, and it will instead let you use Tailscale. We even have a {dash} {dash} Tailscale built in directly. Or, crazy thought, I don't know, kind of
Speaker A
risky here, but we are in the use your tokens poorly section, you can just tell your agent to go set it up. You have a new server, and you're already building a repo that manages your servers, you can tell Claude or Codex, "When you set
Speaker A
up that server with Claude and Codex, you should also probably set it up with T3 code so I can connect remotely over Tailscale." And it will just do it, and it will just work, and it will be just
Speaker A
great. I'm going to frame the stupid prompts a little silly. Have you ever found yourself Google searching, "Where did I leave my keys?" because you got so in the habit of Google searching things that you just default to that? If you
Speaker A
don't find yourself doing that with agents, you're not offloading to them enough. You need to get to the point where you ask a thing to agents that it obviously can't know, and you feel silly for asking. If you don't find the things
Speaker A
you're asking the agents a little bit silly, then you're not asking it silly enough and you're not pushing the limits here yet. Here's a fun example. I have a lot of side projects, and I had a bit of inference to burn the day I sent
Speaker A
this. So I asked, "I want you to go through all my GitHub projects as well as unfinished work on this machine using lots of sub agents to help me figure out what I should be putting more time into.
Speaker A
What are some of my ideas and side projects that seem like they're more in demand now that I might have foregone or that are worth bringing back and finishing. They might be half-baked or projects that have been abandoned for a
Speaker A
while and deserve another pass. Go through all my work on this machine and GitHub and find things I should potentially revive. This is a shitty prompt for a shitty problem. None of this matters, but I feel bad when my
Speaker A
limits reset and they weren't at zero. So, I was looking for things to burn them on so that I could have them all dead and not feel bad when they reset.
Speaker A
And it found a handful of my things that it thought I should prioritize more. I think that is fun and cool. And I already decided what I want to prioritize, so now the next step, snooze. Well, subtle. I don't need to
Speaker A
snooze cuz I don't want it back. Context usage breakdown, I decided what I want to do on this one and what I want to do is kill it. So, we're going to go here and I'm going to close it.
Speaker A
And I'm going to archive it cuz I don't care anymore. Show work tree creation progress. I don't know how far we got with this.
Speaker A
How far along is this work? Is the Tailscale dev server still up? Spin it up if you can. I'd also love for you to file a PR and babysit it to make sure everything is good to go. And then I'll
Speaker A
come back later. Here's one that I forgotten about cuz it was a while ago.
Speaker A
It says three days, but I'm pretty sure I started this way before then. So, I'm just going to ask. I'll be so real, I haven't kept up with this workflow in a bit. I have no idea what the state of
Speaker A
this is or what the value is. Can you give me a rough idea of where this PR is at and why we should merge it? This was going to fail cuz I'm still mostly out of Astro. It might route correctly. I
Speaker A
have a weird routing issue right now. We'll figure it out. These two are still going. That means that they are legitimate prompts doing legitimate stuff. Awesome. I will come back to them later when they say done. I would not have ever clicked these two if
Speaker A
I wasn't making content because I don't care when they are still working. Here's a real example of what my sidebar looks like when I'm actually working. Do you see how many threads I have here? There were even more below.
Speaker A
1 2 3 4 5 6 7 8 9 10 and then a monitor. There's at least like two or three more underneath that. And across my five Claude Code subs, I was able to do all of this and still have some usage left
Speaker A
over. Not bad. And after this, I went to bed and I woke up the next day, I went through them one at a time, made a decision. Do I want to merge this? Do I want to do a follow up?
Speaker A
Do I want to get rid of it? What do I want to do? I went through them one at a time, did all that and after, I noticed some other things I wanted to fix. So, I spun up more threads to fix them and I
Speaker A
went back and looked to see if everything was done and went through and clicked all the ones that were done and did it. I really have been treating this like an email inbox now and it has made me so much more productive. So, I just
Speaker A
said they're confused cuz the threads have only been running or the threads only ran for 20 minutes.
Speaker A
I had spun them all up within the last 20 minutes. Some of them took 30 minutes, some of them took 18 hours.
Speaker A
That was me trying to get everything going before bed. That wasn't work going. Those all said working. I had 12 plus threads that had been working for at least 20 minutes. Most of them kept working after. All of them actually did.
Speaker A
I don't think any of them were done even within the next 10 minutes. So, as silly as the using your tokens poorly section may have been, I do want you to get into that mindset. Especially when you notice
Speaker A
a limits coming up and you haven't burned it, find silly things to have your agents do. You'll learn things from that. You'll learn a lot more than you expect from that. I am still amazed at how good agents are at triaging large
Speaker A
numbers of PR's, at finding things that I should bring in and merge that I might have missed or forgotten about otherwise. You'll be amazed at how much random you can find. And here is where we get into the last section. I've
Speaker A
touched on a good bit of this here, but I want to really, really emphasize this part.
Speaker A
Using your tokens while sleeping. I mean this both literally and metaphorically. If you've ever had the feeling of, "I want to fix this right now, but I have a meeting coming up or I have to leave the office soon, so I'm not going to send
Speaker A
this message to this thread because I have to close my laptop." I get it. I was like this not long ago.
Speaker A
You need to move your dev work remote. You need to pick up a crappy little Linux box. You need to take some old computer that you haven't used in a while, install Ubuntu on it, set it up with Tailscale, install CLI proxy on it
Speaker A
so that it's always coming from your residential IP, and then you use it as a box that your agents can work with. Now, connect it with something like T3 Code, and you can fire off on that box and not have to think about it. I see
Speaker A
people asking if they don't have a computer, what's the most budget option? The most budget option by far is to talk to your friends and family and find somebody with an old laptop or desktop that they'll give you for free that has
Speaker A
at least 8 gigs of RAM and at least what or at least four cores, and you'll be able to do some real work there.
Speaker A
And I'm already seeing some very, very stupid replies. If you think a VPS is the solution here, I might have to get into selling VPSs cuz those sound much more lucrative than bridges nowadays.
Speaker A
Need 32 gigs of RAM and 16 threads? You can get it on Hetzner. It's only $275 a month. That's a great deal. Or you can waste all your money spending $700 on a box that has 32 gigs of RAM and a 1
Speaker A
TB drive. That would be terrible. That's such a waste of money. That's like 2 and 1/2 whole months of renting a worse computer on an IP address that'll get you banned from Claude. Why would you ever buy hardware that you can run on an
Speaker A
IP address that won't get you banned when you can rent something for 2 months that will get you banned? Hopefully you understand sarcasm because there is pretty much no reason to do cloud servers unless your grandfather into a
Speaker A
good deal. Theo, you're wrong here because insert something stupid. It's actually very funny someone said that in chat right after somebody complained about electricity costs for a computer that pulls conservatively 60 to 80 watts of power at worst. And
Speaker A
remember, it's not doing inference. It's doing code. It's going to be using one or two threads most of the time.
Speaker A
Let's go hop into my massively over-speced server to take a look. I have six threads or so running on Alvin right now. And let's take a look at Btop quick. Huh.
Speaker A
You can't see because my face is covering it. This is a box I have at least six threads running on right now.
Speaker A
It's got 32 cores because I got it for a good deal. You might have noticed it's not using them very much. It's rounding to 0%.
Speaker A
So once again, you don't need a lot, you just need enough. Sadly, that box I showed earlier with the GMK tech has gone up in price because my tweet caused them to sell out. You can still hunt find something, but I would highly
Speaker A
recommend you do not do a Mac for this because Mac OS is a show for parallel work. If you are looking at the numbers of sharing here and saying that's not right, I run three agents in code X and my laptop's overheating.
Speaker A
You're right, your MacBook is overheating because your MacBook has a file system and an even shittier security policy that is hard bottlenecking how much you can run at once. Move to a real OS like Linux and you won't have these problems. You
Speaker A
don't need a lot and I bet my ass if you can find your mom or uncle's old PC that's got 4 to 8 gigs of RAM in it and an okay-ish chip and you flash Ubuntu on it and you plug it into your router
Speaker A
somewhere, you're going to have a better experience than you have doing dev work on a Mac like immediately.
Speaker A
And if you live in an area that isn't San Francisco, well, I'll be real. If you live in San Francisco, hopefully you can afford compute and if you can't, get out. The city's too expensive. So if you live
Speaker A
anywhere else, Facebook Marketplace will be your friend. I bet you can find a surprisingly good deal. And any old laptop with any real RAM, you're going to be fine.
Speaker A
And once you make that change, suddenly you're not bottlenecked the same way. There will obviously be edges like if your agents run CI on the machine locally and it's a rust project that saturates all your cores when it
Speaker A
compiles, you'll run into problems. You're an engineer though, and you have tokens. Burn your brain and your tokens to solve the problem. Maybe you move CI to only run on GitHub, or you use one of our awesome partners like Blacksmith or
Speaker A
Depot. Maybe you have a different server that runs the CI. Maybe you set up a queuing system where the CI gets triggered one after another, so you don't have five rust compiles destroying your machine at once. You're engineers,
Speaker A
you can solve those problems. I don't want to just tell you all the solutions to those things, because your problems will be different from mine. Apparently, I succeeded with my goal of moving Maria to the Blacksmith CLI in order to let
Speaker A
her run the CI via CLI instead of it running on her machines directly, cuz she was a bit RAM constrained. And now she's way less RAM constrained. And that's how we end up with posts like this from Stephen that Maria shared.
Speaker A
Request, can you make T3 code less productive in experience so that he burns fewer tokens? This is where you want to be.
Speaker A
And I've one last thing I want to lean into with this bit. This shit's so fun. I know a lot of engineers have felt like the thing they loved died, and that everything has changed too much, and now engineering
Speaker A
isn't fun anymore because you send off a prompt and then you sit there and watch it make a bunch of mistakes and then get annoyed that the code doesn't work and you could have wrote it faster yourself.
Speaker A
I still have that experience. The only difference is I don't watch the thread, I go do something else. And then another thing, and then another thing. Maybe I spin up three threads and then I go to dinner with my friends. Maybe I spin up
Speaker A
seven threads while I am waiting in the lobby to queue in a game. Maybe I go skate while I have some things running and when I'm sitting and recharging, I pull out my phone and check on two of
Speaker A
them quick and kick them off to go do a bit more work. And the fun isn't in those parts, I'll be clear. The fun isn't that I'm checking my phone and sending prompts. The fun is this new type of engineering problem. I can now
Speaker A
justify spending more time micro-optimizing how and when I trigger CI. I'm thinking more about my cores and how they're being used. I'm thinking more about how I can work on the same thing eight times on one box and not
Speaker A
have to think as much after. And I love doing this. I find it so genuinely fun.
Speaker A
And I hope this video helps inspire some more of this fun for y'all because that's my real goal here. All of this chaos has made engineering the most fun I've ever had with it. And if you're not having fun yet, I hope this can help you
Speaker A
get there. And maybe just maybe you have reason to go spin up a couple more accounts and burn a few more tokens. And if you do want to throw some of those tokens our way to make some real
Speaker A
improvements to T3 code, we'll probably ignore them, but we might merge them. So, consider it. Oh, yeah. I did have this DM back and forth with Maria. I could probably scroll and find it, but I don't want to scroll through our DMs
Speaker A
publicly. I set her up to do remote stuff and she got a Linux box so she could do it at her place.
Speaker A
She hits me up in all caps, "Remote work is the best. This is so fun." And I said, "I'm sorry.
Speaker A
You're about to burn so many tokens." And if you've been struggling to hit your limits, I really hope you don't struggle anymore because this is so so fun. I think I've said all I have to on this one. I'm going to go kick off
Speaker A
some more threads and grab another drink. Enjoy the new world of agent maxing and get as many of these tokens out as you can before the subsidization ends, if it ever does.
Speaker A
Let me know if you want a video about that, too, cuz I have a lot of thoughts about how this goes long-term. But this wasn't about that. This is about actually using them and I hope it was helpful. Let me know and until next
Speaker A
time. Peace nerds.
Topics:Claude AIagent threadstoken managementAI automationAPI cost optimizationParallel searchsubscription strategyAnthropic ClaudeAI workflowagent scaling











