Skip to content

How to Build a Software Factory for AI Coding Agents

Explore how AI coding agents transform software factories with in-house and buy/build options, focusing on Boundary's architecture and agentic coding.

Ask about this video. Answers come from its transcript only — with the timestamp, so you can check them.

Generated from the transcript and can be wrong — check the timestamp.

Key Takeaways

  • AI coding agents can effectively replace traditional human-built software factories, offering flexibility between building and buying components.
  • Boundary's software factory architecture demonstrates a layered approach that supports mixing open and closed-source tools for enterprise needs.
  • Codex is currently favored for engineering and execution tasks over other AI models like Claude or Fable.
  • Token usage is a valuable metric to validate actual AI model utilization beyond hype.
  • Extensible and customizable software factories using agentic coding languages like Boundary ML enable advanced automation and orchestration.

What the video covers

  • Introduction to the concept of replacing human software factories with AI coding agents.
  • Discussion on the spectrum of building software factories in-house versus buying managed cloud solutions.
  • Detailed exploration of Boundary's in-house software factory architecture and its four enterprise software factory layers.
  • Comparison between open-source and closed-source products and how they enable mixing compute, dev environments, harness, and control planes.
  • Insights into the programming language Boundary ML designed specifically for agentic coding.
  • Analysis of market trends in AI coding models, including Codex, Claude, Anthropic, and Fable, with a focus on Codex's dominance in engineering tasks.
  • Practical examples of how different AI models are used for various tasks like UI work, code writing, and research.
  • Discussion on token usage as a metric for genuine AI model adoption and performance.
  • Technical details about software factory design patterns, orchestration, lifecycle hooks, and extendable software using TypeScript.
  • Engagement with audience questions and real-world implementation challenges in AI-driven software factories.

Answers

Questions about this video

What is the main difference between building and buying in AI software factories?

Building involves creating software factory components in-house, offering more control and customization, while buying means using fully managed cloud agent solutions. The best approach allows engineers to mix and match open systems flexibly.

Why is Codex preferred over other AI models like Claude or Fable?

Codex excels in engineering and execution tasks due to its ability to read longer contexts and produce more reliable code, making it the favored choice for most coding features despite other models being used for UI or generic knowledge work.

What role does Boundary ML play in AI coding agents?

Boundary ML is a programming language designed specifically for agentic coding, enabling developers to build and orchestrate AI coding agents effectively within software factories.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
You start with the normal software factory where a human builds the thing, and then you swap it out with an agent that builds the thing. There's a couple different options, right? You could like build everything in-house, and then there's like the other side of the
00:11
Speaker A
spectrum, which is like buy everything. In terms of buy versus build, where do you land?
00:16
Speaker A
The real answer is like you should have as an engineer choice and the ability to like work in open systems and plug these things together. So this was a really fun episode where we talked about software factory design patterns. We've
00:27
Speaker A
talked a lot about the broad software factory. And today we zoomed in on this part where it's like agents building the thing and testing the thing and getting feedback. We talked a lot about how the boundary in-house software factory
00:39
Speaker A
works and the architecture of it and contrast it to like sort of full stack cloud agents where you just buy everything and everything is managed for you. Um, so the architecture here was really interesting and then we kind of
00:50
Speaker A
dug into like the four layers of the enterprise software factory, what the stacks are and some actual like implementation examples of like how you can mix and match parts of this stack and how different like open-source and
01:03
Speaker A
closed source products like let you bring your own compute, your own dev environment, build or buy the harness and build or buy the control plane and orchestration. So a lot more detail there. We talked a lot in detail about
01:14
Speaker A
the boundary software factory. Super fun. And then as a little bonus at the start you can see our take on the Financial Times article of hey is OpenAI taking market share from Anthropic and why and why are we also
01:26
Speaker A
codeex build. So super fun episode, excited to get into it. Let's go. Hey folks, this is AI that works. We talk about AI that actually works beyond the demo for real production use cases, enterprise, whatever it is you want
01:39
Speaker A
something that actually works well and is reliable. This is all we do. We do a lot of systems design. We do a lot of context engineering. Um, the very first conversations that led to the term context engineering existing happened on
01:51
Speaker A
this show. I'm going to stop yapping. I'm Dex. I'm the CEO and co-founder of Human Layer. We build a multiplayer coding agent workspace. And Vibb here is I'm Vibb. I'm one of the co-founders at Boundary and we build a programming
02:04
Speaker A
language designed for agentic coding. Amazing. Check it out. boundaryml.com. Is that right? Did you guys not get baml.com? Uh sadly, Bank of America has a slightly bigger war chest than us.
02:17
Speaker A
We'll get it. That's your goal, right? It's like by the end of the We are slightly bigger than them on SEO for Google, though. So, we do beat them in almost every region. Yeah. So, if
02:27
Speaker A
you Google Bam, we're now rank one or rank two and rank three and they are on like page two, I think.
02:32
Speaker A
Um amazing. We have go Google BAML and help their SEO even more. Click the Boundary ML website. Go check it out.
02:39
Speaker A
Programming language for agents. I've been building a bunch of internal tools in it. It's a lot of fun. You can try Human Layer at humanlayer.com. Um, someone asked me today about this article. Did you read the Financial
02:52
Speaker A
Times article about how Fable and Opus are actually like losing market share and Codex is taking over?
03:00
Speaker A
So, like if you think about it, this is kind of interesting. I mean, I'll tell you honestly, like my mom knows nothing about Claude or Anthropic as a company, she only knows ChatGPT.
03:14
Speaker A
Yep. If she ever wants to go ahead and build an agentic stuff, she would only ever use ChatGPT because it's the closest she can get. So, I suspect it's just a relationship of building a consumer brand versus an enterprise brand. Like
03:27
Speaker A
no matter what you do at some point if more people get exposed to your topic the consumer brand will grow bigger and that's I suspect that might be it. I completely disagree. I think really I think I saw the change in like
03:44
Speaker A
December and January of this year where all of the best agent decoders, the ones that are like 6 to 12 months ahead of the like SF meta, they all switched to Codex. So, okay, that's a totally separate thing.
03:58
Speaker A
Like Codex is definitely I mean I switched to Codex, too. Like, and I think No, yeah, because it's better.
04:05
Speaker A
Well, I was actually hesitant for a long time because I could only get Claude to work, but for some reason I kept on seeing I was like I saw Theo switch to Codex, Dylan switch to Codex. A lot of these
04:15
Speaker A
people on Twitter that I was watching only use Codex switch. Peter, well, that's a Peter was early. I had a conversation with Peter Steinberger in October of last year.
04:26
Speaker A
Peter had totally different incentives. This was way before OpenAI or anything. He was like when he was I mean he was working on like 40 side 50 side projects. This is before OpenAI launch but I was like yeah we built this whole
04:37
Speaker A
research plan implement system for Claude and like it makes Claude so much better and he was basically his take was just like this was Codex five maybe, maybe I don't know if Codex is better than Fable back then or back then. No, it was
04:51
Speaker A
Codex back then is not better than Fable now, of course. But he's basically like all that research upfront [expletive] that you have to tell Claude to do, you're just compensating for the fact that you're using a worse model and if you like you
05:01
Speaker A
ask Codex to do something, it will just read stuff for longer, which is exactly what you want.
05:06
Speaker A
Yeah. Um, so have you guys switched fully to Codex? We I so I I will use Opus models for like generic knowledge work type stuff like running the company and like managing pipeline and things like that.
05:20
Speaker A
Um, it's just like a little bit better at like tool use and things like that.
05:24
Speaker A
Like barely noticeable. It's just like what I'm used to. Um, but 90 or 95% of the features I write are written with Terra or Soul.
05:34
Speaker A
You know what's so funny? Yeah. Um, you know what I want to see being done is you know how people are like talking about startups like there's all these people on Twitter's like vibes slopping about how much revenue they
05:44
Speaker A
make and people are like just post the revenue or it doesn't count. What I kind of want to see with people talking about AI models is I kind of want to see people posting their token usage or
05:53
Speaker A
it doesn't count. It's like you can't say you're a Codex man unless you like actually like for example you can't say you're a Codex man until you post your token usage and prove that you're a Codex man because so many people just follow the
06:05
Speaker A
hype without doing it. So I can never tell like for you I know you actually do it like in your case, right? That's Oh yeah, show us.
06:14
Speaker A
I'll show you the token usage. There you go. Let's see. This is cool. Um, let's find Yeah. cost cost by model. I mean, so this is by spend.
06:27
Speaker A
Um, but yeah, we're using a [expletive] ton of Codex. There you go. Um, yeah. It's also interesting. It's not like you have stopped using the Opus models. It sounds like you use it more targeted though.
06:38
Speaker A
It's a more like things that are UI focused will still use Fable or Opus.
06:42
Speaker A
Like okay if I have Fable credits, I will let I will let Soul drive and then if I find a shitty UI, I will I will I will kind of like, okay, throw this out and then we'll we'll have it either invoke a
06:51
Speaker A
Fable sub-agent or I'll just start a new session with Fable. Yeah. I mean, what I find is actually like so I've been doing a lot of UI work myself as well, and for that work, I also end up using like the Opus models
07:02
Speaker A
because like I feel like for whatever reason, Codex sounds a lot more like AI than Claude does when it's writing content when I'm writing things with it. And like or at least maybe I just know how to tweak Codex better to sound more
07:14
Speaker A
like me or Claude to sound more like me than I know how to tweak Codex better.
07:18
Speaker A
Codex is kind of like but for engineering tasks it's not even a question like Codex is execution Codex is better.
07:27
Speaker A
Yeah. Whether it's Codex or
07:42
Speaker A
um I'll be like write the headers and write a little snippet like a a markdown comment about what goes in this section and the reference to where in the transcript it's from. And I'll have it write one little literally like two or
07:55
Speaker A
three sentences at a time. And I'll polish them. And in one session I will have it write the thing. I will change it. I will be like, "Read what I changed it to. Do the next section. Don't [ __ ]
08:06
Speaker A
it." And then like after four or five iterations, you start it starts in one session, you can start to give it a little bit of a sense of a tone in the patterns to follow.
08:15
Speaker A
And then sometimes I'll flush that out to a doc that is like here's all the anti-atterns that you do by default. But it's like my writing process is like I cannot stand like three pages of Claude or or codec slop. Like I just I won't
08:27
Speaker A
read it. I will. So, like I have to do it really small increments at a time, otherwise it's overwhelming.
08:32
Speaker A
Because otherwise, you just don't read. Otherwise, I'm just like, "Oh, I [ __ ] hate this. I might as well write this from scratch." And then I'm like, "Well, I don't have time to do that." Like, how do I get how do I get leverage from AI?
08:41
Speaker A
How do I get it to like kind of like unblock my writer's block without me actually having to like without actually like having to fix every single issue that it that it introduces.
08:50
Speaker A
I found that I can't really have my whole team switch. I still don't have a good way to convert any humans to like force them to go do the like to either model like to codeex or to claude. And the reason is like
09:03
Speaker A
I think a large part of how these AI models work today and these AI workloads work is like they're kind of personal to every human and like for example like obviously like we land on two different spectrums of
09:13
Speaker A
how we like to engineer. I I like to sit and think really hard about the problem for a while and like try and like remove as many roadblocks up front and you do a lot more iterative approach. Neither
09:22
Speaker A
approach is necessarily bad or good, but like clearly like if I and like it's also not like both of us live in one world at all times. We live on the spectrum. We just have a tendency on one
09:31
Speaker A
end or the other. But like if we for if you were forced to always plan upfront and do every single amount of planning up front and I was forced to always be iterative, we would kind of not enjoy the work as
09:44
Speaker A
much. And I think that's the problem with cloud and codeex is you as an engineer got used to a workflow. It's like [ __ ] now I have to change again.
09:51
Speaker A
[ __ ] I have to change again. I have to like get the tweaks and the nuance of the model right every single time.
09:56
Speaker A
Yeah, there's intuition that you have to build up and you need it's almost like a separate mode of being like you're like, "Okay, I'm going to be less productive and the the the benefit is not what I actually produce. It is going to be the
10:10
Speaker A
the intuition I build." Have you ever switched keyboards? Like every time you switch keyboards, like you're just like slightly less productive for a while. Like when I first used the Kinesis, I was like I literally just lost three months of my
10:21
Speaker A
life. Like because I could not type like I used to type. I just feel so dumb for that three months until I finally got used to it. And I think that's kind of what switching models and providers are
10:29
Speaker A
like. And yet I can't convince you to try Vim. Uh Vim is so hard. Vim is I cannot for the life of me understand HJKL. Like what the like this was designed for an era with no mice.
10:45
Speaker A
Can I show you? Can I show you something? Okay, while you do that, let me read some of the questions that are popping up. Um, Matas was asking, "Do you guys use Soul with sub agents or just soul
10:56
Speaker A
alone?" I use Fable for code writing, research, and web search sub agents and it beats Soul by a mile, but within the same harness um uh without using uh sub agents like I don't know how you make anything happen.
11:09
Speaker A
So, here is my most recent workflow for big stuff is like basically like I do some research and create a product outline back and forth and build a canvas that looks uh something like let me see if I can pull
11:22
Speaker A
something up. Uh oh, I updated um let me find a good PRD. Uh show me a workflow.
11:32
Speaker A
V3 task launcher. No, not this one. Actually, this probably has it in there. The Yeah, the PRD.
11:43
Speaker A
So, I will work back and forth and build like a bunch of like interaction mocks with the model. Like, okay, here's how it works. Here's the different parts of this. Here's the visual stuff that we need to add.
11:53
Speaker A
Not fully in like again like these are HTML mockups. They don't have to be faithful. It's just like get the direction and then I will have Soul and Fable implemented overnight. I'll literally queue up a fable session and I'll be
12:04
Speaker A
like implement and then I'll cue another prompt that is like okay review it. Okay, another prompt and it's like use the codeex CLI to have C codeex review it and then compact and then do it again and then do it again and basically like
12:15
Speaker A
queuing up like 15 prompts to let them just argue it out overnight and then I look at what they built in the morning and I look at at a human in the loop like polish okay make this look
12:24
Speaker A
like this this thing is broken and then once that's ready I'll go make the PR and if it's 20,000 lines of code I will actually work back and forth with the models are really good at this like how
12:33
Speaker A
do I break up this work into separate PRs where do the migrations go and then we merge like a one to 3,000 line PR at a time.
12:41
Speaker A
Okay. But I've Okay, so I've got a question for you. When you do this workload um what's your wall time of landing it?
12:50
Speaker A
Like all the work? It's It's hard to say because I am I will have like bursts of like two days of a lot of coding and then I'll have like four days of like doing like CEO [ __ ]
13:04
Speaker A
Um that's fair. The feature that this is about I think This was finished. This part took probably like a day or two on and off with doing other stuff. And then this took like a day. And then I ship the
13:19
Speaker A
first five PRs, maybe like one a day in addition to doing other things. And then Kyle is picking up the last three PRs and like polishing it and getting it ready to like turn on for customers.
13:29
Speaker A
Oh, you're literally the manifestation of a CEO that says, "I vibe coded something. Can my engineer go implement it?" That's what you're doing.
13:37
Speaker A
I implemented the first five pull requests. I laid the I laid all the GR and then we got to PR number five and he's like I don't like how this is architected. I'm like great. You can fix it.
13:46
Speaker A
But it was digestible, right? That PR where we had to like I'm just No, this is the thing though is like prototyping is great and like it's a great way to figure out what you want.
13:54
Speaker A
It's a great way to see what's possible. But like I'm not asking anyone to review a 20k line PR because it's it's like I'm not I'm not gonna like have the same bar of quality. The person who's reviewing
14:05
Speaker A
it is not going to have the same bar of quality and they're going to they're going to resent it. So you got to make these like models are really good at making these things digestible. We don't do this for everything. I'd say 90% of
14:14
Speaker A
what we ship is small enough that you could just do it in one pass. But if it's huge and it's prototype and it's experimental and it's just like I need to create the vision for what this is beyond just those HTML mockups, then
14:25
Speaker A
it's super valuable. Yeah, I actually agree. I find it's very like kind of what I'm hearing you say and I'm going to repeat it back to you and see if it this is what this is what I heard
14:35
Speaker A
which is you basically go through you get an idea you spend a lot of time on like model auto looping to produce a pretty good spec. Then you implement the whole spec end to end and it produces basically a slot PR.
14:48
Speaker A
But the slot PR allows you to decide whether or not the features are exactly what you want them to be and you keep iterating on slot PR.
14:56
Speaker A
Then you basically this is pretty cheap. This this this part is very cheap in terms of human time and we have subs and it's like we use those tokens anyways.
15:05
Speaker A
Um but the engineering happens here. [clears throat] Yeah. Yeah, but once you have the slot PR, then you say, "Okay, this is a slot PR. Use this as the spec and say that you were going to implement this from
15:14
Speaker A
scratch. How would you redo it?" And like what are the faces down into it pieces?
15:20
Speaker A
Yep, that's right. It's like how well it's like how do you break this down into digestible chunks and like what are things that we can incrementally ship to users that maybe add incremental value or uh or at least like easy to digest
15:31
Speaker A
and review. I want to show you something really cool actually. Uh this is kind of something uh I know I know we were talking about a bunch of stuff but I actually um so we've talked about this a little bit
15:44
Speaker A
which is on the idea of like um okay I'm going to screen share. We've talked about like uh our one of the things that we're running into now uh for context is like we're kind of running into this issue where we get a
15:58
Speaker A
bunch of like tiny PRs and like the amount of tiny work is like overwhelmingly large. Uh and each of those tiny work items will improve the system. It just takes a while to go do all of them. So like we've been uh we've
16:09
Speaker A
been trying to make our software factor a little bit better. And I wonder if you've seen this kind of thing before and if you guys have considered this where basically like users report feedback.
16:18
Speaker A
Yep. And every single feedback that comes in, we basically like the idea is you can't trust any feedback ever because all feedback is basically untrustworthy because users are reporting stuff who knows how. So we like immediately check if the feedback
16:31
Speaker A
is already fixed on the latest version. And if it is, we can just notify that user really quickly. If it's not, we basically try and search if it's like an existing repro or new one.
16:40
Speaker A
And these are like basically like we put as like percentages of like how correct is every single subsystem meant to be here. like this this will likely be wrong at some cadence because like you know like you might like maybe this
16:51
Speaker A
feedback report in like you think it's fixed but it's a flaky issue dduping is going to be buggy like I've never seen a good dduping agent for this and the key ones are bad agents are bad if it ddupes half of the things then
17:02
Speaker A
you've saved yourself a lot of time exactly so then you do this then you create like a really really good repro and I think we can spend a lot of engineering to make this repro really really good uh just by like producing
17:14
Speaker A
like n others then you actually create an issue So the idea is like an issue isn't created from feedback like issues are created post feedback and then you just constantly check basically on every single PR every single one is run and we
17:26
Speaker A
just check if every issue is fixed always. It's very cheap. It's like CPU time uh to measure it. We don't have like web browsers or anything to go check. Uh for some stuff we do and that isn't and you just like keep basically like
17:39
Speaker A
prioritizing the work. You assign a human to it. You basically communicate with the system and you either like put a plan together or you like just fix it and then you just like do the whole merging system. But like we tried to I'm
17:50
Speaker A
have you I mean you guys have built tried building something like this before right?
17:54
Speaker A
Where was the bottom? This is I mean like you have at the very top an important piece of text which is for issues with clear code repros.
18:03
Speaker A
Uh where do I where do I have that? Yeah. Yeah. I mean at the very top above feedback reported like if it's easy to reproduce then this whole thing works. If it's hard to reproduce, then uh then you need
18:13
Speaker A
something else or you need you need to send an agent to see if it can create one.
18:18
Speaker A
Yeah. Yeah. And that's kind of what the whole like creator repro is cuz like I don't trust any feedback files by users.
18:24
Speaker A
So we kind of just need to spend a burn our own tokens to try and produce a clear repro and you can't produce for clear repro then you assign a human to it. So like this is the thing we've been
18:34
Speaker A
working on internally and like it's like the next version of agent tries baml but for like human file feedback and I was like I was going to say this this feels shaped very similar to the infrastructure you wrote for agent tries
18:44
Speaker A
BAML and now it's just like cool let's add another like input stream which is human tries BAML or humans agent tries BAML.
18:52
Speaker A
Yeah, it's that one more so than [laughter] anything else. But it's like it's fascinating because like this PR review part is actually the thing that we're really trying to bottleneck and like I actually learned this concept from you when we were talking about this
19:03
Speaker A
in the team. What we were talking about a lot is like how do we push the human time as low as possible.
19:09
Speaker A
So you'll notice like a large part of this is like assign a human to this like a shepherd is like down here and like really this is assignment but then we actually ping the shepherd when this happened in Slack. We're like hey
19:22
Speaker A
here's a bug fix. Here's the repro. Here's the bug fix. go read the PR or this is really hard. Here's a plan. Can you work can you either like download the plan into your local agent or like tell me this plan is good and I'll go
19:34
Speaker A
implement it. What's the um what's the uh accuracy rate of your classify issue difficulty?
19:41
Speaker A
Cuz that seems like an is that a human task? No, we're going to do an agent task and I think we'll actually what we're going to do with this is actually a twofold idea and like we're still playing around
19:51
Speaker A
with this. So, we're building like evals for this actually internally. Yeah. Um, and now we have enough issues as a database that we actually like have eval which is kind of nice. Like our first PR just went out and I think where they go
20:06
Speaker A
like what's cool about this approach that we're doing. By the way, I only see your Excal tab. I don't know if you're trying to show us something else.
20:12
Speaker A
Uh, let me pull this over here. What's cool about this PR for example that we're making now is all these PRs.
20:19
Speaker A
This is the one rule that we have which is like you have to push repros onto here.
20:23
Speaker A
Yep. So you actually have to give us metrics for every single version of this pipeline. So then we can decide whether or not it's good or not. And like this is the dduping PR for example. Like obviously the dduping PR we know is low.
20:34
Speaker A
It's lower than we expected. So like we want to do it even better. Yep.
20:40
Speaker A
Right. But like Yeah. So we don't actually know what the XT rate here is yet. So, we're still going to be playing with this um to figure out what the accuracy rate is, but I'm hoping we can get to at least
20:53
Speaker A
like 60% correct and that is enough. And we can measure it always by lines of code. If the fix is more than 60, if the fix ends up being mclassified as a small issue and ends up being 500 lines of
21:04
Speaker A
code, we just we can just force classify it as the other one. So, it's not that hard.
21:09
Speaker A
Yeah, that makes sense. Um I'm going to show you one. None of this was what today's ep Oh, one more. Let's go.
21:16
Speaker A
So, uh, if you come, we got a little distracted as as you guys could see from the chat. If you type Vimtutor, it is a interactive guide that teaches you Vim. And if you do this for 10 minutes a day, uh, you're welcome.
21:31
Speaker A
No, no, no. Amigos, we're we're we're in the voice era. You don't need Vim. You just need voice.
21:38
Speaker A
Voice tutor. Um, voice t. That's what [clears throat] I need. I need a voice that tells me how I should go talk to talk to other agents.
21:46
Speaker A
Make you more clear. Uh yeah, me and Kyle were joking about like giving the agents a/ bro skill that they can give back to humans of like that was super unclear. I need you to like uh speak more like articulate yourself more
21:59
Speaker A
clearly. Honestly, that would probably fix context problems more than anything else would. It's probably a good idea.
22:07
Speaker A
Yep. Cool. Okay. I want to talk about today our topic which is software factory design patterns. So we've talked a lot about the software factory in various different versions and like cool you start with the normal software factory where a human builds the thing
22:22
Speaker A
and then you swap it out with agent builds the thing and you have all these pieces that like make up your your your software factory and then you make the agent test the thing and then like this part is now fast but the review part is
22:35
Speaker A
still slow and like we've talked about this a bunch of times. Um, I want to come and talk actually a thing that we haven't really done is like zoom in on this part of the factory. Um, which is
22:45
Speaker A
like agent builds the thing. And the way that I think about this is basically like there's a couple different layers.
22:52
Speaker A
We've talked to a lot of different teams and we'll talk about kind of like there's there's a couple different options, right? You could like build everything inhouse, uh, which I think is probably what your team has done. Um and then there's like
23:05
Speaker A
the other side of the spectrum which is like buy everything. So build everything in house might make like what do you guys use for your like um sandbox execution like where where do your cloud sessions run for like agent tries?
23:19
Speaker A
Uh we we buy MacBooks and Mac minis. Okay. So your compute is like agent sessions run on MacBooks and like basically when an agent session runs on demand like it I imagine it sets up some kind of developer environment.
23:38
Speaker A
Uh yeah and stuff like this. Yeah. But it's all kind of already pre-installed on the MacBook. So like part of this is like speed is why we have this cuz like this core stuff is preset up and there's no boot up time in that
23:51
Speaker A
case. Yep. Right. Um, and that's kind of nice. And then the other reason that we use MacBooks is because like we just had some extra MacBooks lying around.
24:00
Speaker A
Yep. And then for the harness, you use cloud code primarily. Cloud code uh cloud code and codeex.
24:08
Speaker A
I don't know. I don't know what the latest one is, but it's agnost. It's like we it's just like CLI you swap out.
24:14
Speaker A
Yeah. So you use the CLI with some sort of some sort of headless mode that like streams out JSON or something, right?
24:20
Speaker A
Yeah. Exactly. Uh we use headless mode that stream.json via the file. It dumps all the chat logs like the chat transcript, you know, like the JSONL file. So we just read that and then Yeah. So then you just tail
24:34
Speaker A
tail the JSONL file or you like you scoop it up once the session is done.
24:39
Speaker A
Yeah, exactly. Um cool. And then like do you have customizations that you add on there? I call it like the outer harness, right?
24:47
Speaker A
So you have like skills you're injecting. Do you inject any MCPS or anything like that?
24:51
Speaker A
We actually have like actual while loops that we have that we've built around this. So like for example to drive things to completion.
24:58
Speaker A
Um Okay, cool. So like like for like we have a pull request bot like code rabbit that we use.
25:04
Speaker A
Uh we want to merge after code rabbit. That's a while loop. That's like a hard-coded Y loop that actually says this. And like if the Y loop hits like max iterations like three, then what we do is like but then we boot
25:13
Speaker A
it out and we ask a human to take a look. Cool. That makes sense. Um uh we don't merge automatically. Uh we do have a human hit merge.
25:28
Speaker A
Yeah. Yeah. Okay. But it's like fix code rabbit until mergeable. Yeah. Like basically do not notify a human until code rabbit is happy.
25:37
Speaker A
Uh and then we have or until three iterations get hit. Yeah. And then uh we have another thing that we're actually about to add now in this new loop that I built which is like our CI/CD takes a long time. So we want
25:50
Speaker A
a human to hit that they're happy and after a human hits that they're happy this agent is supposed to just constantly merge against our main branch and keep it up to date so it can just mer drive it to merge.
26:01
Speaker A
I see. So you're merging from main into the other thing and like basically looping this part.
26:06
Speaker A
Yeah. So after the human is happy it's basically like I I want this to merge.
26:10
Speaker A
just babysit this thing until it merges in cuz we have a merge queue and like sometimes conflicts happen etc.
26:16
Speaker A
So this is the last part that we're adding now and we think this is just going to help. It's basically just a better a it's an agentic merge queue.
26:22
Speaker A
That's how you can think about it. This Yep. Okay. And where does this run?
26:26
Speaker A
Um and actually all this is uh going to be it's the same machine. It's the same thing right?
26:32
Speaker A
Um so does this run like so this is this is separate from from agent tries BAML, right? This is separate from the factory stuff. This is like running on someone's workstation.
26:41
Speaker A
No, this actually we don't want it to run on your workstation. We want it to run on a non on a on a work on a like a sandbox type machine.
26:49
Speaker A
Okay. So, do you run it in GitHub actions or does it run on the same MacBook where the agent runs? Everything is running on the same MacBook.
26:56
Speaker A
Yeah. Okay. So, you just have a pool of MacBooks that run any generic agent task.
27:02
Speaker A
Yeah. the agent loop for example. I bet that one could run as like a pure web layer because like GitHub is automerge.
27:09
Speaker A
So we could just do like automerge and then go do it and then but right now like that's kind of how we're thinking about it.
27:15
Speaker A
Okay. Um and then like my last question is like okay so you have this MacBook and then you have you know your you know Rust tool chain etc set up.
27:26
Speaker A
Yeah. Uh, and then you have your CC/codeex like your like inner harness. You have your like outer harness. And then you basically have like I guess like how does how does this system know that there's a new thing to work on? Is it pulling some
27:48
Speaker A
queue or like there's some there's some way to turn this it runs a web server on it locally and then we use some like herder or something. So, we're running like a web server.
27:59
Speaker A
Okay. So, there's a there's a web server that you can on each MacBook. Yeah. Yeah. And then there's like another way that we I forgot the app we use. We use some app that basically makes that app that makes that thing
28:11
Speaker A
like uh accessible like a tail scale or enrock or something. It's not either of those, but it's something else. I don't remember it off the top of my head. I would um but we use something that makes it basically
28:23
Speaker A
like behave like a remote server that's like deployed somewhere. You have some sort of like privately accessible like thing you can open in a browser.
28:31
Speaker A
Yeah. Well, it's it's a REST API. So like there's no we don't build a UI for it. I mean we could but like we could sort of react off of it but like yeah it's a rest API.
28:40
Speaker A
Okay. So what sends requests to this? Uh now we have yeah well now we have another web server that we actually host that's a pure rest endpoint. So now there's a web app that we build.
28:56
Speaker A
Um yeah, and this web app is the thing that like listens to linear, listens to Slack, listens to kind of everything and that communicates with these.
29:05
Speaker A
Okay, cool. So this is actually like on on the host. Yeah. Yeah, exactly. And then you have basically a cloud web server that receives like web hooks from Slack and Linear.
29:17
Speaker A
Yeah. And GitHub and all the other stuff. Yep. Okay. Yeah. And it also has its like own endpoint that you can communicate with if you want.
29:27
Speaker A
Yeah. So you have this thing as kind of like what I would call your like dispatcher.
29:32
Speaker A
Yeah. Yeah. It's also the thing that has a database and everything on there. Yeah.
29:37
Speaker A
Um so like so yeah this is the this is basically what we call like the orchestration layer. Um, and basically the the the kind of thing I was like want to get across is like you can buy all of this,
29:48
Speaker A
right? If you use like cursor cloud agents or cognition or whatever, you get you get all of that, right?
29:55
Speaker A
Mhm. Uh, yeah. Every single one of these is viable. 100% viable. And so like basically what I what I get into is like what is the what I'm trying to like elaborate here is like what is the um what is the what are the building
30:10
Speaker A
blocks? What are like the core interfaces inside the software factory so that if you decide there's some part of it that's hard and you want to build certain things and buy other things like what does that all look like? And I
30:21
Speaker A
think I've sent you this picture before but I'm going to I'm going to draw it out here for the for the fans.
30:24
Speaker A
Yeah. Well, people here I'll let you draw it out and then I'll just chat.
30:30
Speaker A
Yeah. People are asking them questions. I'll just answer them along the way. What steps you use to run evals?
30:35
Speaker A
Honestly, the eval system is pretty simple and you're seeing that we're making more rigorous in the PR that I showed earlier where now we actually have like benchmarks that we're trying to set of like we download past issues.
30:44
Speaker A
We see if they reproduce and what the rate is. So like now we actually just build tests around them and we have like benchmarks for like what percentage of tests pass and we have like a pass rate and stuff that we're building out. Um
30:56
Speaker A
all the code is actually open source. You'll be able to go see it if you want.
31:00
Speaker A
Uh do you use containers with G Visor? Uh we do not use containers. I mean these are all trusted systems. So the reason for example that you saw earlier in my in my Excala draw that I was showing is like raw feedback coming from
31:12
Speaker A
a user is completely untrustworthy. We run our own prompts to create repros and those repros are things that we trust.
31:18
Speaker A
So we never execute raw code off the prompt up there. We don't want to. So this is why it's like slightly more trustable. So we can just run it directly on the machine.
31:27
Speaker A
Uh, I guess someone could try and Bitcoin mine it, but I think Claude Code and stuff are like smart enough to like not try and do that and we don't really accept remote code.
31:34
Speaker A
Uh, and then do we have a sandbox or Dro? No, we don't use any sandboxes or anything else. We just run it directly on the MacBook because again, like I said, it's trustworthy. So, we don't have any issues on this.
31:44
Speaker A
Yep. And so, like again, you can buy or build that. You could throw it in Daytona. You could throw it in your Kubernetes cluster. You could throw it ETP, Freestyle, whatever it is. Uh, but you need a place for the aven isolates,
31:54
Speaker A
right? Like some people have done this thing with jetpash where like the agent runs it doesn't even have a real Linux file system. It's literally just tools that look like a file system.
32:02
Speaker A
Yeah. And then Simon asked the interesting question like how do you treat issues that are dependent on external sources, data providers and such. The thing is like we don't um lucky for us we don't have that much of
32:13
Speaker A
a problem because we're building a programming language. So a lot of stuff is like either self-reroducable or not and it's very clear when it's not and if it's not we'll just bump it up to a human.
32:23
Speaker A
Like external data source will just go to a human and a human can make some educated decision on that kind of problem.
32:29
Speaker A
Yep. I think I think the folly that a lot of people make is they try and make a system that is fully automatic. That is much harder than getting a system that's 95% automatic.
32:40
Speaker A
So like that's how we're we're aiming for that latter bit. We're not trying to be fully automatic. We sadly, not sadly, I guess luckily will actually keep employing humans for a while because I like people and I like making friends
32:53
Speaker A
if for no other reason. Um, yeah. So, anyways, yeah, the the the comput stuff is basically like you can build it, you can manage your own compute, whether it's EC2 or your own like Kubernetes cluster or literally you
33:05
Speaker A
have a stack of MacBooks in the office like you can buy it and run it in your infrastructure. So I know like people like Daytona will like basically give you like kind of a black box and you get
33:14
Speaker A
an API and just be like hey I need a sandbox and they'll just spit it up for you but it runs in your AWS cloud or in your data center or on your like rack and office or you can buy where it's just like I
33:26
Speaker A
don't want to run any infrastructure. I have like here's my stuff and here's like the vendor stuff and I just send a request and all the execution happens like in their thing.
33:36
Speaker A
Yeah. Like basically it's like you can do as much BOC as you want. Yep.
33:41
Speaker A
Uh and this is like a very well-established domain. And I'm gonna ask you a question.
33:47
Speaker A
How do you decide how much BYOC you want? So this is this is kind of what I want to get at is like if you understand these like four layers or five layers of the stack, you can make the decisions
33:59
Speaker A
based on your needs. Um, and I think the most interesting part of this whole thing that I think is the most controversial is the idea of the dev environment. Um, because the dev environment has many layers, right? It's
34:11
Speaker A
like, okay, like what like language run times do I need to actually be able to uh like compile and test my software? There's like how do I like test it from the outside? So for BAML is very verifiable,
34:28
Speaker A
but if I need to run a web app, I need a way not just to like send work to this workstation, but I need to be able to go like look at what the agent is building and see a preview of it, right?
34:38
Speaker A
You have like a lot of a lot of larger teams have like internal services where like you might run like four layers of your stack when you're developing, but actually the app has like 50, you know, 50 plus services that it needs to run
34:52
Speaker A
and these are all shared. Yep. Uh, you know what's like shared dev? We call it like a dev cloud or something like this.
35:02
Speaker A
You know what's um Okay, I'll I'll let you keep going. Actually, I want to keep I want to see where this goes.
35:08
Speaker A
Yeah. Yeah. Yeah. Yeah. Uh just in the interest of time. Um, basically my take is like once you're going into the vendor's cloud, if if it's running here and you need to test it, like telling telling Daytona or whatever, hey, I need
35:22
Speaker A
to install like I need to use this base Docker image cuz it has Rust pre-installed at a specific version that we like is easy, but when you get to the point where like I need to like poke holes in my in my uh in my like cloud to
35:35
Speaker A
allow this running thing in someone else's compute cloud to access uh the shared things that my stack pack needs cuz we have a 100 repos and I don't want to run all of them in the sandbox. Uh this becomes like quite a bit of
35:49
Speaker A
friction and the the need to basically like debug oh the rust thing is not working on this thing because they're using like a fake Linux kernel that doesn't support this sys call or something. Uh basically like uh it it
36:03
Speaker A
can cause a lot of friction and my one of my thesis is like you are going to want to own this uh unless you're building like tiny little toy nex.js JS apps and you're like, you know, in the
36:13
Speaker A
like majority like in distribution set of stuff, you're going to start to hit friction unless you can own this in a really clean way. Does that make sense?
36:22
Speaker A
I agree. Cool. Um, the next layer is the harness. So, this is like basically and I I think we probably split this into like in yours we kind of split it into like inner hardness versus outer harness. you
36:34
Speaker A
have like the core cloud code or codeex or um this could be like or you can like or amp or devon or factory. You can like buy this from a vendor who builds a harness or you could like build your own
36:47
Speaker A
on top of like open code PI custom things. And basically like you can you can you can kind of say like okay cool there's the inner harness is really thin and we build a bunch of like outer harness
37:01
Speaker A
stuff or you can basically buy an inner harness that comes with a ton of stuff like a browser and testing and stuff. Uh oh my god I can't type.
37:13
Speaker A
And then you can have your outer outer harness be like very thin where it's like just your skills and stuff.
37:20
Speaker A
Does that make sense? Yep. Yeah. You have options here how you do [clears throat] compaction, how you do testing. Yeah.
37:28
Speaker A
Uh what's fascinating here is like I don't know if anyone in the audience has ever worked at a like Google or Facebook, but if you have then you might see a lot of this stack is literally something that you have probably worked
37:39
Speaker A
on before. And Google, for example, they buy you a really nice MacBook, but you never code on your MacBook. You always get what's called a cloud top. And every single time you get a new environment, you actually set up a whole new cloudtop
37:50
Speaker A
and you just set that up repeatedly. And Facebook's the same way. And like that's just how your default coding. And like same thing here, like when no one ever has no one ever has code checked out on their workstation because like they
38:04
Speaker A
basically don't want you to take the code outside the building. Well, that's one part of it, but also it's just like too big. Like you can't have that. And then the other thing that's really fascinating to me is like
38:14
Speaker A
whenever you make like web services and you go do them, you get default URLs that are attached to you. So like my uh uh my I don't I don't even remember what my LDAP was. I'm pretty sure it's like
38:25
Speaker A
what my user ID was. Like I got a subdomain that's like myeldap.c.google or some internal Google domain.com. I could send that around to people. So if I ever build like an internal UI, by default it's sharable in my whole
38:37
Speaker A
company if I want it to be. And you get you can just share that URL and a PM could go look at the web app or whatever it was.
38:44
Speaker A
Exact. And what that means is like any service I want to point to, I can have my own version of the service or I can have a hosted version of the service that I point to by default and it's
38:52
Speaker A
always kind of just works um in that way. And what's funny to me is it feels like we're all trying to reinvent that kind of golden stack where like you go in, you press a button, you get a cloud top issue to you, you can
39:04
Speaker A
start coding on it, it does the work, you can access it from anywhere. You can see all the code on it from the same editor that you're the same editor can just switch to like different machines really really fast.
39:14
Speaker A
And then um on top of that you get all these like nice little web systems where like it just points the right database, right URLs and like sharing and everything is provisioned. You can't share URLs publicly unless you
39:25
Speaker A
explicitly escalate them to a public URL that's sharable. So, it's internet only. Anyway, I just find it so interesting that like we're doing all this stuff and like there are a couple companies that have done this like perfectly already by
39:39
Speaker A
necessity of their work. But it's also it's like there's something to be said for like you might not be Google and you might be able to get away with basically like a full stack cloud agent that does all of this kind of like for you because
39:51
Speaker A
you don't need that flexibility. Yeah. Yeah. Yeah. Exactly. Um cool. The last piece here is I think the most interesting and underserved piece. I'll take the pitch part of this out but is basically the control plan, the orchestration layer. Uh, and so like
40:07
Speaker A
basically the idea here is like you want to be able to dispatch new work. You want to be able to see what's you want to be able to look at the session traces. You want to be able to look at
40:16
Speaker A
the plans and architecture docs that are being made. You want to be able to like schedule things to run every night or on a cron or in response to web hooks. You want to be able to like review and
40:27
Speaker A
iterate on the code itself. uh sort of like a PR shaped thing but like doesn't necessarily need to happen in GitHub.
40:35
Speaker A
There's like permissions and audit and who can talk to what and who can see which services and like managing spend and budgeting. And then there's like this idea of compounding engineering. I we'll start high level. We'll dig into
40:46
Speaker A
like whichever one of these you think is the most interesting. Compounding engineering is we talked about this on a an episode like a month ago of like how do we how do we build a memory system where it's like hey if all your
40:56
Speaker A
engineers across your whole team are yelling at Claude all day and they're all saying the same thing or codecs or whatever it is how do you incorporate that into your uh your outer harness basically.
41:08
Speaker A
Yeah. Um this is I I I mean if I were to go and write the layers of stack I could not come up with a better representation of the stack. Um there is there is one argument to say
41:19
Speaker A
that like the dev environment is actually part of the outer harness and it doesn't it doesn't sit underneath the harness. It's actually part of the harness. But uh I disagree.
41:28
Speaker A
We'll we'll we'll avoid that that semantic right now. I I I disagree. I think it's I think it's better to have the separate layer because it's like when I get a company laptop for example, there's a bunch of crap I have to
41:39
Speaker A
install. I just want that installed. Yeah. There. And like to me like what else goes in the dev environment is like identity provisioning like that all goes into the dev environment not the harness and that's why I think um
41:52
Speaker A
you mean like authenticating to internal services and integrations and stuff. Yeah. Just like identity attached to that environment. There's some sort and that implicitly gives you internal services. It gives you access to certain API keys, gives you access to certain
42:05
Speaker A
scopes like all the stuff you might expect. Um but that belongs in the dev environment layer. It's not in the infrastructure layer. it's not in the harness layer. Um because like I want my harness to not have to think about
42:16
Speaker A
certain provisioning. Yep. Um and one other thing I think is interesting on the dev environment side, then I want to flip it to the audience and hear like which parts of this stack are the most interesting to drill into. Um are you
42:27
Speaker A
familiar with the idea of like pets versus cattle? Okay. While while you describe a new uh concept to me, um I will read the comments. Describe the concept to me.
42:38
Speaker A
So yeah. So I have no idea what this is. basically this idea of like you kind of like have some number of server you like prod web01 you have prod web 02 you have like a set of servers that have like roles in your
42:52
Speaker A
infrastructure okay uh and like if one of them goes down it's like cool I got to go rebuild prod web 03 cuz the disc filled up or whatever it is okay what's cattle and cattle is basically this idea of
43:04
Speaker A
like you don't care if it goes away you've designed your oh yeah it's like web previews for example web previews on PRS or cattle.
43:12
Speaker A
Exactly. So it's basically just like web and then you have some like you know random UID or something. Yeah. Yeah.
43:19
Speaker A
Yeah. And then you just kind of have infinite of these and these are usually cattle is more like okay we have a script that can provision these. No one has to go SSH into the server and set it
43:27
Speaker A
up. Uh and like basically just on demand. So like I would say your MacBook setup today is actually more of a pet setup.
43:34
Speaker A
And I actually think this is a great place to start. I think you can get very far with like basically like if if you bought a new MacBook and you wanted to add it to the stack, you would go
43:45
Speaker A
someone would go like pull it up and manually set it up. You don't have a button you can push that makes it ready to receive like workloads, right?
43:52
Speaker A
Well, there's a repo that you download, you start running uh some start some script and it just provisions it automatically.
43:59
Speaker A
Yeah. Um but you do have to download the repo and go do that. We made it. It's not like someone has to go basically chef or you need the hardware.
44:08
Speaker A
You need the hardware to go run it. Uh which means and we own the hardware.
44:12
Speaker A
Yeah. So when I worked it, we basically like we built a system to turn everybody's workstation in like undergrad in college. We we worked it for the library and we basically built a system to turn all of the workstation
44:24
Speaker A
from pets into cattle cattle. And we had this like literal station in the basement where like whenever someone's machine got corrupted or got a virus or whatever it was, instead of going in and trying to fix what was going on, we would literally
44:37
Speaker A
pop two drives in the in the bay, clone from one drive to the other, bring them a new drive, put it in their machine, and be like, "Cool, go download your [ __ ] from the from the file share."
44:47
Speaker A
Yeah. Yeah. No, I I I agree that dev environment should be more cattle than pro uh than bats. 100%. I actually like when we build our dev environments and like we have a similar stack to you. I run some stuff in EC2. I run some stuff
44:59
Speaker A
in XC.dev. We actually treat it like pets for now because we don't have [ __ ] thousands of work workflows starting up and down.
45:07
Speaker A
I just leave one server running all day and when I need to run stuff on it, I run stuff on it.
45:12
Speaker A
When I need to authenticate to GitHub again, I just log in and reauth to GitHub. It's another great analogy for everyone else is like I would say like AWS Lambda functions or cattle. like an EC2 container you own is more pet than
45:25
Speaker A
cattle. It's it's kind of in the middle. Uh but it's it's kind of in that same kind of workflow.
45:30
Speaker A
Um we've got a couple questions uh in here that I want to go over.
45:34
Speaker A
Uh it's like um it's a weird network cluster. My factory ended up as a tail scale sprawl. Tunnels to extend. Uh tunnels are Yeah, I mean yeah, we don't use tail scale for that reason. Like it's kind of annoying. Um Sedarth has
45:47
Speaker A
some things. It's like uh I sort of understand how software factories can behave in a single repro. How do you deal with multiple repros when you have like multiple things to be completely honest sadly my recommendation is uh
46:00
Speaker A
you're going to be sad because you're not using a monor repo. So the best thing you can do and I know Dexter has helped teams that have tons of like micro repos teams with 80 repos and or 100 repos or
46:13
Speaker A
whatever it is. So like my recommendation would be like simulate a monor repo as much as you can and you can do that in many ways. You can get like a checkout of you can go to a folder where you sim link all the monor
46:24
Speaker A
repos in there. You sim link it and treat it like a monor repo. You can have an actual monor repo that actually like just contains a list of sim links to every single repo that you have and it
46:33
Speaker A
just like with sim links like cloud can traverse your workspace pretty easily. So if I come here and look at the code here we literally this is our template for this. We don't use this anymore cuz this is built into
46:44
Speaker A
human layer now. But uh yeah, we basically say like if you're using especially if you're doing research and you want to go find the stuff, it's like cool, you're in a coordination repo for multiple repos. They're all like one
46:54
Speaker A
level up basically. And so the way the way that you lay things out on disk is basically like source repo one, repo 2 coordination. They don't need to be in the same git tree. You don't need subm modules. You don't need sub trees. You
47:05
Speaker A
don't need sim links. As long as the model knows all of the things that might be in scope and you give it a like part of your workflow that is like go do research and understand what matters to
47:14
Speaker A
this uh this is like it's really not like a product thing. It's literally just like a cool just tell your model where the other [ __ ] is and like you still uh while you say that actually you still don't need one
47:27
Speaker A
like control file that tells you the organization structure and the layout system that you're going to have.
47:32
Speaker A
Yeah. Right. So this is like but like it doesn't have to be sim linked into every repo. You don't need every repo to have like a doubly linked like two-way reference to all the other ones. We just run all of our session you just run all
47:44
Speaker A
of your sessions from one coordinate and this can be a separate repo with just this cloud MD in it or it can be or it can be one you know the repo that everyone uses the most and then sometimes you touch other repos
47:55
Speaker A
well you put all the config there and even if you're not touching that main repo you just you have one canonical place where you start agent work from and then you give it the ability to find those other repos. The sandbox thing
48:07
Speaker A
gets a little trickier if you're doing on demand sandboxes because like do you really want to clone 200 repos onto the box every single time? Well, that's going to make your thing slower and you need more disk and things like this. The
48:17
Speaker A
alternative is like make the human pick which repos are relevant, which works to a point, but as soon as I want to like take advantage of the fact that agents can help me be productive in code bases that I'm not an expert in. Now we're
48:28
Speaker A
back to like, oh, if I missed a thing that was important, then I'm back to like, oh, I need to go talk to the person or know, I need to have that intuition about the codebase to know that. And like we really want to
48:38
Speaker A
enable you to leverage agents to do this. So, this is another time where it's like you could do cattle, but then you're bulk cloning all the repos all the time and have to keep them all up to date. Or
48:47
Speaker A
you could do pets and just have them all checked out. Yeah. Um, what I would the other thing I would recommend is please don't use subm modules.
48:53
Speaker A
Uh, when you use subm modules, subm modules are horrible because they pin the specific commits in every repro.
48:58
Speaker A
You're going to be a sad sad cookie. Unless you're sub trees are good though.
49:02
Speaker A
I know a lot of people who are who are skilled with this stuff that do a lot of work with get sub trees and I've actually played with them and like I think I think it's an option that makes
49:09
Speaker A
I haven't actually used sub trees. So, I'd be I I'll have to go check it out and see how they're different. How are they different? Do you know off the top of your head?
49:16
Speaker A
I I did the tutorial like a month or two ago. I I I can't I I can't do it justice today, but it's it's somewhere it's a it's kind of like a sub a work tree and a subm module had a baby.
49:27
Speaker A
I see. Okay. Yeah. Cuz the biggest problem I have with subm modules, it's just like you you have to like commit in two places every time you update the every time you update the module, the subm module and their actual repo. And
49:38
Speaker A
it's just like not fun. Yeah, sub trees have a little bit of that. But I think I think also I've heard from people who do a lot of this and like who especially I think Mos Mike Hostetler was on an
49:47
Speaker A
episode like six months ago where he was showing his like Ralph loop with git sub trees set up to help people basically do agent work across all their reps.
49:54
Speaker A
You want to pull up your thing really fast? I have a couple other interesting you scaladraw that you had. I had a couple of things I was going to do.
50:01
Speaker A
People have got some other interesting things of like uh where do you see the software factory design pattern evolving in the near in the near term. Honestly, I think the biggest thing that's going to happen is more people like actually
50:13
Speaker A
what's funny is I should actually just paste our our workflow into here because it's basically the same. It's like a more micro version of a part of this stack.
50:20
Speaker A
Yeah. So, I'll put that in here really fast. But like when I think about this, when I think about the evolution, I'm curious about your thoughts on this evolution.
50:29
Speaker A
Yep. Uh let me Sorry, I'm coming to the Excala draw and I'll post it down here.
50:34
Speaker A
There you go. It's at the bottom part. Nice. Um when I when I go think about this um uh when I the way I see this evolving is one it's going to become a lot more frequent.
50:50
Speaker A
Like every company is likely going to build some sort of internal agent that does something like this. And like I personally don't think the whole thing is viable. Uh, and the reason I don't think the whole stack is viable is
51:00
Speaker A
because every company has some weird quirk about their um internal stuff. That is how their setup works and everything. So, it's like you just can't know everything. Like for us, like someone asked earlier, how do you deal with issues with external
51:16
Speaker A
resources? Well, like we're lucky we don't have that problem or it's so rare for us. We don't have to design around that system.
51:23
Speaker A
On the other hand, this is this is why I'm so obsessed with like defining these building blocks. Yeah.
51:30
Speaker A
And like defining really good interfaces between them. So like the interface here is is an interesting one. Like ACP tries to do this uh but ACP is kind of on hit X on Riverside again please.
51:43
Speaker A
I can't see it full screen and I want to see Dexter's diagramming and max potential.
51:47
Speaker A
Yeah. So like between the control plane and the harness you have something called ACP as an option. It's quite narrow. Um I think we need like a thicker version like a wider version of this interface for a like coding harness to talk to a
52:01
Speaker A
UI or an editor or a web app. Basically it's like a protocol for things to talk along. You also have things like AGUI which again is great for broadcasting the UI and the events but actually like the the challenge I have is neither of these
52:15
Speaker A
support hooks which means like something that is like job it is to life cycle a harness and like its events and do special things in response to like I don't know in human layer we have a bunch of hooks of like
52:26
Speaker A
yeah go ahead do you know why though I think it's because every single harness is literally basically its own bespoke UI.
52:33
Speaker A
Yep. and like converting every like claude code and codeex do not have the same hooks and pi has totally different hooks.
52:41
Speaker A
Open code has a plug-in system that's like the hooks are like a completely separate concept.
52:46
Speaker A
Yeah. And I think this is kind of where it like all these things die which is that it's not a there's no standard for this there and and rightfully so like every harness has different trade-offs. Like clawed code
52:58
Speaker A
is much much more like you bring it and it's good like it's guaranteed to be good.
53:04
Speaker A
Yep. Whereas like uh pi is more like you have to configure every little thing and that allows you to have like it comes with control but you have to you have to build more stuff or but also you have to build more stuff.
53:19
Speaker A
Yeah. Exactly. Right. And I think that's why like you can't really do this cuz the abstraction layer in which you represent everything there is just not that good because it and and it by design can't be that good at least in my
53:32
Speaker A
eyes uh not right now but also like you know I used to think about web and then I finally learned about react I was like oh [ __ ] they actually did solve this hard problem by inventing this concept
53:40
Speaker A
of state turns out state is all you need and everything else kind of just devolves from that in a really nice way and like whether you like react or angular doesn't really matter but like web web systems generally are state
53:51
Speaker A
driven engines and that was a novel concept. So like you just got to think about what's the boundary here to think about and I don't know what it is and I think that's why it feels like you guys I know
54:01
Speaker A
do a lot of work of wrapping every single one of these harnesses and making it good to plug into.
54:06
Speaker A
Yep. Yes sir. So um I think that's most of what I had. I mean are there more are there more comments in the chat? I mean like yeah the the interface between the harness and the dev environment is
54:18
Speaker A
something like the Linux kernel and like I don't know the dev environment the infrastructure infrastructure is mostly like how do you go set it up it's like what APIs do you have what control do you have how do you
54:29
Speaker A
go build this there is something called compute SDK that tries to be this which is basically like one library for oh it's become a they made decided to be a benchmark instead of a instead of an SDK but it
54:43
Speaker A
used to be like it doesn't really matter for that layer. Like it's the same reason that you don't like yes it's nice to be able to abstract over like GCP and AWS but thing is once you're a big enough
54:53
Speaker A
company you're basically either a GCP shop or AWS up no one really needs to swap stuff over that often and when you do use the AWS interface is already pretty good.
55:04
Speaker A
Yeah, that's kind of what I mean. Like if you're using Daytona then Daytona interface is good. If you're using like GCP, the GCP SDK is good and the GCP CLIs.
55:14
Speaker A
So like all these things ship CLI. Go ahead. We used to have this rule of like at Sprout, we said don't wrap interfaces with other interfaces or don't wrap APIs with other APIs. It was like okay cool.
55:25
Speaker A
like we have the official Twitter SDK and then I'm going to make our version of the Twitter client which just wraps some of the methods that we need and all the data is just passed through with a little bit of like transformation and
55:36
Speaker A
it's way better to just like just call the official SDK at every call site. You don't need to like put an abstraction in front of it if it all it does is like actually create like okay now I have to like do extra
55:48
Speaker A
work every time I want to add a new call to a new endpoint on Twitter. That's kind anyway that was my spiel about where I think the interface stuff was going. In terms of buy versus build, where do you land? Obviously, I know
56:00
Speaker A
you're biased, but that's okay. I mean, I think the idea is like I think when you use a f fully verticalized cloud agent, some of the challenges you will hit is there's friction in defining your dev environment. There's friction
56:12
Speaker A
in having limited like choices on compute. I think what Devon and Cognition are doing is really interesting in terms of like letting you run what they call like outposts or workers where it's like okay you can run the compute and the outpost will bring
56:24
Speaker A
the harness and the tooling into your level of compute and you and the harness can kind of co-own the dev environment.
56:34
Speaker A
Uh, and so like things like that I think we're going to see more and more of and like the the conversations we're having with a lot of people and like trying to figure out is like what are the right
56:42
Speaker A
interfaces and I think what a lot of people are starting to realize one more time or I can sure most people are either like buy the whole thing or build the whole thing and I think the answer is like you should be
56:56
Speaker A
able to instead of having like it's like it's like composition over inheritance right it's like rather than having to be like okay whatever layer I buy at I have to buy everything below it. And so it's like, oh, if I want orchestration and I
57:08
Speaker A
want to be able to ping agents from Slack, I either have to build the whole freaking thing or I have to buy the whole freaking thing. And like the the real answer is like you should have as an engineer choice and the ability to
57:19
Speaker A
like work in open systems and plug these things together. Uh so like what's interesting for us is like uh our a you're Devon is kind of at these three layers, right?
57:30
Speaker A
Yeah. And they do the infrastructure too. No, they do the cloud sandboxes. Yeah, you can move that.
57:35
Speaker A
Yes, but I think they kind of own they don't want to own all of it. It's like optional for them.
57:40
Speaker A
Uh I I think in most cases they they want to own it. Uh just because it's a better experience. Like it's easier if I can just sign up and like send a Slack message and see work start getting done
57:50
Speaker A
versus like the price. Huh. For enterprise customers, I think they just run the whole thing on prem. Um but yeah, enterprise customers. Yeah.
57:58
Speaker A
So enterprise wants to bring the compute and probably bring parts of the dev environment. Yep.
58:02
Speaker A
Yeah. Like for us, what's interesting is like when we build our agent drives BAML stack, what we're really building is this this kind of we're building like up here and we're kind of like we don't we explicitly do not want to own this
58:13
Speaker A
layer. We don't want to own because you want to use the harness that everyone else is using and like really well I want to make it swappable cuz yeah cuz the harness is going to change like all the time. And really I don't
58:24
Speaker A
really want to own the compute sandbox. I just do this because of like convenience. Like whether we run on a box or not, I don't really care.
58:30
Speaker A
But these are the parts that I I find valuable. It sounds like the layer that you're trying to help is like, hey, you for human layer uh human layer, you just wanted to buy the orchestration, but you wanted to bring your own
58:45
Speaker A
harness. Yeah. This is kind of what you're trying to help plug into. Yeah. and Devon is there and like cloud cloud automations like cloud background stuff is like cool you can do like cloud code as a building block
59:00
Speaker A
or you can bring the entire cloud code st Oh oh that's not good or you can like I got it basically I got I got stop I don't want that one you can have like basically like clawed uh what do they call it like
59:13
Speaker A
automations or I think they call them routines or something but it's like basically like where they'll bring the compute environment and stuff too.
59:22
Speaker A
Yeah. And I'm going to do one more thing really fast. Um, put this up here. And like I would say like there Claude has cron jobs and stuff that you can do and they're trying to like plug into this
59:32
Speaker A
part over here, right? Yeah, exactly. Craw where they're trying to be like so you can if you want to be the full Cloud stack and like uh I know Pearson Marks does this the the Jelly Pod guy.
59:43
Speaker A
He was demoing this at the last thing. like he uses the full cloud stack and like he knows it's separable but he likes cloud code as a harness. He thinks the dev environments there are good enough for him and the compute on demand
59:53
Speaker A
is great and he can schedule automations on it and so he his you know composed stack happens to be all in the cloud ecosystem.
60:01
Speaker A
Exactly. [clears throat] Um I'm going to go read some comments because we got a bunch more really fast.
60:05
Speaker A
Some people are building both. They're like fully multiplayer. It sounds like there's a product there that's kind of cool. if you can drop in the chat if it's interesting. Uh given that we have a choice that it stack would a lampstack
60:15
Speaker A
equivalent emerge I think so I actually think we will get an equivalent to a lamp stack here over time. Um you mean nextJS next.js superbase and uh work.
60:26
Speaker A
No no no no I actually I don't know about that because that's not really what this is about. This is like the harness the dev environment the infrastructure deployments like dev environment is much more closer to docker for me
60:37
Speaker A
like the infrastructure application stack but the yeah the yeah like for this like software factory stack software factory stack yeah exactly uh ATB is like a thing that we call like agent tries BAML which is like our own
60:48
Speaker A
software factory that we built to automatically improve the BAML programming language and then yes these videos are recorded if you want to check them out they're on the YouTube channel I'll drop them off in a bit we post them on Twitter as well.
61:01
Speaker A
Yeah. And then is there an open-source project for control plane orchestration? I think the reason that there isn't an open source project for control plane orchestration right now um is because like one like no one you don't want to use someone else's control
61:17
Speaker A
plane because it's so easy to write code now unless you actually get all the integrations built with it. And right now everyone is everyone that's doing that to some degree wants to make a product and they have some value to
61:29
Speaker A
serve. So like they want to sell it to you and that's not an unreasonable request. There's a lot of good design and intuition that goes into it and like there is um we met the open inspect guy last last week but he's that's open
61:42
Speaker A
source but it's another full stack. His his is like okay cool we bring I think it's like basically like it does all of this and you bring the compute to run it on basically.
61:51
Speaker A
Yeah. Exactly. And like I that's what I mean like there is stuff but it's so easy to build now that I would just build like for us for example like we just decided the two layers of the stack
62:01
Speaker A
we care about and we just plug and play and like you ask me like why would like I actually have thought like should we use human layer stuff for part of our control plane and like the only reason
62:09
Speaker A
we don't is because like they don't have an API surface that I can't like plug into my slack into our GitHub issues into our BAML CLI. So now we have to build this whole layer around this and I think the input control plane
62:22
Speaker A
Oh yeah, but the point is like we still have to build an input control plane to define what we consider an issue. Like maybe we could use human layer for like the good once an issue is good, but for
62:32
Speaker A
that untrustworthy feedback that we get from a another user's agent. Where do we run that? Like that's a and this is the same reason that we don't just use linear pure linear for this anymore because like we have two
62:44
Speaker A
different versions of issues and when you ask like is there an open source project the problem is every company's working style has different definitions.
62:53
Speaker A
Yeah, this is David Cross's thing of like software should be extensible now that agents can write code. Like why are you going to make me plumb some like wissy wig like integration manager of like cool when a linear web hook comes
63:06
Speaker A
in transform it in this way and then do a workflow. It's like just let me write the code for that.
63:11
Speaker A
Yeah, exactly. And Matas is asking like do I see an open source I do personally see an open source solution happening and like I suspect companies like you will probably eventually support something like this because like it's in
63:21
Speaker A
everyone's best interest to go do that and if you have a good prod stack you would build a good open lure solution but I think it's really hard to upfront build an open source solution like part of what I really admire about the work
63:31
Speaker A
that Dra has done is that like they're working with real companies and real companies have real products and instead of optimizing for one company what they can do is they can take the learnings across like 10, 20, 30, 40, 50, 100, a
63:44
Speaker A
thousand companies and build the best interface that really serves across all of these. And then you can imagine like an open source thing being much more u um able to being built like you can't build React the first time the web was
63:59
Speaker A
invented. You [laughter] really have to watch a lot of stuff happen right and wrong before you see that happening.
64:07
Speaker A
Well, and I think step one is the interface layer. like MCP did this really interesting thing where it created this really explosion of ecosystems that let both like agent client and agent harness builders and integration builders like mix and match
64:22
Speaker A
with each other really easily. And I think basically like there's a lot of things that we're working on is like less like what is the what does the product look like at any one of these layers and more like what are the open
64:34
Speaker A
interfaces that enable people to have choice and be able to buy versus build any part of this stack for the parts they want control over without having to uh without having to uh to to go all in on
64:48
Speaker A
like let's say the cloud stack or the the codec stack or the open inspect or build the whole thing themselves.
64:54
Speaker A
And it still is to show like how nice do Claude and Codex want to play with each other.
64:59
Speaker A
Yeah. Like if they don't want to play nice with each other, they will keep breaking APIs and open.
65:04
Speaker A
This is why Claude doesn't support agents MD cuz they're like we want Claude MD to be the Cloud thing and everyone else can use the other file, but we're special.
65:12
Speaker A
Yeah. And like they've just decided that's the case and it's annoying, but like we'll just have to kind of see.
65:17
Speaker A
And yeah, Eric has this great thing. like the browser wars because the thing is like in the past the companies believe owning the browser was worth it and it turns out like it was right. If you own the browser you can make a lot
65:27
Speaker A
of money by doing that. Um and owning mobile was worth a lot of money. I think a lot of people are betting that owning the harness is going to be worth a lot of money.
65:38
Speaker A
You want to see something one last thing that's kind of interesting before uh we close it out for the day.
65:43
Speaker A
You talked about a question from Eric real quick and then yes which is is human layer composable decomposible? So, so human layer owns the uh orchestration layer and it can also we have a harness uh but it's swappable. You can use cloud
65:55
Speaker A
code, you can use our harness, you can use codecs. Um we kind of have noped out of the like we'll eventually have managed cloud sandboxes and stuff but our kind of take is like if you're serious and like the customers that we
66:06
Speaker A
work with all of them would prefer a good interface on top of compute and dev environment that they own basically.
66:12
Speaker A
Yeah. Exactly. Yeah. It's just way easier. like I don't want to deal with setting it up on someone else's infrastructure.
66:18
Speaker A
You got one more thing. I I'll show you extendable software. I was going to make this a whole podcast.
66:21
Speaker A
So if it's interesting, let us know and we can make this a whole thing. Um I've been playing around with this concept a lot.
66:28
Speaker A
Um so when I think about extendable software, I think the best way to extend software is let agents write code.
66:33
Speaker A
It's called code mode. You've probably seen it. So in this case, here's like a string that an agent wrote. It does some code. It does some stuff. And I'm basically just going to start mcking with this string in really
66:44
Speaker A
interesting ways. So I'm going to pass in some code that I'm additional code that I'm going to dynamically generate.
66:50
Speaker A
So like if I do like name two uppercase just run this really fast and I'll run uh dexter. So when I run this I have compilation errors. U let me just do one grap instead.
67:12
Speaker A
H we changed our API to be only reflect. So I had to replace all these when I run this code.
67:24
Speaker A
Oh, never mind. I got nothing good to show. I apologize. Uh this is broken for whatever reason.
67:29
Speaker A
Uh but anyway, we'll do an episode on this about how to run like dynamic code generation, how to actually have your app be extendable from a really extend from a really interesting perspective.
67:38
Speaker A
We I just upgraded BAML version so I can't show this yet I guess. uh I should stop running local dev is what I've really learned. But um I think extendable code is definitely going to be one of the most interesting features
67:48
Speaker A
out uh out ever because like like for example you in human layer how great would it be if people could define their own columns and their own like you could be like hey here's a small little code workflow that I want to run
68:01
Speaker A
for doing this and like they can just go say like oh here's an agent loop that I want to add on because that's kind of what cloud automations is effectively but just done kind of shittily like a cron job a chron Crown job is basically
68:15
Speaker A
like code that says while true at this time run this piece of code. That's like the ultimate crown job.
68:22
Speaker A
Yeah. In a scalable, you know, fault tolerant way where like if the server goes down at that time, it still runs a minute later on another server.
68:29
Speaker A
Yeah. Yeah. So that's the hard part about the execution. But the extendable part is what I really want because if I can only write the workflows that you let me write, I'm kind of sad. But if I can write the
68:39
Speaker A
workflows I can write on your stack and I can buy the reliability from you, then buying the reliability is totally worth paying for, right?
68:48
Speaker A
I mean, I think it comes into like this idea. I mean, in in our own code bases, I know many people have talked about they have like viable zones and they have no vibe zones, right? They have the
68:58
Speaker A
the functional core and then they have an interface and things like pragmatic code can sit on top of it. And there's no reason not to basically create space for users to deploy their cut like instead of user files a bug and if it's
69:12
Speaker A
in the no slop zone our factory builds it chips a feature goes in the next release. It could literally be like in real time it's just like cool these parts of this have you know sandbox execution adapters where your agent
69:23
Speaker A
writes code and that modifies your app in real time. But getting the interfaces right and figuring out okay what is the core versus what is the code and what are the interfaces that are available in that code is is uh it's an interesting
69:36
Speaker A
space that yeah I think I think we've both been exploring a little bit uh and excited to I don't know I I don't have any like cool like demos of that yet but I think Ree from Exeutor actually built
69:44
Speaker A
something like this where like the way you customize and configure exeutor is just by dropping TypeScript files in your directory. Yeah. I I think there's like no other way around this is like long term what you want to do is you
69:54
Speaker A
just want to be like here's my code and every piece of software effectively turns into like a harness.
70:00
Speaker A
Yeah. And open code like the whole point of those things is like hey you want to customize the behavior. Cool. Have your have your agent write code and it live reloads into the agent and now you have new extensions.
70:11
Speaker A
Yeah. Exactly. Because like nothing else I think is like that's like the most interesting stuff in my in my eyes about like how people build apps. It's like the like React built like the uh interactive web and I think we're about
70:23
Speaker A
to have like just like extendable software as like a common place like VS Code amazing because you can add extensions. Tails Tailwind comes out you can have it work in VS Code without VS Code updating and that's like magic.
70:36
Speaker A
Um cool we're way over time. Oh [ __ ] we are. It was a really fun conversation today. I love today's conversation. It's been so long since we've had one of these. Yeah, had a good time.
Topics:AI coding agentssoftware factoryBoundary MLagentic codingCodexopen-source AIenterprise softwaresoftware architectureAI model comparisontoken usage

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →