Skip to content
Get App

Subagents vs Agent Teams? 🧠 Hermes Bots, Goal Loops & Kanban Graphs

Explore the benefits and trade-offs of single agents with sub-agents versus multi-agent teams for efficient AI workflows.

Key Takeaways

  • Delegation to sub-agents reduces token usage and context window overload.
  • Single capable agents with sub-agents often suffice for many tasks without needing full multi-agent teams.
  • Multi-agent teams are beneficial for complex workflows but introduce coordination and cost challenges.
  • Steering sub-agents during execution allows dynamic control and improved output quality.
  • Effective agentic workflows require clear task orchestration, evaluation loops, and sometimes human oversight.

What the video covers

  • The video compares single agents with delegation capabilities to multi-agent teams in AI workflows.
  • Single agents can delegate tasks to sub-agents, reducing token usage and improving efficiency.
  • Multi-agent teams add complexity, coordination challenges, and higher costs but can improve workflow quality for complex tasks.
  • The presenter demonstrates how sub-agents operate in parallel and how instructions can be steered in real-time.
  • Token consumption is significantly lower when using delegation compared to a single agent doing all work alone.
  • The video introduces concepts like goal loops, bot mode with specialized agents, and Conbon for managing complex tasks.
  • A test compares performance and cost between different agentic workflow levels.
  • The importance of balancing delegation and judgment is emphasized to maintain output quality.
  • The video also covers orchestration, task graphs, and human-in-the-loop curation for knowledge workflows.
  • Viewers are encouraged to explore additional resources and memberships for deeper insights.

Answers

Questions about this video

What is the main advantage of using sub-agents with a single parent agent?

Using sub-agents allows the parent agent to delegate tasks, significantly reducing token usage and preventing the parent agent's context window from being overloaded, while still maintaining control over the final judgment.

When should you consider using a multi-agent team instead of a single agent with sub-agents?

Multi-agent teams are more suitable for complex workflows that require specialized skills, tools, or memories, as they can improve quality and efficiency despite added coordination, cost, and latency challenges.

How does the video suggest improving the quality of outputs from single agents?

The video recommends using goal loops to check task completion, steering sub-agents with specific instructions during execution, and incorporating human-in-the-loop curation to enhance output quality.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
That's just a really powerful way for you to significantly reduce the cost. There we go. It just dropped from 45,000 tokens down to 18,000. One agent can already delegate to sub-agents. So, what does making an agent team actually add?
00:13
Speaker A
This is the question I get asked most often about multi-agent workflows. Why build a team if one capable agent can actually get the job done? And honestly, sometimes you shouldn't. One powerful agent with sub-agents can often be
00:24
Speaker A
enough. But other times, a more complex workflow can improve the quality and the efficiency. Hi, my name is Callum, also known as Waterloots, and welcome to today's video on agents, loops, and graphs. When are multi-agent teams worth it? This process also works with agents
00:37
Speaker A
outside of Hermes as well. In today's video, we'll explore the different levels of agentic workflows. One agent with temporary helpers, aka sub-agents, goal loops that check whether the work is actually done, adding specialists with their own tools, skills, and
00:50
Speaker A
memories, aka bot mode, and then Conbon, a way to manage complex agentic tasks. Along the way, we'll check what does each level actually need and when does it make sense for a specific task to add more agents or tooling, including a test
01:02
Speaker A
at the end where I give the same task to each level and see how it does. Now, let's take a look at agents, graphs, and loops.
01:10
Speaker A
So, there are three separate questions that we're going to ask here. Who does the work? How do they check it? And what do they do next? So, today I want to go through an example scenario that applies to almost every type of task working
01:21
Speaker A
with an agent. If information comes in, what actions should I take? How should I take them? And what does done look like for that particular task? For example, an email comes in or I take notes on a new article and I want to add it to my
01:33
Speaker A
LLM wiki. How does the agent process that information in a consistent, repeatable way so that it actually meets the satisfaction that I want for that particular task? But first, before we get into automations, triggers, task graphs, and Conbon, let's start with
01:46
Speaker A
just a single agent to see how it works and orchestrates with sub-agents. We will build up in complexity until we get to a full graph system that lets you orchestrate your tasks. If you find this video helpful, please like, hype, and
01:56
Speaker A
subscribe as I really appreciate your support a lot. If you're looking for more ways to support me, please consider joining my YouTube or Patreon memberships. Members get access to insights, tips, and kits for agentic and knowledge workflows. I've also put
02:07
Speaker A
together a free article that I'll link in the top comment that goes through agents, loops, and graphs in a little bit more detail for beginners as a guide. I'll also include the results of the test that I ran so you can go
02:17
Speaker A
through it in a little bit more depth and hopefully see how this can help you improve the quality of your own workflow.
02:26
Speaker A
All right, so the first example that we're going to go through is using a single agent here, which you're probably already familiar with. This would be just any type of chatbot, but we're going to give this agent the ability to
02:36
Speaker A
have sub-agents. So, we can see in the bottom left here, we have this option for open agents. And we can see that right now at the moment we have no live sub-agents. So when we're working with
02:45
Speaker A
this first example of a single specialized agent, we can think of this as a parent agent that has the ability to create multiple sub-agents or children agents. So if we go over to capabilities on the left here and go over to tools,
02:57
Speaker A
we can search for delegation. And we can see that right now task delegation is turned off. So I'm going to enable this tool right here. And this will enable the parent agent to create a series of sub-agents with specific capabilities. So
03:10
Speaker A
let's get this started. It's going to be able to delegate to solve our particular problem, which is how do I turn incoming information into something useful, an action I can take or knowledge I can reuse, and what is
03:23
Speaker A
the best strategy for that action system? Does it make sense to use a single agent with delegation capabilities or a multi-agent team with different specialized agents? And the idea here is that if we can have one capable agent that's able to meet my
03:35
Speaker A
standards, that's a complete solution on its own. So right now we can see that the agent was able to create all of these sub-agents that are now running in parallel. We have three of them. We can see on the bottom here that we have
03:46
Speaker A
three different workers active at the moment. And what's cool, this is a new update. If we click on the three sub-agents here and click on a specific agent, we're able to give instructions to this sub-agent and steer it. So I
03:57
Speaker A
could get feedback saying, make sure to please include citations. And I would be able to now steer this agent. And that just sent off specific instructions to this particular agent that's looking at concrete end-to-end workflows for inbox information. And I could do that for
04:13
Speaker A
each of the different sub-agents. And now we can see, for example, this particular agent, the one that I just steered, is now using the research grounded citations skill. So I was able to give some instruction to the agent as
04:25
Speaker A
it's running to help control the direction that it's going in. You can click on it here and you can go and see how it's thinking. Oh, we just saw this one finish right now. Now we have two agents instead of three. And we can see
04:35
Speaker A
that one has completed right here and the other two are running. So now the main agent, the parent agent is waiting for the sub-agents to complete its work.
04:43
Speaker A
And then it'll synthesize a concrete recommendation when they return. Okay, here we go. Looks like all three of those just finished. So now the main agent has gone through and it validated a few of the sources and now it is
04:53
Speaker A
creating a report right here. Okay, so this gave what looks like a pretty comprehensive answer. We can see if I scroll all the way to the bottom, there's a lot of information here. So I'd have to go through that a little bit
05:01
Speaker A
more in depth. And what's nice is if we take a look at the bottom here, this only used up 36,000 tokens. And the reason for that is because the parent agent delegated to sub-agents and only had to give definitions for them to go
05:13
Speaker A
off and run. So the sub-agents themselves then used up their own amount of context. And that's one of the major benefits of using a delegation system like this is I didn't have to have this main parent agent blow through its
05:25
Speaker A
context window. Instead, it could delegate to sub-agents of lower intelligence and then take the synthesized reports rather than reading all of the pages itself. So, the overall idea here is that we're delegating the work but not the judgment. And just to
05:37
Speaker A
see how this compares, I'm going to give the same prompt asking it to do research but do not delegate to sub-agent.
05:43
Speaker A
So, the approach where I just used a single agent and told it not to delegate to sub-agents. We can see that used 102,000 tokens whereas the one that delegated to sub-agents only used 37,000 tokens. So keep in mind this is not
05:55
Speaker A
taking into account the sub-agent token usage. So the overall tokens was larger. There was more context used. But that context that made it back into the main higher intelligence agent was significantly compressed relative to the agent that just did it all itself and
06:10
Speaker A
had 102,000. If I ran this turn one more time, I might have to compress the context here and then you lose some quality. So that's the first comparison, just using a single agent versus using a single agent with multiple sub-agents. So
06:21
Speaker A
the biggest difference is that one uses up a lot more context. But this was also a very basic example and the results of both of these are basically the same.
06:29
Speaker A
Build an accountable action steward with delegation capabilities, not a standing team of specialized agents. A permanent multi-agent team adds coordination, cost, latency, and evaluation problems before it h...
06:41
Speaker A
Multi-agent designs work well when the task is genuinely decomposable and parallel. And this is a very similar answer to what we got with the delegation system. Do not begin with a peer team of autonomous agents. Delegate bounded readonly work to specialists
06:54
Speaker A
when that work is genuinely independent. So both of our researchers found effectively the same thing that we should only delegate specialist work when it actually makes sense to do so when the complexity needs it. But all of this is relative. This is specifically
07:06
Speaker A
for a one-off research task that we just did here. As the system starts to scale in complexity, we need to start thinking about other ways of orchestrating. For example, using all the different bots like we've set up in previous videos.
07:17
Speaker A
And one thing I want to mention too is that when we're delegating to sub agents, sub agents know nothing. They start with completely fresh conversations. So they have zero knowledge of the parents conversation history, tool calls, or anything
07:29
Speaker A
discussed before they were created for the delegation. The sub agents only context comes from the goal and the context field that the parent agent populates when it calls delegate task.
07:38
Speaker A
Delegate task is this tool that we enabled at the beginning here. And the context is how much information goes into all of this. But before we get into an actual situation where it makes sense to have multi- aent workflows, there's
07:48
Speaker A
one more thing I want to quickly talk about, which is a way that we can enhance the quality of the output of the single agent, whether it's a parent agent with delegated sub aents or all-in-one. And that actually brings me
08:00
Speaker A
back to the warning of the sub aent system, which is that the context comes from the goal. So what is a goal?
08:10
Speaker A
So to bring up goal, we can type in slash and goal. And this brings up the goal command, which is a way that we can set a standing goal that Hermes works across all turns until it's achieved.
08:21
Speaker A
And basically what that means is before we get into the complexity of setting up all of these specialist agents and figuring out how to coordinate them, there's one more step that you should try first, which is to enable a goal. A
08:31
Speaker A
goal effectively allows you to link multiple prompts together into a bigger picture, a broader task than just one specific research task. For example, instead of managing every step, I hand over a bounded outcome. What success means, how to verify it, and when to ask
08:45
Speaker A
me. The main agent acts as a judge that checks the progress, and the session can continue within its own budget. And the goal isn't endless research. The evidence matters. So before we get into setting up a complex system to see if we
08:58
Speaker A
can get more out of our agent, we can try using this goal command. So let's give our multi- aent a goal here.
09:05
Speaker A
So we can see by using / goal, I now have turned this goal active. I've added a little bit more detail than just saying, hey, can you go research this thing? I've said specifically, I want different streams of information. Please
09:18
Speaker A
use delegated sub aents of lower intelligence to do the research. So each one of these should be using at least a terra medium instead of terra high. I'll talk about model routing in a moment.
09:27
Speaker A
But then I've also said please prepare a synthesized report and then once you are done that corroborate the report against practical workflows from community best practices in Hermes. You also must make sure the report is less than 500 words
09:39
Speaker A
and has clear citations. We can see that the sub agents are all running right now. So what I've effectively done here is I've told the agent what done looks like. So it's taken all of these criteria and it will just keep running.
09:50
Speaker A
We can see turn one of 20. The default limit is 20 turns. The agent is able to loop itself 20 times, checking to make sure that the information and the research and the output meet the quality standard that I've requested. So, this
10:03
Speaker A
is a really great way for you to increase the quality of your work without necessarily introducing a lot of complexity. I honestly use this all the time, especially for coding. So, I can also add criteria on here. For example,
10:13
Speaker A
we must have at least five peer-reviewed sources. And I can add that information to the criteria as the agent continues running. And that now becomes a new definition of what done looks like. So we can see here it took my criteria into
10:28
Speaker A
account because now the criteria is we must have at least five peer-reviewed sources. And if we look here, it only produced eight sources and it doesn't look like they're peer- reviewviewed. So now the agent has started up its next
10:39
Speaker A
task on turn two and turn three and it's delegated to another series of sub aents to now cross-reference the output against peer-reviewed academic sources.
10:48
Speaker A
So it says it's now added an independent organizational coordination and team cognition literature to the evidence set and it's awaiting the peer review before it will work again. As this is working, I can think of more criteria and I can
10:59
Speaker A
add this here and the agent will just keep looping. The more specific you are on what done looks like at the beginning, what the actual goal is, the better this process will work. Depending on the type of task, depending on the
11:09
Speaker A
type of budget you have, depending on your goals, you're going to have different ways of doing this. So, it's worth experimenting to see what actually makes the most sense for your particular situation. So, it's kind of a handsoff
11:20
Speaker A
method of operating. If you've established a good enough goal at the beginning, the agent can run on its own.
11:25
Speaker A
In some systems, I've had my coding agents run for five or six hours straight, working on a coding system, delegating to lower intelligence sub aents, and it keeps automatically testing the code to make sure it works before it presents it back to me for
11:38
Speaker A
approval. And that brings me to an important point because if you're running a goal here and it runs for hours and you have this set to Astra for example or to Fable or whatever your Frontier model is, it's possible that a
11:50
Speaker A
goal could just completely blow through your quota in a very short period of time, especially if you're not sitting there monitoring it. So this brings me to an important question on who should actually be doing the work here.
12:02
Speaker A
So what we can see is the sub agents here. It just says GPT 5.6 Terra. So I can't actually see at the moment if it is medium or lower. So as part of the goal here, I could have specified rather
12:12
Speaker A
than just please use delegated sub aents of lower intelligence. I could have said please use a local model or please use Luna only for the research and then you as the higher intelligence model can synthesize it. So this is where you can
12:24
Speaker A
start being more specific with your prompting and specific with your orchestrating through a single agent delegating to lower intelligence sub agents to significantly preserve your quota. makes a massive difference because keeping an agent working isn't the same thing as using its
12:38
Speaker A
intelligence. Well, I want the orchestrator agent, the main agent to make the judgments, not to carry every source detail or do every little action itself. So, we can use these lightweight helpers for bounded extractions with limited actions. And what I mean by that
12:53
Speaker A
is we can be specific about what type of model is used in which type of situation. So, this can apply to the original goal that I just gave here. I can be specific like I mentioned of what type of model I want to be used. I can
13:05
Speaker A
bring in the agents.mmd file where for this particular profile this particular bot or Hermes agent I can specify inside of the agents.mmd or inside of for example the soul.MD I can say when using sub agent delegation always use lower
13:22
Speaker A
intelligence models. By including it here, it means that every single time this particular bot, this particular agent wants to run a task that involves some sub aent delegation, it will always know to use lower intelligence models.
13:34
Speaker A
And I can go even a step further by going under settings and model and I can choose the auxiliary model for different types of tasks where for example maybe for image analysis I don't want to use my main model. I can change this and use
13:47
Speaker A
a local model or I could use a cheaper model. So you can go in and you can start limiting the specific actions to specific models to help conserve the quota and intelligence of the main agent for when it actually needs it.
14:00
Speaker A
And honestly, this did a really good job. It was able to specifically limit the report to just a couple hundred words instead of the original thousands and thousands of words it gave me off of the first one. So it wasn't just about
14:11
Speaker A
adding that limitation at the beginning. It was framing the goal that I wanted it to go in a particular order for cross referencing and corroborating the report and then make sure that it was 500 words. But honestly, this example here
14:24
Speaker A
is kind of disconnected from my actual workflow. This might be the generic best answer of when delegation beats a specialist team where we want to have an orchestration and delegation for a single evolving state, sequential or predictable work. And in contrast, you
14:39
Speaker A
might want a multi- aent team for breadth first research, different types of due diligence, pipelines with different specialist identities. So this all makes sense, but everything I've just said here is not very specific to me. It doesn't know who I am or what I
14:52
Speaker A
want. And that brings me back to Hermes corroboration where delegate task brings up a fresh context fork join primitive whose children return summaries. So, it fits bounded research, review, and repair work, but it doesn't necessarily make sense for other types of situations
15:08
Speaker A
that require more specialization and more context and more knowledge because like I mentioned, sub agents have zero knowledge from the parents conversation history. They only have what comes from the goal and the context that the parent gives them. But what if we start getting
15:21
Speaker A
into more specific context or more custom situations that involve other types of context outside of what the parent even knows? We can see here the decision test that the answer says is add a specialist if it is independent
15:33
Speaker A
needs different context tools and policy and beats the single agent baseline on quality coverage latency and total review cost. So in other words the answer is that we should add specialists if we want it to keep its own working
15:44
Speaker A
knowledge and have access to different context tools and policies perhaps have a different memory system rather than being spawned as a brand new agent every single time. And I note that the Hermes corroboration here mentions Kambban or Conbon depending on how you want to
15:58
Speaker A
pronounce it. But before we get into conbon, I thought that we could take a quick look at Hermes bot mode actually sets you up for success when you start using conbon because it helps you prepare the specialist profiles. Like we
16:09
Speaker A
can see down here I have orchestrator, researcher, and librarian. So let's take a quick look at bot mode.
16:18
Speaker A
So now we can go from a single agent operating in a loop with delegated sub aents to a graph. That's where profiles come in. Okay. So this is bot mode. So we can see here a bot is just a profile.
16:29
Speaker A
It's the same list of profiles that I have along the bottom here. A bot or a profile can retain its own skills, memory, and working context with a model suited to its role. Each of these different bots here can have a default
16:41
Speaker A
model that's different. For example, maybe my orchestrator has one of the most powerful models available, a frontier model, but maybe the researcher needs a much less powerful model. And then I can use a feature like the group chat to coordinate how they work
16:53
Speaker A
together. So rather than a sub agent starting with zero knowledge, we can configure what type of knowledge and skills and memory we want available for each specialist. And that knowledge and all of that context and capabilities survives the chat conversation. It
17:07
Speaker A
maintains it over time. And since Hermes is a self-learning system, it can also learn more skills as well, getting better at its particular specialized task. And we can take a quick look at a group chat here where we can see, for
17:18
Speaker A
example, I have the librarian, the orchestrator, and the researcher, and they're all working together to solve the original task that I put out. So, you can think of this as a communication graph where we've set up a group chat
17:28
Speaker A
here, and the group chat has particular members, and they work together, communicating back and forth with one another to solve a particular problem.
17:34
Speaker A
So, this is where we go from not just having a loop, but going from a loop to a graph. Now, I've already talked a lot about this in the last two videos. In the first video, I gave an overview on
17:43
Speaker A
what bot mode is, how to get it set up, and how to customize each of these individual profiles. So, I'm not going to touch on that today. And then in the second video, we did a specific research upgrade where I updated the capabilities
17:54
Speaker A
of the researcher to be a much more powerful tool. The main skill that I gave the researcher as an upgrade allowed it to check both academic sources and state-of-the-art sources and then corroborate the two of them. I just
18:04
Speaker A
wanted to quickly show you this because it is a form of graph in the sense that the orchestrator is able to communicate and message back and forth with the different sub aents, but it's not the same as a full task graph like we can
18:14
Speaker A
get into with Conbon. All you need to know at this point is that the orchestrator is the one that I message.
18:19
Speaker A
It's my human proxy, the one that controls the other agents. The researcher is able to investigate. So, in the example of information coming in, the orchestrator would create the task and then give it to the researcher to go
18:30
Speaker A
investigate. And then the librarian works with my existing LLM wiki to take the output research that I've approved as a human and file it into my LLM wiki or my knowledge base in Obsidian. So in a graph, each node can be a single agent
18:43
Speaker A
or action or a bundle of actions. For example, like an agent loop built into the group as part of a bigger macro loop. So in group chats, one of the key limitations is that they only have the ability to go back and forth three times
18:55
Speaker A
before they have to stop and check what you actually want. So before we actually get into running these full control graphs, I just wanted to quickly talk about a couple examples on how this fits together so that you
19:07
Speaker A
can see what the graph might look like before we get into how to set it up just so it makes a little bit more practical sense. So imagine you have an email come in and you want to have a decision.
19:16
Speaker A
Maybe you want to have a draft created, maybe you want to have some research done, maybe both. So the email can come in and it can trigger an agent like the researcher agent to pull in the relevant context and the evidence by doing some
19:28
Speaker A
research. Once it's done, it can then create a brief like a research brief and draft a reply and then hand it off to you for review. If you don't think it's good enough quality or you think it's missing something, you can provide some
19:38
Speaker A
more context, some more commentary, or steer it, and you can have it go back and loop again. So this brings in a loop, but it's a human in the loop. Then if it meets your approval, you can have
19:47
Speaker A
it create the draft or even potentially send it right away if you've already reviewed the final result. It's up to you if you want to have the send capability given to the agent or if you want to click the button. As another
19:57
Speaker A
example, let's say you've come across a great article or a YouTube video or some other resource and you've taken some notes on it. You can drop in the raw source and your notes and you can send it off to the librarian. The librarian
20:09
Speaker A
agent can try and find the fit. How does this fit into your existing knowledge base? It'll check the context connections and draft an entry for you for your review to see is this good enough for it to make it into your
20:21
Speaker A
knowledge base into your curated library. If it's not good enough, you can suggest revisions and the librarian will update the draft entry and then pass it back to you for review. Once you approve it, the librarian can then file
20:32
Speaker A
it into the library and connect it to the rest of your research. And as a more complex system, we can actually start linking these loops together. For example, you can propose a question or a topic to the researcher. The researcher
20:43
Speaker A
can go through a research loop checking sources, checking it against the quality that you've given. This could be a goal loop. And then when it's done all of its research, it can pass it off to you as the human to validate that the
20:54
Speaker A
information is good enough. You can add your own notes, connect it to that cited source and report and then hand it off to the librarian. The librarian can then see how that fits into your existing knowledge base and prepare a draft for
21:05
Speaker A
your review on whether or not it should make it into the library. Again, you can go through a curation loop, but this one has a human in the loop as part of the core loop itself. Once you've approved
21:14
Speaker A
it, the librarian can then go and file this into the library. And then, as a coding example, let's say you're building an app and you want to add the ability to export to Markdown. You can hand this off to your orchestrator and
21:24
Speaker A
explain what you're looking for. And the orchestrator can understand it and then delegate it to the implement. The implementer can then go through a goal loop where it builds it, checks to see if it works, and if it doesn't, then it
21:35
Speaker A
repairs it and then it checks it again. So it can loop as many times as you want until the task is actually done. Once it's done, you can pass it off to a reviewer for an independent review. Kind
21:43
Speaker A
of like a second version of the implementer, but specifically for quality assurance, who can then pass it off to the orchestrator and check to make sure it matches the original goal.
21:51
Speaker A
If it does, then it can hand it off to you for either a pull request or for you to do something with that particular output. In the original research, the Hermes cooperation actually surfaced that conbon is Hermes durable
22:02
Speaker A
alternative with name profiles, shared state, dependencies, retries, and work that survives the lead session. So, what that means is it's a lot easier to start and stop and add more tasks and modify and just it gives you a lot more control
22:15
Speaker A
on the actual task system that's happening here and the way the different agents fit into it. That's why it's called a task graph. But you might be wondering, isn't that kind of similar to what I already showed you where we have
22:25
Speaker A
the orchestrator or the main agent delegating to sub agents? So the Hermes docs has a good breakdown here of what the difference is because they are actually quite different systems. In delegate task, the parent is blocked from doing things until the sub
22:42
Speaker A
agent actually returns. Whereas in Conbon, the orchestrator can just set the agent off, can assign the task, and it can just go and do its own thing, and you don't have to have anything make it back to the orchestrator unless you want
22:52
Speaker A
it to. Rather than having some random anonymous sub agent that starts from scratch, you have a named profile with persistent memory. So that's a pretty important element. Also, for resumability, if the sub agent fails to do its task, it's just done. It just
23:06
Speaker A
comes back and says incomplete. Whereas in conbon, you can set different types of restoration processes. So it can be blocked and unblocked and then rerun again from where it left off. So this is what I mean when I say it's more
23:17
Speaker A
durable. It has the ability to survive errors and keep going. With sub agents, you can steer, you can add some feedback, but it's not necessarily actually adding a human in the loop. But in conbon, we can have specific actions
23:30
Speaker A
or task actually block from happening until the human comes in and says go. And with delegation, the audit trail, all of the context can be lost when the system actually compresses. This is what I was talking about with the context
23:42
Speaker A
window on the bottom here. If this gets compressed, then if there was a sub agent within that compression, you can lose the context associated with it.
23:49
Speaker A
Whereas in conbon, it's stored in a SQLite database. So you get access to that forever. You can go back and take a look. So a good summary here is that delegate task is a tool call. It's just one function. Whereas conbon is a work Q
24:01
Speaker A
where every handoff is a row in a database that any profile or human can see and edit. So there's a lot more transparency that happens here, which is why it starts to get into being a proper control graph or task graph. Thankfully,
24:12
Speaker A
Hermes has a built-in feature called Conbon. So let's take a look at that. So the first thing we need to do for conbon is you may not have this appearing on the side here. So we can go up to capabilities and then go over to
24:23
Speaker A
plugins. And we can see that we have three built-in plugins that come with Hermes. The first is bot mode that comes turned on like we've already looked at and the second is the conbon. So make sure this is turned on here. And then
24:33
Speaker A
you should see this appear on the lefth hand side. So if we go over to conbon here, we are not going to be the ones running things. The agent is. All we need to do is put a card into the ready
24:42
Speaker A
section, assign it to a particular agent, and the agent will pick it up in a minute. And there's even a triage section. So this is where if you're not quite sure exactly what the task is, you can have an agent rewrite the idea into
24:53
Speaker A
a proper task first. So that's actually a really important part of it, and I'll show you how we can change the settings for that in a moment. So triage is for raw ideas where that the specifier is going to flush out the spec and then it
25:04
Speaker A
will move it through the system. To-do means it's waiting on dependencies or unassigned. So we can create task and not assign it and that means the dispatcher won't send it off to go run.
25:14
Speaker A
We can schedule which is we want a task to run at a particular time. So it could be for example check the latest series of emails and see if something has come in that I've been waiting for. If so
25:25
Speaker A
then execute the next task. Ready means that the dependencies are satisfied. it's been assigned to a profile and the dispatcher runs it. Running means it's claimed by a worker and an agent's actually working on it. Block means it's
25:35
Speaker A
waiting for human input. So, this brings that human in the loop. And then review means that there's a review agent specifically checking out the work. So, you can establish which agent you want to operate as the reviewer. And this is
25:45
Speaker A
the completed section. But with all that said, why don't we run a quick test task?
25:50
Speaker A
So, I'm going to leave auto decompose triage task turned on. And I'm going to click new task. I'm just going to call this test task. I'm going to leave this assigned to the Wanderloots tutorials assigne and click create task. So what
26:02
Speaker A
we should see now is we've added this card to triage and the dispatcher is going to recognize this and hand it over to the auto decomposer within a minute.
26:10
Speaker A
And I'll explain all of that in a moment. So triage is a place for raw ideas where the specifier will flesh it out. So we just saw that it went from this raw idea into running. But this task here is clarify the scope because
26:21
Speaker A
it sounds like a test. So that's where putting something in triage will automatically assign it to the default that we have set up and it will run over here. So this should get blocked because this was a task that had no clear
26:33
Speaker A
outcome. So this is where the human in the loop starts to come in because it basically said there's no body. There's no information here. We don't know what they want.
26:41
Speaker A
So the dispatcher basically runs inside of their Hermes gateway inside of the back end. And you don't have to install anything. It just checks every 60cond has there been a new card created? If so, let's move through the conbon loop.
26:54
Speaker A
Let's go through the graph. The idea is that once this task gets created as a card, the dispatcher will see, oh, there's a new card. Let me run the system for you. But if we go back to my
27:03
Speaker A
task for a second, we can see here that we have a rough idea. The specifier will flesh it out. So, we can also see that we have auto decomposed the triage tasks. So, this is where if we go over
27:13
Speaker A
to settings again, and inside of models, we go down to auxiliary models. We can see here that we have two different models set, two different auxiliary models that we can control here. The first is the triage specifier and this
27:24
Speaker A
is the one that will take the idea, the goal that you have and frame it as a better task, which will then get passed on to the conbon decomposer to break that task into a task graph. So you can
27:35
Speaker A
control what you want these models to be here. Perhaps the specifier is a very lightweight model because you're basically just using text to explain text a little bit better. But for the decomposer, you probably want a much more powerful model. And Hermes docs
27:48
Speaker A
actually has a strategy for this where they say decomposing a project, breaking it into well scoped cards takes frontier level judgment using the most intelligence that you have. But executing a card that has a clear goal, context, and handoff doesn't really need
28:02
Speaker A
that same level of intelligence, especially if it's running in a loop. Workers are where the vast majority of tokens are spent. So that's typically where the cost is. So that's what I'm getting at when I talk about this
28:12
Speaker A
decomposer here. We can set this model to whatever we want. I could set this for example to GPT6 Astra if I wanted to. And what would happen is now the conbon decomposer, this decomposer setting that we have here would be using
28:23
Speaker A
Astra to convert whatever task I put inside of this task card into a series of subtasks that then get assigned to the worker profile, which is wonder tutorials that I currently have set at Terra. So that's just a really powerful
28:37
Speaker A
way for you to significantly reduce the cost of your operations by having a powerful planning model and a less powerful execution model or implement or worker agent.
28:50
Speaker A
And we can see up at the top here that we have a default board. So what we can do is we can click on the default. We can rename that board if we want to or create a new one. So I can for example
28:58
Speaker A
create a new board called test create board. And now when I click on this I can switch between different boards. And what's cool too is within that test board, you can assign it to a specific project. So what that means is you can
29:10
Speaker A
have different tasks and different graphs and different configurations and settings and operations that you want for particular boards and all of that can be controlled at the project level.
29:18
Speaker A
So you can create a new project like for example maybe one is your admin project and it deals with everything related to email. Maybe another one is a research project and it deals with all of your research that you're bringing in. Maybe
29:28
Speaker A
another is for a particular coding project or app that you're working on and all of that can maintain the context within this particular project and then you can assign a particular board associated with that project to run these tasks.
29:42
Speaker A
So we can open up the orchestration settings here and what that does is it brings up a list of all of the profiles.
29:47
Speaker A
It brings up all of the different bots or agents that we have available with all of the particular settings and skills and memory and context that we were talking about earlier. Each of these operates as a different specialized agent. So there's a few key
30:01
Speaker A
settings we need to take a look at. The first is the orchestrator profile. So right now it's defaulting to default.
30:06
Speaker A
You can change which agent you want to operate as your orchestrator. This is the orchestrating agent that will decide which agent gets which task. And those tasks get assigned to the assigne which you can set which one you want to be the
30:18
Speaker A
default. Maybe I want to create an implement profile whose sole job is to execute on things. in a similar way that the researcher is the one that researches things and the librarian is the one that files things in my LLM
30:29
Speaker A
wiki. So for this one I'm going to leave on auto decompose just so you can see how this works. I'm going to click new task.
30:39
Speaker A
I'm going to call it multi- aent research and I'm going to include the description of the same prompt that I gave previously. So here we can set the priority. This basically just controls the order if you have multiple tasks
30:48
Speaker A
going in there at once. And then we can have the workspace that we want it to run in. I'm going to leave the current assigne as the waterlude tutorials as the default even though I'm asking for research. So this is where the task
31:00
Speaker A
decomposer should be able to recognize that this is actually a research task, not an implementation task for the tutorials agent. I can list specific skills if I want. I can change the model that I want this to run at or just have
31:11
Speaker A
it run at the profile default. And importantly, I can have this run on goal mode. So by default, each worker agent gets one shot at its card to do the work. But turning goal mode into true for that particular task lets that
31:23
Speaker A
worker run in a goal loop like I've already shown you. And this is its own self-contained goal loop. It doesn't connect to the existing goal that we have elsewhere. And what's cool too is we can also estimate the cost of what
31:33
Speaker A
this task is going to be. So I can click estimate and then it gives me a list of the tokens that it thinks it's going to cost for this particular prompt here. So it took this task and it analyzed it to
31:42
Speaker A
give you an estimate. So it says it's going to be a broad ambiguous research task requiring extensive source review, synthesis, architecture comparison, and likely iterative documentation or prototyping. So now, for example, I can include a limitation here of please
31:54
Speaker A
limit your search to five sources and the output document to 500 words or less. And then I can reestimate it.
31:59
Speaker A
There we go. It just dropped from 45,000 tokens down to 18,000. There's a rough estimate researching up to five sources, comparing this particular task, and then a synthesis under 500 words. And that dropped the complexity of this task down
32:11
Speaker A
to medium. That's a cool way for you to do a lightweight check ahead of time to make sure that your task isn't going to blow through your whole quota. So now let's click create task. And this brings up the board that I was talking about.
32:22
Speaker A
The triage section is where raw ideas. We have a specifier that's going to flesh out the spec. So it's going to be converted from this into a different card. And this is using that auxiliary model that I was talking about. Okay,
32:33
Speaker A
here we go. So it just started. So that triage using the auto decomposer, the dispatcher checked. Oh, it's been a minute. There's something here. Let's break it into two different tasks. So the task decomposer broke it up and
32:43
Speaker A
assigned it to these particular sub aents where each one is now running on its own. And we can click into this and we can see the specific prompt that was given to it created by the auto decomposer using the default model. We
32:54
Speaker A
can estimate the effort of this particular task like I was just showing you. And we can comment. So as these tasks are currently running, we can actually send messages to the worker and if there's a big enough change, we can
33:04
Speaker A
even send a message and then reue it and have it start that task over again. So we get a chance to go through and see how all of this is working, what it's actually doing, how it's analyzing it.
33:12
Speaker A
And we can see this for each card that's running here. So you get a lot of flexibility for controlling these cards here. And what's cool, too, is for example, I can go back to the to-do section, which are the currently blocked
33:23
Speaker A
tasks. So these ones are waiting for the researcher to finish before they run. And we can see this one was actually assigned to default instead of orchestrator. So I'm going to change the assigne from default here back to
33:33
Speaker A
orchestrator. And now we can see that changed the card over here as well. So, I think it's just such a cool way for us to be able to take a look at how all of the tasks are broken up into their
33:42
Speaker A
individual actions and see how they move through the conbon and you can kind of just get a sense on how all of the project is working together. And that's just for one card. I can now add a second card and it can start running
33:54
Speaker A
itself and keep going from here. So, you can really start to plan out maybe you have five or six or 10 or 20 tasks that you want to run. You can put them all in here. You can schedule them. You can let
34:03
Speaker A
the agent know that some are waiting on others. There's a lot that you can do here and it just takes a lot of that sub agent delegation that we were talking about in the first couple examples and makes it a lot more transparent and
34:14
Speaker A
controllable and that's why it's called a control graph or a task graph because we can actually see the graph of how this is operating. What needs to be done first before the next task can be accomplished. So this one is still
34:25
Speaker A
running here but we can see this one finished. We can go through and see the full activity. So we can see that it was created and then promoted to ready to run and then it was claimed by a worker
34:34
Speaker A
which is this researcher agent. It ran through completed it and it created a briefing doc that was saved right here.
34:40
Speaker A
It only reviewed two sources because it was given a two source limit and it maintained everything. Okay. And it looks like the second one just finished as well. And now we can see that the orchestrator task just moved from to-do
34:51
Speaker A
to ready. So once it's going from ready, once the dispatcher ticks in 60 seconds or so, it should get put over to running. And there we go. So basically we had those two research tasks and once they were finished the to-do of
35:04
Speaker A
delivering the report was unblocked because the previous tasks were completed and then once that was unblocked the dispatcher running in the background moved it over to ready. And then once it was ready it shifted it into running. And again we can see all
35:17
Speaker A
of this was done by the auto decomposer. And here's the specific prompt. So it's pretty complex here. This is where it's important to be really strategic with the information and the tools and the memory and the context, the skills and
35:28
Speaker A
everything that you're giving to your specialized agents because you don't want to accidentally have them blow through your entire quota because you gave them more than they needed. So that's one of the dangers of running all of this autonomously is that you might
35:39
Speaker A
be giving these agents more than they need. So it's up to you to go through and experiment and test and just make sure that you're able to be more specific and strategic with how you want to set up these task graphs and how you
35:50
Speaker A
want to run them. So that task just completed right here and now we have multi-agent research ready. So we had the researcher run create its research report pass it off to the orchestrator who then constructed a 500word or less
36:03
Speaker A
recommendation and then pass it back to the original prompt that I had put in here called multi- aent research. And now this should be answered by it passing along the document that was created. Great. There we go. So now the
36:14
Speaker A
dispatcher is running again. The multi- aent research my original question that was created by me. And now it's going to give me the report.
36:22
Speaker A
Great. And everything is now done. And this is the report that it created from incoming information to action and reusable knowledge which is the theme that we've been working with today. So you can see it's 484 words. It is
36:32
Speaker A
limited to five sources and it gave a pretty solid recommendation. Have a single accountable lead agent with bounded delegation is the best default.
36:40
Speaker A
Use a manager style orchestration to keep one owner for state and a final output while allowing scoped specialists. Use a persistent team only when measurement shows sustained independent parallel work. So again, this is basically the same answer as
36:52
Speaker A
what we had before. So now let's quickly break down what just happened here before I explain the alternative option.
36:58
Speaker A
So this is the original prompt that I gave and what happened is you can see this was created by dashboard. So this was created by me. I assigned it to the orchestrator and I gave it this particular prompt. Then we can see that
37:09
Speaker A
the auto decomposer commented and said hey I decomposed I broke up this task into these three subtasks. So these were the two research tasks and then the writing of the article task. The root this question here this task will wake
37:22
Speaker A
once all the children complete. So it took that description it broke it into those three subtasks. It went through and ran the researcher went into running completed its two tasks here. sent it back to the orchestrator to write the
37:34
Speaker A
article, which honestly in the future I would probably have an implement agent or a writing agent who's specialized with this, leaving the orchestrator to just deal with orchestration. So that's one change I'm going to make. And then it gave the summary here. So it accepted
37:46
Speaker A
the deliverable. The final is 485 words. All the sources check out. And here's the simple recommendation. So honestly, that worked pretty well. But maybe we didn't need two researcher agents.
37:55
Speaker A
That's a little bit of overkill for a five source limitation. that probably used up more tokens than what we needed here. So this is where if we close this and we go back to our orchestration settings, we can start moving beyond
38:07
Speaker A
this auto decomposer. So this is where instead we can turn it off and we could let the orchestrator be the one to decompose the triage tasks.
38:19
Speaker A
So why would you want to do that? Well, my orchestrator has a specific soul associated with it. If I go here and click edit soul, I have very specific instructions on how I want the orchestrator to work, how I want it to
38:29
Speaker A
coordinate my team, who I want it to work with. I can put limits, I can connect specific memories. There's more context that can be given here from the orchestrator profile than just the auto decomposer, which doesn't necessarily have that same level of context. This is
38:43
Speaker A
just automatic mode. This is more of a curated mode, and it doesn't even have to go through triage. Instead, it can just go straight into ready. But if we want to use the orchestrator to operate as the task decomposer, we need to give
38:54
Speaker A
it specific conbon skills. So if I go over to orchestrator and then we go to capabilities, we can see here if I go to tools again and then to task, I've turned off task delegation. So basically what that means is I've removed the
39:06
Speaker A
ability of the orchestrator agent to be able to create those delegated sub aents. Instead, I want the orchestrator to run the conbon board. So I can turn on these tools here. And what that does is if we go back over to conbon by
39:18
Speaker A
having auto decompose turned off and the orchestrator turned on as the orchestration profile with the conbon tools the orchestrator will now break apart the tasks and then assign it to the particular profiles defaulting to the default assign here. So automatic
39:32
Speaker A
decomposition buys you convenience. It's really easy. You just type in the task and click go and Hermes will break it down for you. But creating my own planner as the orchestration profile the orchestrator gives me a lot more
39:43
Speaker A
deliberate control over the task design. and the roles in the review. It's not that one necessarily has to be lower cost than the other. It depends on the task, the complexity, and how much control you want to maintain and how the
39:54
Speaker A
system operates. The real benefit is continuity. I can return tomorrow. I can see what's blocked. I can inspect the handoff, and I can resume without reconstructing the conversation or trying to set off more sub agents. So, the orchestrator plans, the dispatcher
40:07
Speaker A
starts the system saying, "We're ready for work." And then the worker executes on the tasks assigned to it. So you can kind of think of each card as work units, not necessarily different agents, but a card can use helpers or contain a
40:18
Speaker A
goal loop in and of itself. So if we were to do that again, we set which orchestration profile we want, what default assign we want, but instead of creating a new card in triage here, new task in triage, it would just sit here
40:30
Speaker A
forever because the dispatcher doesn't have the auto decomposer to break that single task into the series of subtasks that then would be put into ready. I'm not going to go through that orchestration example. The key difference is that rather than dropping
40:43
Speaker A
something into triage, you can either put it into to-do and then wait until you're ready. You can schedule it or if you want to trigger the orchestrator right away, you would drop the new task in ready and let it deal with the task
40:53
Speaker A
decomposition instead of letting the auto decomposer do it. But I actually tested out the different scenarios to see which works best. So why don't we take a quick look at the results of my testing?
41:07
Speaker A
So my thought was there's a lot of different coordination methods here. which one actually makes the most sense.
41:12
Speaker A
The results showed that a single agent with task delegation for a simple task at least and running the orchestrator without the auto decomposer actually had pretty similar token usage and pretty similar quality. But using the auto decomposer and adding the goal loop
41:27
Speaker A
actually added more overhead for this particular task. So this is obviously a very simple example. It's not by any means a proper benchmark. I just wanted to see how this worked for a particular research task for me. And one thing to
41:38
Speaker A
keep in mind is that even if the workflow in the conbon takes more tokens, it might be worth it because you can go through and take a look at what was actually done. So it gives you more transparency. The actual right choice
41:49
Speaker A
just depends on what your work really needs. So I'm going to keep exploring this more. I'm going to try and do some more complex conbon systems with coding so I can actually see what makes sense for my own workflows, how we can
42:00
Speaker A
leverage the orchestration and control graphs that Hermes has built into it. So, do you need an agent team? Not necessarily. Start with the simplest level that actually meets your needs and then add coordination and complexity as you find yourself actually needing more.
42:18
Speaker A
You want to find a specific problem that that extra coordination and effort is going to actually solve. Now, today's example was a very simple one. Next, we'll get more into how sub aents, loops, and graphs can fit together to
42:30
Speaker A
solve more complex workflows. What workflow would be the most helpful for you? Please let me know in the comments and I'm happy to explore these topics more in future videos. If you found this video helpful, please like and subscribe
42:39
Speaker A
as I really appreciate it. And please consider joining my memberships. My members enable me to continue making free videos like this one. Thanks again for watching and I will see you in the next video.
Topics:AI agentsmulti-agent systemssub-agentstask delegationHermes botsgoal loopsKanban graphsworkflow automationknowledge managementagent orchestration

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries and speaker detection. 30 minutes free.

Or transcribe another YouTube video here →