This video explains why companies should avoid overspending on cloud AI and highlights Nvidia's Switchyard and Neotron 3.5 Lightning for cost-effective, customizable AI tooling.
Ask about this video. Answers come from its transcript only — with the timestamp, so you can check them.
Generated from the transcript and can be wrong — check the timestamp.
Key Takeaways
- Avoid using expensive frontier AI models for every task; customize smaller models instead.
- Human supervision and structured processes are essential for effective AI adoption.
- Nvidia’s Switchyard and Neotron 3.5 Lightning enable cost-effective, customizable AI workflows.
- Observability and data-driven feedback loops improve AI performance and reduce costs.
- Local AI infrastructure can significantly cut cloud AI spending while maintaining high utility.
What the video covers
- Companies are overspending on cloud AI tooling by relying too heavily on expensive frontier models for all tasks.
- The key problem is turning over judgment to AI without human supervision, leading to wasted resources and poor outcomes.
- Nvidia's Switchyard is introduced as a supervision architecture that routes AI tasks to appropriate models with human oversight.
- Neotron 3.5 Lightning is a 30 billion parameter model designed for easy customization and efficient use in organizational AI workflows.
- Customizing smaller models based on user data can outperform large frontier models on specific tasks, reducing costs significantly.
- Building an organizational structure around AI usage data creates a feedback loop for continuous improvement and cost control.
- The video emphasizes the importance of observability, human-in-the-loop supervision, and structured AI adoption processes.
- A powerful local workstation setup with dual RTX Pro 6000 GPUs is demonstrated as a practical platform for running these AI tools.
- Nvidia’s approach includes logging, evaluation, fine-tuning, and reinforcement learning to optimize AI deployment.
- The video warns that indiscriminate use of expensive cloud AI models is financially unsustainable and unnecessary for many tasks.
Chapters
- 00:00Introduction: The Problem with Cloud AI Spending
- 01:40The Need for Human Supervision in AI Tooling
- 03:14Nvidia Switchyard and Neotron 3.5 Lightning Overview
- 04:36Building Data Sets and Organizational Learning
- 06:01Hardware Setup: HP Z8 Fury with Dual RTX Pro 6000 GPUs
- 07:27AI Application Instrumentation and Process Architecture
- 09:04Training, Fine-tuning, and Reinforcement Learning Details
- 10:51Cost Efficiency and Practical AI Deployment Advice
- 12:35Summary and Final Thoughts on AI Spending and Supervision
Full Transcript — Download SRT & Markdown
Speaker A
Your company is burning money. It's literally setting money on fire on your AI tooling. You know it. I know it.
Speaker A
Your boss probably doesn't know it. Maybe you should send him this video. It's kind of an oblique and subtle suggestion that maybe they pump the brakes on the AI stuff. The trap isn't just paying $25 for only a million
Speaker A
output tokens. I mean, that should have cost pennies. The bigger trap is turning the judgment over to the AI because increasingly what the tooling encourages you to do is to just turn it loose and let it do the thing. Can you tell he's
Speaker A
not happy with the codec updates? Uh, more on that in a minute. [music] As software engineers, of course, of course we want labor augmentation. There has never been enough labor to do everything that we want to do in
Speaker A
software engineering. I mean, give me another programmer who can read every file and work all night and never gets bored and costs almost nothing long as you're not sending it to the cloud. And we probably, you know, we probably do
Speaker A
get roped in at least a little bit helping the rest of the company adopt this AI thing as software engineers for the company. And that probably turns into, well, just go buy this thing and turn it loose. No, no, no, no, no.
Speaker A
That's the aforementioned burning money for warmth that I'm warning you about. So, yeah, turn over nothing. We devs, we still want to supervise. We want to structure. We want to be confident that in our brains, we have built an
Speaker A
understanding so that we can understand. We want to know what problem we're trying to solve. We want to do work that ultimately others value. I mean, folks, this is very, very basic stuff.
Speaker A
It applies to AI. And I've seen some disasters where almost everyone has lost sight of this really basic stuff. Now, for AI tooling, I want to know what it did. I want to know what reasoning it had, what led to the next step, what
Speaker A
information it used when it failed. I want to know what circumstances led to that failure and then what happened next. What corrective actions did the humans take to help the AI? That is a process that we have to build. That is
Speaker A
not a product. But today some of the most important software components that let anyone build a robust type of process like that around AI adoption in their company, that is getting a major upgrade and that is what you in the
Speaker A
janitorial and IT forces, the janitorial core should know. Switchyard. Switchyard is a key component of that from Nvidia. It's a supervision architecture for routing user requests to different AI models. It
Speaker A
is kind of like task router, specialized local worker or a specialized worker or a customized worker or a small model or something, an observable result tool, trace or an evaluator, and then you can accept or escalate that to a stronger model. It's human supervision is a big component of
Speaker A
this and that is important to me and I think that's probably important to you too. That's probably why you're watching this video. Nvidia is also launching Neotron 3.5 Lightning, which is a 30 billion parameter model that's meant to
Speaker A
help companies with this kind of observational structural flywheel. Or at least I found that it fits really well here. The model isn't the news, though.
Speaker A
It's the software tooling that Nvidia is providing that companies can take and customize and run with this. And you can customize the model. It's never been easier. The model that's customized that's small can beat a Frontier model because the Frontier model is never
Speaker A
going to be omniscient. I promise you. So what I'm saying is you take the small model, you customize it based on how your users are using it or the data or whatever and then that is going to
Speaker A
beat the frontier model every time because you can't expect the frontier model to be omniscient. That's not how that works. So when I say you need an organizational structure here to incorporate this and something clicks for you immediately, you spot three
Speaker A
advantages. Cost: don't spend Opus money on a task that Quinn can do. Observability: you know what component did what where in your organization. And organizational learning: every routing decision, every escalation, every interaction that your users have or your
Speaker A
other fellow devs have, this is a secondary data set that you capture to understand how your folks are using AI.
Speaker A
And this, in my humble opinion, is the only path to sanity. I really can't underscore how important the easy customizability aspect of this is because with this tool you're building a data set. This is kind of secondary to the task. But the
Speaker A
data set about how your folks are using AI really, it's the only path to sanity.
Speaker A
And with the easy customizability of local models and the data set that you're building based on how people actually use stuff, something magical is going to happen for your organization in terms of cloud token cost
Speaker A
and spend. Oh, incidentally, I think we're getting really close to the high water mark where the finance bros pump the brakes a little bit on this AI spending, which may lead to relief in other sorts of scenarios because
Speaker A
finally, your company does not need to spend all of this money on frontier models for just every little task. I mean, it makes sense intrinsically when you say it like that. Yeah, helping Bob with his PowerPoint presentation, do we
Speaker A
really need a $25 per million token model helping with that? No. No, we don't. But how do we manage that? Goes for software engineering tasks, too. You already know it's lunacy that we would use the frontier model for everything.
Speaker A
But let me give you the context and vocabulary to really drive that point home with this new stuff. Now, [snorts] the cloud tokens, they probably get spent on other things ultimately, robots, designer drugs, cancer research. Maybe
Speaker A
that part doesn't slow down. But the days of Opus helping you remember the order that the arguments should go in, those days should be over. And it's important to your corporation's bottom line that those days should be over. So,
Speaker A
let's take a look at this at a piece at a time. This machine behind me is running a dual RTX Pro 6000. This is the HP Z8 Fury that I reviewed previously.
Speaker A
There's some forum write-ups and how-tos, and there's been fun stuff we've been doing with this platform. This platform can accommodate up to four high-end GPUs. Right now, it's got two RTX Pro 6000s in it, 96 gigs of VRAM apiece,
Speaker A
but this platform will do 384 gigs of GPU memory right in the box. Check out my other videos and the forum posts if you're more curious about our HPZ workstation and you want to go down that rabbit hole. Right now though, this
Speaker A
machine is a literal flywheel connecting all of these pieces together. It's running the model router Switchyard from Nvidia. It's running two AI models, a small model and a much larger model. Uh, but these models are customizable without full retraining. And that's kind
Speaker A
of becoming table stakes for these open models. And that's part of how a small model can beat a frontier model at specific tasks. And I've got the research papers to back that up in the forum thread. You should check that out.
Speaker A
Conceptually, this thing plugs into something that Nvidia has been working on here for a while. The data flywheel.
Speaker A
It's a structural concept. It's not even really about Nvidia or Nvidia products or anything like that, but more about how you can structure process changes with the new advanced fancy pants AI tools in your company, but
Speaker A
without the setting money on fire for warmth aspect of that. Their architecture is almost comically explicit about this. Instrument the AI application, which means log the production traffic, build evaluation and fine-tuning data sets from those logs, and then evaluate smaller models,
Speaker A
customize them, and promote the ones that work. The original blueprint implementation is now deprecated for new production deployment on GitHub, but it's still useful conceptually, like for me mentally, working on this and then building my own thing w
Speaker A
kind of how I got here. But that's not the part that's running on the workstation behind me anyway. I mean the workstation also like the router has still has access to the Frontier model subscription that's in the mix. The
Speaker A
router here can reach it but also if my team springs for some giant GB300 class system somewhere else you know bigger local AI fine the router that's on here can reach that too. It's still local to the organization. It's just not local on
Speaker A
this machine but that router can reach all of those components wherever they are. the application. Ideally, our users and what they're doing don't really care about where the intelligence came from.
Speaker A
It's all routed through here. In my case, Nvidia has clearly seen the structural implications of this coming and I think probably also the economic implications, the ones that I was kind of hinting about about demand and AI and
Speaker A
all this other sort of thing. That's why I'm strongly encouraging you to tool up on this and learn to self-service. Even if you you just want to throw money at cloud AI, it would be irresponsible not to build something like this just to get
Speaker A
visibility into what your users are doing. And if that structure can also save you a ton of money with no extra effort really on top of that, I think that speaks to the mind of the seuite does it not? Now Nvidia also has emotron
Speaker A
orchestrator 8B which was running on here too. I was using that previously that's older but there's a lot of useful diagrams and explanation with the their tool exh orchestrator and you should check that out to understand more about
Speaker A
how this kind of thing running locally like you can just download this and run this on your own hardware. This is not cloud stuff. This thing is 8 billion parameters. It's tiny. It'll run on a laptop. Not super fast but it definitely
Speaker A
will run on a laptop. And so even though it's 8 billion parameters, we're not asking the 8 billion parameters to know everything. It's an orchestrator. It can call search. It can call code tools. It can call other specialist AI models. It
Speaker A
can call giant general purpose models including GPT5 and claude opus. On uh humanity's last exam, which is a benchmark, Nvidia reports the orchestrated system scored a 37.1% versus 35.1% on GPT5. Now keep in mind this is 2025 but at 30% of the cost and
Speaker A
two and a half times faster. So that little model was better at the task because it was in a tool calling ecosystem like is on this. It's not trying to be omnisient like the next version of GPT or claw. Now the orchestrator 8 billion
Speaker A
in 2025. It's still useful but now we have Neotron 3.5 lightning that is 30ish billion parameters. It's fast. It's fast like the 8 billion parameter model. It's local. And according to Nvidia's pre-release material, and I got to play
Speaker A
with it a little bit ahead of time. They're not just throwing weights over the wall. What they're doing here is they're giving you the full LoRa, the the supervised fine-tuning, reinforcement learning. It's recipes and training data. This is the machinery
Speaker A
that you need to turn that generic model into your model with customizations without full retraining. Now, Nvidia didn't mention Orchestrator 8B in their release or, you know, I may be off base in here, like connecting those in my
Speaker A
brain, but because I was trying to build practical useful things for business, I sort of connected the dots here. And for business use cases where I'm doing a little bit of customization and tool calling in my early testing, Lightning
Speaker A
3.5 behaves much more like what I want here. Speed and, you know, a relatively sophisticated model, but that is capable for routing tasks, route it to the appropriate model and tool calling for the work that I want to do. And so you
Speaker A
add in retrieval, augmented generation, and a vector database, which are things that I can't really get into in this video, but all of that will live happily here on our HP workstation. And the generic model, you know, it doesn't
Speaker A
magically know your your current internal documentation or your inventory or customer state or whatever changed yesterday. But we give the system retrieval and tools so that it can get the authoritative information that it needs.
Speaker A
And uh even if there were a path to do that with a frontier model, do I really want Open AI to have access to all of that or or anthropic? No. No. No one wants that. So on this local machine,
Speaker A
this is an enormous intelligence and utility multiplier over simply handing every problem to one generic model and expecting it to somehow have gone full veger and learned everything that was learnable. No, this is a system of systems and that system assembles the
Speaker A
resources it needs to serve the request to find the answer and anything else is akin to asking that gigantic cloud model to cosplay as an omnisient intelligence.
Speaker A
Um I guess that's actually what the current seuite implementation strategy with codeex CLI or cloud desktop might actually be. I mean that's certainly that's the impression that I get from some emails I've been copied on. I mean is that is that what you think folks
Speaker A
engage below? So that can become the power of model routing and this is where sophisticated buyers are moving in the industry and even outside of Nvidia stack take Quinn 3.6 how smart is the 27B Quinn 3.6 six, you know, how's the
Speaker A
35B A3B stack up? Can it code? Yeah. Yeah, that's all fine, but no, that's not what I mean. I mean, it's good, but but also no. That's not the part that I want to look at and squeal with
Speaker A
excitement about Quinn 3.6. Quen 3.6 is disruptive because it's insanely well documented for customization. Far under reportported is how approachable the customization ecosystem has become with modern open models like this. And Nvidia is leading in some respects and uh
Speaker A
embracing the openness in other respects which is exactly what most businesses need to embrace this and run you know run with custom like run with cool stuff. I mean you got full fine tuning Laura Qura SFT GRPO Megatron you
Speaker A
know multiGPU reusable data formats runnable examples. The Quinn documentation literally gives you the data set structure and starts walking you through commands, you know, MS Swift supports full parameter tuning, Laura, Q, Laura, you know, distributed training, reinforcement learning
Speaker A
methods, including uh GRPO, DPO, PO. Nvidia has done exactly the same kind of thing with their playbooks and the models that they've released recently.
Speaker A
We We don't need a machine learning PhD for AI customization projects anymore. I can do it myself right here. We, the unwashed IT masses, the lowly computer janitors can reach this level of functionality for our organizations on a
Speaker A
workstation like this sitting on a desk. And that's important. I mean, my go-to has always been building observability.
Speaker A
I mean, that's kind of fundamentally an IT problem. That's fundamentally where it starts. The CIOS and the CTO's are worried about that. We need to understand what the users are actually doing. How does it, you know, service the customer or make a better customer
Speaker A
experience. So now we can take observability on how the AI is actually being used, evaluate it and reinccorporate it into the system serving those same requests. Suppose on day one I have to send 80% of the requests through the model router here
Speaker A
to a frontier model and pay $25 per million tokens. Okay, fine. But if the workloads contain repeatable patterns and I'm capturing and evaluating those prop those patterns properly, I should be able to start moving those classes of requests downstack so that we're not
Speaker A
sending all of those to the frontier model, maybe a different cloud model, maybe a different local model. Where does the frontier model succeed where my local model fails? That's potentially a training example. Human corrects the result. Another useful signal. Uh the
Speaker A
same task that users are asking for 10,000 times. Maybe the generic model shouldn't keep relearning how our company does stuff from a system prompt and we should customize it. We can evaluate. We can curate. We can fine-tune. We can deploy. We can measure
Speaker A
again because you know what's the saying? It's like you can't improve anything you're not measuring. That's what's going on here. That's kind of the flywheel of intelligence. And you also get free token savings or you get token savings out of that because it's always
Speaker A
going to be cheaper for the smaller, less intelligent local model to be able to do things. So this is what our process becomes. And the outcomes are twofold. Like I say, economic the frontier inference becomes an exception instead of the default. And
Speaker A
organizationally the company stops renting all of its intelligence but also starts accumulating some of its own. I mean the the finance bros are maybe starting to get clued into that. And you it folks can probably focus on the
Speaker A
economic aspect of what I'm saying because you retain capital. um and capture institutional knowledge if you implement this process where you weren't before. I mean, this is a this is a win-win that really will speak to your seuite. Now, I promised we'd get into
Speaker A
the brass tax of how the customization happens and uh that sort of stuff. So, check this out. This is Laura. Think of Laura as basically a small patch for a big AI model. That's really all it is here. here and the model is set up to
Speaker A
handle this. Neatron 3.5 Lightning is already I mean it's a 30 billion parameter model mixture of experts. It only has three billion parameters active at any given time. Nvidia is explicitly pitching it as a model that you're supposed to customize. Smaller models
Speaker A
fine-tune faster, cheaper, and on much more modest hardware. I mean, I know it's expensive, but it's not, you know, a quarter of a million dollars. Come on.
Speaker A
You freeze the base model and you train a comparatively tiny set of additional weights to teach the small model your particular job. And operationally, Nvidia has already built the other half of this. So, NIM can keep one base model
Speaker A
resident and dynamically load and unload the Laura adapters while it's running and serve multiple specialized adapters from that same model that conserves resources and produces a better user experience. So maybe there's one capable base model here and accounting gets the
Speaker A
accounting adapter and software engineering gets the code review adapter and maybe support you know end user support or customer support gets the support adapter. That weird internal product that hasn't been documented since 2017. It gets the adapter that
Speaker A
contains the dark knowledge known only to Gary. Gary can finally go on vacation. Maybe the interesting part is the system around the model around switchyard and all of the componentry that's feeding switchyard. It's the secondary effect you get from that. The data set that you
Speaker A
accumulate showing how people actually use the system is ultimately probably the most important thing. It's it's not I don't even think it's accurate to call it AI telemetry. It's institutional knowledge and it shows what people are actually asking the system for because
Speaker A
their expectations in reality uh probably not great. And I guarantee you a vast majority of those those requests users are making do not need to be served by a model that costs $25 per million tokens. and the IT folks, you
Speaker A
get a clear accounting of where users get stuck, what workflows repeat, what the organization thinks is simple and they're like trying to do with AI, but isn't because of reasons that you'll have insight into. But most importantly, which processes are actually producing
Speaker A
useful results for staff and customers. That's the lowhanging fruit. All this internal corporate knowledge is, you know, at almost every company is usually uh captured very badly, if it is captured at all. But this facility like building this in as a process to capture
Speaker A
this that is what I'm telling you to do because now you've basically got all of the components right here to build that paint by numbers style. So this blog post from February also helps explain and fill in some of the gaps a little
Speaker A
bit more. It's a complete pipeline for taking small amounts of domain information, generating structured synthetic training examples and automatically evaluating them. you could filter them down and produce a data set that's ready for fine-tuning or distillation. NVIDIA's explicit pitch is
Speaker A
that the workflow is making the model specialization accessible without requiring enormous data sets or machine learning PhDs working on this. This is literally the flywheel. The structure here saves tokens and actually tells you, you know, what model your
Speaker A
organization actually needs for the problem at hand instead of just shoveling dollars into the AI black hole and hoping for a good outcome. just, you know, spam the the giant AI with a million context with all the stuff.
Speaker A
Another way I know I'm right here is that there's been a recent CodeCli update. That's what I was complaining about in the beginning. In my own use, I'm seeing much less useful insight into why the agent is doing particular
Speaker A
things. I'm not alone in noticing this. Uh there's a visibility problem here. There are open issues on the Codeex repository where reasoning summaries that were present in the session log uh are not being surfaced to the UI.
Speaker A
they're they're not showing up and so it's like okay well maybe I need to capture the session log and I'm not even sure that the session log always captures everything that I'm expecting.
Speaker A
Now I have a hypothesis about why that's happening and it's kind of a dark pattern. Highquality model outputs and task traces are extraordinarily valuable if you're trying to train a cheaper specialized model which is exactly what everybody that spends money on cloud
Speaker A
tokens should be trying to do at this point. That's that part's not controversial. OpenAI has a product called model distillation that's built around taking outputs from Frontier models and using them to fine-tune cheaper models for specific tasks.
Speaker A
[laughter] You know, those Halcon days of like a year and a half ago, they were encouraging customers to do that. So, I think there's at least an economic incentive for Frontier providers not to make every useful internal signal
Speaker A
trivially exportable to somebody building this kind of flywheel where we want to run stuff locally and maximize value. And again for the seauite listening here the path to ruination is to abdicate your you your human reasoning your wetwware reasoning and
Speaker A
just offload all of that to the cloud. The smartest folks in your organization want the execution to be observable.
Speaker A
They want to know what resources were used. They want to know what was decided, how the users corrected after a failure. And they want to know ultimately what worked. And you need to build something like this to capture that kind of information. I
Speaker A
mean, besides being the responsible way to supervise automated labor, the information here will feed the system, which makes the next version of the system better, which will make, you know, your company a more functional company. I've got a full write up to go
Speaker A
with this video with references to what I remember reading over the last few months and the new launches and everything else on the forum that that is linked below because there are academic papers that you know suggest things like if we you know like the LLM
Speaker A
router paper and some other things to back up what I'm saying and also expand on it and give you some ammunition for the the boardroom or or uh the seuite or the rest of your tech team depending on
Speaker A
what it is that you're you're trying to build for for integration. I think that with these pieces, this structure is how you can responsibly integrate this, but with also you know like there is a real danger I think of
Speaker A
of AI psychosis in the seauite if I'm not overstating that for comedic effect and building this kind of structure gives you some observability and that's really the bottom line here. It's like you can't fix what you don't measure.
Speaker A
And it's like, well, this this answer from this system feels really good, but then if you don't do an accounting of like, okay, well, we tried to take that answer and we took it apart and it fell apart under scrutiny and reinccorporate
Speaker A
that into the system, then you you're not learning, you're just spinning your wheels. So very excited why the switchyard updates and Nemo Neotron 3 lightning, but more importantly that the um state-of-the-art for local and open models is changing in terms of both
Speaker A
documentation and implementation to make it easier for non-machine learning PhDs to pick it up and customize them and modify them and get a useful result from a model that is under a 100red billion parameters. So it'll it'll basically run
Speaker A
on anything and that's exciting. I wonder this level one. It's been a a bit of a ramble, but also I think that, you know, this might be the high water mark and like now those cloud tokens may be
Speaker A
used for robots and other things, but I think this is going to be the high watermark for where we see those tokens being used in a cloud context because it literally does not make sense for these kinds of tasks to run in a cloud
Speaker A
context. Yeah. Well, we can chat about that on forum. I'm wonderless level one. I'm signing out. I'll see you there. [music]
Topics:AI toolingNvidia SwitchyardNeotron 3.5 Lightningcustom AI modelscloud AI costsAI supervisionorganizational AIAI observabilitylocal AI infrastructuresoftware engineering AI











