Cost-Efficient Enterprise AI Depends on Architecture
Our engineers found smarter architecture that cut costs by nearly 90%.
Priyanka Mudgal
Principal, AI Engineering
Workday
Our engineers found smarter architecture that cut costs by nearly 90%.
Priyanka Mudgal
Principal, AI Engineering
Workday
As AI agents take on more work across HR, finance, and IT, the architecture behind these agents increasingly determines how well they scale. Two systems can do the exact same work and produce the exact same output, yet run very differently under the hood.
Most conversations about AI performance focus on model quality and instructions, but architecture is the variable that quietly decides whether an agent stays efficient as it takes on more complex work.
Workday’s engineering team found that with the right architectural choices, teams can cut AI operating costs by nearly 90% without any loss in quality, accuracy, or capability.
Report
Breaking every task into small, careful steps and checking progress after each one is the instinct good engineering teams have relied on for decades, and for traditional software, it still holds up.
But there’s a different problem when applying that same instinct to an AI agent. Every check-in is a full round-trip to the model. And with each round-trip, the model drags along everything it needs to see: tools available, entire conversation history, the results of everything already done.
The model re-reads every part of it, every single time. So a workflow built from 20 small steps doesn't cost 20 times one big step—it costs closer to 20 squared, because each new step re-pays for all the history stacked up behind it.
The real lever on efficiency isn't the size of any one step. It's how often the system stops to check in at all.
With the right architectural choices, teams can cut AI operating costs by nearly 90% without any loss in quality, accuracy, or capability.
Our engineering team studied a set of AI workflows that had grown organically over time. Each workflow consisted of many small, individually reasonable tool calls, wrapped in a planning layer that kept the agent's checklist up to date as it worked.
It revealed a pattern where systems were checking in with themselves far more often than necessary, after nearly every micro action, rather than at natural milestones in the work. Each of those check-ins carried the accumulated cost of everything before it.
The fix wasn't a smarter set of instructions, but reducing how often the system paused to check in: from constant micro-check-ins down to a handful of meaningful milestones per run.
Applied across four real production systems, the results were substantial: AI check-ins dropped by roughly a third on average, and by as much as 84% in the most expensive system.
Notably, the system that cost the most to begin with saw the biggest savings, which is a useful rule of thumb: fix your most expensive process first, because that's where improvements pay off fastest.
We then layered on a second technique: a provider-level discount available when parts of a request stay identical between steps (think of it as a "loyalty discount" for repeated information) on top of the leaner design. That added a further 66–76% reduction.
Combined, the two approaches led to total cost reductions across all four systems ranging from 77% to 95%, with an average of 89%.
When costs get out of hand, the instinct is to shorten instructions given to the AI but that rarely helps. Instructions are usually a small fraction of the real cost. The bigger levers are structural:
How often does the system need the AI model's advice? Every unnecessary consultation adds cost and that cost compounds for the rest of the run.
How much irrelevant information is being shown each time? Unnecessary capabilities for the current step cost money just by being available.
How much old output needs to be kept around? Results that no longer matter are still being paid for if they're left sitting in view.
Leaders should reach for lower-level cost tricks like caching discounts, prompt compression, context pruning, etc. only after addressing these questions.
Fix your most expensive process first, because that's where improvements pay off fastest.
We've distilled these lessons into a short list of questions that can help any organization think more deliberately about AI costs:
1. Should this even run? The cheapest optimization is work that's correctly stopped before it starts, which is a duplicate or out-of-scope request caught early.
2. Does this step need real judgment, or is it just routine processing? Steps that are simple, predictable, and deterministic don't need the AI's judgment at all—plain logic handles them just as well, at a fraction of the cost.
3. What does the system need to see right now? Only show the AI what's relevant to the current step, not everything it's capable of. For example, a tool the model doesn’t need this turn shouldn't be visible this turn.
4. What should the output look like? Return compact summaries instead of full raw details when possible, while keeping the ability to pull up full details later if needed.
5. What does the system need to remember? Compress completed work into a short summary rather than carrying the entire history forward.
6. What can be made cheaper to send? This is where discounts, using smaller/faster models for simple tasks, and similar techniques come in with real savings. But this should only occur once the design itself is already lean.
Ultimately, a workflow your organization can't afford to run at scale isn't one it can truly depend on.
The lesson here goes beyond any single technique. As AI agents take on more responsibility across HR, finance, and IT, efficiency and predictability matter just as much as raw capability. The discussions above show that leaner, well-governed systems can be just as powerful as their costlier counterparts.
Ultimately, a workflow your organization can't afford to run at scale isn't one it can truly depend on. Treating cost as a governance metric, not an after-the-fact surprise, is what turns AI into a durable, dependable part of how your business runs.
Report