AI Accountability, Not Capability, Is Redefining Enterprise Trust
Trust has always been the basis of any enterprise relationship, but tenure and scale are no longer enough.
Gerrit Kazmaier
President, Product and Technology
Workday
Trust has always been the basis of any enterprise relationship, but tenure and scale are no longer enough.
Gerrit Kazmaier
President, Product and Technology
Workday
When generative AI first flooded the market, it didn't spark immediate investment; it created paralysis. Enterprise leaders were bombarded with promises but rightly hesitated to hand over their most critical HR and financial workflows to unproven black boxes. Most of them got stuck in a “PoC hell” of countless promising prototypes that never saw production light of day, due to hallucinations, erratic failures and undermined business value.
Trust has always been the basis of any enterprise relationship—and Workday has been earning the trust of our 11,500 customers for over 20 years. The old test of enterprise trust was tenure and scale, but that’s no longer enough. The new test for enterprise trust is accountability. Can you prove quickly and explainably that the system did the right thing, for the right reason, and produced a real and positive business outcome?
As core business decisions are made faster and increasingly by machines, accountability is not optional. Our CTO Gabe Monroy emphasizes an agent isn’t just a model; it’s an identity, with permissions, guardrails, and governance all working together alongside provable efficacy with defined accuracy. Agents must be lawful to be useful. An agent that acts quickly but can't be explained, audited, and trusted is more dangerous than one that's merely slow.
With Workday, every step is checked against the security model and compliance logic before an agent acts. It’s not a prompt and a model, but an actual AI system that is based on deterministic business processes and probabilistic reasoning. As CEO Aneel Bhusri has pointed out, it's this combination that unlocks AI for the enterprise.
This is what we call the World Model for Work. Today’s AI models are built primarily on the world's digitized footprint, but the systems and people they manage are poorly represented there. Workday has incredible insight into the world of work, and how people and finances work across the globe. Some domains in HR and Finance even differ significantly from country to country, industry to industry, and even customer to customer. That dataset is ultimately our World Model for Work, and it is the fast track to trusted AI that gets work done without breaking the rules.
Report
In the early days of AI, novelty and the promise of transformation was enough to inspire belief. Now, as enterprise leaders, we are rightly asking harder questions. Features and use are not enough. We have to measure accountable outcomes.
We are rightly asking harder questions. Features and relevance are not enough. We have to measure accountable outcomes.
Frontier models are good at a large number of tasks that require complex reasoning, but they sit on the sidelines of the operational system. They are detached from the business semantics, working off downloaded spreadsheets, and unable to close the loop back to where the work actually happens.
Given the choice between accuracy and speed, business leaders choose accuracy every time, especially in the core functions of the business. This caution is understandable, but it limits speed and hems in the incredible potential of AI.
In fact, only 27% of organizations are using AI that is embedded in their core enterprise system. In a workplace where agents work alongside people, we can’t expect our employees to copy and paste data between systems. The two have to be one and the same.
Recent Workday AI research showed the danger of relying on confident but unaccountable AI outputs. Our researchers provided information to agents, and then watched what happened to their responses (and the AI explanations of the responses) when they made minor changes to the underlying data:
A human might reasonably mistake a long, detailed, confident AI response as being accurate, but our research challenges this assumption. If the underlying context shifts, so might the accuracy of the response. In fact, even when the underlying change to the source material should not materially change the outcome, it can more often than you might expect.
Since confidence isn't proof of accuracy, the instinct might be to reach for a bigger, smarter model to close the gap, but that is the wrong instinct.
The real answer isn't a single frontier model asked to do everything. Instead, better outcomes come from combining small models such as classifiers, open-weight models, and large frontier models—matching each to the job. It’s about building AI systems that consider failure modes, retry strategies, consistency checks, and task-specialized domain models in a secure, auditable and governed process framework. The architecture has to pair the right model for the task with deterministic guardrails to check its work. That is what will deliver true confidence, backed by lawful systems.
The architecture has to pair the right model for the task with deterministic guardrails to check its work. That is what will deliver true confidence, backed by lawful systems.
Enterprise decision makers no longer evaluate technology on capability lists, flashy demos or prototypes; they evaluate it on speed to a provable outcome. Value has to be demonstrated immediately as a solution.
Workday’s AI agents already deliver on this promise to our customers. In our second quarter, we drove nearly $600 million in Annual Recurring Revenue (ARR) from our AI SKUs, up more than 200% year-over-year. And when we rolled out Sana Enterprise internally, our own Workmates built nearly 22,000 custom agents in just three weeks.
Workday’s AI agents already deliver on this promise to our customers. In our second quarter, we drove nearly $600 million in ARR from our AI SKUs, up more than 200% year-over-year.
This is what AI looks like when it is embedded into real work, combining immediate value and speed to redefine enterprise trust.
Report