Whitepaper

Governing enterprise AI
in procurement and finance

Why procurement and finance need a vertical, governed AI platform

It's a near certainty that dozens, if not hundreds, of people at your company are using unsanctioned AI tools to manage spend and suppliers.

With the easy and widespread access to an increasing number of AI products promising "streamlined workflows" via AI agents, both in desktop tools and now baked into the search capabilities within every browser tab, it's almost harder to avoid AI than it is to access it.

Ungoverned shadow AI exists within the walls of every company. According to our 2026 State of AI in Spend report, 57% of business leaders we polled say they're actively using AI tools their employer never sanctioned.[1] Just 47% say AI purchasing runs through formal procurement or IT review. Governance on paper and behavior in practice are very different things.

We believe this is a regulatory exposure crisis in the making. Work that would never survive a SOX audit is becoming routine in procurement and finance workflows which keep no record of who or what made each decision, be it AI or human.

The exposure is financial too. Uber spent its 2026 AI budget in four months and now caps engineers at $1,500 per month per coding tool.[2] Most companies can't produce that number at all, because AI spend arrives through consumption-based contracts, corporate cards, and tools that never went through review. Some companies are even experimenting with 'tokenmaxxing,' pushing employees to consume as many AI tokens as possible on the theory that usage signals productivity. The ungoverned usage that breaks the audit trail also breaks the budget.

But the root cause runs deeper than access to the tooling. Employees reach for shadow AI because company pilots, or new capabilities from incumbent vendors, fail to deliver their promised, or even necessary, value. The same survey found that 62% of leaders use AI multiple times a day, the demand is real, and it reaches the highest levels of the organization.[1]

Shadow AI and stalled transformation look like different problems, but they arise from the same core issue. LLM-based systems can fail when the intelligence requirement has no governed structure to complete the work in a trustworthy manner. The quality of models released by frontier AI labs like OpenAI and Anthropic is increasing fast, but in enterprise software these models are raw materials; necessary infrastructure but not sufficient on their own.

The answer is a vertical, governed AI platform designed for the work of procurement and finance.

In this paper, we build the case in three steps: first, why AI is difficult to govern in procurement and accounts payable workflows; second, what a platform must do to clear that bar; and third, how to decide between building and buying, where for almost every team the right move is to buy the governed foundation and build your differentiation on top.

Why procurement and finance are the hard case for AI

Not all business processes carry the same weight. When a sales rep uses AI to draft an email for cold outreach, hallucinations can lead to rewrites or, at worst, minor embarrassment. When AI judges that a $200,000 software contract should be approved and clears it without the right authority, the business has authorized spend it should never have approved, accepted terms that put it at risk, and created the exact control failure that a SOX audit is built to catch. Both activities use AI, but the cost of errors differs by orders of magnitude.

Procurement and finance live in the second category: a domain where decisions are regulated, mistakes are expensive, and lax governance can end in jail time.

In most functions, governing AI means keeping it on-brand and roughly accurate. In procurement and finance, governing the work means something far stricter. Every action has to comply with the policy that governs that specific transaction, which can vary by region and spend type. Every approval has to run through the right authority for its threshold, while every action has to leave a record an auditor can reconstruct: who acted, what they did, and how and why the AI made the recommendation it did. All of this is close to the opposite of what a general-purpose LLM provides. These systems are probabilistic by design, hold no memory of your approval matrix, and produce zero audit trail unless workflows are painstakingly built to capture one.

But this issue is not only a risk problem. The same governed foundation that keeps a transaction compliant is what lets AI work it well: surfacing a supplier's track record, catching an auto-renewal before it hits, capturing the savings a generic assistant leaves on the table.

Whether for risk or performance, the model needs a governed place to do the work, and almost none of that place exists inside the model.

Teams are already running into this wall. The Hackett Group's 2026 Procurement Agenda study found that although 56% of procurement organizations have deployed agentic AI in pilot or at scale, early productivity and cycle-time gains are still in the single digits, 9.7% and 9.3% respectively.[3] Deployment without a governed foundation rarely converts into measurable gains.

The work demands a level of governance that off-the-shelf models don't provide and internal teams can't easily build or maintain. Governance is the missing link.

The governance layer: where enterprise AI earns trust

Picture the enterprise AI stack as five layers.

At the bottom are the systems of record: the ERP, the CLM, the supplier and materials masters where you find the source-of-truth for spend. Each business structures this data differently, shaped by its industry, regulations, and processes. At the top are the systems of intake, where requests enter and stakeholders interact. As AI finds its way onto every employee's desktop, the top layer becomes the first touchpoint for many enterprise processes, and increasingly where policy and process are first surfaced to users. But generic models alone can't be trusted to surface, apply, and enforce those policies.

Three layers sit between them and do the actual work. The context layer reads the systems of record and resolves what a given transaction means: which supplier, which policy, which budget, which prior commitments. The orchestration layer plans and carries the work across systems and human approvers. The governance layer enforces policy at the moment of action and records what happened. Together they turn a capable model and a system of record into automation you can trust.

A governed platform matters because agents can break the assumptions enterprise software is built on. Traditionally a person sat in the loop at every step, with defined permissions, a paper trail, and policies to follow. Agents without these layers sidestep those safeguards: they choose their own tools, leave no trace of their reasoning, and reset context between sessions.

A model can reason about the work but cannot enforce a policy or remember your approval matrix. A system of record can store the result but cannot plan or act. Neither was built to run an agent safely. Everything that makes an agent safe to run lives in the three layers between them.

What the governance layer must do

Context and orchestration are what let the middle layer do the work. Three controls are what make that work governable, and they are the three requirements to put in a buying specification.

Requirement
What it does
What to ask the vendor
Guardrails
Preventive control at the moment of action: policy and permission enforcement on every step, an identity and a delegated spend authority for each agent bounded by category and threshold, segregation of duties, and human review where risk demands it.
How are an agent's access rights and spending authority defined and enforced? Show me segregation of duties for a non-human actor, and how we set human-in-the-loop thresholds, as well as tighten or loosen them over time.
Observability
Detective control: a complete, queryable record of what every agent did, which tools it called, what data it touched, what it consumed to do the work, and how and why it reached a decision, plus the evaluations and monitoring to catch degradation before it becomes an incident.
Show me the audit trail for a single agent decision end to end (reasoning, data touched, approvals granted). Show me what our agents consumed last month by workflow. What evaluations ship for our domain, and how do you detect and alert on issues before they become an incident?
Auditability
The compliance record. Every AI-assisted action is logged in a durable, tamper-evident form tied to the identity that acted, the policy that applied, and the authority that cleared it, retained and attributable to the standard an external audit requires. Where observability lets your team watch the work as it runs, auditability lets an auditor reconstruct it years later.
Show me the full record of a single AI-approved transaction the way an auditor would see it: who or what acted, under which policy, with what spend authority, and how the decision was reached. How long is it retained, can it be altered, and how do you prove it hasn't been?

These controls accomplish multiple goals; guardrails stop the transaction that shouldn't clear, observability shows your team what the agents actually did, and auditability proves it to someone outside the company. The record that satisfies an auditor is also what tells you which workflows are working well enough to widen. In Zip, that record includes what the AI consumed, with real-time visibility into token usage across Zip AI so procurement can run agents on a budget it can predict.

This ensures both governance and impact, both visible and measurable.

Paths to deliver governed AI to procurement and finance

Knowing the requirements of the governance layer turns the path forward from a technical question into a procurement decision. Many businesses are wrestling with this problem: build the governed middle layer, or buy it? Early adopters have found it isn't the either/or it looks like.

Capable organizations are increasingly tempted to build. Simpler use cases have made this seem easy: wire an agent to a model API, give it a few tools, and you have a working prototype of a procurement assistant. What the demo obscures is the reality of governance.

The governance layer is a set of systems an internal IT team will need to author and maintain forever: the approval matrices and supplier context it reasons over, the workflows and integrations that carry the work, and the guardrails, observability, and audit trail that keep it compliant. Most of that work sits below the demo surface, and it takes teams far longer to build than they estimate. Worse yet, most procurement functions aren't staffed to help design these agentic workflows, map integrations that shift underneath you, or define policies for how agents interact with human stakeholders (including suppliers). Neither IT nor procurement is staffed to do this work, and the delay compounds.

The clearest evidence sits in the deployment data. MIT's Project NANDA studied enterprise GenAI in 2025 and found that externally sourced tools reach production about twice as often as internal builds, roughly a 67% success rate against about a third for in-house efforts.[4] The teams with the deepest engineering benches are not exempt from that pattern.

None of this means building is always the wrong call. When a workflow is a real competitive differentiator and no vendor serves it, building is exactly right, and the test later in this section shows how to spot those cases. The mistake is building the governed foundation itself, the part every procurement team needs and none competes on.

Buying a standalone, generalized AI platform fails from the opposite direction. A general-purpose enterprise AI assistant can put powerful tools on every employee's desktop quickly, but it replicates the shadow AI problem in a sanctioned form: no purchasing policy context, no orchestration, no governance, and no audit trail. For company-specific employee experiences it can create genuine efficiencies, like answering questions from an approved internal knowledge base. But when these systems take on regulated activities, the work that matters most in procurement and finance, they fall short.

The path through this decision is to avoid the false choice. Businesses do not need to either build or buy. The answer is to buy the governed foundation, and build what differentiates your company on top of it.

The future of enterprise AI is buy plus build

In every business there are activities that create real competitive advantage, and a few that define how the company works. In these situations, custom development creates tools that exist only for that company.

Businesses also have required but non-differentiating activities: they must pay their employees according to local labor law, run proper due diligence checks on third parties they work with, and manage accounting practices according to regulated standards. Procurement and finance processes generally fall into this second category, though parts of them can be genuinely differentiating.

The middle layer (the context, orchestration, guardrails, observability, and audit trail every procurement organization needs and none should rebuild) is best purchased as a foundation you build on for the future. Reserve your own engineering for the workflows and agents that encode how your business specifically wins. That reframes the decision from build vs. buy to buy plus build.

This reframe gives businesses two practical tests to guide their architectural decisions. The first is what to build versus buy: build only where a use case is both a true differentiator and unserved by any commercial platform, and buy everything else. Plot your use cases on those two axes and the build list shrinks to a handful (which is the point).

The second test is how to choose the platform you buy, and here your own requirements are the scorecard. The three controls above are what the platform must enforce; scoring a vendor means testing whether those capabilities add up to governance that holds. That comes down to four dimensions, the same four our Enterprise AI Risk Assessment measures. Ask one question per dimension:

For context, can it encode your policies, approval matrices, and supplier data and apply them automatically, so the rules live in the system rather than in a handbook?

For visibility, can it show you every AI tool and agent running across procurement and finance, and which of them can take actions rather than only generate text?

For guardrails, does it enforce identity, least-privilege access, and spend authority at the moment of action, holding any agent below a threshold for human sign-off?

For auditability, can it reconstruct any AI-assisted decision end to end, with the same rigor as a human approval, and flag drift before it becomes an incident?

The Enterprise AI Risk Assessment takes each dimension further, with sixteen scored questions that show where your own governance would pass or fail an audit today.

A vendor that clears all four isn't really a single purchase decision. It is a choice about which foundation your procurement and finance functions will run on as they move toward autonomy, the platform you will still want to be on in 2030.

Governance today, autonomy tomorrow

Right now, somewhere in your business, an employee is still running a contract through a tool no one approved. They aren't trying to be reckless but do what everyone has always done throughout history: move faster and get more out of paradigm-shifting new tools.

Shadow AI and stalled pilots teach the same lesson: both expose the value AI hasn't yet delivered. In procurement and finance, governance is the precondition for that value. It is not a safeguard bolted on after; rather the governed middle layer is what turns a probabilistic model into work you can trust to run on its own. And indeed, only trusted work compounds into what everyone is after: faster cycle times, captured savings, decisions that hold up in an audit, and eventually autonomous processes.

The potential of AI in procurement and finance is real, but only for the teams that lay the governed foundation first. Get the middle layer right, and the compounding effect will follow.

About Zip

Zip is an AI procurement orchestration platform, built to be the governed middle layer this paper describes: the context, orchestration, guardrails, and observability that let procurement and finance teams run AI on regulated work.

Sources

[1] Zip, State of AI in Spend 2026. Survey of 1,050 procurement, finance, IT, and operations leaders at companies with 300 or more employees, fielded April to May 2026. All figures are self-reported. Available at zip.com/resources/state-of-ai-in-spend-2026.

[2] Bloomberg, Uber Caps Usage of AI Tools Like Claude Code to Manage Costs, June 2, 2026. Uber exhausted its annual 2026 AI budget in four months and capped employees at $1,500 per month per agentic coding tool.

[3] The Hackett Group, 2026 Procurement Agenda & Key Issues Study. Among procurement organizations deploying AI, reported gains were 9.7% in productivity and 9.3% in cycle-time reduction; 56% had deployed agentic AI in pilot or at scale. thehackettgroup.com/insights/2026-procurement-key-issues-2601/

[4] Challapally, A., Pease, C., Raskar, R., et al. The GenAI Divide: State of AI in Business 2025. MIT Project NANDA, July 2025. Externally sourced tools reached deployment at roughly twice the rate of internal builds (about 67% success versus roughly one-third for in-house efforts).

Written By
Nick Heinzmann
Head of Research
Nick Heinzmann is the Head of Research at Zip, the world's leading procurement orchestration platform. With a deep understanding of procurement trends and a knack for uncovering actionable insights, Nick helps leaders navigate the evolving procurement landscape. His expertise fuels Zip’s research initiatives and thought leadership content.

AI procurement orchestration, from intake to pay

Enter your business email to keep reading