AI entered most companies as a twenty dollar ChatGPT Plus charge on one employee’s corporate card. Then a Claude Pro charge from someone else. Then a team plan, a few GitHub Copilot seats for engineering, a Perplexity subscription a department head approved on their own. The technology arrived one line item at a time, and finance is only now treating it as a category worth governing.
AI behaves like a blend of software, labor, and infrastructure, and the disciplines that already apply to those categories apply here.
Audits are dull, and they are still the right place to start. Before approving anything new, account for what is already running.
AI waste tends to hide in three places:
Standalone subscriptions, the ChatGPT, Claude, and Perplexity seats paid by card or pushed through expense reimbursements.
AI built into software the company already owns and already pays for: Microsoft 365 Copilot riding on existing Office licenses, Gemini bundled into Google Workspace, plus the AI add-ons in Notion, Slack, and Salesforce that get switched on and almost never measured.
Employee tooling that is free today or personally paid, the Cursor seat or the second ChatGPT account someone runs out of pocket until a manager eventually agrees to expense it.
The audit should also surface duplication. Three teams may each be paying for their own research layer, one on Perplexity, one on Glean, one routing the same questions through ChatGPT. Engineering may run Cursor while GitHub Copilot already comes bundled in the GitHub Enterprise seats the company pays for, so the same developer is covered twice.
One error is to treat every AI charge as innovation spend. Some of it is shadow IT with better branding. The opposite error is to treat all shadow AI as a problem to be stamped out. Some of those unofficial tools are doing real work.
This is why the audit has to be part financial and part ethnographic. Finance systems will tell you what is being paid for. They will not tell you what matters. For that you talk to the power users. Ask what they reach for every day, what they tried and abandoned, what they would fight to keep, and where slow approvals pushed them into workarounds. Employees can tell you which tools are toys and which aren’t
A demo shows you the tool at its best, in conditions the vendor controls. A slick walkthrough of a support agent like Sierra or Decagon clearing tickets, or of GitHub Copilot scaffolding a feature, proves the tool works in the vendor’s hands. It does not prove the tool works on your data, your edge cases, and your team. A pilot does, and that is the only thing you are buying. The pilot is also where the build versus buy question gets answered, because running something for thirty days tells you how much of it you could have assembled from a foundation model and a weekend.
A serious pilot runs at least thirty days. It starts from a baseline, because savings claims are meaningless without one. How long does the task take today. What does it cost. What quality is acceptable. What error rate can you tolerate. What human review does the work require before it ships.
With that baseline in place, the pilot should measure four things:
Usage - Does your team continue to use it
Productivity - Does it save your team time
Quality - Is it improving quickly
Cost- What are you paying for intelligence
Ordinary software improves on a predictable curve. You buy a version, it gets patches, it gets better slowly. AI does not behave that way. The underlying model can change, the pricing can change, the leading vendor can change, and your own use case can change, sometimes inside a single contract term.
A tool that looked excellent at purchase can become redundant because a foundation model absorbed its core feature. Standalone AI copywriters like Jasper and Copy.ai built businesses on a capability that GPT-class models now do natively for the price of a chat subscription. Grammarly spent years as the default writing checker, watched general assistants close most of the quality gap for free, and responded by acquiring Coda and the Superhuman email client and rebranding the entire company around a broader bet. A workflow that seemed too risky last year can become viable once better controls ship. What you are buying keeps changing under you, so the cadence of your review should match the cadence of the change.
Agentic systems move the buying decision from “can it answer” to “can it act.” A tool that drafts a reply for a human to send is one kind of purchase. Superhuman Go pulling your calendar to schedule a meeting, Salesforce Agentforce working a support case end to end, a Cursor agent committing code, or a Ramp agent approving an expense is another kind entirely.
Finance does not need perfect technical fluency before it can govern AI. Waiting for that fluency is how the spending got ahead of the policy in the first place. The core disciplines transfer directly. Know what you spend. Require an owner for every tool. Test before you scale. Define the return you expect and check whether you got it. Cap the downside. Avoid lock-in that removes your ability to switch. Review on a schedule. Kill what does not work.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.