AI in Accounts Payable: Separating Real Automation from Marketing
- What does AI actually do in accounts payable today?
- Which AI claims hold up, and which are rebranded rules and OCR?
- What should you ask a vendor to test an AI claim?
- Where does AI measurably change AP performance, and where does it not?
- What are the risks of putting a model in the payment path?
- Where Corpay AP automation and Corpay AI fit in your process
AI in accounts payable means using machine-learning models instead of fixed rules to read invoices, propose GL coding, flag anomalies, and answer questions about spend. What ships under that label varies enormously from one vendor to the next.
The gap has a name. Gartner's June 2025 research on agentic AI coined "agent washing" for the practice of relabeling old software, and put a count on it. Of the thousands of vendors claiming agentic capability, Gartner judged roughly 130 to be genuine. The rest were rebranded chatbots, RPA scripts, and assistants with a new coat of paint.
Sorting one from the other is a matter of knowing which questions produce a falsifiable answer in a demo, and which produce a slide.
Key Takeaways
Real model-driven capability in AP sits in four places today, including document extraction, coding suggestion, anomaly detection, and plain-language query over spend data. The rest of the workflow still runs on rules.
Gartner's agent-washing finding is the right default assumption for any claim involving autonomy, because most products marketed as agents are assistants with an approval step.
The test for learned extraction is a supplier the system has never seen, run against your own invoices during the demo.
A vendor who cannot produce a straight-through rate, an exception rate, and a confidence-score distribution has given you an answer.
Automation moves cost and cycle time. Approval judgment, supplier relationships, and the audit trail obligation stay exactly where they are.
Once a model touches the payment path, remittance and vendor banking fields turn into a security exposure that accuracy metrics never capture.
What does AI actually do in accounts payable today?
Four AP tasks have genuine model-driven capability in production. A model reads unstructured invoice documents, suggests GL coding, flags transactions that look unusual for a given supplier, and answers plain-language questions about spend data.
Adoption has moved quickly enough that ignoring the category is no longer a defensible position. According to Ardent Partners' 2025 report Accounts Payable Metrics That Matter in 2025, 44% of AP teams already use AI in some form and 75% expect to within 12 months.
Everything around those four tasks is still rules and workflow. Approval routing follows a matrix somebody configured, payment-method selection follows policy, and duplicate checks compare fields the way they always have. When a demo presents everything AP automation covers as AI-driven, what sits underneath is usually a conventional workflow engine with a model attached to the capture step. Walk the AP process end to end and the model touches two or three stations on a line of a dozen.
Where the model's output lands matters as much as the output itself. A coding suggestion that never reaches the general ledger cleanly is a suggestion your team will re-key by hand. Teams running AP automation against NetSuite tend to find this out during the first month-end close, when the sync either holds or it doesn't.
What is the difference between AI, machine learning, and a rules engine?
A rules engine applies conditions a person wrote. A machine-learning model infers patterns from examples and shifts as it sees more of them. AI is the umbrella over both, which is why the term works so well in marketing copy and so badly in a requirements document.
The practical difference shows up at the edges. A rule that routes every invoice over $10,000 to the controller behaves identically on invoice one and invoice one million. A model trained on your invoice history behaves differently as that history grows, which is both the upside and the audit problem.
Rules are predictable and blind. Models adapt and resist explanation. Most working AP stacks run both, and a straight-talking vendor will tell you which layer handles what.
What are AI agents in accounts payable?
An AI agent is software that plans and executes a multi-step task with limited supervision, instead of answering one prompt at a time. In AP, that would mean a system that receives an invoice, resolves a PO discrepancy with the supplier, adjusts the coding, and releases payment with no person in the loop.
Very little on the market does that today. According to Gartner's June 2025 press release Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, escalating costs, unclear business value, and inadequate risk controls will kill more than four in ten of these projects before the decade turns.
Ask what an agent in your AP process would need permission to do. It needs authority to change a payment amount, alter a vendor master record, and release funds. Most finance teams, hearing that list read aloud, decide they want an assistant with a human approval gate. That is what most products described as agents already are.
Which AP tasks still have no working AI?
Several, and they cluster around information that never made it into a system in the first place. Supplier enrollment is a phone-and-email problem involving a person at another company who has no particular incentive to answer you today. Dispute resolution depends on knowing which relationships tolerate pressure and which don't.
The tasks that resist automation share a shape:
Supplier enrollment and banking-detail collection, where the bottleneck sits at another company.
Dispute resolution, where the correct answer depends on the relationship rather than the document.
One-off exceptions justified by a conversation that nobody logged.
Judgment calls a reviewer has to defend later, including most of the fraud patterns a model gets asked to catch.
None of that is a knock on the technology. It's a reminder that the messy middle of AP was always the expensive part, and models were trained on the tidy part.
Which AI claims hold up, and which are rebranded rules and OCR?
Extraction and anomaly scoring are usually real. Coding intelligence is partly real. Anything described as autonomous is almost always an assistant with an approval step, and the agent-washing pattern Gartner measured is reason enough to assume so until a vendor proves otherwise.
The failure rate at the ambitious end of the market is well documented. MIT Media Lab's Project NANDA study The GenAI Divide: State of AI in Business 2025 found that 95% of enterprise generative-AI pilots deliver no measurable P&L return, drawing on 150 executive interviews, 350 employee surveys, and 300 public deployments.
Extraction mechanics deserve their own treatment, and how invoice capture reads a document covers the detail. What matters for evaluation is narrower. You need to know whether the reading generalizes to a supplier the system has never processed.
Claim as vendors phrase it | What's usually underneath | The question that tells them apart |
"AI-powered invoice capture" | Template matching for configured suppliers, with a model covering the remainder | Can it read an invoice from a supplier we have never configured, live, right now? |
"Smart GL coding" | A lookup of the last coding applied to that vendor | Show me the confidence score behind this suggestion, then code a vendor whose spend splits across cost centers |
"AI fraud detection" | Duplicate checks and threshold alerts with a dashboard | What has it flagged that a rule would have missed? |
"Touchless processing" | A straight-through rate measured on the cleanest subset of invoices | What is the rate across our full mix, including scanned and multi-line documents? |
"Agentic AP" | An assistant that drafts an action a person approves | What can it execute without human approval, and what is the permission model? |
"Continuous learning" | A model retrained on a vendor-wide schedule | Does it learn from our corrections, and how long before that shows up? |
Sort every claimed capability into real, partly real, or rebranded before you compare prices.
How can you tell learned extraction from template matching?
Bring a supplier the system has never seen. Template-based capture needs a configured layout per supplier, so a new one either fails outright or lands in a setup queue. Learned extraction generalizes on the first pass, imperfectly, and improves from there.
The demo version of this test takes four minutes. Hand over a PDF from a supplier you onboarded last quarter and watch the screen. If the response involves configuring something during implementation, you have your classification, and no amount of roadmap discussion changes it.
Is the GL coding suggested by a model or copied from last time?
Ask for the confidence score. A model produces one and a lookup does not, and the difference is visible inside a minute.
Smart coding is frequently a recall of the last coding applied to that vendor, which works beautifully until a vendor's spend splits across cost centers or departments. Pull one of those vendors from your own ledger and watch what the suggestion does. If it returns the majority coding every time and never abstains, the system is recalling history and presenting it as a prediction.
Does anomaly detection find new patterns or replay old rules?
Duplicate-invoice checks and threshold alerts are rules with a dashboard on top. A model finds the invoice that is unusual for this supplier at this time, whether that's an amount outside their normal range, a sudden change in terms, or first-time bank details on an old relationship.
The stakes justify the scrutiny. The Association for Financial Professionals' 2025 AFP Payments Fraud and Control Survey Report found that 79% of organizations were targets of attempted or actual payments fraud in 2024.
Duplicate detection is the fair benchmark for any anomaly claim, because the duplicate-payment problem is well understood and rule-based systems have handled it for years. If a vendor's anomaly story amounts to catching duplicates, they have sold you a feature that shipped two decades ago.
What should you ask a vendor to test an AI claim?
Ask questions that produce a falsifiable answer in the room. The general evaluation ground of payment methods, security, ERP fit, and exception handling is already covered by the seven questions worth asking any AP automation vendor. What that list doesn't reach is the AI layer, where the claims are newest and the proof is thinnest.
Protect cash flow with modern AP
Modernize AP to cut costs, speed approvals, and mitigate payment risk — gaining the real-time visibility to protect cash flow and scale with confidence.
Download the whitepaperWhat has to be in the demo for it to mean anything?
Your invoices, not theirs. A curated demo set proves that the software works on curated invoices, which nobody was disputing.
A supplier the system has never processed, added live while you watch.
A scanned document sitting alongside a clean digital PDF.
A multi-line invoice with a table that wraps across pages.
An invoice with a missing PO number, where extraction quality usually shows itself.
One genuinely ugly invoice, the one your AP team complains about by name.
Send them ahead and ask that they be loaded cold. If the vendor needs a week to get ready, the preparation time is itself the answer to how much configuration the software demands before it performs.
Which numbers should a vendor produce on request?
Four numbers, and a vendor who tracks them will have them on hand.
Straight-through rate measured on an invoice mix that resembles yours in supplier count and document quality.
Exception rate, with their definition of an exception in writing, because definitions vary wildly.
Confidence-score distribution, which reveals whether the model abstains when it should.
Time to reach that straight-through rate at a new customer, measured in weeks.
A vendor who can't produce these is telling you the number isn't tracked, which is a signal about what the roadmap prioritizes.
Who works the exception when the model is unsure?
Someone does, and the answer decides whether automation reduces your workload or relocates it. Software that flags an exception has finished its job the moment the item appears in a queue. A service that works the exception contacts the supplier, chases the missing PO, re-keys what needs re-keying, and closes the item.
The most common complaint from AP teams who have lived through a rollout is that the work changed shape without shrinking. That happens when nobody owns the exception queue, and it is the most predictable failure mode in this category. Ask who staffs it, where they sit, and what their volume looks like on a bad week.
Where does AI measurably change AP performance, and where does it not?
Cost per invoice and cycle time move. Judgment, relationships, and control obligations do not. The pressure to chase the first without acknowledging the second is genuine, and more than half of the CFOs surveyed in Deloitte's 2026 CFO Signals Spotlight 1Q26 said their CEOs have asked them to focus on managing and reducing costs.
What changes about cost and cycle time?
Processing cost, mostly. According to Ardent Partners' 2025 report The State of ePayables 2025, the average cost to process a single invoice is $9.84.
APQC's 2024-2025 accounts payable benchmarking cycle puts median full-process cost per invoice at $21.40, with the top quartile reaching $10.18. The spread between those two numbers is wide enough that the distance between an average AP function and a good one is worth more than most software decisions on the table.
Time follows cost. According to the Institute of Finance & Management's 2025 AP benchmarking research, manual tasks consume 84% of the average AP practitioner's time, and that is the pool automation draws down. A model that reliably extracts and codes the bulk of your invoice volume takes real hours out of the week.
What does AI not change?
Controls, segregation of duties, and the person who signs. A suggestion is an input to a decision somebody still owns and still has to defend.
Auditors care that the approval happened, that the approver had authority, and that evidence exists, whichever layer produced the coding. What an AP audit examines doesn't shift because a model made the first pass. If anything it gets harder, because you now have to explain a decision the software can't fully articulate.
Why do so many AI pilots stall before production?
The reasons are organizational more often than technical. Pilots stall on data access, on workflow fit, and on the absence of anyone whose job description includes the exception queue. The model usually works. The surrounding operation usually isn't ready.
Anyone who has watched how AP automation evolved to this point recognizes the pattern, because it predates the current wave by two decades. The same holds for the myths that attach to payment automation, which have simply acquired a newer vocabulary.
What are the risks of putting a model in the payment path?
Remittance and vendor banking data carry the exposure. An extraction error on an invoice line is an accounting problem you catch at reconciliation. An extraction error on a bank detail is money leaving the building.
The FBI Internet Crime Complaint Center's 2025 Internet Crime Report recorded $3,046,598,558 in reported business email compromise losses across 24,768 complaints, with 86% of those losses moving by wire or ACH. Vendor-banking fields are the target, and any system that writes to them earns the scrutiny you would give a wire approval.
The rail underneath is enormous. Nacha's January 2025 release Same Day ACH Passes Major Milestone in 2024 as the ACH Network Shows Higher Growth reports that the ACH Network carried 33.6 billion payments worth $86.2 trillion in 2024, up 6.7% in volume year over year. B2B payments on the network reached 7.3 billion in the same year, up 11.6%.
Most of the fraud reduction credited to automation comes from control design, not model intelligence. Removing manual touchpoints from invoice processing closes more exposure than any detection feature a vendor demos.
What happens when a model is confidently wrong?
It routes for approval like anything else, which is precisely the problem. A confidence score earns its keep only when the system abstains below a threshold you control, and a model that never abstains is worth failing a vendor over.
Ask where the threshold lives, who can change it, and what happens to items that fall below it. If everything routes to approval regardless of confidence, the score is decoration.
How does AI interact with fraud controls and segregation of duties?
A suggestion is not an approval, and that boundary has to live in software where it can be enforced. The model can propose coding, propose a match, and propose a payment date. Approval authority stays with named humans, and vendor-bank-detail changes stay outside the model's reach entirely.
Designing the approval workflow a model routes into matters more once suggestions enter the queue, because a reviewer who sees a confident recommendation approves faster and looks less carefully. Three-way matching is the control that extraction feeds, and it only works when PO and receipt data are as clean as the invoice data.
What should security review before you sign?
Three questions about the model, plus the usual attestations about the vendor.
Where model training data lives, and whether your invoices train a shared model that serves other customers.
How the vendor documents an AI-influenced decision for an auditor, in a form the auditor will accept.
What happens when the vendor changes the model, and whether you are notified before it reaches production.
On the attestation side, ask for the report, not the badge. Corpay is SOC 2 Type II compliant, and Comdata, a Corpay company, is a PCI DSS Level 1 service provider. Scope matters more than the label in either case, which is why the underlying report is the document worth reading.
Where Corpay AP automation and Corpay AI fit in your process
The hole in most AI pitches is the exception queue. A model that extracts well still produces items it can't resolve, and somebody has to work them. That is the gap software alone doesn't close, and it's the reason a managed service sits alongside the automation rather than behind it.
Corpay AI is a conversational intelligence layer built into every Corpay platform. You ask questions in plain language and get answers you can act on. It summarizes vendor spend across categories and highlights the trends underneath, points to approval bottlenecks and outstanding items, and surfaces opportunities around payment timing and terms. For an AP manager it behaves like a concierge for invoices, approvals, and vendor questions.
The automation underneath has to work first. Corpay AP automation captures invoice data automatically, routes it to approvers, and removes manual entry, with in-field auto-coding and sync back to your ERP. That sync spans 180+ ERP integrations, including NetSuite, Sage Intacct, Microsoft Dynamics 365, and Acumatica, with QuickBooks supported as well. Corpay AI sits on top of that data, and invoice automation is where capture and coding happen.
See how Corpay handles your invoice mix in a walkthrough built on your own documents.
Frequently Asked Questions
What is AI used for in accounts payable?
AI in AP is used for four things today, including reading invoice documents, suggesting GL coding, flagging transactions that look unusual for a supplier, and answering plain-language questions about spend. Approval routing and payment execution remain rules-driven workflow.
What are AI agents in accounts payable automation?
An AI agent plans and executes a multi-step AP task with limited supervision, such as resolving a PO discrepancy and releasing payment. Most products marketed as agents today are assistants that draft an action for a human to approve.
How does AI reduce errors in accounts payable?
Machine-learning extraction reduces keying errors by pulling invoice fields directly from the document, and anomaly scoring surfaces transactions that deviate from a supplier's normal pattern. Errors that come from bad approval design or dirty vendor master data are unaffected.
What are the benefits of AI in accounts payable?
The measurable benefits are lower processing cost per invoice and shorter cycle times, driven by less manual keying and faster exception triage. Control obligations, approval judgment, and supplier relationship work stay with your team.
How is AI being used in accounts payable today?
Most production use sits in document capture and coding suggestion, with anomaly detection and natural-language spend queries gaining ground. Fully autonomous invoice-to-payment execution is rare, and the projects attempting it have a high cancellation rate.
How can you tell if a vendor's AI claim is real?
Test it on a supplier the system has never processed, using your own invoices, and ask for the confidence score behind a coding suggestion. Then request a straight-through rate, an exception rate, and a confidence distribution in writing.
- What does AI actually do in accounts payable today?
- Which AI claims hold up, and which are rebranded rules and OCR?
- What should you ask a vendor to test an AI claim?
- Where does AI measurably change AP performance, and where does it not?
- What are the risks of putting a model in the payment path?
- Where Corpay AP automation and Corpay AI fit in your process
Switch to Corpay
Discover how making the move to Corpay streamlines payments and strengthens your business.
Talk to an ExpertSmarter payments. Stronger growth. Keep business moving.
Corpay powers payments for 800,000+ businesses worldwide. Let’s build what’s next for yours.