AI in Payments: Where Ledgers Fit in the Agentic Stack
By Formance Staff
Payments
Ledger
Loading article content
Loading table of contents
Loading author
_LEDGER/
Post your first agent-stamped transaction
Clone the open-source Formance Ledger on GitHub, run it locally, and post your first agent-stamped transaction against it before any of this lives in production.
AI payment systems sit behind procurement bots, treasury sweeps, marketplace payouts, and usage-billing services, but the ledgers underneath them may still assume a person clicking checkout once.
Consider a supplier payout request that times out at the payment service provider (PSP), prompting the AI agent to retry three times in 400 milliseconds. The engineering team opens the ledger and cannot say whether one, two, or three payments were recorded. The agent's log shows three attempts and one success reply, but a success reply does not prove only one payment was recorded. The ledger has to be the component that knows, and in most AI payment stacks, it does not.
The core problem is a missing boundary between layers. Agentic payment protocols like AP2, ACP, and x402 stop at authorization and leave the recording work to the implementer. A ledger built for human checkout has no rule for an AI agent that retries at machine speed.
The problem scales with the market: agentic commerce transaction value is projected to reach $3.5 trillion by 2031, so every ledger built on the human-checkout assumption inherits the same failure mode at that volume.
This guide shows where the ledger fits in the agentic payment stack (agent, policy, ledger, execution adapter, reconciliation), why agent requests must stay separate from ledger records, and what to put in an agent's command contract before the first live payout. Everything about how AI agents move money follows from keeping the agent and the ledger apart.
Why agentic payment protocols stop at authorization
Agentic payment protocols stop at authorization because their job ends once the agent has proven permission to pay. Settlement, recording, and reconciliation fall to the implementer: the team building the ledger, orchestration, and reconciliation layer on top of the protocol.
Three protocols draw the same boundary from three different situations.
1. AP2 stops at the signed mandate
AP2 defines Checkout and Payment Mandates as signed SD-JWTs, verified cryptographically by the Merchant, Credential Provider, and Network. The current specification explicitly excludes deterministic signature schemes like Ed25519 and requires schemes like ECDSA instead.
Once those mandates are signed, the AP2's job is done. The specification says nothing about where the resulting ledger record lives or how it is stored.
2. ACP stops at the Shared Payment Token
ACP, an open standard created by Stripe, OpenAI, and Meta, defines how an agent completes a purchase without ever holding card details.
A Shared Payment Token passes from buyer to agent to merchant, scoped to a specific merchant and cart total. The merchant stays the merchant of record and charges the token through its existing payment provider. ACP hands off at the token and not at the ledger.
3. x402 stops at on-chain settlement
x402 defines a request-and-settlement cycle over HTTP. A 402 response, a signed payment payload, and a facilitator that confirms on-chain settlement. Nothing in the cycle specifies protocol-level holds or mandatory human approval. Both live in the surrounding architecture and not in the x402 protocol.
Five responsibilities the implementer owns
Five responsibilities past the protocol boundary fall to the implementer:
1Idempotency key lifetime
2Agent identity written into every record
3Fund holding and release
4Human approval as a persisted ledger state
5Reconciliation across payment rails with different settlement speeds
Ledger ownership is a money movement architecture decision, and AP2, ACP, and x402 leave that decision to the implementer.
How AI payments break a ledger built for human checkout
AI payments break a human-checkout ledger through two machine-speed failures that human-paced controls cannot catch: duplicated records from retries and out-of-scope records from AI drift. An AI planner retries and drifts faster than any control designed around a person's pause between clicks, so both failure modes surface at machine speed.
Duplicate records from retries
Duplicate records happen when the agent retries a request the ledger has already accepted, and no idempotency check catches the repeat.
The duplicate-record failure is shown by a scheduled transaction that runs and produces duplicate debits and downstream fees. An idempotency check at the ledger layer prevents duplication.
Out-of-scope records from AI drift
Out-of-scope records happen when the command the agent sends does not match the payment type or amount that was authorized. Nothing at the write boundary checks the difference between what was authorized and what was sent.
The out-of-scope failure is shown by the Citibank/Revlon wire error. On August 11, 2020, Citibank sent approximately $900 million, including principal, on a Revlon loan when only an interest payment had been authorized. Three reviewers approved the transfer before it went out.
The Citibank/Revlon wire error shows why implementations should check the payment type and amount against what was authorized before recording it.
Why agent speed triggers these failures
Agent speed makes both duplicate records and out-of-scope records more likely when controls assume a person is pausing.
An agent that times out may retry immediately and repeatedly without checking the account, and an AI planner may send a command whose amount or recipient drifts from the mandate it was given.
Over 40% of agentic AI projects are forecast to be canceled by the end of 2027, with weak risk controls among the reasons. A person may wait, refresh, and check the statement before clicking again, but an agent removes the pause, leaving the ledger's write-time checks as the only remaining safeguard.
Where the ledger fits in the agentic payment stack
The ledger sits at the third of five layers in the agentic payment stack (agent, policy, ledger, execution adapter, and reconciliation) as the only component allowed to change balances.
At Formance, we've built the open-source ledger at this third layer for teams running agent-initiated payments. It has four key properties: read-only agent access enforced at the tool surface, idempotency in the policy layer, all-or-nothing writes, and reconciliation as a separate feedback loop.
Layer
Input
Output
Not allowed to own
Agent
Task, mandate, model context
Proposed command with mandate ID, agent ID, model version, and idempotency key
Balances, records, and finality
Policy
Proposed command, with budget, allowlist, and threshold rules
Balanced transaction template, or a pending-approval state
The decision to skip a threshold
Ledger
Transaction template after a saved idempotency check
Recorded or rejected transaction as one atomic unit
Rail state and agent reasoning
Execution adapter
Committed hold
Rail instruction, and asynchronous confirmation or rejection
1. Agent layer: read-only proposals via credential-scoped tools
The agent proposes commands and never writes to the ledger, and the read-only boundary is enforced at the tool surface.
Every field the agent emits (mandate ID, agent ID, model version, and idempotency key) is meant for the policy and ledger layers to validate. The agent keeps task state, conversation context, and its own attempt log, but no balance.
Block agent write access at the tool surface so the model never holds a credential that can change a balance. Read access can include balances, transactions, and metadata.
2. Policy layer: money logic as code the finance team can read
The policy layer runs budget and allowlist checks outside the model, reads a funded account for the budget check, and confirms the destination is a known supplier.
It also applies the threshold that decides whether the command posts now or waits for a signed human approval, along with supervisory dashboards and overrides that make agent-initiated payments supervisable.
3. Ledger layer: all-or-nothing writes with idempotency short-circuits upstream
Each transaction is either recorded in full or not at all. When the agent retries, the policy layer catches the repeat and returns the original record instead of writing a new transaction.
Idempotency in the policy layer is what makes retry safety hold under agent-speed retries, and securing ledger integrity at the write boundary keeps each accepted transaction all-or-nothing.
The policy layer compiles the agent's proposed command into a Numscript transaction (Formance's purpose-built language for describing financial transactions) before the Formance Ledger sees it.
4. Execution adapters: async rail I/O that becomes the next ledger record
An execution adapter takes a committed hold, sends the rail instruction, and later returns a confirmation, a rejection, or nothing until a timeout fires. Each outcome becomes the next ledger record.
The adapter sits outside the ledger whether it fronts a PSP, a bank, a custodian, or an on-chain facilitator. Keeping the adapter out of the ledger means rail-specific quirks (batch windows, chain reorgs, and PSP retry semantics) never leak into the balance-writing layer.
5. Reconciliation layer: drift exceptions for operators, never the agent
The reconciliation layer flags drift as an exception that an operator or workflow, and never the agent, resolves.
Drift is treated as a signal, so every gap between ledger and provider produces a traceable record that an operator must sign off on.
How a mandate becomes a record through hold, settle, and release
An agent mandate maps to three ledger events (hold, settle, and release), with rail finality windows setting how long each hold may wait before it commits or reverses.
The policy layer's approval creates a hold against a funded budget account, rail confirmation creates a settlement record, and a timeout, cancellation, or rejection creates a release.
The hold is the only event the agent triggers directly and the only one that needs the agent's identity stamped on it, because settlement and release react to adapter events.
A procurement fleet reserving a $184,000 supplier payout
A procurement fleet is funded with $2,500,000 for the quarter from @platform:treasury:operating.
The fleet's budget lives in @platform:procurement:fleets:fleetA:budget, and mandate 7f3c uses @platform:procurement:fleets:fleetA:mandates:7f3c:hold for its hold.
The agent wants to reserve a $184,000 supplier payout under mandate 7f3c. The hold moves the amount from the budget account into the hold account tied to the mandate, with the agent's identity stamped in the transaction metadata.
// AGENT_PAYMENT_HOLD// Event: reserve USD 184,000 from the procurement fleet budget for mandate 7f3csend [USD/2 18400000] ( source = @platform:procurement:fleets:fleetA:budget destination = @platform:procurement:fleets:fleetA:mandates:7f3c:hold)set_tx_meta("event_type", "agent_payment_hold")set_tx_meta("mandate_id", "7f3c")set_tx_meta("agent_id", "proc-fleet-a-worker-07")set_tx_meta("model_version", "planner-2026.08.2")set_tx_meta("idempotency_key", "mandate_7f3c:hold")
If the budget account cannot cover the send, the hold fails, and nothing is recorded. A ledger-enforced budget stops the hold at write time rather than tripping a counter after money has already moved.
The submitter or policy layer saves the idempotency key (mandate ID plus a stable action name, like mandate_7f3c: hold), returns the previously recorded transaction on an identical retry, and rejects reuse of the key with different parameters.
Tie the key's lifetime to the mandate's expiry, so a retry after the mandate expires fails rather than opening a new hold.
A settlement confirmation from the adapter moves the held amount from @platform:procurement:fleets:fleetA:mandates:7f3c:hold to @counterparties:psps:pspA.
A rejected approval or an expired mandate returns the amount to the budget account as a reversing record, and the original hold record stays in the ledger's append-only log.
The adapter reports rail events like an x402 on-chain confirmation. The agent must not treat a rail event as a committed ledger record, because the ledger only changes state when the settlement record is written.
Rail finality sets the timeout, and the agent reads only four states
The hold-settle-release cycle above assumes the rail commits within the hold's timeout. Different payment rails become final at different times, so the maximum hold wait should match its rail's finality window.
Once a transaction is final on the rail, errors become difficult or impossible to reverse, so the ledger must support the reversal path before the first live payout. Understanding transaction finality on each rail sets the timeout.
The agent sees four states, and only these four: pending is the hold, committed is the settlement record, failed is a rejected hold, and reversed is a compensating record written after the fact.
Anything else the agent thinks it knows about the payment (rail confirmations, in-flight retries, and adapter telemetry) is not a state the ledger recognizes, so the agent should not act on it. Keeping the agent's read model this narrow is what stops it from double-spending against a hold that hasn't yet resolved.
How reconciliation corrects agent-assumed spend
Reconciliation corrects what the agent thinks it has spent by comparing the agent's state against the ledger and the provider's settlement file, then flagging any gap as a drift exception.
Compare ledger balances against provider balances per asset, so @counterparties:psps:pspA in the ledger is checked against the PSP's reported USD balance.
The agent cannot clear the drift exception because the agent has no path to write the correction. An operator resolves a hold the PSP never saw, or a settlement the ledger never recorded, or a reversing record does.
After an authorized correction is written, the agent's next budget read returns the corrected balance.
A ledger for AI agents that doesn't utilize the reconciliation loop lets the agent's attempt log stand in for the truth. The account reconciliation patterns that work for a human-paced platform apply unchanged, at a shorter interval (minutes to hours instead of daily).
Why AI payments need a ledger built for agents
An AI payment stack without a real ledger will hold up under demos and single-agent tests until the first machine-speed failure lands: a retry storm, a drifted mandate, or a rail confirmation the agent trusted too soon.
The question becomes operational: who authorized what, when, and against which budget. If the answer lives only in the agent's log, it's not an answer.
The agentic stack must put the ledger in the middle and keep the agent read-only above it. Idempotency and human-approval thresholds run in a policy layer the finance team can read, while rails sit behind execution adapters, isolated from the balance-writing layer. Reconciliation runs as a separate feedback loop, where drift becomes a signal rather than a silent correction.