On Tuesday, September 22, the institutions that actually move the money filed their opening brief in the agentic commerce wars. Six banks on three continents: ASB Bank, Bank of America, Capital One, Commonwealth Bank of Australia, ING and NatWest. Their joint paper, Building Trust in Agentic Commerce, is voluntary, nonbinding, and contains no timetable. It is also the clearest signal yet that the payment system’s custodians have looked at the 2026 shopping stack and decided it needs adult supervision.

The banks’ posture is neither hype nor refusal. Agentic commerce, they write, could become “a mainstream way in which consumers and merchants transact,” and in the paper’s own words quoted by Gizmodo: “We, too, are excited by the promise of agentic commerce and are eager to work with customers, industry and stakeholders to enable its future.” The excitement comes with a risk register attached, spanning “transparency, safety, privacy & data, choice, and interoperability,” which grows “as greater autonomy is given to AI agents.”

Read the paper next to the week’s headlines and the urgency explains itself. But the most interesting thing in the document is not what it warns about. It is the thing its framework quietly cannot do, and which nobody at the payments table currently owns.

What the paper actually proposes

Five principles: transparency, safety, privacy and data, choice, and interoperability. Beneath them sits one concrete mechanism, the paper’s centerpiece, detailed by PYMNTS: providers should preserve an auditable record of what a customer asked an agent to do, what authority the customer granted, and what happened before and after payment. Evidence of consumer instructions, authentication, intent, transaction decisions and outcomes. Records of warnings or interventions. The explicit purpose: giving participants a basis to investigate scams, recover money and resolve disputes.

The problem being solved is that the payment chain has gone blind in real time. Issuers and acquirers may lack access to an agent’s identity, the merchant of record, the customer’s intent and the purchase details. Customers may not know whom to contact if an agent buys the wrong item, overspends or falls for a scam. Merchants fear disputes and chargebacks arising from decisions they did not control.

On safety, the paper is unusually specific about the malpractice it has already observed in the wild: agents asking users for card details and typing them directly into websites, and agents favoring payment methods with weaker consumer protections, because a payment rail with fewer safeguards is friction-free, and friction is the enemy of a fast checkout. On transparency, the banks want consumers, merchants and payment providers to know when an agent is involved and on whose behalf it acts, and, in the paper’s most commercially pointed line, agents should explain how they prioritize products and payment methods, including sponsored options. An agent that quietly ranks a higher-commission product first is the agentic version of an undisclosed paid placement, and the banks have just called it what it is.

Mark Monaco, Bank of America’s head of global payments solutions, framed the mandate in one passage: “As agentic commerce continues to evolve, establishing trust and confidence across the ecosystem will be critical to its long-term success,” adding that “building confidence among consumers, merchants and financial institutions will require thoughtful approaches to identity, authorization, fraud prevention, liability management and customer protection.”

A follow-up paper will translate the principles into protocols and standards. What arrives now is the framework, and frameworks tell you what their authors are worried about. This one is worried about money movement: who authorized it, who executed it, who pays when it goes wrong.

Why the banks moved this week

The banking sector does not publish six-institution joint papers prophylactically. It publishes them when the telemetry on its own rails starts to change shape.

The telemetry is loud. Meta’s Muse, launched September 8, passed 2.5 million downloads in two weeks, per Sensor Tower data shared with CBS News, enough to displace ChatGPT as the top free iPhone app. Amazon blocked Muse from its store on September 20 over agent identification and credential concerns. And Meta disclosed and patched a zero-day vulnerability in the assistant that could have let an attacker hijack Muse and inherit every permission the user had granted it: browsing, buying, paying. That last detail is the one that should keep a fraud team awake. An agent with delegated purchasing authority is not a chatbot. It is a standing payment instruction with a browser.

The consumer base beneath all this is already enormous and already nervous. PYMNTS Intelligence counts 50 percent of Americans having made a retail purchase with AI’s help, and 22 percent beginning their product research with an AI tool. IBM’s agentic commerce research puts 41 percent of consumers using AI assistants to research products and 31 percent using them to hunt deals. But the willingness collapses exactly where the banks’ paper focuses: only 24 percent of consumers say they would let an AI agent both shop and pay. In PYMNTS’ phrasing, “the change stops as the agent gets closer to the money,” because “the consumer seems comfortable delegating research, but much less comfortable delegating identity, payment choice or an irreversible decision.”

The banks quote the same hesitation in their own language: “Consumers are unclear if AI agents will act in their interests. They are concerned that AI agents may buy the wrong thing or spend too much, or even worse, lose their money to scams and fraud. They are not sure whether they will be protected or who they will need to go to if things go wrong.”

That last sentence is a banking problem. “Who they will need to go to” is literally the dispute-resolution architecture, and the paper admits it is currently broken across the value chain: “When things go wrong, there is unclear and inefficient allocation of liability, and disputes processes do not involve all relevant parties.”

What the banks got right: liability follows the entry point of error

Buried in the paper’s dispute principle is the most consequential sentence in the document, easy to miss: liability should reflect “where an error or risk entered the transaction.”

This is a genuinely important idea, and it deserves more attention than the paper’s release got. Today’s chargeback regime was built for card-present and card-not-clicked-by-a-bot commerce: the issuer, the merchant and the network argue among themselves, with the consumer holding a right of reversal that functionally pre-assigns blame to the merchant. Agentic commerce breaks that allocation in every direction. If an agent bought the wrong item, is that the customer’s fault for writing a vague instruction, the agent operator’s fault for a sloppy interpretation, the merchant’s fault for a misleading listing, or the platform’s fault for a manipulated ranking? The banks’ answer, that liability should attach to where the error entered, is the only structurally fair one. It is also, as we will see, unenforceable with the data the banks’ own audit trail collects.

The blind spot: the audit trail certifies the plumbing, not the product

Here is the limitation nobody is discussing. The proposed audit trail records the human side of the transaction with forensic precision: the instruction, the delegation of authority, the authentication, the payment decision, the outcome. It reconstructs what the customer asked for and what the agent did.

It certifies nothing about what the agent knew.

Consider the failure mode the banks themselves name first: the agent “buys the wrong thing.” Trace it with the proposed audit trail. The customer asked for a safe stroller under $300. The agent searched, ranked, shortlisted, purchased. Every step logged, every authorization timestamped, the reconstruction immaculate. And yet the answer to the dispute remains invisible to the record, because it was decided before the payment chain was ever invoked: the stroller topped the agent’s shortlist because it showed 4.9 stars and 12,000 reviews, and some meaningful fraction of those reviews were purchased, incentivized, or machine-written. Amazon’s own enforcement reporting puts suspected fake reviews blocked in the hundreds of millions per year, and independent measurement keeps finding AI-generated review text on bestseller pages. The audit trail will faithfully reconstruct a decision that was built on contaminated evidence. It records the recipe. It says nothing about the ingredients.

This is not a marginal gap. It is the difference between transactional trust and product trust, and the paper’s five principles govern only the first. Banks can certify that money moved with consent. Networks can certify that identity was verified. Neither can certify that the 4.9 stars were earned, because neither party observes the review corpus, and no party in the payment chain has any incentive to build that observation layer. The merchant will not audit its own reviews. The agent operator ranks whatever corpus it ingests. The bank sees an authorized transaction. The error entered through a door none of them was watching.

Where the error entered: the vacant seat at the dispute table

Return to the paper’s best idea: liability follows the point of entry. Now ask the operational question a dispute analyst would ask. When a customer disputes an agent purchase on the grounds that the product was misrepresented by its review record, who proves where the error entered?

The agent operator points to the marketplace: we ranked what the corpus said. The marketplace points to user-generated content: reviews are the words of customers, not ours. The merchant points to both. The bank holds a clean record of authorization. Everyone is telling the truth from where they sit, which is the structural signature of a missing party: an independent evaluator whose product-quality evidence predates the transaction and can be compared against what the agent consumed at decision time.

That party does not exist inside the payment chain, which is precisely why the banks’ paper cannot name it. It has to exist outside the transaction, answering to none of the parties whose incentives contaminate the inputs. This is the design space GoBuy occupies, and the mapping onto the banks’ framework is nearly one-to-one:

  • Filtered evidence before scoring. Smart Score, 0 to 100, is computed on review quality after manipulated and low-information reviews are removed. The score an agent consults is built on a cleaned corpus, not the raw one the marketplace displays.
  • Curation that cannot be sponsored. Top seven verified products per category, identical for every caller, with no paid placement. The shortlist was not purchased into existence, which answers the banks’ sponsored-priorization concern structurally rather than by disclosure.
  • Durability as the trust test. The GoBuy Verified badge requires 80 or above sustained for 90 days. A commission budget can buy a week of ranking. It cannot easily buy a quarter of verified quality.
  • Auditable by construction. Delivered over MCP at gobuy.ai/api/mcp, the trust consultation is itself a timestamped, machine-readable event. An agent that checks independent product evidence before purchase adds that check to the very audit trail the six banks just proposed: instruction, authorization, verification, payment, outcome.

That last point deserves emphasis, because it completes the banks’ own architecture. Their paper wants the path from instruction to payment reconstructed when things go wrong. An independent evidence layer makes the reconstruction meaningful. When the dispute is “the agent bought the wrong thing,” the record can show that the wrong thing scored 62 on filtered evidence and carried no verification, while the right thing scored 88 with a 90-day badge: the error entered at the ranking layer, and liability can follow it there. Without that layer, the same dispute is four parties shrugging at a clean payment log.

The counterargument from retail’s old guard

Not everyone believes the agentic moment is structural. Ron Johnson, the former Apple executive who built the original Apple Stores, told TechCrunch that AI “will improve the online shopping experience” without changing “the way we shop,” because “AI will never be able to have you physically experience a product. They’ll just become more informed shoppers when they come to the store.”

Notice what even the skeptic concedes: informed shoppers. The entire agentic stack, from Muse’s deal-hunting to the banks’ audit trail, is an information game layered over commerce. And an information game is only as good as its evidence. Johnson’s objection actually sharpens the point rather than blunting it. If agents produce more informed shoppers, the value concentrates in whoever supplies the informing, and the suppliers with clean inputs win the agents. Contaminated review corpora are not a fixed feature of the landscape. They are an unclaimed liability, and the first actor that absorbs them, at the evidence layer, becomes infrastructure.

What to watch

Four signals over the next quarter:

  1. The follow-up paper. The six banks promise a sequel translating principles into protocols and standards. Watch whether product-quality signals appear in it, or whether the framework stays confined to identity, authorization and payment. If it stays confined, the evidence gap gets solved by outside actors, which is already happening.
  2. Network rules. Mastercard and Visa are both building agentic payment capabilities. The moment audit-trail-style requirements enter network rules rather than voluntary papers, they stop being suggestions. Whether those rules touch ranking inputs is the tell.
  3. The first holiday-scale agentic dispute. Muse is at 2.5 million downloads and climbing into a shopping season where, per Coresight, 30 percent of consumers would already use AI to compare holiday prices. The first viral dispute over an agent-bought product will stress-test the liability vacuum in public.
  4. Regulatory convergence. The FTC’s personalized pricing comment window closes September 25, two days after the banks’ paper landed. Disclosure regimes built for human screens and machine-mediated shelves are on a collision course, and the audit trail concept gives regulators a ready-made enforcement instrument.

The banks are right that agentic commerce’s constraint is trust. Their paper builds the skeleton: instruction, authorization, payment, dispute. What it cannot supply, from inside the payment chain, is the part of the record that says the product was worth buying. That layer is being built outside the chain, deliberately independent of every party with an incentive to shade it.

Check what your agent is about to buy, before it buys it, at gobuy.ai. And if you build agents, wire them to product evidence the payment chain can audit: gobuy.ai/agent-docs.