Agentic commerce has spent two years arguing about architecture. Open agents that roam the whole web versus closed agents that live inside one retailer’s app. Protocol maximalists who want checkout inside the chatbot versus the everything-through-Visa camp. On September 8 and 9, 2026, the industry stopped arguing and ran the experiment instead, because two of the biggest companies on earth shipped shopping agents within 24 hours of each other, built on opposite philosophies.
Meta launched Muse on Tuesday, calling it “the world’s first personal AI agent built for everyone.” Instacart launched Clementine on Wednesday, an AI grocery assistant that turns a conversation, a list, or a recipe into a ready-to-buy cart for most U.S. and Canadian customers.
One of them works. The other one could not buy toilet paper. And the reasons why tell you exactly where agentic commerce is going, and exactly what it is still missing.
The Errand Test: What Happened to Muse
PYMNTS did the industry a favor this week and tested Muse the way an actual human would: three ordinary errands. Reorder six rolls of toilet paper on Amazon. Order a pizza from Domino’s. Book a restaurant reservation on Resy.
Muse struggled with all three. None of the three transactions completed.
The failure mode matters more than the failure. It was not reasoning. Muse could plan the task fine. The friction was everything required before the task: Amazon credentials had to be shared through a secure link before the agent could even see order history. Gmail and Calendar each demanded their own separate authorization flow. When payment time came, Muse routed through Link, Stripe’s checkout system, and Meta noted Muse is the first agent covered by Link’s purchase protections. Decent infrastructure. Still zero completed transactions.
For calibration, PYMNTS timed the same toilet paper reorder done the old way, directly on Amazon: under 30 seconds, no credential sharing, no app connections, no failed checkout.
The security architecture is genuinely serious, to be clear. Muse runs on a dedicated Muse Secure VM, a separate cloud computer housing the agent and the user’s connected data, with a “Sentinel” agent that must approve anything Muse sends out to the internet. That is real containment engineering, and it is more than most agent launches ship. But containment is not capability. An agent that needs three authorization ceremonies to attempt an errand a human finishes in 30 seconds is, as PYMNTS put it, an intern who needs detailed instructions, explicit account access, and supervision at every step. Delegation only helps if the handoff is faster than the task.
Mark Zuckerberg’s August vision essay promised that “everyone will have an exceptionally capable personal agent” whose agent “will work 24/7 on your behalf.” He also told the Sources podcast he expects Muse to eventually pay for itself, with Meta taking “a very small cut of whatever the transaction is.” Hold that business model in mind. It matters later.
Why Clementine Works: The Garden Owns Everything
Instacart’s Clementine is the counter-argument, and it is winning on the evidence.
Type “I need a week of budget-friendly kid lunches,” “high-protein easy dinners for two,” “everything for a football watch party for 20 people,” or simply “order my usuals,” and Clementine assembles the cart. Because it has access to Instacart’s lifetime orders, catalog, and daily inventory signals, it knows what is currently on the shelves of your chosen store, what is on sale there, and what you buy week after week. It generates recipes, applies your preferences, surfaces deals, and reorders from history. Then you check out inside the same app, with the same saved payment you have used for years.
CEO Chris Rogers framed the moat plainly: “We’ve spent nearly 15 years learning how families shop and eat, and Clementine puts that knowledge to work.” In May he put a number on it: “With over 1.6 billion lifetime orders, we have a unique and deep understanding of the grocery journey, and we’re using that to build the gold standard of agentic grocery AI.”
The Q1 2026 results behind that confidence: $10.29 billion in quarterly gross transaction value, up 13 percent, crossing $10 billion for the first time alongside $1.02 billion in total revenue. Ninety-one million orders. And Cart Assistant, the white-label enterprise version of the same technology, is already available to about 25 percent of U.S. customers, with Food Bazaar, Heritage Grocers Group, and Woodman’s live and Aldi U.S., Harmons, Save Mart, and Stew Leonard’s on the way. Instacart has also integrated with ChatGPT and Claude, so the garden is hedging its bets on where the conversation starts.
Notice what makes the two launches diverge. Muse fails because it must negotiate access to systems it does not own, one credential handshake at a time. Clementine succeeds because the catalog, the inventory, the order history, the payment method, and the checkout all belong to the same entity. The universal agent needs permission for everything; the garden agent needs permission for nothing. In agentic commerce, 2026’s version of the mobile-native versus mobile-web fight, the garden just took the round.
The Demand Side Is No Longer Hypothetical
Neither company is building into a vacuum. The PYMNTS Intelligence report “The 50 Million Consumer Migration” found that 49.6 million U.S. adults now begin retail product research with AI, including 39 million who have stopped relying primarily on the traditional search channel where they once started. AI recommendations have introduced 10.8 million consumers to a retailer they had not used before and moved 15.9 million toward a different brand.
The money follows. Consumers who tried a new retailer after using AI spent an average of $1,430 on AI-assisted purchases over three months, against $931 for AI retail researchers who made purchases, and $1,564 for those who made an unplanned purchase. Among those AI-influenced shoppers, 43 percent found a better price, 27 percent chose a different brand, and 26 percent bought a different product than originally intended.
Sixty-four percent of shoppers expect to shop with AI agents at least occasionally within the next two years, and 30 percent expect to do so frequently or almost always, per the Global Digital Shopping Index. Forty-eight percent of online shoppers already used AI to research their most recent purchase.
That last set of numbers contains the strategic asymmetry defining this whole market: 56 percent of consumers would let an AI agent search and compare products, but only 37 percent would let one authorize payments, and 35 percent would give one access to saved payment methods. Discovery is delegated. Money is not.
The Trust Ceiling, in One Word
Visa CEO Ryan McInerney named the ceiling at the Goldman Sachs Communacopia + Technology Conference on September 8, the same day Muse launched:
“We are seeing adoption for shopping, but not yet for autonomous payments.”
And then, the sentence every agentic commerce roadmap should be printed under: “The barrier to that, if I had to describe it in one word, would be trust.”
Three-quarters of consumers Visa surveyed do not trust agentic platforms to make autonomous payments with their money and financial information. Sixty-one percent said they would trust an agent to pay if Visa were involved, a figure that rises above 70 percent among people who use large language models at least weekly. Visa is leaning on it, pairing the shopping-side adoption with its $2.4 billion BioCatch acquisition to push fraud prevention upstream into identity.
Mastercard moved the same day Instacart did. On September 9 it expanded Agent Suite for Merchants and debuted Agent Connect, a single integration meant to connect merchants, AI agents, digital platforms, and payment providers. The framing deserves a close read: Agent Suite makes “merchant-defined information” including product details, pricing, availability, and fulfillment options “accessible across models, platforms and agent experiences while maintaining control over how that information is used, represented and acted upon.”
Payment trust has the card networks building agentic rails. Behavioral trust has Meta’s Sentinel VM and Anthropic’s approval gates. Read Mastercard’s language again and notice what every layer of this stack has in common: the seller defines the information, and the seller keeps control over how it is represented. The merchant’s interests are represented at the protocol level. The shopper’s are assumed to follow.
Merchants, for their part, are hedging too: 46 percent said pricing and the final price paid are the functions they are least willing to let AI agents handle, 31 percent plan to invest within a year in automated product search and comparison versus 26 percent for AI completing checkout on a customer’s behalf, and nearly half put autonomous checkout, multi-merchant bundling, and dispute management in the “later or never” category.
The Gardener Sells the Shelf
Which brings us to the part of Clementine’s launch that got the least scrutiny: the gardener monetizes the garden, and the garden’s fastest-growing revenue line is advertising.
Instacart runs three engines: the consumer marketplace, the enterprise platform, and an advertising ecosystem for brands. In Q1 2026, advertising and other revenue hit $286 million, up 16 percent year over year, the fastest growth since Q3 2023, spanning more than 9,000 brands and more than 310 Carrot Ads retail-media partners. Every new enterprise client becomes a Carrot Ads partner. The ads business is not a side hustle; it is the growth story.
Now put Clementine next to that number. The launch announcement says Clementine “surfaces relevant deals.” In May, Rogers described the new generative recommendation system that powers this world: it uses real-time cart context to predict what a shopper actually needs, adding flour and eggs triggers vanilla extract and cinnamon instead of cookies, and “early data shows higher engagement and better advertiser results.”
Higher engagement, fine. But “better advertiser results” is the tell. The system that decides what jumps into your cart when you say “order my usuals” is being evaluated, in part, on how well it performs for the brands paying to be recommended. When a human shopped the virtual shelf, sponsored placement had to fight for attention against a scan of the page. When Clementine shops for you, there is no page. There is one cart, built by one agent, whose employer books $286 million a quarter helping 9,000 brands get into it. Substitutions, deal-surfacing, “budget-friendly” rankings: every one of those is now a merchandising decision made invisibly, inside a conversation, by an entity paid by both sides of it.
This is the FTC’s Amazon case transplanted to a conversational interface, except with less visibility. On a marketplace website you can at least see the sponsored badge. Inside an agent-built cart, you see groceries. Meta, for its part, plans to take “a very small cut of whatever the transaction is.” Mastercard hands merchants control over how their products are “represented.” Every actor in the stack has a revenue reason to shape the agent’s judgment, and none of them has a revenue reason to check it.
And it extends past grocery. The same PYMNTS data shows 27 percent of AI-influenced shoppers chose a different brand and 26 percent bought a different product. What changes a model’s brand judgment? The corpus it reads, and the cheapest thing to manufacture in that corpus is still the fake review, in the volumes Amazon already blocks by the hundreds of millions per year. Garden agents inherit the garden’s advertising incentives. Open agents inherit the web’s contaminated reviews. Different architecture, same missing layer.
The Layer Nobody Shipped
A trust stack with payment trust and behavioral trust but no product trust will optimize exactly what it measures: authorized transactions and well-behaved agents, filled with whatever the garden or the corpus put there. The fix is an independent layer whose incentives run the other way:
- Score the product, not the placement. GoBuy’s Smart Score runs 0 to 100 on review quality and authenticity after manipulated and low-information reviews are filtered out, not on review volume or seller-stated averages, the two things ad money buys most easily.
- Require persistence, not a lucky window. GoBuy Verified demands a filtered score of 80 or above held across 90 days, the difference between a stable property of a product and a purchased spike.
- Curate instead of drowning. Top seven verified products per category rather than a sponsored-first firehose, which removes the placement auction from the discovery path entirely.
- Meet agents where they decide. Any agent, garden-variety or open-web, can query filtered trust data over MCP at gobuy.ai/api/mcp before it puts anything in a cart. Docs live at gobuy.ai/agent-docs. Humans get the same signal injected directly on Amazon product pages via the Chrome extension.
Clementine deciding between two brands of vanilla extract, Muse eventually reordering your toilet paper, Mastercard Connect brokering a merchant’s self-description to a hundred agents: all of them are making product-trust decisions with seller-supplied evidence. The agent layer is being built this quarter. The verification layer is still mostly unbuilt.
What to Watch
First, whether Instacart discloses how Clementine weighs sponsored products inside agent-built carts, and whether regulators extend the disclosure logic from the FTC’s existing fake-review and sponsored-placement rules to conversational placement. Second, whether Muse’s completion rate improves enough by holiday season to re-run the errand test, because the open-agent path dying on setup friction would concentrate even more commerce inside walled gardens. Third, whether Agent Connect’s “merchant-defined information” standard gets any independent-verification counterpart, or whether seller self-attestation becomes the de facto product data layer for every connected agent. Fourth, Cart Assistant’s rollout numbers past that 25 percent mark, because white-labeling the garden to Aldi and Stew Leonard’s is how a closed architecture quietly becomes the default.
The experiment ran. The garden won. Now the only question left is whether anyone checks what the gardener is planting.
Before you, or your agent, trust a cart: verify what’s actually in it. Start at gobuy.ai, and wire independent trust scoring into any agent at gobuy.ai/agent-docs.