Agentic shopping has quietly crossed the threshold from demo to infrastructure. Bain & Company’s retail practice, in its May 2026 analysis of autonomous shopping, reports that around 30% of US consumers now use generative AI for product comparison and recommendations, that shopping referrals from ChatGPT have more than doubled across the US, France, the UK, and Germany over the past year according to Similarweb estimates, and that Amazon’s on-site agent Rufus generated nearly $12 billion in incremental annualized sales, with monthly active users up 115% year over year. For some retailers, AI now accounts for up to a quarter of referral traffic.
The machine is now a customer. Which raises the question a team of researchers from Columbia Business School and Yale asked directly in a paper titled “What Is Your AI Agent Buying?”: when the machine chooses, what actually drives the choice?
Their answer should unsettle anyone building commerce on top of these systems. Using ACES, a provider-agnostic simulator that runs randomized controlled trials on shopping agents, the researchers audited six frontier models: Claude Sonnet 4, GPT-4.1, and Gemini 2.5 Flash as of August 2025, and their successors Claude Opus 4.5, GPT-5.1, and Gemini 3.0 Pro Preview as of December 2025. The agents browsed mock storefronts where the researchers randomized everything: product position, prices, ratings, review counts, and the platform badges attached to each listing.
Because everything was randomized, every effect they measured is causal. And the causal picture is this: AI agents obey badges, ratings, and position far more than they evaluate products. Those are precisely the signals most vulnerable to purchase and manipulation in modern e-commerce.
Finding 1: Platform Endorsements Are Rocket Fuel for Agents
The single largest lever in the entire study was not price, not review quality, not brand. It was the platform endorsement badge.
When ACES attached an “Overall Pick” badge to a product with a 10% baseline selection probability, that probability jumped to 24.3% under Claude Sonnet 4, 19.9% under GPT-4.1, and 42.6% under Gemini 2.5 Flash. That is a 2x to 4x lift from a badge the marketplace assigns itself. The researchers’ conclusion: “the strong positive response to overall pick implies that platform endorsements are treated as credible signals,” echoing how human buyers behave.
Sponsored tags, by contrast, got penalized. The same 10% baseline product fell to 8.9% under Claude Sonnet 4, 8.0% under GPT-4.1, and 7.9% under Gemini 2.5 Flash when labeled “Sponsored.” Agents have apparently absorbed enough training data about advertising to discount it.
Here is the uncomfortable synthesis. Agents discount advertising but reward endorsement, and endorsement is a lever the platform controls unilaterally. Why would a platform ever buy an ad slot inside an agent-mediated world when it can assign a badge that agents reward with 2x to 4x more force, and which cannot be discounted because it does not look like an ad?
Bain’s data shows exactly where this goes. Amazon is already serving sponsored ads inside Rufus chats, Google is surfacing sponsored products in its AI Overviews, and OpenAI is testing ads in ChatGPT in partnership with retailers such as Target. The monetization of the recommendation channel is not a future risk. It is a shipped feature. The ACES results suggest the durable form of that monetization will not look like an ad. It will look like a “credible signal.”
Finding 2: Agents Obey Star Ratings, and Star Ratings Are Compromised
Across all six models, the researchers found agents “react positively to higher average ratings and to a larger number of ratings.” Bain flags the same result from the practitioner side, warning that agents “heavily weigh review counts and average ratings, creating a new optimization challenge for retailers and brands.”
Read those two sentences together with what we know about the review economy. The FTC’s Rule on the Use of Consumer Reviews and Testimonials, in force since October 2024, exists because fake and manipulated reviews had become an industrial supply chain: purchased five-star campaigns, review suppression, incentivized testimonials, insider reviews that never disclose themselves. The agency has spent the past year bringing enforcement actions against the brokers who sell this machinery.
Now install that compromised signal as a direct input into automated purchasing. A rating delta the researchers describe as tiny, +0.1 stars, tripped GPT-4o on a pure dominance test 71.7% of the time in the August 2025 snapshot, meaning the older model failed to consistently identify the objectively higher-rated product. The December 2025 models nearly stopped failing these one-dimensional tests, which is genuine progress. But sensitivity is not the same as judgment. The new models reliably detect which number is bigger. They still cannot detect whether the number is real. No model can, from inside the listing.
That is the core defect: an agent that obeys ratings with high precision converts every fake review in the corpus into a direct instruction. Manipulated input plus obedient processor equals manipulated output, executed at machine speed and presented with machine confidence. Fake reviews used to have to fool a skimming human. Now they only have to fool a parser.
Finding 3: There Is No Stable “Top” Product, Only Model-Dependent Winners
The most commercially explosive finding concerns market shares. When the researchers let each agent select from an eight-product assortment under identical prompts, different models produced wildly different winners.
In the fitness watch category, Claude Sonnet 4 chose the Fitbit Inspire 45% of the time, while GPT-4.1 and Gemini 2.5 Flash chose it only about 25% of the time. Then the model upgrades landed, and the market flipped. With Claude Opus 4.5, the Fitbit Inspire’s share jumped from 45% to 77%. With GPT-5.1, it collapsed from 25% to 6%. The same product, the same shelf, the same prompt: two concurrent frontier agents diverged by 71 percentage points on one SKU.
Position effects were equally unstable. Listing position caused large changes in selection rates, and “not just the magnitude but even the direction of the bias differs widely across models.” GPT-4.1 and GPT-5.1 exhibited nearly opposite position preferences: the preferred position of one was the least preferred position of the other. These biases persist even in headless, text-only API interfaces, which means structured product feeds do not fix them.
There is also a concentration problem. In the stapler category, Amazon Basics dominated while the Arrow brand was never selected by any model. The researchers describe agents “concentrating demand on a few modal products while ignoring others entirely,” raising what they call market dominance questions: agentic demand may suppress the natural dispersion of consumer preference and harden around whichever SKUs the current model weights favor.
For anyone selling online, the implication is brutal. Your agentic market share is not a function of your product. It is a function of which model version your customers happen to run, where you sit in the feed, and which badges the platform assigned this quarter. When OpenAI or Anthropic ships an upgrade, demand shocks arrive with the changelog.
Finding 4: Sellers Will Optimize for the Machine, and the Machine Is Volatile
The researchers then ran the obvious next experiment: they gave the seller side an AI agent too, and let it rewrite product descriptions using competitor sales data.
In 67% of category-model pairs, a one-shot AI-optimized description produced no statistically significant change. In the remaining 33%, the gains were large: +25.8 percentage points of market share for an office lamp under Claude Sonnet 4, +7.1 points for GPT-4.1 in another category. And in several cases, the AI-optimized description actively backfired and destroyed share, because a tweak that helps with one model hurts with another.
This is AI-SEO in its native form: sellers tuning listing text, attributes, and review signals not for human eyes but for agent parsers. The study’s own framing is that “sellers can respond,” and Bain’s retail clients are already being briefed to “structure their product catalogs and data so that agents surface their content accurately.” The description-optimization arms race has a name now, and the ACES results say its returns are real but wildly unstable across models. Combine unstable returns with model-update demand shocks and you get the paper’s summary verdict: “agentic markets are volatile and fundamentally different from human-centric commerce.”
The Trust Math Nobody Is Doing
Step back and assemble the full picture from these two datasets.
First, adoption is real: 30% of US consumers already use generative AI for product comparison, Rufus alone is credited with $12 billion in incremental sales, and Bain says roughly half of consumers are not yet comfortable with fully autonomous checkout, which means the volume still has room to multiply as comfort catches up.
Second, the decision layer inside that volume is driven by badges, ratings, review counts, and position: signals that are either assigned by the monetizing platform (badges, position) or notoriously manipulable at scale (ratings, review counts).
Third, the platforms are already monetizing the agent channel through sponsored placements, and the ACES data implies endorsement badges are an even more effective monetization lever than ads, because agents reward them instead of discounting them.
Fourth, consumers tell Bain they trust retailer-owned agents three times more than third-party agents. The agents humans trust most are the ones operated by the marketplaces with the strongest incentives and the widest levers.
Every layer of this stack is being standardized and capitalized: payments through the Agentic Commerce Protocol and Google’s Universal Commerce Protocol, financing inside the conversation, discovery through Rufus and ChatGPT integrations. The one input that remains unverified is the one everything else compounds on: whether the product the agent selects is actually good, as measured by something other than the marketplace’s own badges and the marketplace’s own review corpus.
What a Real Trust Layer Needs to Look Like
The ACES authors call for “continuous auditing” of agent decision-making. That is necessary but insufficient. Auditing tells you the agent obeyed the badge. It does not tell you the badge was earned. What agentic commerce needs is an independent, machine-readable trust signal that satisfies four conditions the current stack does not:
- Independent of the marketplace’s levers. A score that no platform assigns, no badge system controls, and no sponsored budget can move. ACES proves agents reward endorsement; the answer is not better badges but signals the platform cannot mint.
- Built on review quality, not review quantity. Agents weigh review counts heavily, and review counts are the cheapest signal to buy in bulk. Filtering manipulated reviews before scoring, and weighting authentic, verified-purchase feedback upward, removes the poisoned input rather than re-weighting it.
- Scored on sustained performance. A 0-100 score computed over a long evaluation window, like GoBuy’s Smart Score, which requires a product to hold 80 or above across 90 days to earn the GoBuy Verified badge, is structurally harder to game than a snapshot star rating that moved 0.1 stars last Tuesday.
- Curated, not exhaustive. The ACES concentration findings show agents already collapse choice onto a few SKUs. The useful trust layer does not hand the agent ten thousand sorted listings. It answers the actual question: what are the top 7 products in this category that survive filtering? Fewer, better inputs reduce the surface area for position bias and badge capture.
This is precisely the gap GoBuy was built to fill. GoBuy’s Smart Score is computed from review authenticity and quality rather than raw volume, manipulated reviews are filtered out before they can inflate anything, and the output is a short verified list rather than an infinite shelf. And because the consumer of a trust signal is increasingly not a human but an agent, the whole system is exposed via MCP at gobuy.ai/api/mcp, so any shopping agent can consult independent trust data before it commits a purchase, instead of inheriting the marketplace’s badge-and-rating complex blind. For humans still browsing the old way, the Chrome extension injects the same trust panel directly onto Amazon product pages.
The Bottom Line
The Columbia and Yale researchers set out to ask what our AI agent is buying. The answer turned out to be: whatever the current model version weighs most, positioned where that version prefers it, wearing whatever badge the platform sold or assigned this cycle. Until agents consult trust signals that none of those parties control, agentic commerce is a machine that executes the marketplace’s incentives at machine speed, with a human’s money.
The infrastructure layer is solved. The financing layer is solved. The verification layer is the entire ballgame now.
See what independent trust looks like before your next purchase, or before your agent’s next one: gobuy.ai. Developers building shopping agents can connect the trust layer directly at gobuy.ai/agent-docs.