Two documents, published five months apart, describe the same marketplace. One is an enforcement report. The other is a measurement. Read together, they explain why product trust in 2026 feels the way it does: an enormous machine is winning the blocking war, and the corpus is losing anyway.
The first document is Amazon’s first Trustworthy Shopping Experience Report, released in April, which expands the company’s annual Brand Protection Report into a broader account of how the store polices itself. It is full of numbers that deserve respect. The second is a study from AI detection lab Pangram, which scraped 30,000 front-page reviews across 500 Amazon best-sellers and checked how many were written by machines. It is full of numbers that deserve attention.
Start with the enforcement side, because it is genuinely formidable. Then look at what the measurement found sitting on the front pages anyway.
What Amazon’s Report Actually Says
The Trustworthy Shopping Experience Report covers four pillars: proactive controls, anticipatory tools, accountability, and consumer education. The review integrity numbers sit inside the first pillar.
On reviews specifically, the report says Amazon’s systems “analyze thousands of data points across billions of reviews before a review appears in the store, drawing on review data that dates back to 1995 to help detect fake or abusive content.” The headline result: “In 2025, we proactively blocked hundreds of millions suspected fake reviews from our store.”
That is one year of blocked reviews. Not flagged for human review. Blocked, before publication.
The surrounding enforcement stack is similarly scaled:
- 32,000+ bad actors pursued by the Counterfeit Crimes Unit since 2020, through litigation and criminal referrals across 14 countries
- 15 million+ counterfeit products identified, seized, and disposed of in 2025
- More than 100 websites that brokered fake reviews and scams shut down through legal action in 2025
- 99.9 percent of suspected infringing listings blocked before a brand owner ever had to report them
- Omniscan, a machine learning system that photographed and checked more than 12 million products in fulfillment centers across the US, Canada, UK, Türkiye, Saudi Arabia, and Europe
- An early warning system that in one 2025 case blocked infringing listings of a viral branded product eight days before the brand owner even shared its IP with Amazon
Amazon also frames the problem honestly in one respect: “Creating the most trustworthy shopping experience in the world is a journey,” the report concludes. “This report is our most transparent account yet of the progress we’ve made, the challenges that remain, and the commitment that drives us forward.”
A journey is the right word. Because the second document measures where the journey currently stands, on the exact surface shoppers see first.
The Delta: What Slipped Past the Filter
Pangram’s methodology is simple enough to audit. The lab scraped 30,000 front-page reviews across 500 of Amazon’s best-selling products, spanning ten categories from baby products to laptops to medical devices, recorded star ratings and Verified Purchase status, and ran everything through its AI text detector.
The findings, in the lab’s own words: “We found that 3% of the total reviews studied - 909 total reviews - were AI-generated with high confidence.”
Ninety percent of those reviews are not marginal cases. They are machine-written text that survived Amazon’s full filtering stack, which analyzes thousands of signals per review against 30 years of historical data, and landed on the front page of the best-selling products in the store. These are not the reviews Amazon blocked hundreds of millions of. These are the ones that got through, on the listings with the most traffic, where the marginal purchase decision actually happens.
The star distribution is where it becomes an integrity problem rather than a curiosity. AI-written reviews gave five stars 74 percent of the time. Human reviews on the same products gave five stars 59 percent of the time. Run the reverse direction and the asymmetry doubles: humans left one-star reviews at 22 percent, AI at 10 percent. The pollution is not neutral noise. It is directionally inflationary wherever it lands.
And it does not land uniformly. A seller paying a content farm does not spray 909 reviews across 500 products. They concentrate fire on their own listing, where Pangram’s underlying rate means a targeted injection can shift a front page materially: enough five-star volume, arriving fast enough, to change the displayed average before organic reviews can dilute it. The aggregate 3 percent understates the effect on any single contested listing, which is the only listing a shopper or an agent is looking at.
The Day the Verified Purchase Badge Died
The single most consequential number in the study is not the 3 percent. It is this: “93% of the first-page AI-generated reviews had the ‘Verified Purchase’ badge, showing that this badge alone is no longer a reliable symbol of trust.”
Sit with the mechanics of that. The Verified Purchase badge exists because it was supposed to be the one signal a shopper could not be fooled by: whatever else might be fabricated, at least this person bought the thing. If 93 percent of machine-written front-page reviews carry it, then either the reviews are attached to real transactions the seller controls, or the badge is being applied more loosely than anyone assumed.
Both paths exist and both are documented. The first is brushing, the scheme the FTC describes in its consumer guidance: a seller ships cheap or empty packages to real addresses, creating genuine-looking transactions that generate the purchase credential, then reviews flow from those transactions. The second is the refund-outside-the-platform arrangement, where compensated reviewers actually buy the product, get reimbursed through a side channel, and leave the five stars. The transaction is real. The review is theater. The badge cannot tell the difference, because the badge verifies a payment event, not an opinion’s provenance.
This matters beyond Amazon because the Verified Purchase pattern is the trust architecture the whole industry copied. “Verified buyer” badges, “confirmed purchase” labels, purchase-gated review prompts across every major marketplace all descend from the same assumption: that a transaction record authenticates the sentiment attached to it. When the sentiment is manufactured but the transaction is real, the architecture authenticates nothing. Pangram’s conclusion is blunt: “Unless platforms take action, it will soon be hard to trust any reviews on the internet.”
Humans Are Coin Flips Now
If the badge cannot carry the load and the corpus is polluted, the last line of defense is the reader’s own judgment. Research published this year measured it.
A study from the University of Hawai’i, “Genuine or Fake? Explaining Consumers’ Perception and Detection of AI-Generated Fake Reviews,” ran 151 consumers through 906 review classifications. The result: “humans cannot reliably distinguish between genuine and AI-generated fake reviews (accuracy = 53.2%).” Worse, participants were “especially worse at detecting negative AI-generated fake reviews.” A coin flip is 50 percent. The trained intuition of the review-reading public is worth 3.2 points over randomness, and the gap closes to noise on negative reviews, where a fabricated one-star attack on a competitor is effectively invisible.
The machines are doing better. Separate detection research published in May, covering a multimodal model that reads text, images, and reviewer behavior together, achieved 93 percent accuracy on Amazon review data and 91 percent on Yelp, outperforming traditional text-only methods.
Read those two numbers as a system design constraint. Human screening is now statistically worthless against AI-generated reviews. Machine screening works well but only exists where it is deliberately deployed, which today means inside the platforms policing their own corpora and inside a handful of detection labs. The shopper has no access to the first and does not know the second exists.
Why Blocking at Scale Cannot Fix the Corpus
Amazon would reasonably object that hundreds of millions blocked versus 909 found on front pages is a stellar catch rate, and on its own terms it is. But the frame hides the structural problem, which is that the economics of the arms race run one direction.
The cost of generating a plausible review has collapsed to approximately zero. The cost of catching one is thousands of data points of forensic analysis against three decades of historical data. The attacker iterates in seconds; the defender’s rule set updates on a slower loop. Every blocked batch teaches the broker what tripped it. Meanwhile the FTC’s rule banning fake reviews and testimonials, final in 2024, made buying, selling, and suppressing reviews subject to civil penalties, which raises the stakes but not the detection budget of the platforms being gamed. Amazon itself now participates in a cross-industry coalition with Booking.com, Expedia, Glassdoor, TripAdvisor, and Trustpilot, stating that “sophisticated tools are being used to flag suspicious review histories and remove AI-generated reviews before a customer encounters one.” Pangram’s front-page sample is what that claim looks like when measured from outside.
The conclusion is not that Amazon is failing. It is that blocking is a rate, and trust is a stock. The blocked reviews never touch the corpus, but the ones that slip through become permanent residents of the product page, compounding into the average rating forever. A filter, however good, only controls the inflow. It does nothing to discount the pollution already priced into the 4.7 you are looking at.
The Agent Inherits the Pollution
Now add the layer this column has tracked all month: default-on agent checkout across roughly a million Shopify storefronts, payment rails from Mastercard and Stripe, and roughly 60 percent of AI-assisted purchases still closing on Amazon, per Forkast’s reporting.
An agent building a shortlist does not read reviews the way a wary human limply attempts to. It ingests the star average as structured data: a float, a count, a confidence. It weights the rating, ranks against alternatives, and hands the output to a checkout rail that was shipped this month. The 53.2 percent human defense does not even apply, because there is no human reading at all. The pollution propagates from corpus to recommendation to payment without a single skeptical glance anywhere in the pipeline, at machine speed, with pre-authorized credentials.
The industry built identity, authorization, and recourse rails in September. The input layer still runs on a signal that is 3 percent fabricated on the highest-traffic pages, directionally inflated toward five stars, and wearing a purchase badge 93 times out of 100.
Scoring Must Assume Pollution
The design answer to an adversarial corpus is not better blocking. It is scoring that assumes pollution exists and discounts for it before presenting anything.
That is the specific bet behind GoBuy:
- Filter first, then score. Smart Score 0-100 is computed on the review corpus after authenticity analysis, not before. A rating built on manufactured praise and a rating earned over years are different objects and get different scores, regardless of how similar their averages look.
- Quality over quantity. A 4.7 from 40 verified-durable reviews can outrank a 4.8 from 15,000 that include coordinated patterns. Volume is not evidence. Durability is.
- Curation with no seller base. We show the top 7 products per category, ranked by filtered evidence. No marketplace take rate, no advertising auction, no incentive for the shortlist to contain anything other than what survives scrutiny.
- Time-bound verification. GoBuy Verified requires holding 80 or above across 90 days, which filters launch-week rating theater from sustained quality, exactly the distinction a fast-injection attack is designed to blur.
- Agent-native delivery. The same filtered evidence is exposed via MCP at gobuy.ai/api/mcp, so an agent ranking products consults a signal that already assumed the reviews were polluted. The Chrome extension puts the trust panel directly on the Amazon page for humans still shopping by hand.
Amazon’s report closes by calling trust a journey. Fair. But shoppers and agents do not purchase the journey, they purchase against today’s corpus, and today’s corpus contains 909 machine-written reviews on the best-sellers’ front pages, most of them badge-carrying, most of them five stars. Verify what the reviews are hiding before you, or your agent, believe the average: gobuy.ai. Building a shopping agent? Wire in pollution-proof product evidence in minutes: gobuy.ai/agent-docs.