Amazon’s first Trustworthy Shopping Experience Report, released in April 2026, contains a sentence that should stop anyone who builds product discovery systems. Describing its review defenses, the company wrote: “In 2025, we proactively blocked hundreds of millions suspected fake reviews from our store.” The systems behind that number “analyze thousands of data points across billions of reviews before a review appears in the store, drawing on review data that dates back to 1995,” and Amazon’s legal actions shut down more than 100 websites that year that were brokering fake reviews and scams aimed at the store.

Hundreds of millions blocked. Billions analyzed. Three decades of data. And yet every experienced shopper still squints at a 4.9 with 40 reviews and wonders. A paper published on August 25, 2026 in the arXiv economics category explains, with formal proofs, why that squint is rational, why the suspicion should be strongest at the top of the scale, and why even the impressive enforcement numbers above cannot tell us whether manipulation is shrinking. The paper is “Rating Manipulation: Credibility Inversion and Audit Leakage” by Van-Quy Nguyen of the Faculty of Fundamental Sciences, College of Technology, National Economics University, Hanoi. It is theory, not empirics, which is precisely its value: it isolates the logic that field data can only show us through a glass darkly.

One Number, Two Layers

The paper’s organizing move is to split what buyers see from what sellers do. A marketplace displays a single average rating, call it z. Behind it sits a vector: the authentic reviews the seller earned, plus whatever fake or induced reviews the seller chose to add, each carrying a chosen score. Buyers observe only z. Enforcement acts on the hidden vector. Nearly everything interesting in the paper falls out of that asymmetry.

The first result is about composition. Suppose a low-quality seller with an authentic 4.1 average wants to display a 4.8. What is the cheapest way to get there? The theorem says: use only fake scores above the target. A fake 4-star review pulls the average in the wrong direction and would require still more fakes to undo the damage. So a seller engineering a 4.8 buys fives, possibly some fours if fives are expensive or heavily policed, and never anything below the target. Among the scores it does use, it equalizes marginal cost per unit of what the author calls score leverage: how strongly one additional fake at that score moves the average. The paper describes this as an Euler-equation analog, the same kind of marginal-balance condition that governs how firms split spending across inputs.

Read that again as a product-trust claim. Manipulation is not noise smeared randomly across a rating distribution. It is construction, concentrated above the target, allocated by leverage. The upper tail of a rating histogram is where purchased reputation lives. The bottom of the distribution, the genuine complaints, is the part a manipulator cannot cheaply touch.

Credibility Inversion: Why the Top of the Scale Is the Least Informative Place

Buyers in the model are Bayesian. Seeing a displayed rating, they ask: how characteristic is this number of a genuinely high-quality seller, relative to a low-quality seller who might have manufactured it? Credibility is a likelihood ratio, not a level.

Here is the inversion. If low-quality sellers are especially likely to occupy the very top of the scale, then a near-perfect rating is more characteristic of fraud than of quality, and Bayesian buyers trust it less than a slightly lower rating. The paper calls this “credible imperfection”: “It does not mean that buyers prefer lower numbers. It means that an almost-perfect rating can become suspicious when low-quality sellers are especially likely to manufacture the top of the scale.”

The paper’s worked example is worth memorizing. Suppose genuinely high-quality sellers cluster tightly around 4.60 with little dispersion, while low-quality sellers are scattered, averaging 4.10 with much wider spread, because manipulation intensity varies from seller to seller. Under those distributions, credibility peaks at approximately 4.66. Below that, higher is better news. Above it, each step toward 5.0 makes a rating more likely to be manufactured, so each step is worse news. A 4.7 outsells a 4.9 not because shoppers enjoy imperfection but because the 4.9 is statistically more likely to be a costume.

The author is careful to kill the pop version of the claim: “The familiar statement that ‘four beats five’ is therefore only an illustration; the general result compares any two exact ratings through the information they carry.” The theorem is not “four good, five bad.” It is that the credibility ordering of ratings need not match their numerical ordering, and will not match it wherever manipulation concentrates above the credibility frontier.

This is not exotic. The empirical literature the paper builds on has documented versions of it for a decade: Mayzlin, Dover, and Chevalier (2014) found promotional review patterns on Amazon tied to competition; Luca and Zervas (2016) estimated that roughly one in five Yelp reviews in their sample was suspicious, concentrated among struggling restaurants; Filippas, Horton, and Golden (2022) showed reputation inflation eroding the information content of ratings as everything drifts toward five stars. Nguyen’s contribution is to derive the mechanism cleanly: inflation is not just a nuisance, it flips the sign of what the top of the scale means.

The Retreat from Perfection: Smart Frauds Display Less and Sell More

The second result is the one sellers will find darkest. What happens when a manipulative seller starts caring more about actual sales, say because margins rise or lifetime customer value grows?

Intuition says: buy more fakes, push the rating higher. The model says the opposite can hold. Once buyers are suspicious of the top of the scale, pushing toward 5.0 costs money and destroys demand. A seller whose sales matter most rationally retreats to a lower, more credible displayed rating, one chosen near the peak of buyer belief. The paper’s phrasing is exact: along this branch “the displayed rating declines while credibility and sales rise,” and the seller “may display less and sell more.”

Pause on the implication for detection. The sophisticated manipulators, the ones optimizing for revenue rather than for a vanity badge, are not the ones with 4.9s and suspiciously uniform five-star bursts. They are sitting just below the credibility frontier, indistinguishable in the histogram from honest sellers whose genuine record is merely very good. The crude frauds get caught by pattern filters. The optimized frauds look like the honest upper middle of the market. Any trust system that scores the displayed average, or even the shape of the visible distribution, is scoring a signal the adversary has already learned to shape.

Audit Leakage: Why “Reviews Blocked” Cannot Measure Success

Now the enforcement half, which lands directly on Amazon’s report and on every platform transparency disclosure in the industry.

Suppose the platform cracks down specifically on five-star fakes, making that score more costly to use. Globally, the theorem shows, use of the audited score falls. That sounds like victory. But hold the displayed rating fixed and look at the whole vector: manipulation shifts toward the other scores above the target. More fake fours, or more fake 4.5s where half-star scales exist, cheaper levers now relatively attractive. The displayed average does not move, so buyers see exactly what they saw before. The paper names this “audit leakage”: “the review mix changes even though buyers continue to see the same displayed rating.”

When the seller is also free to re-target, the audited score falls and the displayed average drops slightly, but the response of every other score is ambiguous: substitution pushes them up, the lower target pushes total manipulation down. The author’s conclusion is the operational one: “A fall in the audited category is therefore not enough to measure the overall success of enforcement.”

This is why the transparency numbers platforms publish, including Amazon’s genuinely large ones, underdetermine the question everyone actually cares about. “Hundreds of millions blocked” tells you enforcement is active. It cannot tell you the surviving mix is honest, because the model’s whole point is that the mix can rot while the number you display and the number you announce both look stable. What would measure success is disclosure of composition shifts: the surviving distribution of scores per product over time, cross-checked against post-purchase signals like returns, refunds, and review editing. No major marketplace publishes that. The same reporting asymmetry appears at the FTC level: the agency’s fake-review rule, which we covered alongside the Yelp enforcement data earlier this summer, penalizes the practice, but public enforcement statistics still count actions taken, not substitution prevented.

The Ranking Theorem No Marketplace Will Implement

The paper’s final section turns to search ordering, and this is where theory stops being academic. Platforms allocate attention: position one gets more clicks, more consideration, more purchases than position five, a fact established empirically by Ursu (2018). Nguyen proves the allocation rule a buyer-oriented platform should use: rank sellers by posterior-adjusted expected buyer value, where the posterior is exactly the Bayesian belief from the credibility analysis. At equal prices, “a seller with a lower displayed rating should receive more attention precisely because that rating is more credible,” and the punchline follows immediately: “A ranking based only on raw scores can, therefore, direct the most attention toward the less credible seller.”

Sit that theorem next to the FTC’s August 31 lawsuit alleging Amazon converted its ad auctions into a mechanism that extracted over $20 billion while concealing the change from advertisers, which we analyzed this week, and the picture completes itself. The ranking function is not merely un-posterior-adjusted. It is monetized. It is the platform’s most valuable inventory. Expecting the operator of the shelf to also correct the credibility of the labels on the shelf asks it to tax its own sponsored results, since the raw-score signals that sponsored placements amplify are exactly the signals credibility inversion discounts. The paper frames the fix as a design recommendation for “a buyer-oriented platform.” The structural news of 2026 is that the buyer-oriented platform does not exist, and cannot be expected to.

The Agentic Amplifier

If human shoppers are bad at internalizing likelihood ratios, AI shopping agents are worse, because they are maximally literal. We covered the Columbia-Yale ACES audit earlier this August: frontier models acting as shopping agents obey star ratings, review counts, and platform badges with mechanical directness, delivering selection lifts of up to 4x for badged products. An agent does not squint at a 4.9 with 40 reviews. It reads the field, sorts by the number, and cites “high rating” as its reason.

Credibility inversion turns that behavior from a heuristic into a vulnerability. A population of agents consuming displayed averages programmatically is a population of ideal marks for upper-tail manipulation: they respond to the number instantly, at scale, with perfect consistency, and they never see the composition the number was built from. Worse, the manipulator does not need to fool a human’s judgment, only the sort key. The Nguyen model assumes buyers eventually learn the market-wide relationship between ratings and quality, which disciplines manipulation over time. Agents that hard-code “higher is better” into their pipelines short-circuit that discipline on behalf of whoever can cheapestly move the number. The economics of fake reviews have always been good. The economics of fake reviews for machines are better.

What a Posterior-Aware Trust Layer Looks Like

The paper’s own limitations section concedes it does not solve the audit budget or the optimal ranking rule. But its logic points at the architecture, and the architecture is not a better average. It is scoring that consumes the composition, not the compression:

  • Filter before scoring. A trust score should be computed from the review record after manipulated and low-information reviews are removed, not from the displayed average. This is the core of GoBuy’s Smart Score: quality of the surviving evidence, not quantity of the submitted evidence, precisely because quantity is the cheapest thing for a manipulator to buy.
  • Score against the frontier, not the ceiling. Credibility inversion says the informative signal is where an honest record plausibly sits, which is why a 4.6 built from filtered, heterogeneous, occasionally critical reviews can outrank a 4.9 built from uniform praise. Trust systems should be indifferent, even mildly skeptical, toward the top of the raw scale.
  • Require sustained composition, not burst achievement. A seller optimizing Nguyen-style buys the minimum manipulation that clears a threshold. A badge earned by holding a filtered score of 80-plus across 90 days, as GoBuy Verified requires, prices in persistence that burst campaigns must repeatedly re-fund and re-hide, raising the expected enforcement cost the model centers on.
  • Rank by adjusted value, then stop expanding the list. The ranking theorem says attention is the scarce resource. Returning only the top seven verified products per category is attention allocation under a posterior-adjusted cutoff, not a search box with a filter applied.
  • Deliver it where the decision happens. Agents cannot squint, but they can query. GoBuy’s trust layer is exposed over MCP at gobuy.ai/api/mcp, so an agent comparing products can pull the filtered score and verification status instead of the marketplace’s displayed number, and developers can wire it in at gobuy.ai/agent-docs.

None of this abolishes manipulation. The theory guarantees the adversary adapts, substituting across scores and tactics. What independent verification does is move the target: from a public number any seller can purchase, to a computation over evidence the seller does not control, run by a party whose ranking is not for sale.

What to Watch

Three disclosures would signal the industry taking composition seriously. First, whether any marketplace begins publishing per-product score distributions over time rather than averages and aggregate enforcement counts, the only data format that makes audit leakage visible. Second, whether the AI agent ecosystem standardizes on any posterior-aware trust signal before the first documented case of an agent fleet farmed by a rating campaign, a story that now has a working business model and a proof of concept. Third, whether regulators, having banned fake reviews, start asking platforms about substitution: not how many reviews were blocked, but what the surviving mix looks like.

The star rating survived twenty years of the internet because it compresses beautifully. The economics of 2026 keep arriving at the same conclusion from different directions: the compression is the vulnerability. The number is not the knowledge.

Stop trusting the average your marketplace chose to show you. Check the filtered score at gobuy.ai, or give your agents the posterior-adjusted signal at gobuy.ai/agent-docs.