Ask a room of commerce executives why AI shopping agents matter and you will hear about convenience, conversion, and checkout automation. Ask an economist the same question and you get a stranger answer: delegation moves the bottleneck. The shopper’s problem was never really finding products. It was articulating what she wants in a space of many hard-to-specify dimensions, and then trusting that what the intermediary shows her resembles the truth.

A working paper published August 9 on arXiv, “From Product Search to Preference Articulation: The Economics of Agentic Commerce”, by Lingxiu Dong, Kaiwen Luo, and Fasheng Xu of Washington University in St. Louis’s Olin Business School and the University of Connecticut’s School of Business, is the cleanest formal treatment of that shift we have seen. It is pure theory, current version dated August 2026, which is exactly why it deserves attention now: it isolates incentives that the industry’s product announcements structurally cannot admit.

The paper’s three results map precisely onto the three questions every operator in agentic commerce is currently fumbling through with A/B tests: when do agents beat browsing, who wants that transition first, and what happens to recommendation quality once the platform controls it. The third answer is the one that should be screenshotted and pinned above every trust and safety team’s desk.

The Setup: Two Ways to Shop, One Attention Budget

The model compares two canonical regimes. Manual search: the consumer inspects a sample of products herself, evaluating each with perfect fidelity but covering only a handful of listings. Agentic search: the agent screens the entire catalog at negligible marginal cost, but through noisy representations of both sides of the match, the consumer’s latent preferences and the products’ true attributes.

The paper’s organizing concept is preference complexity: the number of satisfaction-relevant dimensions that are hard to articulate before search but easy to judge on inspection. Its running example is home-office furniture. Price, dimensions, and material you can specify in advance. Whether the chair stays comfortable through a nine-hour workday, whether the desk looks professional on camera, whether the finish actually fits the room: these are known to the shopper only in the evaluating, not the asking. The consumer holds a finite attention budget and chooses not just the regime but the intensity: how many products to inspect manually, or how many rounds of dialogue to spend refining the agent’s model of her preferences. Refinement narrows the search but costs friction.

That last detail is the paper’s core move, and it is what makes the results feel true to life. Using an agent does not abolish effort. It converts effort from inspecting products into articulating yourself. The authors state the headline plainly: agentic commerce “shifts scarcity from product inspection to preference articulation.”

Result One: Manual Search Has a Finite Death Threshold

The first theorem pair contrasts how the two regimes age as complexity grows.

Manual search dies at a cutoff. The mechanism is the curse of dimensionality, stated cleanly: under uniform sampling, the probability that a random product lies within distance r of the consumer’s ideal point is r to the power k, which decays exponentially in the number of dimensions. Each additional inspection costs effort and buys a shrinking expected improvement. Beyond a finite complexity level, which the paper calls the cutoff in its Theorem 4.4, the consumer stops inspecting entirely: mismatch returns to the no-search benchmark and, starkly, “platform revenue falls to zero.” Not declines. Falls to zero. In sufficiently complex categories, browsing is not merely inefficient. It is economically nonexistent.

Agentic search never collapses. Complexity raises the marginal value of refinement, so once talking to the agent becomes worthwhile, it stays worthwhile at every higher complexity level. Mismatch may creep up and articulation effort with it, but “mismatch remains below the no-search benchmark and revenue remains positive” at every finite complexity. The paper calls this complexity attenuation without collapse.

The commercial reading: agents do not win by being nicer to use. They win category by category, in exactly the places where preference complexity is high and product-by-product evaluation is costly, furniture, electronics, health products, anything with feel, fit, and reliability dimensions that spec sheets do not carry. In simple categories, the model implies, humans and agents are competitive. In complex ones, it is agents or nothing, because nothing is what manual search rationally delivers.

Result Two: The Adoption Lag

The second theorem explains the adoption paradox the industry keeps bumping into. The platform ranks regimes by conversion revenue. The consumer ranks them by total loss: mismatch plus her own search expenditure, including the dialogue effort of refinement. Two different scorecards, and the paper proves they cross at different complexity levels.

When manual inspection is sufficiently cheap, the platform’s revenue threshold sits below the consumer’s adoption threshold. Between them lies what the authors name the adoption lag: a range of product categories where agentic search would generate more platform revenue if used, yet consumers rationally keep browsing. The paper’s own summary is the sentence every agentic commerce roadmap should be audited against: “higher platform revenue from agentic search is not sufficient to induce consumers to adopt it voluntarily.”

This is not a hypothetical. The paper cites the survey record: Worldpay finds 44 percent of US consumers would let an AI assistant browse for them but only 6 percent would relinquish complete control over the purchase; Contentsquare finds 30 percent would let an agent complete a purchase; Adobe had 39 percent using generative AI for online shopping by early 2025, mostly for research rather than execution. The funnel we have covered before, heavy AI use at the comparison stage, thin delegation at checkout, is the adoption lag measured in the wild. During the lag, expect platforms to discount, bundle, and default-switch users into agent modes, because the model says that is where the revenue is. Expect consumers to keep one hand on the wheel, because the model says that is rational too.

The paper even documents the mirror case. When inspection costs are high, the thresholds can reverse: consumers would prefer the agent over a range where the platform still earns more from manual search, and “consumer demand for the agent may remain latent.” Anyone wondering why a marketplace with world-class AI labs on retainer ships conservative shopping features should sit with that sentence. Withholding the agent can be the revenue-maximizing move.

Result Three: Inverted Fidelity, the Uncomfortable Centerpiece

Then the paper does something rare in this literature: it hands the platform a dial and watches what it does with it. Conditional on the consumer using agentic search, the platform chooses sigma, the representation noise level, trading accuracy cost against conversion.

The finding, Proposition 7.2, is the one to remember. Attention-rich consumers, those with larger budgets for refinement dialogue, can absorb more representation noise by simply talking to the agent longer. So the revenue cost of degrading fidelity is smaller for them, and the profit-maximizing platform “assigns weakly lower fidelity to the attention-rich segment, while that segment remains weakly more profitable to serve.”

Read that again slowly. The platform’s most engaged, most patient, most valuable customers receive the noisiest recommendations, because their patience functions as a free substitute for platform-provided accuracy. The paper is explicit that this inverts half a century of quality-versioning results from Mussa and Rosen onward: fidelity is ordered by attention budget, not willingness to pay. Your stoicism becomes the product feature that funds the platform’s cost savings. In the exponential-refinement benchmark, the inversion becomes strict discrimination when three conditions hold jointly: interaction friction is low, the feasible noise range is wide enough to exhaust the constrained segment’s budget, and accuracy costs are intermediate. The authors’ summary of platform logic is chilling in its cleanliness: “accuracy is most valuable where consumer effort cannot substitute for it.”

Now overlay the deployed world. The same model that treats sigma as a Gaussian engineering parameter is, in production, controllable through ranking, sponsorship, metadata, and whatever the agent reads about the product. A platform that can profitably serve noisier representations to its best customers can also, in less careful hands, sell access to the noise. This is where theory meets the enforcement record we have been tracking all summer: the FTC and 22 states are currently litigating allegations that the first screen of Amazon results is the output of a disguised first-price auction, and the formal economics of credibility inversion show the top of the review scale is now the least informative place in the store. The paper models product-side representation error as mean-zero. The single most important fact about real product-side error in 2026 is that it has a mean, the mean points toward whoever paid, and nothing in the platform’s objective function, as modeled, penalizes that direction.

The Layer the Model Points At

The authors flag their own boundary. Their extensions section concedes the baseline leaves open “third-party agents with advertising or steering incentives,” which “could make agent ownership consequential.” That is an economist’s understated way of saying the paper’s fidelity dial is currently being sold to the highest bidder in every deployed system, a risk the SkillShift research on covert skill steering demonstrated empirically two weeks ago. And the evaluation literature keeps confirming that final outcomes alone cannot tell you whether an agent chose well or was aimed: the ACWorld benchmark paper from August showed that “final state alone can miss evaluated errors,” requiring process-level evidence to audit agent behavior across even its 785,022-listing catalog.

Put the three results together and the architecture requirement falls out. If scarcity has moved to preference articulation, and if platform-side fidelity is a choice variable with an inverted incentive, then the consumer side of the match needs effort and candor, but the product side needs verification that is not a dial on the platform’s console. Independent product trust is best understood as exactly that: exogenous fidelity. Sigma that cannot be raised for the profitable segment because it is not owned by the platform at all.

That is the design logic behind GoBuy, and the paper sharpens each element of it:

  • Score the evidence, not the representation. GoBuy’s Smart Score is computed from review quality after fake and low-information reviews are filtered out, not from the displayed average the platform or seller presents. In the model’s terms, it reduces product-side noise at the source rather than trusting the intermediary’s reading of the product.
  • Persistence as fidelity proof. GoBuy Verified requires a filtered score of 80-plus held across 90 days. Degraded or purchased representations are cheap to flash and expensive to sustain, so sustained quality is the observable that survives.
  • Decision-shaped output. The model’s consumers are attention-constrained; drowning them in an infinite grid re-imposes the inspection burden agents were supposed to absorb. Returning only the top seven verified products per category is an attention-budget format, not a marketing flourish.
  • Queryable, not readable. Verification delivered over MCP at gobuy.ai/api/mcp means an agent pulls structured trust data in one lookup instead of browsing adversarial pages, which is also the cheapest way to keep product-side fidelity high without spending the consumer’s dialogue budget compensating for it. Integration docs for any shopping agent stack are at gobuy.ai/agent-docs.

What to Watch

Three empirical markers would tell us the paper’s predictions are live. First, fidelity disclosures: whether any platform publishes representation-quality metrics by segment, or whether the first hard evidence arrives instead through litigation, the pattern that gave us the ad auction complaint. Second, whether agent-mode defaults and switch-friction tactics intensify in exactly the high-complexity categories the model predicts, over the objections of users who keep opting back out. Third, whether third-party trust signals get wired into agent stacks as defaults before the first documented case of a fidelity gradient exploited commercially, the same race condition between audit and incident we have tracked across the ACES and SkillShift results.

The paper closes by reframing the consumer as the load-bearing element of the new system: her willingness and ability to articulate preferences is what voluntary adoption, refinement, and even platform profitability now hinge on. Fair. But a consumer can only refine what she can articulate, and she can only articulate preferences over products whose representations are honest. Preference articulation is the new scarcity. Product verification is the old one, and the economics of agentic commerce, modeled honestly, say it is not going away. Test the representations your agents consume at gobuy.ai, and wire the trust layer into your stack at gobuy.ai/agent-docs.