Monday was a celebration. At the first MCP Dev Summit in Toronto, the Linux Foundation-hosted protocol community marked 500 million monthly SDK downloads, up from 97 million in March, across nearly 16,000 public servers. TikTok used the same news cycle to announce its full agentic commerce stack, with MCP connectors letting agents from Claude, Perplexity, and six other platforms manage ad campaigns and a Shopping Assistant mediating purchases inside the For You feed. The message of the week was that the plumbing is finished and the shopping can begin.
Tuesday, two independent investigations examined the plumbing. What they found should reset every assumption about what happens when an agent consults a server before it spends your money.
The first investigation, published by OX Security as “15,465 MCP Servers, 0 Governance,” looked at what people actually install. The second, documented by Ars Technica, followed one independent researcher’s proof that the most basic trust mistake in web security appears independently in MCP servers written by a hyperscaler, a global bank, a database company, and two governments. Neither finding is about a single bad actor. Both are about a substrate that was adopted before it was inspected.
A Marketplace With No Bouncer
The OX Security team analyzed 15,465 publicly indexed MCP servers across five registries, deduplicated to 5,095 unique hostnames. Their headline finding is not any single statistic. It is the absence of a mechanism. “We found no guardrails and no review,” wrote Moshe Siman Tov Bustan, Security Research Team Lead at OX. “Security is a recommendation, not a policy.”
The frame the report uses is instructive: in 2012, Google ran Bouncer, an automated scanner that checked Android apps for malware before they reached users. It was imperfect; researchers slipped malware past it within months. But it existed. MCP marketplaces in 2026 have no equivalent. Anyone can write a server, push it, and publish it. And even if a review process existed, OX notes a deeper problem: remote MCP servers can run backend code that differs entirely from what their public repository shows. “Code review tells you what the developer published, not what the server runs.”
Then the geography. Of 5,095 unique hostnames:
- 15.6 percent resolve to infrastructure outside the United States, including 19 in China and 18 in Russia. An agent connected to these servers may send data to jurisdictions the security team never approved.
- 0.45 percent route traffic through consumer tunneling services, mainly ngrok-free. These publicly listed servers run from personal machines and, likely, home networks.
- 2.3 percent no longer resolve at all. Six sit on expired domains registrable for $4 to $12 a year.
That last bullet deserves a slow read, because it is the cheapest attack on agent infrastructure anyone has priced this year. In OX’s words, “a new owner would inherit an established server identity, along with requests from any agent still configured to call it.” Not a zero-day. Not a supply chain compromise in the traditional sense. A domain renewal somebody forgot, and a $4 registration that inherits the trust of every agent configuration that still points at it.
Location, the report adds, can also change after the fact: an operator could launch a server on a clean US IP address and later route traffic somewhere else. This is not hypothetical hygiene. OX’s earlier work this year traced critical vulnerabilities in Anthropic’s own MCP reference code, downloaded more than 150 million times. The official foundation is audited by volunteers; the ecosystem around it is audited by no one.
Protocol Pivoting: The Hallway Nobody Watches
The second investigation is smaller in scale and larger in implication. Independent researcher Syed Anas Mohiuddin, working with no institutional affiliation, has spent the last five months testing a hypothesis: that the security gap in MCP is structural, not incidental. In May he had one example and an argument. In an October update, he has confirmation from five security teams that share nothing except the protocol.
The pattern, at each site, is server-side request forgery. An MCP server exposes tools, an agent calls them with arguments, and some of those arguments are URLs, paths, or endpoints. If the server builds an outbound request from that value without checking where it resolves, the agent effectively decides what the server’s network identity talks to. Agents read untrusted content and act on it, so the agent becomes a channel for anyone who can get text in front of it.
The confirmed roster:
- Google. The MCP Toolbox for Databases (googleapis/mcp-toolbox) initialized its HTTP client without a redirect policy and without validating target IP addresses. A crafted path parameter could make the toolbox follow a redirect to an internal endpoint and send requests on the attacker’s behalf. Google assigned CVE-2026-14540 with a CVSS score of 8.0.
- JPMorgan Chase. A documentation-search MCP server in the bank’s open-source repository contained two sibling tools. One applied a domain allowlist before fetching. The other,
related(), fetched any caller-supplied URL with no restriction. The instructive detail: the component was forked from an AWS project that never dereferenced the URL at all. JPMorgan’s rewrite added the fetch and left out the allowlist. Two competent teams, one routine fork, and a hole neither codebase had on its own. - Weaviate. A configurable endpoint value decided where the server sent outbound requests. The fix constrained it to the Google API hosts the feature actually needs.
- The French government. DINUM, the interministerial digital directorate, maintains the official MCP server for the national open-data platform. A URL field supplied by data producers was fetched server-side in a way open to DNS rebinding and able to reach cloud metadata.
- Tangerang City, Indonesia. A security-operations MCP server whose web-checking tool advertised SSRF protection, but only rejected literal IP addresses. Any DNS name pointing at a private or metadata address sailed through.
Read the list again with the industries removed. A hyperscaler, a bank, a database company, a national government, and a city government, on three continents, wrote MCP servers independently, and the same mistake appeared in all five. As the researcher put it: “That is not individual carelessness. MCP makes it very easy to expose a function to an agent, and nothing in that path prompts the developer to ask who controls each argument, or what happens to the data coming back.”
There is more in the queue. Sixteen published GitHub security advisories now credit him as reporter, covering command injection, authentication gaps, session hijack, credential leaks, and bypasses of earlier fixes, including a 1Password MCP tool that leaked its own service-account token and a CryptoAPIs hub that let unauthenticated callers spend the operator’s API key. Five US federal MCP servers, reported September 2, are still in triage, including a Department of Veterans Affairs benefits server that logs full upstream error responses without redaction, bodies that can contain a veteran’s name, Social Security number, date of birth, and address, triggered by routine validation failures. Japan’s Digital Agency runs a grants MCP server that, as reported to JPCERT on September 1, had no authentication at all; the pull request fixing it remains unmerged.
Why Every Scanner Misses This
The most quietly damning section of the research explains why the standard enterprise tooling, the software composition analysis and dependency scanners every large buyer already pays for, sees none of this. Those tools look for known-vulnerable packages and follow call graphs through code. In an MCP server, the dangerous input does not come through a code path those tools model. It comes over the transport, as a tool argument the model chose, described by a tool manifest the scanner never reads.
“The call graph stops at the transport boundary,” as the writeup puts it. Google’s missing redirect policy, JPMorgan’s allowlist gap, and the VA’s unredacted logs all sit in code a dependency scanner would pass. None of the dependencies are vulnerable. The vulnerability is in what the server trusts.
Ars Technica’s framing, built on interviews with Rapid7’s Douglas McKee, lands on the same structural point from the defense side: “Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between.” McKee’s operating rule deserves to be pinned above every agent developer’s desk: “Anything passed from an LLM to your tool should be treated like input from a stranger on the Internet, because in a prompt injection scenario that’s exactly what it is.”
What This Means When the Agent Is Shopping
Now connect this to Monday’s commerce news, because the connection is the story. TikTok’s Shopping Assistant, ChatGPT’s Instant Checkout, Shopify’s agent-ready rails: all of them assume the agent can consult external context mid-purchase, and the consult mechanism of record is MCP. When a shopping agent calls a product-data server, it inherits two claims at once: that the server is what it claims to be, and that the data it returns is true.
Tuesday’s audits demolish the first claim. Six expired domains, 23 servers on home connections, 19 hostnames in China and 18 in Russia, and no marketplace review anywhere. For a commerce agent, the expired-domain finding is not an abstract supply chain risk, it is a product recommendation hijack priced at $4. Register the domain, inherit the server identity, and every agent still configured to call it now consults you before checkout. The attacker does not need to hack the agent. The agent was built to trust the pipe.
But here is the uncomfortable part: even a perfectly vetted, perfectly hosted, perfectly signed server does not solve the second claim, because the everyday threat to product data is not a hostile server. It is a legitimate server faithfully returning a contaminated corpus. This column walked the evidence last week: a Singapore enforcement action against a provider that sold machine-written five-star reviews with a rating calculator and a replacement warranty, gig listings still live two days later at S$5 a post, and Trustpilot’s own accounting of 4.5 million fake reviews removed from one platform in 2024 alone. A well-behaved MCP server that scrapes a marketplace and serves those reviews with a 4.8 average is infrastructure doing exactly what it was built to do. The poison is upstream of the pipe, in the corpus itself.
That is why the two audits, important as they are, only cover half the trust problem in agentic commerce. Server identity is a solvable infrastructure problem, and OX names the solution set: vetting, code signing, origin verification. Delegation trust is a solvable architecture problem, and the answer is the zero-trust model the industry abandoned in its rush to agents, plus McKee’s stranger-on-the-Internet rule for every tool argument. Data merit is the problem nobody has assigned, and it is the one that decides whether the purchase is any good.
The Three Layers a Shopping Agent Actually Needs
For anyone building or operating commerce agents, the requirements now stack cleanly.
Vetted servers. Treat every MCP server the way you treat a vendor: verify who operates it, where it runs, whether the published code matches the running code, and whether the domain is actually owned by the operator. Until marketplaces add bouncers, that diligence is the operator’s job. A $12 expired domain should never outrank a due diligence checklist.
Zero-trust delegation. Assume any value crossing a protocol boundary, MCP, A2A, or whatever comes next, is attacker-controlled until validated at the point of use. Validate resolved addresses, not hostnames. Constrain every fetching tool, not just the first one written. Redact upstream bodies before they reach logs.
Independent product evidence. The merit question, is this actually a good product, cannot be answered by the seller’s own data, the platform’s own assistant, or an unvetted scraper. It needs a source with no stake in the transaction: filtering fake and incentivized reviews before any score is computed, weighting authentic ones, and scoring review quality rather than review count, because count is precisely what the S$5-per-post economy manufactures. GoBuy’s Smart Score, 0 to 100, is built on exactly that filtered foundation, the GoBuy Verified badge requires a score above 80 sustained across 90 days, the time window a burst campaign cannot fake, and only the top seven products per category survive the cut. For agents, the same evidence arrives as one structured call to gobuy.ai/api/mcp; for humans, the Chrome extension injects the trust panel directly onto the Amazon page.
The protocol community will fix its layer. The researcher presents the cross-vendor pattern at MCPCon on October 23, and the fixes he documents, Google’s startup-time URL rejection, JPMorgan’s deployed allowlist, are neither exotic nor expensive. Marketplaces will eventually hire bouncers. What neither fix touches is the corpus, and the corpus is where the shopping happens. The industry spent Monday wiring agents into feeds and checkouts. It spent Tuesday learning nobody vets the pipes. Wednesday’s job is admitting that even vetted pipes carry unfiltered water. Check what you are about to buy, and what your agent is about to buy for you, at gobuy.ai, and if you build shopping agents, wire them to filtered, independent product evidence at gobuy.ai/agent-docs.