Before you connect a new MCP server to Claude, ChatGPT, or any other agent, three things are checkable in under five minutes: who actually published it, what access it's asking for, and whether it's pinned to a specific version rather than whatever the maintainer ships next. None of that is optional, because nothing about the Model Context Protocol itself checks a server's code for malicious behavior before your agent runs it -- that job falls entirely on you or your client.
Anthropic open-sourced MCP in November 2024 to solve a real integration problem: connecting many different AI apps to many different tools and data sources used to mean custom code for every single pairing. Anthropic donated the protocol to a new vendor-neutral body, the Agentic AI Foundation, on December 9, 2025 -- the same foundation that now also hosts Google's rival agent-to-agent protocol alongside it. Claude, ChatGPT, Microsoft Copilot, and Gemini all support MCP for connecting to outside tools. None of that governance touches server safety: it decides who steers the standard, not who's allowed to publish a server that implements it.
What actually goes wrong when nobody checks
The clearest illustration is a single disclosure from May 2025, when researchers found that GitHub's own official MCP server could be turned against its user without any bug in the server's code at all.
The GitHub MCP "toxic agent flow", May 2025
- Posts a public GitHub issue containing hidden instructions, written to look like ordinary text.
- Reads the issue while doing the task it was asked to do: check open issues on the public repo.
- Follows the hidden instruction instead of the developer's -- it has no way to tell issue text from a command.
- Uses its own access token -- often scoped to every repository the developer can see, not just the public one -- to read a private repository.
- Publishes what it found into a pull request on the public repo, where the attacker can read it.
Invariant Labs, the security firm that found this, was explicit that nothing malfunctioned: the server read a public issue it was allowed to read, and opened a pull request it was allowed to open. The fix isn't a patch to GitHub's server code -- it's never handing an agent one token that spans both public and private repositories in the first place.
Three more incidents, and one still-open argument
That May 2025 disclosure wasn't an isolated case. The incidents below show the same underlying problem from different angles -- a trusted approval that doesn't mean what you think, and a clean-looking package that turns malicious only after it's earned your trust.
MCP, launch to open dispute
- Nov 25, 2024 — Anthropic open-sources MCP, the connector standard now behind this whole category of tool.
- May 26, 2025 — Invariant Labs discloses the GitHub MCP "toxic agent flow" -- a scoped-token fix, not a server patch.
- Jul 29, 2025 — Cursor ships a fix for MCPoison (CVE-2025-54136) after Check Point shows an approved MCP server can be silently swapped for a malicious one.
- Sep 15, 2025 — The postmark-mcp npm package ships version 1.0.16 -- its first with a backdoor that BCCs every email it sends.
- Dec 9, 2025 — Anthropic donates MCP to the newly formed Agentic AI Foundation, alongside OpenAI and Block.
- Apr 15, 2026 — OX Security discloses a systemic command-injection flaw in MCP's default server-launch design; Anthropic calls the behavior expected, not a bug.
The postmark-mcp case is the one to notice if you've ever installed anything from an open registry without reading the diff first: the package shipped fifteen ordinary, functioning versions specifically to earn trust before version 1.0.16 quietly added the line of code that BCC'd a copy of every email it touched to an outside address. Koi Security caught it after the package had been downloaded 1,643 times -- not because MCP's registry flagged it, but because someone happened to look.
Two things people assume about MCP that aren't quite true
In April 2026, security firm OX Security published research alleging a systemic flaw in how MCP servers are launched by default -- one that could, in the worst case, reach across as many as 200,000 running instances, and that produced ten official CVE numbers, nine of them rated critical. Anthropic didn't dispute the technical mechanism the researchers found, but declined to change MCP's reference implementation, saying the flagged behavior -- the STDIO execution model -- is a secure default whose input sanitization is the responsibility of whoever builds on top of it. That leaves a genuine, unresolved dispute rather than a settled fact, alongside a second, quieter misconception about what a registry listing actually proves:
Two claims worth separating from the facts
- MCP's default server-launch design is a critical, systemic vulnerability putting up to 200,000 servers at risk.
- A server listed in the official MCP Registry has had its code checked for malicious behavior.
The five-minute vet before you connect a new server
None of this means treating every MCP server as guilty until proven innocent. It means running the same five checks every time, regardless of how official a server looks:
Vet an MCP server before your agent runs it
- The official MCP Registry ties a server's name to a verified GitHub account or domain through namespace authentication -- a name like io.github.username/server means that specific GitHub user published it.
- A server that only needs to read one calendar shouldn't be handed a token that reaches every repository or inbox you own -- a scoped token instead of a blanket one is exactly what would have closed the May 2025 GitHub MCP gap.
- postmark-mcp shipped fifteen clean releases before its backdoor landed in version 1.0.16; a floating install would have picked that version up automatically.
- MCPoison (CVE-2025-54136) exploited the opposite assumption in Cursor: once approved, edits to a server's command or arguments ran silently. Cursor fixed this in version 1.3, released July 29, 2025.
- Feed it a file, issue, or message containing an obvious fake instruction and confirm the server -- and your agent -- treat it as inert content, not a command.
Running those five checks once closes the specific gaps above. The ways this quietly gets skipped anyway are the same four every time.
Four ways this check gets skipped without anyone noticing
None of this requires distrust of MCP as a standard -- it's doing exactly the integration job it was built for, and the incidents above are the ordinary growing pains of a fast-adopted connector standard, not a reason to avoid it. It requires treating a new MCP server the same way you'd treat a new CI dependency running with production access or a new inbox you're about to hand an agent -- worth five minutes of checking before it's running with your credentials, not after.
- Anyone can publish an MCP server; registry listing verifies the publisher, not the code.
- A May 2025 GitHub MCP flaw let one poisoned issue leak private repository data.
- A backdoored npm package quietly BCC'd emails for weeks before Koi Security caught it.
- Check Point's MCPoison let an approved server be silently swapped for a malicious one.
- Anthropic disputes a 2026 researcher claim that its own default design is a critical flaw.