The State of MCP Server Security

This is what happens when every Model Context Protocol server in the MCP Trust Registry is read by one deterministic engine instead of being ranked by stars: 31,300 npm and PyPI packages, each scanned from its real published code, each graded by the same auditable rules. Nothing below is a popularity chart, an activity feed or a survey — every number is a scan result, and the same input would produce it again.

Packages scanned
31,300
one scan result per published package
Products listed
24,571
forks of one server collapse into a single catalog card
Findings recorded
52,324
across 19,190 packages that raised at least one
No findings at all
38.7%
scanned clean, not one scored issue
Toxic flows
0
cross-tool exfiltration chains — a measured zero, not a gap

Trust grade distribution

95.5% of scanned packages hold an A. That is the honest headline, and it is less reassuring than it sounds: the story worth reading is the 1,396 packages that do not — they are the ones carrying the prompt injection, the tool poisoning and the unreviewable install scripts.

Trust grade distribution across the catalogGrade A: 29,904 (95.5%)Grade A95.5%29,904Grade B: 807 (2.6%)Grade B2.6%807Grade C: 360 (1.2%)Grade C1.2%360Grade D: 126 (0.4%)Grade D0.4%126Grade F: 103 (0.3%)Grade F0.3%103

A grade describes the exact version that was scanned, never the project in general. Percentages are of 31,300 scanned packages.

Score distribution

Every package starts at 100 and loses points for what the engine actually finds. The result is not a bell curve — it is a wall at the top and a long, thin tail, which is exactly what a deterministic rule set does to a population where most packages are boring and a few are not.

Trust score distribution, 0 to 1000-9: 440-910-19: 3310-1920-29: 3320-2930-39: 6630-3940-49: 242440-4950-59: 636350-5960-69: 12612660-6970-79: 36036070-7980-89: 80780780-8990-99: 13,11313,11390-99100: 16,79116,791100

100 is kept as its own column: “perfect” and “nearly perfect” are different claims.

Blast radius and what the scan could see

The grade answers “is this dangerous”. These two answer the questions that decide how much the grade is worth: how much damage could this server do if it went rogue, and how much of the package was actually readable. A high capability is not a finding — it is the reason a finding would matter.

Capability blast radius
Capability blast radiusminimal: 21,144 (67.6%)minimal67.6%21,144moderate: 1,028 (3.3%)moderate3.3%1,028high: 9,128 (29.2%)high29.2%9,128critical: 0 (0%)critical0%0

  • minimal — no file, network or shell reach worth naming
  • moderate — reaches one of file, network or shell
  • high — broad reach — file and network, or shell
  • critical — unrestricted reach

Scan coverage
Scan coveragesource: 29,970 (95.8%)source95.8%29,970metadata: 1,330 (4.2%)metadata4.2%1,330

  • source — the published code was read
  • metadata — no readable source shipped — manifest only

What actually fires

Ranked by how many packages each rule flagged, not by how alarming it sounds. The bars are packages; the table also gives raw findings, because one rule can fire several times inside one package. Rule ids and their thresholds are part of the published engine — nothing here is a human judgement call.

Most frequently firing rulesMTC-SUP-011: 11,293 (low)MTC-SUP-011low11,293MTC-SRC-002: 8,662 (high)MTC-SRC-002high8,662MTC-SUP-012: 4,857 (info)MTC-SUP-012info4,857MTC-SRC-003: 1,726 (medium)MTC-SRC-003medium1,726MTC-SRC-001: 1,363 (high)MTC-SRC-001high1,363MTC-SRC-009: 1,250 (medium)MTC-SRC-009medium1,250MTC-SRC-005: 1,204 (medium)MTC-SRC-005medium1,204MTC-SUP-010: 1,091 (medium)MTC-SUP-010medium1,091MTC-SUP-006: 976 (medium)MTC-SUP-006medium976MTC-SRC-006: 772 (high)MTC-SRC-006high772MTC-SRC-007: 498 (medium)MTC-SRC-007medium498MTC-SRC-010: 446 (high)MTC-SRC-010high446

Bar length is packages affected; colour is the rule’s severity.

RuleSeverityPackagesFindings
MTC-SUP-011 low 11,293 11,293
MTC-SRC-002 high 8,662 22,659
MTC-SUP-012 info 4,857 4,857
MTC-SRC-003 medium 1,726 2,889
MTC-SRC-001 high 1,363 2,559
MTC-SRC-009 medium 1,250 1,979
MTC-SRC-005 medium 1,204 1,915
MTC-SUP-010 medium 1,091 1,091
MTC-SUP-006 medium 976 3,886
MTC-SRC-006 high 772 1,406
MTC-SRC-007 medium 498 969
MTC-SRC-010 high 446 726

Findings by severity

This is the reconciliation between the two numbers above: most packages carry something, but the bulk of it is low-severity supply-chain hygiene, which is why the grade distribution still leans hard on A. info notes are counted here for completeness and are never scored.

Findings by severitycritical: 0critical0high: 28,189high28,189medium: 12,842medium12,842low: 11,293low11,293info: 4,919info4,919

Counts are individual findings, so one package can appear in several rows.

Popularity is not safety

Each bar is normalised to full width, so these compare mix, not size. If downloads predicted safety, the mix would shift visibly from bottom to top. Read it for that, and for nothing else.

Covers 3,545 packages — 11.3% of the catalog. Only packages carrying an npm weekly-download figure can appear here; the rest publish no such number and are excluded rather than silently counted as zero.

Grade mix by weekly download volumeABCDF<100 downloads/week1,860 packages<100 downloads/week · A: 1,798 (97%)A 97%<100 downloads/week · B: 38 (2%)<100 downloads/week · C: 17 (1%)<100 downloads/week · D: 2 (0%)<100 downloads/week · F: 5 (0%)100-1k downloads/week1,043 packages100-1k downloads/week · A: 977 (94%)A 94%100-1k downloads/week · B: 37 (4%)4%100-1k downloads/week · C: 13 (1%)100-1k downloads/week · D: 9 (1%)100-1k downloads/week · F: 7 (1%)1k-10k downloads/week466 packages1k-10k downloads/week · A: 426 (91%)A 91%1k-10k downloads/week · B: 26 (6%)B 6%1k-10k downloads/week · C: 10 (2%)1k-10k downloads/week · D: 3 (1%)1k-10k downloads/week · F: 1 (0%)10k-100k downloads/week136 packages10k-100k downloads/week · A: 129 (95%)A 95%10k-100k downloads/week · B: 6 (4%)4%10k-100k downloads/week · D: 1 (1%)100k+ downloads/week40 packages100k+ downloads/week · A: 32 (80%)A 80%100k+ downloads/week · B: 5 (13%)B 13%100k+ downloads/week · C: 1 (3%)100k+ downloads/week · D: 2 (5%)D 5%

Bucket counts are on the right of each bar; segment labels are that grade’s share of the bucket.

Grade drift

A score belongs to a version, not to a project, so the interesting question is what happens when a package ships a new one. 3,564 packages have been scanned more than once and 244 of them changed grade — the drop column is the rug-pull signal worth watching.

Score movement, latest rescan
Score movement on the latest rescanimproved: 249unchanged: 3,307dropped: 8↑ 249 improved→ 3,307 unchanged↓ 8 dropped

Only the latest rescan against the one before it — one comparison per package, so a single flapping package cannot dominate the summary. A drop means the new version scored lower than the version before it.

Most recent grade changes
ServerGrade moveDetected
Thetacog FB 2026-07-22
N8n FD 2026-07-22
Patchwork BA 2026-07-22
Appstore Connect CA 2026-07-22
Vitalmcp BA 2026-07-22
Typegraph CB 2026-07-22
Auto Install CB 2026-07-22
4da CB 2026-07-22

The full feed, with versions and the engine build behind each move, is on what changed.

What is in the catalog

Context for everything above: which ecosystems these packages come from, whether the publisher could be verified, and how the top categories are weighted. These count packages, so they add up to 31,300 — the catalog’s own category pages count products and read lower on purpose.

Top categories
Packages by categoryDeveloper Tools: 9,363Developer Tools9,363Web, Search & Scraping: 2,027Web, Search & Scra…2,027Finance & Commerce: 1,830Finance & Commerce1,830AI & Agents: 1,768AI & Agents1,768Productivity & Workflow: 1,688Productivity & Work…1,688Other: 1,620Other1,620Design & Media: 1,244Design & Media1,244AI, Memory & Reasoning: 1,163AI, Memory & Reaso…1,163Learning & Documentation: 1,140Learning & Docume…1,140Business & CRM: 1,083Business & CRM1,083Data Science & ML: 1,059Data Science & ML1,059Security & Testing: 1,059Security & Testing1,059

Top 12 by package count.

Registry and publisher verification
Packages by registry and verificationnpm: 23,676npm23,676PyPI: 7,624PyPI7,624verification: vendor: 140verification: vendor140verification: source: 3,609verification: source3,609verification: none: 27,551verification: none27,551

“vendor” means the package is published by the vendor it claims to integrate; “source” means the repository was confirmed.

How fresh this is

Scanning runs on a queue, not on page load. These figures are recomputed at most once a minute from the stored scans, and this snapshot was built 2026-07-23 09:08 UTC.

27,736 packages scanned in the last 24h 31,300 in the last 7 days oldest scan on record 2026-07-22 newest 2026-07-23 engine 1.5.0 (31,300)

A package scanned on an older engine build carries that build's rules until it is re-read, so part of any difference between two packages can be our methodology rather than their code. That is why every scan records the engine version that produced it, and why what changed marks the rows where the engine moved. Re-check any package yourself with the free API or the CLI — same input, same score, no account needed.