The State of MCP Server Security

This is what happens when every Model Context Protocol server in the MCP Trust Registry is read by one deterministic engine instead of being ranked by stars: 30,351 npm and PyPI packages, each scanned from its real published code, each graded by the same auditable rules. Nothing below is a popularity chart, an activity feed or a survey — every number is a scan result, and the same input would produce it again.

Packages scanned
30,351
one scan result per published package
Products listed
24,071
forks of one server collapse into a single catalog card
Findings recorded
39,729
across 11,880 packages that raised at least one
No findings at all
60.9%
scanned clean, not one scored issue
Toxic flows
1,218
cross-tool exfiltration chains — a measured zero, not a gap

Trust grade distribution

94.4% of scanned packages hold an A. That is the honest headline, and it is less reassuring than it sounds: the story worth reading is the 1,707 packages that do not — they are the ones carrying the prompt injection, the tool poisoning and the unreviewable install scripts.

Trust grade distribution across the catalogGrade A: 28,644 (94.4%)Grade A94.4%28,644Grade B: 1,605 (5.3%)Grade B5.3%1,605Grade C: 75 (0.2%)Grade C0.2%75Grade D: 17 (0.1%)Grade D0.1%17Grade F: 10 (0%)Grade F0%10

A grade describes the exact version that was scanned, never the project in general. Percentages are of 30,351 scanned packages.

Score distribution

Every package starts at 100 and loses points for what the engine actually finds. The result is not a bell curve — it is a wall at the top and a long, thin tail, which is exactly what a deterministic rule set does to a population where most packages are boring and a few are not.

Trust score distribution, 0 to 1000-9: 000-910-19: 0010-1920-29: 0020-2930-39: 0030-3940-49: 3340-4950-59: 7750-5960-69: 171760-6970-79: 757570-7980-89: 1,6051,60580-8990-99: 26,62126,62190-99100: 2,0232,023100

100 is kept as its own column: “perfect” and “nearly perfect” are different claims.

Blast radius and what the scan could see

The grade answers “is this dangerous”. These two answer the questions that decide how much the grade is worth: how much damage could this server do if it went rogue, and how much of the package was actually readable. A high capability is not a finding — it is the reason a finding would matter.

Capability blast radius
Capability blast radiusminimal: 16,519 (54.4%)minimal54.4%16,519moderate: 5,040 (16.6%)moderate16.6%5,040high: 8,548 (28.2%)high28.2%8,548critical: 244 (0.8%)critical0.8%244

  • minimal — no file, network or shell reach worth naming
  • moderate — reaches one of file, network or shell
  • high — broad reach — file and network, or shell
  • critical — unrestricted reach

Scan coverage
Scan coveragesource: 28,690 (94.5%)source94.5%28,690metadata: 802 (2.6%)metadata2.6%802live: 859 (2.8%)live2.8%859

  • source — the published code was read
  • metadata — no readable source shipped — manifest only
  • live — coverage recorded by the scan

What actually fires

Ranked by how many packages each rule flagged, not by how alarming it sounds. The bars are packages; the table also gives raw findings, because one rule can fire several times inside one package. Rule ids and their thresholds are part of the published engine — nothing here is a human judgement call.

Most frequently firing rulesMTC-SRC-002: 7,588 (high)MTC-SRC-002high7,588MTC-SUP-011: 5,566 (info)MTC-SUP-011info5,566MTC-SUP-012: 4,432 (info)MTC-SUP-012info4,432MTC-SRC-002: 2,110 (low)MTC-SRC-002low2,110MTC-SRC-003: 1,580 (medium)MTC-SRC-003medium1,580MTC-SRC-001: 1,464 (high)MTC-SRC-001high1,464MTC-CAP-005: 1,100 (low)MTC-CAP-005low1,100MTC-SUP-010: 1,041 (low)MTC-SUP-010low1,041MTC-SRC-005: 1,032 (medium)MTC-SRC-005medium1,032MTC-SRC-009: 874 (medium)MTC-SRC-009medium874MTC-NET-005: 859 (info)MTC-NET-005info859MTC-CAP-001: 819 (high)MTC-CAP-001high819

Bar length is packages affected; colour is the rule’s severity.

RuleSeverityPackagesFindings
MTC-SRC-002 high 7,588 17,124
MTC-SUP-011 info 5,566 5,566
MTC-SUP-012 info 4,432 4,432
MTC-SRC-002 low 2,110 4,259
MTC-SRC-003 medium 1,580 2,552
MTC-SRC-001 high 1,464 1,985
MTC-CAP-005 low 1,100 1,957
MTC-SUP-010 low 1,041 1,041
MTC-SRC-005 medium 1,032 1,556
MTC-SRC-009 medium 874 1,369
MTC-NET-005 info 859 859
MTC-CAP-001 high 819 1,326

Findings by severity

This is the reconciliation between the two numbers above: most packages carry something, but the bulk of it is low-severity supply-chain hygiene, which is why the grade distribution still leans hard on A. info notes are counted here for completeness and are never scored.

Findings by severitycritical: 266critical266high: 21,806high21,806medium: 8,810medium8,810low: 8,846low8,846info: 10,918info10,918

Counts are individual findings, so one package can appear in several rows.

Popularity is not safety

Each bar is normalised to full width, so these compare mix, not size. If downloads predicted safety, the mix would shift visibly from bottom to top. Read it for that, and for nothing else.

Covers 22,027 packages — 72.6% of the catalog. Only packages carrying an npm weekly-download figure can appear here; the rest publish no such number and are excluded rather than silently counted as zero.

Grade mix by weekly download volumeABCDF<100 downloads/week17,862 packages<100 downloads/week · A: 17,049 (95%)A 95%<100 downloads/week · B: 763 (4%)4%<100 downloads/week · C: 40 (0%)<100 downloads/week · D: 8 (0%)<100 downloads/week · F: 2 (0%)100-1k downloads/week3,295 packages100-1k downloads/week · A: 3,097 (94%)A 94%100-1k downloads/week · B: 182 (6%)B 6%100-1k downloads/week · C: 9 (0%)100-1k downloads/week · D: 5 (0%)100-1k downloads/week · F: 2 (0%)1k-10k downloads/week665 packages1k-10k downloads/week · A: 611 (92%)A 92%1k-10k downloads/week · B: 47 (7%)B 7%1k-10k downloads/week · C: 5 (1%)1k-10k downloads/week · D: 2 (0%)10k-100k downloads/week141 packages10k-100k downloads/week · A: 131 (93%)A 93%10k-100k downloads/week · B: 8 (6%)B 6%10k-100k downloads/week · C: 2 (1%)100k+ downloads/week64 packages100k+ downloads/week · A: 59 (92%)A 92%100k+ downloads/week · B: 4 (6%)B 6%100k+ downloads/week · C: 1 (2%)

Bucket counts are on the right of each bar; segment labels are that grade’s share of the bucket.

Grade drift

A score belongs to a version, not to a project, so the interesting question is what happens when a package ships a new one. 30,332 packages have been scanned more than once and 6 of them changed grade — the drop column is the rug-pull signal worth watching.

Score movement, latest rescan
Score movement on the latest rescanimproved: 21unchanged: 30,287dropped: 24↑ 21 improved→ 30,287 unchanged↓ 24 dropped

Only the latest rescan against the one before it — one comparison per package, so a single flapping package cannot dominate the summary. A drop means the new version scored lower than the version before it.

Most recent grade changes
ServerGrade moveDetected
Phren Cli BA 2026-09-07
Memoir Cli CA 2026-09-07
Sinataoke Cn AB 2026-09-07
Brave Real Browser AC 2026-09-07
Trace AB 2026-09-07
Phren Cli AB 2026-09-06
Par5 AB 2026-09-06
Brave Real Browser CA 2026-09-06

The full feed, with versions and the engine build behind each move, is on what changed.

What is in the catalog

Context for everything above: which ecosystems these packages come from, whether the publisher could be verified, and how the top categories are weighted. These count packages, so they add up to 30,351 — the catalog’s own category pages count products and read lower on purpose.

Top categories
Packages by categoryDeveloper Tools: 8,559Developer Tools8,559Web, Search & Scraping: 2,028Web, Search & Scra…2,028Finance & Commerce: 1,864Finance & Commerce1,864AI & Agents: 1,723AI & Agents1,723Other: 1,669Other1,669Productivity & Workflow: 1,612Productivity & Work…1,612Design & Media: 1,238Design & Media1,238Learning & Documentation: 1,133Learning & Docume…1,133AI, Memory & Reasoning: 1,093AI, Memory & Reaso…1,093Business & CRM: 1,070Business & CRM1,070Data Science & ML: 1,059Data Science & ML1,059Security & Testing: 1,044Security & Testing1,044

Top 12 by package count.

Registry and publisher verification
Packages by registry and verificationnpm: 22,517npm22,517PyPI: 6,975PyPI6,975npm: 859npm859verification: vendor: 99verification: vendor99verification: source: 3,756verification: source3,756verification: none: 6,887verification: none6,887verification: repo: 19,609verification: repo19,609

“vendor” means the package is published by the vendor it claims to integrate; “source” means the repository was confirmed.

How fresh this is

Scanning runs on a queue, not on page load. These figures are recomputed at most once a minute from the stored scans, and this snapshot was built 2026-09-07 15:43 UTC.

22,213 packages scanned in the last 24h 23,227 in the last 7 days oldest scan on record 2026-07-24 newest 2026-09-07 engine 1.13.0 (30,310) · 1.12.1 (35) · 1.10.0 (6)

A package scanned on an older engine build carries that build's rules until it is re-read, so part of any difference between two packages can be our methodology rather than their code. That is why every scan records the engine version that produced it, and why what changed marks the rows where the engine moved. Re-check any package yourself with the free API or the CLI — same input, same score, no account needed.