Web Scraper Toolkit MCP Server

web-scraper-toolkit PyPI v0.3.5

Published by imyourboyroy — no publish provenance, so origin is unverified, but the source is public: the repository link below is self-declared yet readable, so you can inspect the code before adopting it.

Trust grade
A
93/100
Last scanned get badge →
Trust
A · 93/100
Adoption risk for you: the threat score, then adjusted down for blast radius, publisher verification and how much the scan could see. Deterministic; every point is auditable.
Capability
High
Blast radius if it went rogue — what the server’s tools could reach. Independent of trust.
Coverage
Source
How much the scan could actually inspect. Shallow coverage is stated, never hidden.
Share this Trust Score
𝕏 Share LinkedIn Reddit
A Why this grade threat 100 − adoption risk = 93/100

The grade answers one question — how safe is this server for you to adopt — so it is computed in two auditable stages. Nothing below is an opinion or an LLM's guess; every line is a real term the deterministic engine applied, and the same input always yields the same number.

1. Threat score — 100 − 0 = 100. What the published surface and source actually contain:

The deterministic scan raised no scored threat in the surface it inspected — the threat score stayed at 100. Capability observations and advisory notes are recorded but never lower it.

2. Client adoption risk — 100 − 7 = 93. Three small, subtract-only factors that reflect your risk in adopting it — a clean scan proves less on a powerful, unverified or barely-inspectable package, so the grade says so plainly:

PointsAdoption-risk factor
−6 capability blast radius (high) — client exposure if the model is manipulated
−1 publisher verification (public source) — no provenance, but the source is public and inspectable

Capability observations and info notes are shown under Findings but never scored. Open any row's finding below for the file, line and evidence behind a deduction.

Findings 3

high Shell/command execution in server code (src/web_scraper_toolkit/core/script_diagnostics.py)MTC-SRC-002

In the server's implementation (`src/web_scraper_toolkit/core/script_diagnostics.py:259`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: r() completed = subprocess.run( command, cwd=str(self.project_root), text=Tr

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server src/web_scraper_toolkit/core/script_diagnostics.py

low Hardcoded egress to an external endpoint in packaging/dev tooling (tests/test_playwright_manager.py)MTC-SRC-003

In a packaging/dev/install script (shipped, but not the server runtime) (`tests/test_playwright_manager.py:301`): A hardcoded outbound call to a fixed external host inside server code is a classic exfiltration/telemetry channel — especially paired with reads of local data. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: e( pm.smart_fetch("https://pressrelease.com/files/example.docx") ) self.assertIsNone(conten

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server tests/test_playwright_manager.py

low Shell/command execution in packaging/dev tooling (tests/test_challenge_fixture_replay.py)MTC-SRC-002

In a packaging/dev/install script (shipped, but not the server runtime) (`tests/test_challenge_fixture_replay.py:46`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: , ] completed = subprocess.run( command, cwd=str(PROJECT_ROOT), text=True, captu

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server tests/test_challenge_fixture_replay.py

Tools 62

Each tool and what it can reach — statically extracted from the published source.

  • batch_scrapeingests untrusted input
  • crawl_siteingests untrusted input
  • deep_researchingests untrusted input
  • download_fileingests untrusted input
  • get_token_countreads sensitive data
  • scrape_urlingests untrusted input
  • search_webingests untrusted input
  • batch_contactsno sensitive capability
  • browser_accessibility_treeno sensitive capability
  • browser_clickno sensitive capability
Show 52 more tools ↓
  • browser_closeno sensitive capability
  • browser_evaluateno sensitive capability
  • browser_get_elementsno sensitive capability
  • browser_get_interaction_mapno sensitive capability
  • browser_hoverno sensitive capability
  • browser_navigateno sensitive capability
  • browser_press_keyno sensitive capability
  • browser_read_pageno sensitive capability
  • browser_screenshotno sensitive capability
  • browser_scrollno sensitive capability
  • browser_solve_challengeno sensitive capability
  • browser_typeno sensitive capability
  • browser_wait_forno sensitive capability
  • cancel_jobno sensitive capability
  • chunk_textno sensitive capability
  • clear_cacheno sensitive capability
  • clear_historyno sensitive capability
  • clear_host_profileno sensitive capability
  • clear_sessionno sensitive capability
  • click_elementno sensitive capability
  • configure_host_learningno sensitive capability
  • configure_retryno sensitive capability
  • configure_runtimeno sensitive capability
  • configure_scraperno sensitive capability
  • configure_stealthno sensitive capability
  • detect_content_typeno sensitive capability
  • extract_contactsno sensitive capability
  • extract_linksno sensitive capability
  • extract_tablesno sensitive capability
  • fill_formno sensitive capability
  • get_cache_statsno sensitive capability
  • get_configno sensitive capability
  • get_historyno sensitive capability
  • get_host_profilesno sensitive capability
  • get_metadatano sensitive capability
  • get_sitemapno sensitive capability
  • health_checkno sensitive capability
  • list_jobsno sensitive capability
  • list_sessionsno sensitive capability
  • new_sessionno sensitive capability
  • poll_jobno sensitive capability
  • reload_runtime_configno sensitive capability
  • run_bot_surface_diagnosticno sensitive capability
  • run_browser_info_diagnostic_toolno sensitive capability
  • run_challenge_diagnosticno sensitive capability
  • run_playbookno sensitive capability
  • save_pdfno sensitive capability
  • screenshotno sensitive capability
  • set_host_profileno sensitive capability
  • start_jobno sensitive capability
  • truncate_textno sensitive capability
  • validate_urlno sensitive capability

What this scan could not see

Versions 1

Scan history per published version. The engine is deterministic — the same version always yields the same score, so a changed score means the package itself changed.

VersionScoreFindingsEngineScanned
v0.3.5 latest A 93/100 3 1.9.0 2026-07-23

Embed this score

Show this server's live Trust Score in your README, docs or website. The badge is served straight from the registry and updates automatically after every rescan — no API key needed. It links back to this page, so anyone who sees the grade can also read the findings behind it instead of taking a number on faith.

MCP Trust Score: A · 93/100
Markdown (GitHub README)
[![MCP Trust Score](https://mcptrustchecker.com/registry/web-scraper-toolkit/badge.svg)](https://mcptrustchecker.com/registry/web-scraper-toolkit)
HTML
<a href="https://mcptrustchecker.com/registry/web-scraper-toolkit"><img src="https://mcptrustchecker.com/registry/web-scraper-toolkit/badge.svg" alt="MCP Trust Score" height="20"></a>
Prefer shields.io styling? Point it at https://mcptrustchecker.com/registry/web-scraper-toolkit/badge.json via https://img.shields.io/endpoint?url=…

Verify this score yourself

The score above is reproducible: the same package version always yields the same result. Run it locally or over the free API — no account, no LLM, fully deterministic.

npx mcptrustchecker scan web-scraper-toolkit --online --registry pypi

Use the free API → How scoring works

More in Developer Tools