Data Science MCP Server

data-science-mcp PyPI v1.2.0

Published by an unidentified publisher — no publish provenance and no public repository, so the publisher could not be verified and the source cannot be independently located.

Data Science MCP Server — Model training, evaluation, and evolution tools for agentic ML workflows. Integrates with agent-utilities IModelEvolver (CONCEPT:AHE-3.15).

Trust grade
A
92/100
Last scanned get badge →
Trust
A · 92/100
Adoption risk for you: the threat score, then adjusted down for blast radius, publisher verification and how much the scan could see. Deterministic; every point is auditable.
Capability
High
Blast radius if it went rogue — what the server’s tools could reach. Independent of trust.
Coverage
Source
How much the scan could actually inspect. Shallow coverage is stated, never hidden.
Share this Trust Score
𝕏 Share LinkedIn Reddit
A Why this grade threat 100 − adoption risk = 92/100

The grade answers one question — how safe is this server for you to adopt — so it is computed in two auditable stages. Nothing below is an opinion or an LLM's guess; every line is a real term the deterministic engine applied, and the same input always yields the same number.

1. Threat score — 100 − 0 = 100. What the published surface and source actually contain:

The deterministic scan raised no scored threat in the surface it inspected — the threat score stayed at 100. Capability observations and advisory notes are recorded but never lower it.

2. Client adoption risk — 100 − 8 = 92. Three small, subtract-only factors that reflect your risk in adopting it — a clean scan proves less on a powerful, unverified or barely-inspectable package, so the grade says so plainly:

PointsAdoption-risk factor
−6 capability blast radius (high) — client exposure if the model is manipulated
−2 publisher verification (unlinked) — no provenance/repo link, but the shipped source was fully read

Capability observations and info notes are shown under Findings but never scored. Open any row's finding below for the file, line and evidence behind a deduction.

Findings 8

high Dynamic code execution in server code (data_science_mcp/inference/base.py)MTC-SRC-001

In the server's implementation (`data_science_mcp/inference/base.py:6`): Evaluating strings as code is the most direct RCE primitive; if any tool input reaches it, the server executes attacker-chosen code. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: post-train reliability eval (one completion per case). Both vLLM and SGLang expose the **same OpenAI-compatible HTTP pr

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server data_science_mcp/inference/base.py

high Dynamic code execution in server code (data_science_mcp/kernels/_runner.py)MTC-SRC-001

In the server's implementation (`data_science_mcp/kernels/_runner.py:49`): Evaluating strings as code is the most direct RCE primitive; if any tool input reaches it, the server executes attacker-chosen code. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: ] = {} try: exec(compile(src, "<candidate>", "exec"), namespace) # noqa: S102 — sandboxed subprocess ex

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server data_science_mcp/kernels/_runner.py

high Dynamic code execution in server code (data_science_mcp/training_pipeline.py)MTC-SRC-001

In the server's implementation (`data_science_mcp/training_pipeline.py:238`): Evaluating strings as code is the most direct RCE primitive; if any tool input reaches it, the server executes attacker-chosen code. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: fn for the reliability eval (defaults to a no-op echo when omitted so the pipeline still completes on CPU).

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server data_science_mcp/training_pipeline.py

high Shell/command execution in server code (data_science_mcp/kernels/_runner.py)MTC-SRC-002

In the server's implementation (`data_science_mcp/kernels/_runner.py:49`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: ] = {} try: exec(compile(src, "<candidate>", "exec"), namespace) # noqa: S102 — sandboxed subprocess ex

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server data_science_mcp/kernels/_runner.py

high Shell/command execution in server code (data_science_mcp/kernels/kernel_verifier.py)MTC-SRC-002

In the server's implementation (`data_science_mcp/kernels/kernel_verifier.py:53`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: proc = subprocess.run( [self.python_exe, "-m", "data_science_mcp.kernels._runner",

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server data_science_mcp/kernels/kernel_verifier.py

low Dynamic code execution in packaging/dev tooling (tests/test_launch.py)MTC-SRC-001

In a packaging/dev/install script (shipped, but not the server runtime) (`tests/test_launch.py:2`): Evaluating strings as code is the most direct RCE primitive; if any tool input reaches it, the server executes attacker-chosen code. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: ed launcher + benchmark eval (CONCEPT:DS-AHE.trainer.concept-4/006). Config builders and the ``accelerate launch`` argv

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server tests/test_launch.py

low Shell/command execution in packaging/dev tooling (scripts/security_sanitizer.py)MTC-SRC-002

In a packaging/dev/install script (shipped, but not the server runtime) (`scripts/security_sanitizer.py:136`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: try: result = subprocess.run( ["git", "ls-files", "--cached", "--others", "--exclude-standard"],

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server scripts/security_sanitizer.py

low Shell/command execution in packaging/dev tooling (tests/conftest.py)MTC-SRC-002

In a packaging/dev/install script (shipped, but not the server runtime) (`tests/conftest.py:81`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: ngine.sock") proc = subprocess.Popen( [binary, "--socket-path", sock], stdout=subprocess.PIPE,

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server tests/conftest.py

Tools 38

Each tool and what it can reach — statically extracted from the published source.

  • build_training_datasetno sensitive capability
  • compose_rewardno sensitive capability
  • cross_validateno sensitive capability
  • curate_corpusno sensitive capability
  • dataset_lineageno sensitive capability
  • decontaminate_corpusno sensitive capability
  • dedup_corpusno sensitive capability
  • deep_train_predictno sensitive capability
  • describe_datasetno sensitive capability
  • ds_specialize_kernelno sensitive capability
Show 28 more tools ↓
  • evaluate_modelno sensitive capability
  • evolve_model_classno sensitive capability
  • fit_modelno sensitive capability
  • generate_interpretability_testsno sensitive capability
  • get_pareto_frontierno sensitive capability
  • grade_responseno sensitive capability
  • load_datasetno sensitive capability
  • merge_adapters_tiesno sensitive capability
  • predictno sensitive capability
  • prepare_pretrain_datano sensitive capability
  • pretrain_modelno sensitive capability
  • quant_derivativesno sensitive capability
  • quant_forensicno sensitive capability
  • quant_market_makingno sensitive capability
  • quant_microstructureno sensitive capability
  • quant_signalsno sensitive capability
  • quant_sizingno sensitive capability
  • quant_statespaceno sensitive capability
  • quant_validationno sensitive capability
  • rank_modelsno sensitive capability
  • run_interpretability_suiteno sensitive capability
  • split_datasetno sensitive capability
  • train_dpono sensitive capability
  • train_grpono sensitive capability
  • train_ppono sensitive capability
  • train_rewardno sensitive capability
  • train_sftno sensitive capability
  • train_tokenizerno sensitive capability

What this scan could not see

Versions 1

Scan history per published version. The engine is deterministic — the same version always yields the same score, so a changed score means the package itself changed.

VersionScoreFindingsEngineScanned
v1.2.0 latest A 92/100 8 1.13.0 2026-08-25

Embed this score

Show this server's live Trust Score in your README, docs or website. The badge is served straight from the registry and updates automatically after every rescan — no API key needed. It links back to this page, so anyone who sees the grade can also read the findings behind it instead of taking a number on faith.

MCP Trust Score: A · 92/100
Markdown (GitHub README)
[![MCP Trust Score](https://mcptrustchecker.com/registry/data-science-mcp/badge.svg)](https://mcptrustchecker.com/registry/data-science-mcp)
HTML
<a href="https://mcptrustchecker.com/registry/data-science-mcp"><img src="https://mcptrustchecker.com/registry/data-science-mcp/badge.svg" alt="MCP Trust Score" height="20"></a>
Prefer shields.io styling? Point it at https://mcptrustchecker.com/registry/data-science-mcp/badge.json via https://img.shields.io/endpoint?url=…

Verify this score yourself

The score above is reproducible: the same package version always yields the same result. Run it locally or over the free API — no account, no LLM, fully deterministic.

npx mcptrustchecker scan data-science-mcp --online --registry pypi

Use the free API → How scoring works

More in Developer Tools