Elite Reasoning MCP Server

elite-reasoning-mcp PyPI v1.2.1

Published by snehgabani — no publish provenance, so origin is unverified, but the source is public: the repository link below is self-declared yet readable, so you can inspect the code before adopting it.

Trust grade
A
93/100
Last scanned get badge →
Trust
A · 93/100
Adoption risk for you: the threat score, then adjusted down for blast radius, publisher verification and how much the scan could see. Deterministic; every point is auditable.
Capability
High
Blast radius if it went rogue — what the server’s tools could reach. Independent of trust.
Coverage
Source
How much the scan could actually inspect. Shallow coverage is stated, never hidden.
A Why this grade threat 100 − adoption risk = 93/100

The grade answers one question — how safe is this server for you to adopt — so it is computed in two auditable stages. Nothing below is an opinion or an LLM's guess; every line is a real term the deterministic engine applied, and the same input always yields the same number.

1. Threat score — 100 − 0 = 100. What the published surface and source actually contain:

The deterministic scan raised no scored threat in the surface it inspected — the threat score stayed at 100. Capability observations and advisory notes are recorded but never lower it.

2. Client adoption risk — 100 − 7 = 93. Three small, subtract-only factors that reflect your risk in adopting it — a clean scan proves less on a powerful, unverified or barely-inspectable package, so the grade says so plainly:

PointsAdoption-risk factor
−6 capability blast radius (high) — client exposure if the model is manipulated
−1 publisher verification (public source) — no provenance, but the source is public and inspectable

Capability observations and info notes are shown under Findings but never scored. Open any row's finding below for the file, line and evidence behind a deduction.

Findings 8

high Tool "run_elite_eval_suite" exposes command/code executionMTC-CAP-001

Tool "run_elite_eval_suite" appears to run shell commands or evaluate code (keyword "eval" in tool name). Arbitrary execution driven by model input is one of the most dangerous MCP capabilities; combined with any untrusted input it becomes RCE.

Fix: Sandbox execution, allowlist commands/arguments, and never pass model output to a shell unescaped.

Location: tool run_elite_eval_suite

high Tool "export_eval_harness" exposes command/code executionMTC-CAP-001

Tool "export_eval_harness" appears to run shell commands or evaluate code (keyword "eval" in tool name). Arbitrary execution driven by model input is one of the most dangerous MCP capabilities; combined with any untrusted input it becomes RCE.

Fix: Sandbox execution, allowlist commands/arguments, and never pass model output to a shell unescaped.

Location: tool export_eval_harness

high Shell/command execution in server code (test_quality_gate.py)MTC-SRC-002

In the server's implementation (`test_quality_gate.py:22`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: ubprocess process = subprocess.Popen( [sys.executable, "-m", "uvicorn", "core.integration.sync_server:app",

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server test_quality_gate.py

high Shell/command execution in server code (test_team_sync.py)MTC-SRC-002

In the server's implementation (`test_team_sync.py:23`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: nd server_process = subprocess.Popen( [sys.executable, "core/integration/sync_server.py"], stdout=su

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server test_team_sync.py

medium Untrusted input can drive an external actionMTC-FLOW-005

Untrusted-input tools ([validate_predictions, browse_tool_usage, introspect]) co-exist with external-action tools ([run_elite_eval_suite, export_eval_harness]). A prompt injection could cause unwanted external actions, though no direct sensitive-data leak path was found.

Evidence: untrusted [validate_predictions, browse_tool_usage, introspect] → sinks [run_elite_eval_suite, export_eval_harness]

Fix: Require confirmation for state-changing/egress actions triggered after processing untrusted content.

Location: flow validate_predictions → run_elite_eval_suite

low Mutating tool "run_elite_eval_suite" declares no destructiveHintMTC-CAP-005

Tool "run_elite_eval_suite" can mutate/egress but declares no destructiveHint. Clients that don't default to spec-safe behavior may not prompt before running it.

Fix: Declare accurate annotations, and gate destructive tools on user confirmation regardless.

Location: tool run_elite_eval_suite

low Mutating tool "export_eval_harness" declares no destructiveHintMTC-CAP-005

Tool "export_eval_harness" can mutate/egress but declares no destructiveHint. Clients that don't default to spec-safe behavior may not prompt before running it.

Fix: Declare accurate annotations, and gate destructive tools on user confirmation regardless.

Location: tool export_eval_harness

low Shell/command execution in packaging/dev tooling (scripts/release_check.py)MTC-SRC-002

In a packaging/dev/install script (shipped, but not the server runtime) (`scripts/release_check.py:27`): Spawning a shell/process is command-execution capability; with unsanitized tool input it is command injection / RCE. This is read from the code itself — not from the tool description — so a poisoned server cannot hide it behind honest-looking metadata.

Evidence: > {name}") result = subprocess.run(command, cwd=ROOT) if result.returncode != 0: raise SystemExit(result

Fix: Review this call path: confirm it never receives unsanitized tool input, constrain it, or remove it. Treat a server whose code reaches these sinks as high-capability regardless of what its tools claim.

Location: server scripts/release_check.py

Tools 92

Each tool and what it can reach — statically extracted from the published source.

  • browse_tool_usageingests untrusted input
  • export_eval_harnessruns code / shell
  • introspectingests untrusted input
  • run_elite_eval_suiteruns code / shell
  • validate_predictionsingests untrusted input
  • adopt_vs_buildno sensitive capability
  • after_action_reviewno sensitive capability
  • analyzeno sensitive capability
  • analyze_prompt_sequenceno sensitive capability
  • archive_goalno sensitive capability
Show 82 more tools ↓
  • assess_confidenceno sensitive capability
  • auditno sensitive capability
  • autonomous_scanno sensitive capability
  • bayesian_updateno sensitive capability
  • benchmark_trackno sensitive capability
  • bias_scanno sensitive capability
  • build_experiment_treeno sensitive capability
  • calculate_expected_valueno sensitive capability
  • calibration_predictno sensitive capability
  • calibration_resolveno sensitive capability
  • calibration_scoreno sensitive capability
  • check_anti_patternsno sensitive capability
  • check_goalsno sensitive capability
  • compound_growthno sensitive capability
  • decision_council_reviewno sensitive capability
  • delete_goalno sensitive capability
  • delete_prevention_ruleno sensitive capability
  • elite_doctorno sensitive capability
  • elite_doctor_jsonno sensitive capability
  • elite_outcome_scorecardno sensitive capability
  • five_whysno sensitive capability
  • fmea_analysisno sensitive capability
  • fmea_risk_gateno sensitive capability
  • generate_autonomous_goalsno sensitive capability
  • get_autonomous_statusno sensitive capability
  • get_elite_workflowno sensitive capability
  • get_prompt_quality_trendno sensitive capability
  • get_quality_trendno sensitive capability
  • get_tool_usage_statsno sensitive capability
  • get_user_profileno sensitive capability
  • get_user_thinking_modelno sensitive capability
  • ingest_contextno sensitive capability
  • learnno sensitive capability
  • list_prevention_rulesno sensitive capability
  • list_team_usersno sensitive capability
  • memory_context_packno sensitive capability
  • memory_search_contextno sensitive capability
  • memory_sync_decisionsno sensitive capability
  • memory_sync_mistakesno sensitive capability
  • memory_sync_rulesno sensitive capability
  • nuclear_prompt_breakdownno sensitive capability
  • orchestrate_request_toolno sensitive capability
  • planno sensitive capability
  • polish_promptno sensitive capability
  • pre_commit_auditno sensitive capability
  • predictno sensitive capability
  • predictive_preventionno sensitive capability
  • query_temporal_graphno sensitive capability
  • reasoning_preflightno sensitive capability
  • recommend_open_source_integrationsno sensitive capability
  • record_decisionno sensitive capability
  • record_hypothesisno sensitive capability
  • record_missed_detectionno sensitive capability
  • record_mistakeno sensitive capability
  • record_prompt_intentno sensitive capability
  • record_prospective_failureno sensitive capability
  • record_quality_scoreno sensitive capability
  • register_prevention_ruleno sensitive capability
  • rememberno sensitive capability
  • remember_contextno sensitive capability
  • research_benchmark_catalogno sensitive capability
  • resolve_hypothesisno sensitive capability
  • resolve_prospective_failureno sensitive capability
  • roi_tool_budgetno sensitive capability
  • search_decisionsno sensitive capability
  • search_thinking_patternsno sensitive capability
  • select_reasoning_protocolno sensitive capability
  • self_diagnoseno sensitive capability
  • set_goalno sensitive capability
  • share_skillno sensitive capability
  • simulate_future_regretsno sensitive capability
  • smoke_test_gateno sensitive capability
  • socratic_challengeno sensitive capability
  • swiss_cheese_auditno sensitive capability
  • sync_team_memoryno sensitive capability
  • update_goalno sensitive capability
  • update_thinking_patternno sensitive capability
  • update_user_configno sensitive capability
  • verify_capabilities_toolno sensitive capability
  • workflow_runno sensitive capability
  • workflow_statusno sensitive capability
  • workflow_update_stepno sensitive capability

Toxic flows 1

Cross-tool combinations that form a data-exfiltration primitive (untrusted input → sensitive source → external sink).

What this scan could not see

Versions 1

Scan history per published version. The engine is deterministic — the same version always yields the same score, so a changed score means the package itself changed.

VersionScoreFindingsEngineScanned
v1.2.1 latest A 93/100 8 1.9.0 2026-07-23

Embed this score

Show this server's live Trust Score in your README, docs or website. The badge is served straight from the registry and updates automatically after every rescan — no API key needed. It links back to this page, so anyone who sees the grade can also read the findings behind it instead of taking a number on faith.

MCP Trust Score: A · 93/100
Markdown (GitHub README)
[![MCP Trust Score](https://mcptrustchecker.com/registry/elite-reasoning-mcp/badge.svg)](https://mcptrustchecker.com/registry/elite-reasoning-mcp)
HTML
<a href="https://mcptrustchecker.com/registry/elite-reasoning-mcp"><img src="https://mcptrustchecker.com/registry/elite-reasoning-mcp/badge.svg" alt="MCP Trust Score" height="20"></a>
Prefer shields.io styling? Point it at https://mcptrustchecker.com/registry/elite-reasoning-mcp/badge.json via https://img.shields.io/endpoint?url=…

Verify this score yourself

The score above is reproducible: the same package version always yields the same result. Run it locally or over the free API — no account, no LLM, fully deterministic.

npx mcptrustchecker scan elite-reasoning-mcp --online --registry pypi

Use the free API → How scoring works

More in Developer Tools