Model Context Protocol (MCP)#
Use scikitplot.mcp to make documentation or a custom corpus searchable
from an MCP-compatible assistant, developer tool, local process, container, or
service.
The normal user workflow is intentionally small:
choose the sources that should count as reference evidence;
use simple lexical retrieval, or add corpus-backed semantic/hybrid retrieval;
verify the selected corpus and retriever with known queries;
choose how the MCP server should run; and
let the client call the read-only
search_docstool and inspect citations.
You do not need to understand the retrieval internals to get started. Retrieval finds and ranks evidence; it does not make the underlying data true.
Scientific grounding: evidence, not absolute truth#
scikitplot.mcp is designed for source-grounded answers: retrieve relevant
passages, preserve where they came from, and let the client reason from visible
evidence instead of unsupported recall.
A grounded answer is traceable, not automatically correct. Source documents can be incomplete, stale, biased, contradictory, or simply wrong. Retrieval can also miss relevant material.
Important
Treat retrieved data as evidence for or against a hypothesis, not as absolute truth. A retrieval score measures relevance for ranking; it is not a probability that a statement is correct.
A useful scientific response should therefore prefer language such as:
“the retrieved sources support …”;
“the available evidence is consistent with …”;
“the sources disagree …”;
“this remains uncertain …”; or
“no supporting passage was retrieved.”
Prefer evidence that is current, attributable, reproducible, and independently checkable. When sources conflict, preserve the disagreement instead of forcing a single confident answer.
How the pieces fit#
The three modules have different jobs. You can use scikitplot.mcp by itself
for a small corpus, then add the other layers only when your retrieval needs grow.
Component |
Main job |
What it does not mean |
|---|---|---|
Read, chunk, normalize, enrich, embed, index, and preserve source metadata for custom corpora. |
Corpus content is not automatically verified or true. |
|
Provide fast local approximate-nearest-neighbor search over vectors for semantic retrieval. |
Vector similarity is not factual correctness or scientific confidence. |
|
Expose a bounded, read-only retrieval contract with citations to MCP-compatible clients. |
MCP does not retrain the model or validate every claim in the corpus. |
Common paths are intentionally composable:
Small / simple
source JSONL -> lexical retrieval -> scikitplot.mcp -> MCP client
Semantic
sources -> scikitplot.corpus -> embeddings -> scikitplot.annoy
-> CorpusAnnoyRetriever -> scikitplot.mcp -> MCP client
Hybrid
sources -> scikitplot.corpus -> lexical + semantic retrieval
-> rank fusion -> scikitplot.mcp -> MCP client
In every path, the final responsibility is the same: retrieve useful evidence, keep provenance attached, and communicate uncertainty honestly.
At a glance#
Main tool:
search_docsDefault transport: local
stdioHTTP transport: Streamable HTTP
Default HTTP endpoint:
http://127.0.0.1:8000/mcpDefault health endpoint:
http://127.0.0.1:8000/healthzSmall custom corpus: optional UTF-8 JSONL via
--docs-jsonlCorpus pipeline:
scikitplot.corpusfor richer ingestion, chunking, embeddings, storage, and provenanceSemantic retrieval:
CorpusAnnoyRetrievercan composescikitplot.corpuswithscikitplot.annoyHybrid retrieval:
HybridRetrievercan combine lexical and semantic ranked resultsDefault HTTP behavior: stateless
Evidence model: retrieval scores rank relevance; they do not certify truth
Safety model: returned document text is explicitly untrusted reference data
Choose your scenario#
I want to… |
Start with… |
Best fit |
|---|---|---|
See whether MCP works |
|
No client and no network required. |
Connect a local assistant |
default |
The MCP client starts the Python process. |
Search a small set of my own docs |
|
Simple local or CI use. |
Connect through localhost HTTP |
|
Local services and HTTP-capable clients. |
Run in Docker |
|
Container-friendly HTTP defaults. |
Add a CI/readiness check |
|
Detect missing, stale, or wrong documentation indexes. |
Serve a team or remote client |
HTTP behind production controls |
Authentication, TLS, network policy, and quotas are required. |
Build a curated multi-format corpus |
Ingest, chunk, normalize, embed, index, and retain provenance before serving. |
|
Search concepts and paraphrases |
|
Semantic retrieval using corpus embeddings and |
Balance exact terms and semantic meaning |
|
Fuse lexical and dense rankings while keeping one MCP tool contract. |
Search a larger or richer corpus |
|
BM25, vector, or hybrid retrieval without changing the MCP tool contract. |
Quick start#
Install the optional MCP server dependency in the environment containing scikit-plots:
python -m pip install "mcp>=2,<3"
First verify the backend without opening a port:
python -m scikitplot.mcp --self-test --self-test-query transport
Then start the default local server:
python -m scikitplot.mcp
Note
The default command uses stdio. It normally waits for an MCP client.
It does not open a browser page, and waiting for the client is expected.
Scenario 1: connect a local assistant#
Use stdio when your MCP client can start a local command for its tools.
This is the simplest mode and does not expose a network listener.
Use these values in the client’s MCP server configuration:
command: python
arguments: -m scikitplot.mcp
The client starts the process and communicates with it over standard input and output.
Tip
Run --self-test from the same Python environment first. This catches an
import, dependency, or documentation-loading problem before the MCP client is
involved.
Scenario 2: test before connecting anything#
The self-test loads the selected retrieval backend, performs one read-only search, validates the output contract, prints JSON, and exits.
python -m scikitplot.mcp \
--self-test \
--self-test-query transport
Require the query to return at least one result:
python -m scikitplot.mcp \
--self-test \
--self-test-query transport \
--self-test-require-match
This is useful for local troubleshooting, Docker image validation, CI, and readiness gates.
Scenario 3: search your own documentation#
For a small local corpus, create a UTF-8 JSON Lines file with one document per
line. doc_id and text are required. title, source_uri, and
anchor are optional but recommended because they improve citations.
{"doc_id":"install","title":"Installation","source_uri":"/install.html","anchor":"install","text":"Install scikit-plots with pip or conda."}
{"doc_id":"plots","title":"Plotting","source_uri":"/plots.html","anchor":"examples","text":"Use plotting helpers to visualize model results."}
Validate the corpus before serving it:
python -m scikitplot.mcp \
--docs-jsonl docs.jsonl \
--self-test \
--self-test-query plotting \
--self-test-require-match
Then start the normal local stdio server with the same corpus:
python -m scikitplot.mcp --docs-jsonl docs.jsonl
Good document IDs are stable and simple, for example:
getting-started, api.metrics.roc, or guide:deployment.
Important
JSONL mode is intentionally bounded and best suited to small and moderate
local corpora. For larger documentation collections, use a production
retrieval backend through scikitplot.mcp rather than continually
growing one in-memory file.
Scenario 4: build a curated corpus#
Use scikitplot.corpus when your evidence comes from more than a small
hand-written JSONL file or when you need a repeatable ingestion pipeline.
A corpus can read and transform source material, preserve provenance, create embeddings, and build a similarity index before MCP is involved:
from scikitplot.corpus import BuilderConfig, CorpusBuilder
builder = CorpusBuilder(
BuilderConfig(
chunker="paragraph",
normalize=True,
enrich=True,
embed=True,
build_index=True,
)
)
result = builder.build("./data/")
hits = builder.search("what evidence supports this claim?")
mcp_result = builder.to_mcp_tool_result(
"what evidence supports this claim?"
)
This is useful because corpus quality can be tested independently from transport or the AI client. Inspect representative chunks, source metadata, and search results before serving them.
Note
Curating a corpus improves control and traceability; it does not convert the corpus into ground truth. Keep source dates, authorship, versions, and other provenance needed to judge the evidence later.
Scenario 5: add semantic retrieval with Annoy#
Use semantic retrieval when users may ask for the same concept with different
words. CorpusAnnoyRetriever composes scikitplot.corpus embeddings with
an approximate vector index provided by scikitplot.annoy.
from scikitplot.mcp import CorpusAnnoyRetriever
retriever = CorpusAnnoyRetriever.from_corpus_annoy(
"./docs/",
metric="angular",
n_trees=10,
)
hits = retriever.search(
"connect a local assistant without opening a network port",
k=5,
)
The same embedding model is used for corpus and query vectors so they share one vector space. Annoy then finds nearby vectors efficiently.
Important
Annoy performs approximate nearest-neighbor search. Its parameters trade retrieval recall, index size, build cost, and query speed. A closer vector is a stronger semantic match, not stronger evidence that the passage is true.
For scientific or safety-sensitive use, evaluate retrieval on a held-out set of known questions and relevant passages instead of choosing tuning parameters by feel alone.
Scenario 6: combine lexical and semantic evidence#
Exact terms and semantic meaning fail in different ways. API names, flags, identifiers, and error strings often favor lexical/BM25 search, while paraphrases and synonyms often favor dense semantic search.
HybridRetriever can fuse already-configured retrieval legs with Reciprocal
Rank Fusion (RRF):
from scikitplot.mcp import HybridRetriever
retriever = HybridRetriever(
[lexical_retriever, semantic_retriever],
weights=[1.0, 1.0],
)
hits = retriever.search("why is the HTTP service reachable but search stale?")
RRF combines rank positions rather than pretending BM25 and cosine scores are the same quantity. A result supported by more than one retrieval leg may rank higher, but fusion still measures retrieval usefulness, not factual certainty.
Use hybrid retrieval when your corpus contains both precise technical language and natural-language explanations.
Scenario 7: make CI detect the wrong corpus#
A health endpoint can tell you that a process is alive. It cannot prove that the expected documentation was loaded.
Add a unique canary document to the corpus, then require that exact document in a self-test:
python -m scikitplot.mcp \
--docs-jsonl docs.jsonl \
--self-test \
--self-test-query MCP_CANARY_7F3A91C2 \
--self-test-require-match \
--self-test-expected-doc-id scikitplot-canary-001
This helps detect:
a stale image;
the wrong mounted corpus;
an incomplete build;
a broken index; or
a deployment that is healthy at HTTP level but functionally wrong.
For an immutable corpus and deterministic backend, repeated self-tests with the same input should also produce stable results.
Scenario 8: run a local HTTP endpoint#
Use Streamable HTTP when a local service or client connects over HTTP:
python -m scikitplot.mcp --transport streamable-http
Default local endpoints:
MCP:
http://127.0.0.1:8000/mcpHealth:
http://127.0.0.1:8000/healthz
Probe the running health endpoint:
python -m scikitplot.mcp \
--transport streamable-http \
--probe
The health check confirms that the HTTP service is reachable and returns the
expected minimal health response. Use --self-test when you also need to
check corpus loading and search behavior.
Scenario 9: run in Docker#
Use the explicit Docker profile:
python -m scikitplot.mcp --docker
It selects container-friendly defaults, including Streamable HTTP, a container network bind, and the health route.
For local development, publish the container port only on the host loopback interface when possible:
docker run --rm \
-p 127.0.0.1:8000:8000 \
scikitplot-mcp:latest
Then use:
MCP: http://127.0.0.1:8000/mcp
Health: http://127.0.0.1:8000/healthz
Important
--docker changes deployment defaults. It does not add identity,
authentication, authorization, TLS, firewall rules, or per-user quotas.
For production containers, install a built wheel in a clean runtime image
instead of using an editable -e installation. This avoids build-system work
on import and gives more deterministic startup and rollback behavior.
Scenario 10: serve remote or team clients safely#
Treat remote MCP deployment like any other network service.
A safer production shape is:
MCP client
|
v
authenticated TLS proxy / gateway
|
v
private scikitplot.mcp HTTP service
|
v
read-only documentation retriever
Use deployment controls appropriate to your environment, including:
authenticated client identity;
authorization for the intended users or service accounts;
TLS at the trusted network boundary;
firewall or private-network restrictions;
request and concurrency limits;
logging and monitoring without leaking secrets; and
an explicit update and rollback process for the corpus and runtime image.
The CLI refuses an ordinary unauthenticated non-local HTTP bind unless you make
that choice explicit. Do not use --allow-unauthenticated-remote as a
shortcut for a production security layer.
How a documentation search works#
A client discovers search_docs and calls it with:
queryThe documentation question or search text.
kThe maximum number of passages to return. The default tool value is
5.
The server then:
validates the tool arguments;
bounds the request;
asks the configured read-only retriever for matching documentation;
cleans and bounds returned text;
validates citation links;
returns passages and machine-readable citations; and
marks retrieved material as untrusted reference data.
A normal result therefore carries both useful text and provenance instead of returning an uncited block of generated prose.
When nothing matches, the result is a normal empty search result with a clear message rather than a fabricated documentation passage.
Read retrieval results scientifically#
Do not collapse retrieval quality, source quality, and factual correctness into one number. They answer different questions.
Signal |
Useful interpretation |
Do not interpret it as |
|---|---|---|
Retrieval score / rank |
“This passage appears relevant to the query.” |
Probability that the passage or final answer is true. |
Citation / |
“This is where the retrieved claim came from.” |
Proof that the source is authoritative or current. |
Agreement across sources |
Corroborating evidence worth investigating. |
Automatic independence or consensus. |
Missing result |
The retriever did not find supporting material in the searched corpus. |
Proof that the claim is false. |
Conflicting passages |
Evidence that uncertainty, version drift, or genuine disagreement exists. |
A reason to silently choose the highest-ranked passage. |
A practical evidence loop is:
question
-> retrieve candidate evidence
-> inspect provenance and dates
-> compare independent sources
-> note contradictions / missing evidence
-> state the best-supported hypothesis
-> cite what supports it
-> keep uncertainty visible
This makes the system data-driven without becoming data-obedient: data gets a voice, but evidence is still tested, compared, and interpreted.
Security by design#
The user-facing tool is intentionally narrow: search documentation and return references. The server does not need write access to perform this workflow.
Protection |
What it means for you |
|---|---|
Read-only, idempotent search tool |
Repeating a documentation search should not modify your project or corpus. |
Untrusted-content label |
Retrieved pages are context, not commands for the model to obey. |
Bounded query and result sizes |
One request cannot ask the tool to emit an unbounded amount of documentation. |
Bounded passage text |
Large source chunks are truncated before they enter the MCP result. |
Closed tool arguments |
Unknown or misspelled |
Citation URL validation |
Unsafe URL forms such as script/data schemes and credential-bearing links are not emitted as trusted citations. |
Bounded concurrency |
The search service limits simultaneous retrieval work and can report that it is busy. |
HTTP request-size limit |
Streamable HTTP accepts bounded request bodies rather than unlimited input. |
Remote-bind guard |
Accidental unauthenticated non-local HTTP exposure is rejected by default. |
Minimal health response |
The public health route does not need to reveal paths, indexes, environment values, or secrets. |
Important
These controls reduce risk but do not make arbitrary retrieved text trustworthy. A malicious or compromised documentation page can still contain semantic prompt-injection text. The client/model must continue treating retrieved passages only as reference material.
Reliable operation#
Use three different checks for three different questions:
Check |
Answers |
Use it for |
|---|---|---|
|
“What configuration will actually run?” |
CLI/environment debugging before startup. |
|
“Can this corpus load and answer a valid search?” |
Builds, CI, deployments, and functional readiness. |
|
“Is the running HTTP service alive?” |
Runtime/container health checks. |
Inspect configuration without starting the server:
python -m scikitplot.mcp --print-effective-config
For Docker-resolved settings:
python -m scikitplot.mcp --docker --print-effective-config
This is especially useful when CLI options and SCIKITPLOT_MCP_* environment
variables are mixed.
Useful runtime controls#
Most users can keep the defaults. These controls are available when deployment requirements change:
Option |
Default |
Purpose |
|---|---|---|
|
|
HTTP bind address. |
|
|
HTTP port. |
|
|
MCP HTTP endpoint. |
|
|
Lightweight HTTP health route. |
|
|
Bound simultaneous searches. |
|
|
Bound HTTP request-body size. |
|
off |
Keep HTTP sessions only when the client/deployment requires them. |
|
|
Control operational logging detail. |
For automated deployments, the same settings can be supplied with
SCIKITPLOT_MCP_* environment variables.
Common problems#
The command appears to hang#
If you ran:
python -m scikitplot.mcp
this is usually normal. stdio mode is waiting for an MCP client. Run
--self-test if you only want to check that the server works.
The health check passes but search is wrong#
/healthz is intentionally a lightweight process/HTTP check. Run a
--self-test against the real corpus, preferably with a canary and
--self-test-expected-doc-id.
My JSONL file is rejected#
Check that:
the file is UTF-8 JSON Lines, not one large JSON array;
every non-empty line is a JSON object;
every document has a non-empty
doc_idandtext;document IDs are unique; and
the corpus stays within the bounded JSONL limits.
No documentation matches#
An empty result is not automatically a server failure. Try a shorter or more
specific term, verify the intended corpus with --self-test, and confirm the
expected document is actually present.
The top result is relevant but the claim is still wrong#
This is possible and should be expected in real corpora. Retrieval ranking asks “what best matches the query?” rather than “what is certainly true?”
Check the citation, source date, document version, and surrounding context. Search for independent supporting or contradicting passages, try a lexical/semantic hybrid query, and update or quarantine stale corpus material when necessary.
A non-local HTTP bind is refused#
This is a safety guard. Prefer localhost, an isolated container/private network, or a properly authenticated production gateway. Make remote exposure an intentional deployment decision rather than disabling the guard by default.
A production container rebuilds on import#
Do not ship a Meson editable installation in the runtime image. Build a wheel in a builder stage and install that wheel into a clean runtime stage.
Designed to grow without changing the basic workflow#
The simple user contract stays the same even when the retrieval system becomes more capable:
user question
-> search_docs
-> DocsRetriever
-> cited, bounded, untrusted passages
The retriever behind that contract can evolve independently.
- Small local corpus
InMemoryBm25Retrieverprovides a deterministic, dependency-light path for examples, tests, and bounded JSONL documentation.- Keyword-focused production search
Bm25Retrievercan adapt a full-text/BM25 backend for exact API names, flags, identifiers, and error messages.- Semantic search
CorpusAnnoyRetrievercan connect corpus embeddings and an ANN index for concept and paraphrase matching.- Hybrid search
HybridRetrievercan combine multiple retrieval legs with Reciprocal Rank Fusion. A non-strict hybrid configuration can continue using healthy legs if one retrieval backend fails.
This separation is useful for future changes: transport, indexing strategy, embedding model, storage engine, or ranking can evolve without requiring every MCP client to learn a new documentation tool.
Measure each retrieval backend on representative queries before promoting it.
For semantic or hybrid search, track retrieval metrics such as recall at k
or manually reviewed relevance alongside latency and memory. Do not optimize a
single score as if it measured truth.
Keep the user guide simple#
For normal use, remember this sequence:
1. Choose and curate the evidence sources
2. Self-test the corpus and retrieval behavior
3. Choose stdio or HTTP
4. Connect the client
5. Search and inspect citations
6. State conclusions in proportion to the evidence
Everything after that is an optimization or deployment concern.
Where to go next#
Most users can stop here.
For corpus construction and provenance-aware processing, see
scikitplot.corpus. For local approximate vector search and its
speed/recall trade-offs, see scikitplot.annoy.
For custom retrieval or programmatic integrations, see scikitplot.mcp.
The public API includes DocsRetriever, RetrievedChunk,
InMemoryBm25Retriever, Bm25Retriever, CorpusAnnoyRetriever,
HybridRetriever, and create_server.
For CLI options available in your installed version, use:
python -m scikitplot.mcp --help