scikitplot.corpus#

scikitplot.corpus#

Tools for turning files, URLs, media, and text sources into canonical CorpusDocument evidence that can be transformed, embedded, stored, searched, adapted, and exported.

Choose the API that matches the job#

CorpusPipeline

Direct stage control for one source or an explicit batch.

CorpusBuilder

High-level build/search/adapt convenience for end-to-end corpus workflows.

FluentCorpus

Immutable, reusable configuration. Chained setter order describes what is configured; it does not define execution order.

RuntimeCorpus

The operational form of a validated Fluent plan. Materialization constructs runtime components; source processing starts only with run() or add().

RetrievalIndex / VectorIndexBackend

Lower-level retrieval and vector-backend extension APIs.

The common processing picture is:

Source
  -> Read
  -> Chunk / Normalize / Enrich
  -> Embed
  -> Store
  -> Index
  -> Retrieve
  -> Adapt / Export

Examples

Direct pipeline control:

>>> from pathlib import Path
>>> from scikitplot.corpus import CorpusPipeline, ParagraphChunker
>>> pipeline = CorpusPipeline(chunker=ParagraphChunker())
>>> result = pipeline.run(Path("article.txt"))
>>> print(f"{result.n_documents} chunks from {result.source}")

The dependency-free default sentence backend is REGEX. Passing a spaCy model name explicitly selects the spaCy shorthand instead:

>>> from scikitplot.corpus import SentenceChunker
>>> portable = SentenceChunker()
>>> spacy_chunker = SentenceChunker("en_core_web_sm")

High-level builder:

>>> from scikitplot.corpus import CorpusBuilder, BuilderConfig
>>> builder = CorpusBuilder(
...     BuilderConfig(
...         chunker="paragraph",
...         normalize=True,
...         enrich=True,
...         build_index=True,
...     )
... )
>>> result = builder.build("./data/")
>>> results = builder.search("quantum computing")

Reusable declarative configuration:

>>> from scikitplot.corpus import FluentCorpus
>>> fluent = FluentCorpus().chunker("paragraph").storage("memory")
>>> fluent.validate()
[]

Materialization is explicit and performs no source read by itself:

>>> from scikitplot.corpus import RuntimePolicy
>>> with fluent.materialize(policy=RuntimePolicy(allow_network=False)) as runtime:
...     len(runtime.documents)
0

A configured source becomes operational only when run() is called:

>>> runtime_fluent = (
...     FluentCorpus().source("article.txt").chunker("paragraph").storage("memory")
... )
>>> with runtime_fluent.materialize() as runtime:
...     result = runtime.run()

Network and optional-capability examples:

Live URLs, OCR, ASR, spaCy, NLTK resource-backed NLP, model embeddings, and native vector backends depend on the corresponding environment capability. User-facing examples should not fabricate results when a capability is absent. Documentation/gallery examples should either use a portable executed path or report a clear skip while keeping the optional configuration visible.

URL ingestion:

>>> result = pipeline.run_url("https://en.wikipedia.org/wiki/Python")

YouTube transcript:

>>> result = pipeline.run(
...     "https://www.youtube.com/watch?v=rwPISgZcYIk"
... )

Image OCR:

>>> from scikitplot.corpus import DocumentReader
>>> reader = DocumentReader.create(Path("scan.png"))
>>> docs = list(reader.get_documents())

With model embeddings:

>>> from scikitplot.corpus import EmbeddingEngine
>>> engine = EmbeddingEngine(backend="sentence_transformers")

Dependency-free local helpers:

>>> from scikitplot.corpus import HAMLET_TEXT, HashEmbedder, SimpleEnricherSpec
>>> HashEmbedder(dimension=32)([HAMLET_TEXT[:120]]).shape
(1, 32)
>>> FluentCorpus().enricher(SimpleEnricherSpec()).validate()
[]

HashEmbedder is a deterministic lexical hashing baseline, not a learned semantic model. HAMLET_TEXT is bundled convenience sample data rather than an authoritative scholarly edition.

See scikitplot/corpus/README.md for the user-oriented API map, runtime policy boundary, retrieval modes, and optional-capability guidance.

User guide. See the Corpus User Guide section for further details.

Adapter layer#

Class inheritance

Inheritance diagram of LangChainCorpusRetriever, MCPCorpusServer

to_langchain_documents

Convert CorpusDocument instances to LangChain Document.

to_langgraph_state

Convert documents to a LangGraph-compatible state dict.

to_mcp_resources

Convert documents to MCP resources/read response format.

to_mcp_tool_result

Format documents as an MCP tools/call response.

to_huggingface_dataset

Convert documents to a HuggingFace Dataset.

to_rag_tuples

Convert documents to (text, metadata, embedding) tuples.

to_jsonl

Yield documents as newline-delimited JSON strings.

to_numpy_arrays

Convert documents to a dict of NumPy arrays suitable for batch ML.

to_tensorflow_dataset

Convert documents to a tf.data.Dataset.

to_torch_dataloader

Convert documents to a torch.utils.data.DataLoader.

LangChainCorpusRetriever

LangChain-compatible retriever backed by RetrievalIndex.

MCPCorpusServer

MCP server adapter for corpus search.

Archive-within-archive#

extract_archive

Extract an archive to a destination directory.

is_archive

Check if a file path has a supported archive extension.

Base Classes#

Class inheritance

Inheritance diagram of DocumentReader, DummyReader

ChunkerBase

Abstract base class for all text chunkers.

DefaultFilter

Standard noise filter ported and improved from remarx's include_sentence.

DocumentReader

Abstract base class for all format-specific document readers.

DummyReader

A no-op reader that validates source existence and accessibility.

FilterBase

Abstract base class for corpus document filters.

PipelineGuard

Wrap any document stream with resilience, deduplication, and checkpointing.

Chunkers#

ChunkerBridge

Adapter that wraps a new-style chunker as a ChunkerBase- compatible object.

FixedWindowChunkerBridge

Bridge for FixedWindowChunker → ChunkerBase contract.

ParagraphChunkerBridge

Bridge for ParagraphChunker → ChunkerBase contract.

SentenceChunkerBridge

Bridge for SentenceChunker → ChunkerBase contract.

WordChunkerBridge

Bridge for WordChunker → ChunkerBase contract.

bridge_chunker

Wrap chunker in a bridge if it is a new-style chunker.

register_bridge

Register a custom bridge for a user-defined chunker class.

unregister_bridge

Remove a previously registered bridge for chunker_class.

TokenizerProtocol

Structural protocol for word tokenizers.

SentenceSplitterProtocol

Structural protocol for sentence segmenters.

StemmerProtocol

Structural protocol for word stemmers.

LemmatizerProtocol

Structural protocol for word lemmatizers.

FunctionTokenizer

Wrap any Callable[[str], list[str]] as a TokenizerProtocol.

FunctionSentenceSplitter

Wrap any Callable[[str], list[str]] as a SentenceSplitterProtocol.

FunctionStemmer

Wrap any Callable[[str], str] as a StemmerProtocol.

FunctionLemmatizer

Wrap any Callable[[str, Optional[str]], str] as a LemmatizerProtocol.

CustomTokenizerRegistry

Thread-safe module-level registry for named custom components.

register_tokenizer

Register a named TokenizerProtocol implementation.

get_tokenizer

Retrieve a registered tokenizer by name.

register_sentence_splitter

Register a named SentenceSplitterProtocol implementation.

get_sentence_splitter

Retrieve a registered sentence splitter by name.

register_stemmer

Register a named StemmerProtocol implementation.

get_stemmer

Retrieve a registered stemmer by name.

register_lemmatizer

Register a named LemmatizerProtocol implementation.

get_lemmatizer

Retrieve a registered lemmatizer by name.

ScriptType

Dominant Unicode script detected in a text sample.

detect_script

Detect the dominant Unicode script in text.

is_cjk_char

Return True if ch is a CJK / Japanese / Korean character.

is_rtl_char

Return True if ch belongs to a right-to-left script.

split_cjk_chars

Split text into individual CJK character tokens.

MULTI_SCRIPT_SENTENCE_RE_PATTERN

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

FixedWindowChunker

Produce fixed-size sliding-window chunks over a document.

FixedWindowChunkerConfig

Configuration for FixedWindowChunker.

WindowUnit

Unit of measurement for window size and step.

ISO_TO_NLTK

ISO_TO_NAME

NLTK_TO_ISO

NLTK_STOPWORD_LANGUAGES

frozenset() -> empty frozenset object frozenset(iterable) -> frozenset object

BUILTIN_LANG_STOPWORDS

coerce_language

Normalise any language specifier into a list of canonical NLTK names.

resolve_stopwords

Return a frozenset of stopwords for one or more languages.

iso_to_nltk

Resolve an ISO 639-1/639-3 code to a canonical NLTK language name.

nltk_to_iso

Resolve a canonical NLTK language name to its primary ISO 639-1 code.

ParagraphChunker

Split a document into paragraph-level Chunk objects.

ParagraphChunkerConfig

Configuration for ParagraphChunker.

SentenceBackend

Supported sentence-splitting backends.

SentenceChunker

Split a document into sentence-level Chunk objects.

SentenceChunkerConfig

Configuration for SentenceChunker.

LemmatizationBackend

Lemmatization backend.

StemmingBackend

Stemming algorithm.

StopwordSource

Stopword list source.

TokenizerBackend

Word tokenisation backend.

WordChunker

Process a document at word level, producing normalised token chunks.

WordChunkerConfig

Configuration for WordChunker.

Corpus Builder#

Class inheritance

Inheritance diagram of CorpusBuilder

BuildResult

Result of a corpus build operation.

BuilderConfig

Configuration for CorpusBuilder.

CorpusBuilder

Unified corpus builder — end-to-end pipeline orchestrator.

Custom Hooks#

BuilderFactories

Component factory callables for FactoryCorpusBuilder.

CustomChunker

Wrap any callable as a ChunkerBase.

CustomEnricherConfig

Custom backend callables for CustomNLPEnricher.

CustomFilter

Wrap any callable as a FilterBase.

CustomNLPEnricher

NLPEnricher extended with fully-replaceable NLP backends.

CustomNormalizer

Wrap any callable as a NormalizerBase.

CustomRetrievalIndex

RetrievalIndex extended with a fully-replaceable custom scorer callable.

FactoryCorpusBuilder

CorpusBuilder extended with pluggable component factories.

HookableCorpusPipeline

CorpusPipeline extended with per-stage lifecycle hooks.

PipelineHooks

Lifecycle callbacks for HookableCorpusPipeline.

Downloader#

Class inheritance

Inheritance diagram of BaseDownloader, AnyDownloader

AnyDownloader

Auto-dispatching downloader with multi-URL and per-parameter list support.

BaseDownloader

Abstract base class for all format-specific URL downloaders.

CustomDownloader

Wraps a user-supplied callable as a BaseDownloader.

DownloadResult

Immutable result object returned by every BaseDownloader.

GitHubDownloader

GitHub URL downloader with automatic blob → raw normalisation.

GoogleDriveDownloader

Google Drive share-link downloader.

WebDownloader

Generic HTTP/HTTPS file downloader.

YouTubeDownloader

YouTube content downloader.

Embeddings#

DEFAULT_CACHE_DIR

Path subclass for non-Windows systems.

DEFAULT_MODEL

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

EmbeddingEngine

Multi-backend sentence embedding engine with SHA-256 file caching.

DEFAULT_AUDIO_MODEL

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

DEFAULT_IMAGE_MODEL

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

DEFAULT_TEXT_MODEL

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

LLMTrainingExporter

Export a corpus with embeddings to LLM training formats.

MultimodalEmbeddingEngine

Unified embedding engine for any CorpusDocument modality — text, image, audio, video, or multimodal.

Enricher#

BUILTIN_STOPWORDS

frozenset() -> empty frozenset object frozenset(iterable) -> frozenset object

EnricherConfig

Configuration for NLPEnricher.

NLPEnricher

Pipeline component that populates NLP enrichment fields on CorpusDocument.

Export#

export_documents

Export a list of documents to output_path in the given format.

load_documents

Load CorpusDocument instances from a previously exported file.

Metadata#

CollectionManifest

Descriptor for a named corpus collection.

CorpusStats

Aggregate statistics over a CorpusDocument collection.

compute_stats

Compute aggregate statistics over a document collection.

provenance_from_filename

Extract provenance metadata from a source filename using heuristics.

Normalizers#

DedupLinesNormalizer

Remove exact duplicate lines while preserving first-occurrence order.

HTMLStripNormalizer

Remove HTML and XML tags from the document text.

LanguageDetectionNormalizer

Detect document language and set CorpusDocument.language.

LowercaseNormalizer

Convert the document text to lowercase.

NormalizationPipeline

Apply a sequence of normalisers in order.

NormalizerBase

Abstract base class for all text normalisers.

UnicodeNormalizer

Apply Unicode normalisation (NFC, NFD, NFKC, or NFKD).

WhitespaceNormalizer

Collapse runs of whitespace and optionally strip leading/trailing space.

TextNormalizer

Pipeline component that populates normalized_text on CorpusDocument instances.

TextNormalizerConfig

Configuration for TextNormalizer.

normalize_text

Normalise text according to config.

Pipeline#

Class inheritance

Inheritance diagram of CorpusPipeline, PipelineResult

CorpusPipeline

Orchestrates the full corpus ingestion pipeline.

PipelineResult

Immutable summary of a single pipeline run.

create_corpus

Create and export a corpus from a single source file.

Readers#

Class inheritance

Inheritance diagram of MarkdownReader, CustomReader

ALTOReader

ALTO XML reader for scanned document archives.

AudioReader

Text extraction from audio files via companion transcript/lyrics parsing, Whisper ASR, and optional audio classification.

CustomReader

Fully user-customizable reader for any file extension and resource type.

normalize_extractor_output

Coerce an extractor return value to a list of raw chunk dicts.

ImageReader

OCR-based text extraction from raster image files.

PDFReader

PDF document reader with pdfminer.six → pypdf cascade.

MarkdownReader

Markdown document reader.

ReSTReader

reStructuredText document reader.

TextReader

Plain-text document reader.

VideoReader

Text extraction from video files via subtitle parsing and/or automatic speech recognition.

WebReader

Fetch a web page and extract structured text via BeautifulSoup.

YouTubeReader

Extract the transcript of a YouTube video using youtube-transcript-api.

TEIReader

TEI/XML document reader with dramatic structure extraction.

XMLReader

Generic XML document reader with configurable XPath.

ZipReader

Generic ZIP archive reader — dispatches each member to its natural reader.

Registry#

ComponentRegistry

Central look-up table for corpus pipeline components.

registry

Central look-up table for corpus pipeline components.

Similarity#

RetrievalConfig

Configuration for similarity search.

RetrievalHit

A single search result.

RetrievalIndex

Multi-mode similarity index over CorpusDocument collections.

Schema#

SectionType

Semantic label for the role of a text chunk within its source document.

ChunkingStrategy

Describes how a CorpusDocument was segmented from raw text.

ExportFormat

Supported serialisation targets for a completed corpus.

SourceType

Semantic label for the kind of source from which a document was read.

MatchMode

Search mode for intertextual matching queries against a corpus index.

Modality

Primary content modality of a CorpusDocument.

ErrorPolicy

Per-document error handling behaviour for PipelineGuard.

CorpusDocument

Canonical representation of a single text chunk in a processed corpus.

_PROMOTED_RAW_KEYS

frozenset() -> empty frozenset object frozenset(iterable) -> frozenset object

documents_to_pandas

Convert a list of CorpusDocument instances to a pandas.DataFrame.

documents_to_polars

Convert a list of CorpusDocument instances to a polars.DataFrame.

Source#

CorpusSource

Declarative descriptor for one or more document sources.

SourceEntry

A single resolved source entry yielded by CorpusSource.iter_entries.

SourceKind

Discriminant for the kind of source an entry represents.

Storage#

Class inheritance

Inheritance diagram of InMemoryStorage, SQLiteStorage

InMemoryStorage

Thread-safe in-memory dict store.

JSONLStorage

Append-friendly JSONL (newline-delimited JSON) flat-file store.

QueryResult

Result container returned by StorageBase.query.

SQLiteStorage

SQLite-backed corpus store with FTS5 full-text search.

StorageBase

Abstract base class for all corpus storage backends.

StorageQuery

Query parameters for StorageBase.query.

URL#

URLKind

Classification of a URL for routing to the correct handler.

classify_url

Classify a URL into one of the known URLKind categories.

download_url

Download a URL to a local file.

infer_extension

Infer a file extension from HTTP response headers and URL path.

probe_url_kind

Probe a URL with a HEAD request to classify by Content-Type.

resolve_url

Resolve a provider-specific URL to a direct-download URL.