Corpus#
Examples for corpus are ordered as a learning path rather
than by implementation detail.
# 💡 corpus Need additionals packages
curl -O https://raw.githubusercontent.com/scikit-plots/scikit-plots/main/requirements/corpus.txt
pip install -r requirements/corpus.txt
pip install scikit-plots[corpus]
# (Recommended)
# !pip install datasets transformers
# !pip install nltk gensim langdetect faster-whisper openai-whisper pytesseract youtube-transcript-api
# sudo apt-get install tesseract-ocr
See also
Start here#
Configure Corpus declaratively — learn
FluentCorpus, immutable plans, validation, branching, fingerprints, and thematerialize()boundary.Build and search a real Hamlet corpus — use
RuntimeCorpusend to end:run(),add(), storage, retrieval, export, and lifecycle.Compare chunking strategies — compare sentence, word, fixed-window, and morphological semantic chunking on the same OCR text.
Process an MP3 — learn audio provenance and companion-transcript precedence without requiring Whisper in the normal gallery path.
Process a mixed-media ZIP — inspect archive-member routing,
archive.zip/member.extprovenance, and per-extension reader settings.Process a YouTube transcript — execute a deterministic local proxy, configure the real YouTube reader, and keep the live transcript request explicit and optional.
Build a multi-source WHO corpus — see the explicit stage-by-stage integration path, partial source success, keyword retrieval, adapters, and where
CorpusBuilderfits.
Which API should I use?#
Goal |
Start with |
|---|---|
Process one source with direct stage control |
|
Build/search heterogeneous sources with partial-success reporting |
|
Create immutable, reusable, branchable configuration |
|
Execute a Fluent plan and manage runtime state/lifecycle |
|
Extend vector indexing/retrieval directly |
|
Capability matrix#
The normal gallery path prefers deterministic local execution. Optional capabilities are either preflighted and skipped when unavailable, or shown as configuration-only examples.
Example |
Normal path |
Optional capability |
Behavior when unavailable |
|---|---|---|---|
FluentCorpus basics |
local/core |
none |
not applicable |
Hamlet RuntimeCorpus |
local/core + NumPy |
native Annoy branch |
configuration only; not built |
OCR chunking comparison |
local image |
Tesseract, NLTK |
explicit |
MP3 ingestion |
MP3 + local SRT companion |
NLTK, Whisper |
optional sections |
Mixed-media ZIP |
local archive |
PDF/OCR/Whisper readers |
individual optional member capability may produce no documents; archive-security failures still fail |
YouTube transcript |
local synthetic proxy |
youtube-transcript-api + network, NLTK |
live/optional sections |
WHO multi-source integration |
local sidecars only |
PDF/OCR/Whisper |
each unavailable source reports |
Gallery reliability rule#
The examples distinguish optional capability absence from real defects:
missing optional package/resource/native capability/network opt-inReport a visible, specific
SKIPand continue when the example can remain truthful.invalid public API / security-policy failure / installed-backend defectFail visibly. The gallery must not convert a real regression into a skip.
A missing local sidecar never silently enables public-network access.
Install only what you need#
The core text/runtime examples use the normal Corpus installation. Media and
NLP examples may additionally use packages such as NLTK, an OCR backend,
Whisper, or youtube-transcript-api. System tools such as Tesseract may also
be required for the corresponding optional path.
Do not install every optional dependency merely to read the gallery. The portable path is designed to remain useful when those capabilities are absent.
Browser / WASM note#
Declarative configuration, local text processing, and portable brute-force retrieval are the strongest browser/WASM candidates. OCR, Whisper, native ANN backends, and live external services depend on the actual JupyterLite/xeus runtime and should not be assumed available until verified in that target environment.
Build and Search a Real Hamlet Corpus with FluentCorpus
Build and Search a Real Hamlet Corpus with FluentCorpus
Build and Search a Real Hamlet Corpus with FluentCorpus