2.3 KiB
corpus-mcp
corpus-mcp builds a portable SQLite full-text index from Markdown and text
files, then exposes ranked document excerpts through a local MCP server.
Installation
Install the command in an isolated environment with either tool:
pipx install .
uv tool install .
For development:
uv sync
uv run pytest
The Python runtime's SQLite library must include FTS5. The command reports a clear error at indexing time when FTS5 is unavailable.
Build A Corpus
Index all .txt and .md files below a directory, at any depth:
corpus-mcp index ./documents \
--database ./corpus.sqlite \
--description "Internal product documentation"
For a longer description, use --description-file description.txt. Source
files are read as UTF-8 (with or without a BOM) by default. Use --encoding
when a corpus uses a different encoding.
Indexing is an atomic full rebuild. If reading or indexing fails, an existing database remains unchanged. Directory and file symlinks are not followed.
Run The MCP Server
corpus-mcp serve /absolute/path/to/corpus.sqlite
The server uses stdio, so configure it as a subprocess in the MCP client. For example:
{
"mcpServers": {
"corpus": {
"command": "corpus-mcp",
"args": ["serve", "/absolute/path/to/corpus.sqlite"]
}
}
}
Restart a running MCP server after rebuilding its database.
MCP Tools
search_corpus(query, limit=10, excerpts_per_document=2, syntax="plain")
returns ranked, highlighted excerpts grouped by document. Plain searches match
all whitespace-separated terms without interpreting operators. Set syntax to
fts5 for phrases, OR, NOT, and prefix expressions. Each result includes
the containing chunk's character offsets, which can be passed to
read_document_excerpt to retrieve more surrounding context. Search ranks a
bounded pool of chunk matches so broad queries remain responsive; as a result,
a document with many highly ranked passages may occupy multiple candidates.
read_document_excerpt(path, offset=0, max_chars=4000) reads a bounded range
and returns previous and next offsets. Responses are capped at 20,000
characters so a large document cannot accidentally fill the model context.
corpus_info() returns the persisted dataset description and build metadata.