Files
corpus-mcp/README.md
T
2026-07-13 18:49:24 +02:00

2.3 KiB

corpus-mcp

corpus-mcp builds a portable SQLite full-text index from Markdown and text files, then exposes ranked document excerpts through a local MCP server.

Installation

Install the command in an isolated environment with either tool:

pipx install .
uv tool install .

For development:

uv sync
uv run pytest

The Python runtime's SQLite library must include FTS5. The command reports a clear error at indexing time when FTS5 is unavailable.

Build A Corpus

Index all .txt and .md files below a directory, at any depth:

corpus-mcp index ./documents \
  --database ./corpus.sqlite \
  --description "Internal product documentation"

For a longer description, use --description-file description.txt. Source files are read as UTF-8 (with or without a BOM) by default. Use --encoding when a corpus uses a different encoding.

Indexing is an atomic full rebuild. If reading or indexing fails, an existing database remains unchanged. Directory and file symlinks are not followed.

Run The MCP Server

corpus-mcp serve /absolute/path/to/corpus.sqlite

The server uses stdio, so configure it as a subprocess in the MCP client. For example:

{
  "mcpServers": {
    "corpus": {
      "command": "corpus-mcp",
      "args": ["serve", "/absolute/path/to/corpus.sqlite"]
    }
  }
}

Restart a running MCP server after rebuilding its database.

MCP Tools

search_corpus(query, limit=10, excerpts_per_document=2, syntax="plain") returns ranked, highlighted excerpts grouped by document. Plain searches match all whitespace-separated terms without interpreting operators. Set syntax to fts5 for phrases, OR, NOT, and prefix expressions. Each result includes the containing chunk's character offsets, which can be passed to read_document_excerpt to retrieve more surrounding context. Search ranks a bounded pool of chunk matches so broad queries remain responsive; as a result, a document with many highly ranked passages may occupy multiple candidates.

read_document_excerpt(path, offset=0, max_chars=4000) reads a bounded range and returns previous and next offsets. Responses are capped at 20,000 characters so a large document cannot accidentally fill the model context.

corpus_info() returns the persisted dataset description and build metadata.