Document → Clean Markdown / Chunks for RAG pipelines, vector databases, and AI agents.
DocForge is an open-source document ingest platform that converts PDF, HTML, DOCX, and plain text into normalized Markdown plus deterministic chunked JSON with token estimates — ready for Pinecone, Weaviate, pgvector, LangChain, LlamaIndex, or any embedding pipeline.
Live demo: https://docforge.konsole.one
Repository: https://github.com/bdeeps/docforge
- Why DocForge?
- Features
- Who is this for?
- Quick start
- Architecture
- API reference
- Chunking strategies
- AI agents & MCP
- Documentation index
- Deployment (Railway)
- Docker
- Environment variables
- Project structure
- Development
- FAQ
- License
Every RAG and agent pipeline needs the same first step: turn messy documents into clean, chunkable text. DocForge solves that with:
| Problem | DocForge solution |
|---|---|
| PDFs with broken layout | PyMuPDF text extraction → structured Markdown |
| HTML noise (scripts, styles) | BeautifulSoup + markdownify cleanup |
| Word docs (.docx) | python-docx → Markdown with headings & tables |
| Inconsistent chunk boundaries | Deterministic chunking (same input → same chunks) |
| Unknown token counts | tiktoken estimates per chunk and document |
| Agent integration | Self-registration, MCP tools, machine-readable manifest |
| Production deploy | Railway Postgres, Docker, OpenAPI |
Search terms: RAG document ingest, PDF to markdown API, document chunking, MCP document server, agent self-registration, vector store preprocessing, deterministic chunking, tiktoken chunk size.
- Multi-format ingest — PDF · HTML · HTM · DOCX · TXT · Markdown
- Clean Markdown output — normalized whitespace, preserved structure
- Deterministic chunking
headings— split on H1–H6 withheading_pathmetadatasize— fixed token windows with configurable overlap
- Stable chunk IDs — SHA256-derived, idempotent vector upserts
- Token estimates —
tiktoken(cl100k_basedefault) - Agent self-registration — REST + MCP, API keys (
X-DocForge-Key) - MCP server —
docforge-mcpfor Cursor, Claude Desktop, agent tooling - Machine-readable discovery —
AGENTS.md,llms.txt,/api/v1/agent-manifest - AEO FAQ — 24 Q&A pairs at
docs/faq.md - Web UI — upload, chunk config, Markdown/JSON preview
- OpenAPI — interactive docs at
/api/docs - PostgreSQL — Railway-ready persistent agent registry
- ML / platform engineers building RAG ingestion pipelines
- AI agent developers needing MCP-native document tools
- DevOps teams deploying document preprocessing on Railway/Docker
- Cursor / Claude users who want one command to ingest docs into a vector store
- Answer engines & crawlers — structured docs, FAQ, JSON-LD schema
git clone https://github.com/bdeeps/docforge.git
cd docforge
chmod +x scripts/dev.sh
./scripts/dev.shpython3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
cd web && npm install && npm run build && cd ..
docforge-apicurl -X POST http://127.0.0.1:8787/api/v1/ingest \
-F "file=@document.pdf" \
-F "strategy=headings" \
-F "max_tokens=512"Response includes markdown, chunks[], token_estimate, and document_id.
flowchart LR
subgraph Input
PDF[PDF]
HTML[HTML]
DOCX[DOCX]
TXT[Text]
end
subgraph DocForge
API[FastAPI REST]
MCP[docforge-mcp]
CONV[Converters]
CHK[Chunking Engine]
AGT[Agent Registry]
DB[(SQLite / Postgres)]
end
subgraph Output
MD[Clean Markdown]
JSON[Chunked JSON]
TOK[Token Estimates]
end
PDF --> CONV
HTML --> CONV
DOCX --> CONV
TXT --> CONV
CONV --> CHK
CHK --> MD
CHK --> JSON
CHK --> TOK
API --> CONV
MCP --> API
AGT --> DB
API --> AGT
Stack: Python 3.11+ · FastAPI · SQLAlchemy (async) · PyMuPDF · tiktoken · React + Vite + Tailwind
Base path: /api/v1 · OpenAPI: /api/docs
| Method | Endpoint | Auth | Description |
|---|---|---|---|
GET |
/health |
— | Health + DB status |
GET |
/agent-manifest |
— | Machine-readable agent discovery JSON |
POST |
/ingest |
Optional key | File → Markdown + chunks (full pipeline) |
POST |
/convert |
Optional key | File → clean Markdown only |
POST |
/chunk |
Optional key | Markdown → chunked JSON |
POST |
/agents/register |
— | Agent self-registration (API key once) |
GET |
/agents |
— | List registered agents |
GET |
/agents/me |
Required key | Authenticated agent profile |
curl -sS -X POST 'http://127.0.0.1:8787/api/v1/ingest' \
-H 'X-DocForge-Key: df_YOUR_KEY' \
-F 'file=@report.pdf' \
-F 'strategy=headings' \
-F 'max_tokens=512' \
-F 'overlap_tokens=64'{
"id": "a1b2c3d4e5f67890",
"index": 0,
"content": "## Section\n\nParagraph text…",
"token_estimate": 142,
"char_count": 580,
"heading_path": ["Introduction", "Overview"],
"metadata": { "strategy": "headings", "section_index": 0 }
}Full REST reference: docs/agents/rest.md
| Strategy | Best for | Key params |
|---|---|---|
headings |
Manuals, wikis, structured reports | min_heading_level, max_heading_level, max_tokens |
size |
Logs, transcripts, fixed embedding windows | max_tokens, overlap_tokens, token_model |
Both strategies are deterministic — identical input and config produce identical chunk IDs and boundaries.
Guide: docs/agents/chunking.md
DocForge is built for agent-native document ingest. Agents discover, self-register, and ingest without human UI.
| Resource | Local | Deployed |
|---|---|---|
| Agent guide | AGENTS.md |
https://YOUR_HOST/AGENTS.md |
| JSON manifest | /api/v1/agent-manifest |
https://YOUR_HOST/api/v1/agent-manifest |
| LLM index | llms.txt |
https://YOUR_HOST/llms.txt |
| FAQ (AEO) | docs/faq.md |
https://YOUR_HOST/docs/faq.md |
| Cursor skill | .cursor/skills/docforge/SKILL.md |
https://YOUR_HOST/docs/skill.md |
curl -sS -X POST 'https://docforge.konsole.one/api/v1/agents/register' \
-H 'content-type: application/json' \
-d '{"name":"my-rag-agent","capabilities":["ingest","chunk","convert"]}'Save api_key, cursor_config, and mcp_env from the response (shown once).
{
"mcpServers": {
"docforge": {
"command": "docforge-mcp",
"env": {
"DOCFORGE_API_URL": "https://docforge.konsole.one",
"DOCFORGE_API_KEY": "df_YOUR_KEY"
}
}
}
}MCP tools: get_agent_manifest · register_agent · list_agents · convert_document · chunk_markdown · ingest_document · agent_whoami · health_check
See mcp-config.example.json and docs/agents/mcp.md.
| Document | Description |
|---|---|
| README.md | This file — overview & quick start |
| AGENTS.md | Primary guide for AI agents |
| docs/faq.md | 24-question FAQ (AEO-optimized) |
| docs/agents/rest.md | REST API reference |
| docs/agents/mcp.md | MCP tool reference |
| docs/agents/chunking.md | Chunking strategies & RAG tips |
| docs/skill.md | Cursor Agent Skill (HTTP copy) |
| CONTRIBUTING.md | How to contribute |
| SECURITY.md | Security policy |
DocForge deploys to Railway with PostgreSQL for persistent agent registrations.
- Fork / connect github.com/bdeeps/docforge
- Create Railway project → Deploy from GitHub
- Add PostgreSQL plugin →
DATABASE_URLauto-set - Set service env:
| Variable | Value |
|---|---|
DOCFORGE_HOST |
0.0.0.0 |
PORT |
(Railway auto) |
DATABASE_URL |
(Postgres plugin auto) |
- Deploy using included
Dockerfile
DocForge normalizes postgres:// → postgresql+asyncpg:// automatically.
Health check: GET /api/v1/health (returns database: "connected" when Postgres is reachable)
docker compose up --buildOptional local Postgres:
docker compose --profile postgres up --buildCopy .env.example → .env
| Variable | Default | Description |
|---|---|---|
DOCFORGE_HOST |
0.0.0.0 |
Bind address |
DOCFORGE_PORT / PORT |
8787 |
Server port (Railway sets PORT) |
DOCFORGE_APP_ROOT |
. (cwd) |
Root for web/dist and agent docs |
DOCFORGE_DATABASE_URL / DATABASE_URL |
SQLite file | Postgres on Railway |
DOCFORGE_MAX_UPLOAD_MB |
50 |
Max upload size |
DOCFORGE_API_KEY_HEADER |
X-DocForge-Key |
Agent auth header |
docforge/
├── AGENTS.md # AI agent onboarding guide
├── llms.txt # LLM / answer-engine discovery index
├── docs/
│ ├── faq.md # AEO FAQ (24 Q&A)
│ ├── skill.md # Cursor skill (HTTP)
│ └── agents/ # MCP, REST, chunking references
├── src/docforge/
│ ├── api/ # FastAPI routes + docs serving
│ ├── mcp/ # docforge-mcp stdio server
│ ├── converters.py # PDF/HTML/DOCX → Markdown
│ ├── chunking.py # Deterministic chunk engine
│ └── agent_manifest.py # Machine-readable discovery JSON
├── web/ # React UI (Vite + Tailwind)
├── Dockerfile
├── docker-compose.yml
└── .cursor/skills/docforge/SKILL.md
pip install -e ".[dev]"
pytest # run tests
ruff check src tests # lint
cd web && npm run dev # UI dev server (proxies /api)Common questions: docs/faq.md — formats, chunking, MCP setup, Railway, troubleshooting.
Quick links when deployed:
Suggested repository topics:
rag · document-processing · markdown · pdf · chunking · fastapi · mcp · model-context-protocol · vector-database · llm · ai-agents · document-ingestion · tiktoken · langchain · embeddings · railway · python
MIT — Copyright (c) 2026 DocForge contributors
DocForge — turn documents into RAG-ready Markdown and chunks.
Star on GitHub ·
Live demo ·
Agent guide ·
FAQ