Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
name: CI

on:
pull_request:
push:
branches: [main]

jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: 20
registry-url: https://registry.npmjs.org

- run: npm ci

- run: npm run build
58 changes: 58 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# CLAUDE.md

Guidance for Claude Code when working in the **context-engine** repo.

## What this is

`@q1k-oss/context-engine` — a TypeScript **library** (plus an optional standalone HTTP server) for the customer-knowledge layer: document extraction, knowledge-graph building, and prioritized-context retrieval.

**Primary consumption (ADR-037): as a library.** `q1k-controlplane`'s Temporal worker imports this package and calls its pure functions **in-process** inside activities — there is no deployed context-engine HTTP service in the document-ingestion path. context-engine owns the *domain logic* (Docling/Gemini extraction, entity extraction, MINT mapping, chunking); durability/retry/concurrency/persistence live in controlplane's Temporal workflow + tenant Postgres. This mirrors the btree (logic) + Temporal (durability) split.

The Express server (`src/server.ts`, `createApp`) still exists for standalone chat/graph use, but the **file-upload route and async `processFile` orchestration were removed** (ADR-037) — ingestion is controlplane's job now.

## The library surface (what controlplane imports)

Pure, side-effect-free functions exported from `src/index.ts`:

- `doclingClientService.extractFromFile({ storagePath, mimeType })` → `ExtractedContent` (spawns the Python Docling adapter; falls back to Gemini on failure).
- `structureToMint(structure, metadata)` / `toMintDocument(...)` → token-efficient MINT encoding (`@q1k-oss/mint-format` ≥ 1.1.0).
- `chunkDocument(structure, opts)` → `DocumentChunk[]` — §-boundary chunks with overlap + `{ text, sectionRef, pageRef, chunkIndex }`. Deterministic (stable `chunkIndex` = idempotency key for the persist activity).
- `claudeClientService` / `geminiClientService` — LLM clients (extraction only).
- Types from `src/types/` (`ExtractedContent`, `DocumentStructure`, `DocumentChunk`).

None of these touch the DB or filesystem (beyond Docling reading the file path it's handed). Persistence (`knowledge_chunks`, the `kg_*` tables) is controlplane's.

## Commands

```bash
npm run build # tsc -> dist/
npm test # vitest run
npm run dev # tsx watch src/server.ts (standalone server)
npm run db:push # apply schema (standalone server only; controlplane owns its own tenant schemas)
```

## Document extraction (Docling)

`doclingClientService.extractFromFile()` spawns `python/docling_extract.py` (Docling). On any failure it falls back to Gemini, so a missing Python/Docling runtime degrades gracefully.

- **Device:** the layout model (RT-DETR-v2) requests float64, unsupported on Apple-Silicon MPS. The script pins the accelerator via **`DOCLING_DEVICE`** (default `cpu`, verified stable on macOS). Set `DOCLING_DEVICE=cuda`/`mps`/`auto` on a capable host.
- **Python command:** override `uv run python` via **`DOCLING_CMD`** (e.g. `python3`, or `uv run --project <ce-repo> python` when consuming this package from another repo whose cwd lacks the Docling env).
- **Packaging:** `python/` + `pyproject.toml` ship in the npm package (`files` field) so the script travels with the library. The consuming worker still needs a Python+Docling runtime reachable via `DOCLING_CMD`.

Local Docling setup:
```bash
uv sync # installs Docling (first run pulls ML models, ~hundreds of MB)
npx tsx scripts/smoke-docling.ts # MINT + chunk smoke (no Docling needed)
DOCLING_CMD="uv run python" npx tsx scripts/smoke-docling.ts file.pdf # real Docling extraction + chunking
```

## MINT usage

MINT (`@q1k-oss/mint-format`) packs data into LLM prompts token-efficiently: graph nodes/edges (`claude-client.service.ts`) and parsed documents (`structureToMint`/`encodeDocument`). Requires `@q1k-oss/mint-format` ≥ 1.1.0.

## Conventions

- ESM, NodeNext — **import with `.js` extensions** even from `.ts`.
- Subpath imports break `require()` of `package.json` — use the declared `exports`.
- Schema in `src/db/schema/*.ts`. The `files` table is now **dead** (orchestration removed); left in place rather than surgically dropped.
- Standalone server routes use a trusted-caller pattern; for the library path, auth/tenant-isolation are the consuming worker's concern.
12 changes: 6 additions & 6 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

16 changes: 12 additions & 4 deletions package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@q1k-oss/context-engine",
"version": "0.2.0",
"version": "0.3.0",
"description": "AI-powered knowledge graph engine that extracts and structures domain knowledge from conversations",
"type": "module",
"license": "MIT",
Expand All @@ -25,7 +25,9 @@
"types": "./dist/index.d.ts",
"files": [
"dist",
"!dist/__tests__"
"!dist/__tests__",
"python",
"pyproject.toml"
],
"exports": {
".": {
Expand Down Expand Up @@ -67,6 +69,10 @@
"./tools": {
"types": "./dist/tools/index.d.ts",
"import": "./dist/tools/index.js"
},
"./extraction": {
"types": "./dist/extraction.d.ts",
"import": "./dist/extraction.js"
}
},
"engines": {
Expand All @@ -81,12 +87,14 @@
"db:generate": "drizzle-kit generate",
"db:migrate": "drizzle-kit migrate",
"db:push": "drizzle-kit push",
"db:studio": "drizzle-kit studio"
"db:studio": "drizzle-kit studio",
"test": "vitest run",
"test:watch": "vitest"
},
"dependencies": {
"@anthropic-ai/sdk": "^0.39.0",
"@google/generative-ai": "^0.21.0",
"@q1k-oss/mint-format": "^1.0.2",
"@q1k-oss/mint-format": "^1.1.0",
"cors": "^2.8.5",
"drizzle-orm": "^0.38.3",
"express": "^4.21.2",
Expand Down
11 changes: 11 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
[project]
name = "context-engine-docling"
version = "0.1.0"
description = "Docling extraction adapter for context-engine file ingestion"
requires-python = ">=3.11"
dependencies = [
"docling>=2.0.0",
]

[tool.uv]
package = false
157 changes: 157 additions & 0 deletions python/docling_extract.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
#!/usr/bin/env python3
"""Docling extraction adapter for context-engine.

Reads a single file path (argv[1]), parses it with Docling, and prints a JSON
object to stdout matching the TypeScript `ExtractedContent` shape:

{
"textContent": str,
"structure": {
"title": str | None,
"sections": [{"heading": str, "content": str, "level": int, "pageRef": int | None}],
"tables": [{"caption": str | None, "headers": [str], "rows": [[str]]}],
"lists": [{"type": "ordered" | "unordered", "items": [str]}],
"figures": [{"caption": str | None, "pageRef": int | None}]
},
"metadata": {"pageCount": int, "wordCount": int, "characterCount": int}
}

The TS caller (`docling-client.service.ts`) spawns this and falls back to the
Gemini extractor on a non-zero exit, so this script must fail loudly (exit 1 +
stderr) rather than emit partial JSON.

Run: `uv run python python/docling_extract.py <file>`

Device: the layout model (RT-DETR-v2) requests float64, which Apple-Silicon MPS
does not support — so we pin the accelerator via `DOCLING_DEVICE` (default `CPU`,
verified stable on macOS). Set `DOCLING_DEVICE=CUDA`/`MPS`/`AUTO` on a capable host.
"""

import json
import os
import sys


def _page_of(item) -> "int | None":
prov = getattr(item, "prov", None)
if prov:
first = prov[0] if isinstance(prov, (list, tuple)) else prov
return getattr(first, "page_no", None)
return None


def extract(path: str) -> dict:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import (
AcceleratorOptions,
PdfPipelineOptions,
)

# Pin the accelerator device (lowercase string: auto/cpu/mps/cuda/cuda:N).
# Default cpu avoids the Apple-Silicon MPS float64 crash in RT-DETR-v2;
# override with DOCLING_DEVICE on a capable host.
device = os.getenv("DOCLING_DEVICE", "cpu").lower()
pipeline_options = PdfPipelineOptions(
accelerator_options=AcceleratorOptions(device=device)
)
converter = DocumentConverter(
format_options={
InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
}
)

doc = converter.convert(path).document

text_content = doc.export_to_markdown()

sections: list[dict] = []
lists: list[dict] = []
current_list: "dict | None" = None

for item in getattr(doc, "texts", []) or []:
label = str(getattr(item, "label", "")).lower()
text = (getattr(item, "text", "") or "").strip()
if not text:
continue

if "list_item" in label:
if current_list is None:
current_list = {"type": "unordered", "items": []}
lists.append(current_list)
current_list["items"].append(text)
continue
current_list = None # any non-list item closes the run

if "section_header" in label or "title" in label:
sections.append(
{
"heading": text,
"content": "",
"level": int(getattr(item, "level", 1) or 1),
"pageRef": _page_of(item),
}
)
else:
if not sections:
sections.append({"heading": "", "content": "", "level": 1, "pageRef": _page_of(item)})
sections[-1]["content"] = (sections[-1]["content"] + "\n" + text).strip()

tables: list[dict] = []
for tbl in getattr(doc, "tables", []) or []:
try:
df = tbl.export_to_dataframe()
headers = [str(c) for c in df.columns.tolist()]
rows = [[("" if v is None else str(v)) for v in row] for row in df.values.tolist()]
except Exception:
headers, rows = [], []
caption = None
try:
caption = tbl.caption_text(doc) or None
except Exception:
pass
tables.append({"caption": caption, "headers": headers, "rows": rows})

figures: list[dict] = []
for pic in getattr(doc, "pictures", []) or []:
caption = None
try:
caption = pic.caption_text(doc) or None
except Exception:
pass
figures.append({"caption": caption, "pageRef": _page_of(pic)})

title = sections[0]["heading"] if sections and sections[0]["heading"] else None

return {
"textContent": text_content,
"structure": {
"title": title,
"sections": sections,
"tables": tables,
"lists": lists,
"figures": figures,
},
"metadata": {
"pageCount": len(getattr(doc, "pages", []) or []),
"wordCount": len(text_content.split()),
"characterCount": len(text_content),
},
}


def main() -> int:
if len(sys.argv) < 2:
print("usage: docling_extract.py <file>", file=sys.stderr)
return 2
try:
result = extract(sys.argv[1])
except Exception as exc: # noqa: BLE001 — surface any failure to the TS caller
print(f"docling extraction failed: {exc}", file=sys.stderr)
return 1
json.dump(result, sys.stdout, ensure_ascii=False)
return 0


if __name__ == "__main__":
sys.exit(main())
Loading
Loading