Skip to content

About

Compare RAG pipelines stage by stage — branch at any node, run the variants on the same document, and see which combination actually answers better.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

RAG preview

Pick an engine at each stage of your RAG pipeline, run it, and compare the results side by side.

RAG preview — the pipeline overview

▶ Open the live demo — a real session, nothing to install

Why

Most RAG stacks lock in their first guess. You pick a parser, a chunk size, an embedding model — and when the answers come back wrong, finding out which of those decisions did the damage means rebuilding the whole thing.

Here every stage is a node you can branch from. Fork at the chunker, swap the retriever, and the two pipelines grow side by side over the same document. The evaluate stage then ranks them, so "which combination actually works" becomes a table you read instead of a hunch you defend.

How a pipeline works

A pipeline runs in two phases. Indexing happens once per document and is cached as a tree; Query runs per question on top of whichever index you built. Stages marked ◇ are optional — skip them and the pipeline still runs.

Indexing

# Stage What it does Engines
1 Upload Renders the file to page images, so you can see what the parser saw —
2 Parse Pulls text and layout out of the file — OCR, vision LLMs, or native extractors 16
3 ◇ Clean Strips repeated headers and footers, rejoins hyphenated line breaks, normalizes formulas 1
4 Chunk Cuts the text into retrievable pieces — by document structure, by meaning, or by fixed size 6
5 ◇ Metadata Attaches facts about each chunk (section, page, keywords) without touching its text 2
6 ◇ Contextual Gives each chunk back the context it lost when it was cut out 4
7 Embed Turns each chunk into a vector 7
8 Index Stores the vectors so they can be searched 1

Query

# Stage What it does Engines
9 ◇ Query rewrite Reshapes the question before searching — expand it, step back, or draft a hypothetical answer (HyDE) 5
10 Retrieve Pulls candidate chunks for the question 3
11 ◇ Rerank Reorders those candidates by how well they actually answer it 3
12 Generate Writes the answer from the retrieved chunks 4
13 ◇ Evaluate Scores the answer — grounded in its context, relevant, matching the reference 1

53 engines in total. Korean documents are a first-class target: KURE, KoE5 and arctic-ko embeddings and Korean rerankers ship as built-in options, and every engine has a page under /document explaining how it works, with diagrams.

One thing to know before you click in: the pipeline overview is in English, but the stage workbenches and the engine docs are written in Korean.

The demo

The live demo is one real session exported to static HTML: a 10-page paper, 186 nodes, every stage page and every node's input and output. It is read-only — you can click through everything, but not run anything new.

The leaderboard is the part worth looking at. Pipelines that answered the same question are ranked together, and expanding a row shows the lineage that produced it. Dropping top_k to 1 barely moves the score; cutting the same document into 64-token chunks collapses it at the same top_k. The cause was chunk quality, not how much was retrieved — which is the kind of thing this tool exists to make visible.

Examples

Three of the thirteen pipeline stages, and the tree they all live in — every stage has a workbench like these. Each image links to that exact page in the live demo.

Upload stage — the document rendered to page images Upload. The document rendered to page images up front, so every later stage can be checked against what was actually on the page. Parse stage — page image with layout boxes beside the extracted markdown Parse. The page as the model saw it, with the layout blocks it detected, next to the markdown it produced — figures lifted out, algorithms and formulas kept as math. Switch engines on the left and compare the same page.
Embed stage — cosine similarity heatmap across chunks Embed. Cosine similarity between chunks as a heatmap. A model that pulls unrelated chunks together shows up here as a bright grid rather than a bright diagonal. Node tree fanning out from one document All of it at once. Every node you ran, laid out by stage. Seven parsers feed six cleaners feed twelve chunkers — each path is a pipeline you can follow to its score.

Getting started

You need Docker, the Compose plugin, and an OpenRouter API key — every LLM and embedding call goes through OpenRouter.

git clone https://github.com/legojeon/rag-preview.git
cd rag-preview
cp .env.example .env

Fill in at least these:

Variable Value
POSTGRES_PASSWORD any random string
BETTER_AUTH_SECRET openssl rand -base64 32
BETTER_AUTH_URL the address you'll actually open (e.g. http://localhost:20016)
OPENROUTER_API_KEY your OpenRouter key
AUTORAG_HOST_REPO_DIR $PWD
DOCKER_GID getent group docker | cut -d: -f3
docker compose up -d --wait postgres
docker compose up -d

Frontend on http://localhost:20016, API on http://localhost:20017.

Sign-up is gated by admin approval. To approve the first account — your own — directly:

docker compose exec postgres psql -U rag -d jeus-rag \
  -c "update \"user\" set role='admin', approved=true where email='<your-email>';"

Running the GPU engines

Local GPU parsers and embedders run as sibling containers the worker starts per job. Those images aren't in this repo — build them with deploy/runpod/build.sh, push them to your own registry, and point AUTORAG_IMAGE_REPO at it. See deploy/README.md and deploy/runpod/.

You don't need a GPU. Leave AUTORAG_IMAGE_REPO empty and the Docker path switches off; everything that goes through OpenRouter — the vision-LLM parsers, OpenAI embeddings, all generation and evaluation — keeps working. If you'd rather run the GPU engines in the cloud, there's a RunPod serverless path (RUNPOD_*).

Layout

backend/     FastAPI + worker. engines/ holds one implementation per stage engine
frontend/    Next.js 14 App Router — the pipeline canvas and the per-stage workbench
config/      engines.yaml — engine registration and metadata
deploy/      compose operations, GPU engine image builds, RunPod endpoints
demo/        the session dump the static demo reads

State is split in two: PostgreSQL holds nodes and sessions, workspace/ holds parse output, vectors and images.

Tests

pytest backend/tests          # 676 passed
cd frontend && npx vitest run # 557 passed

The backend tests need DATABASE_URL and isolate themselves into a separate jeus-rag_test database, so they won't touch your development data.

License

Code is MIT.

The document in the demo is Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Albert Gu, Tri Dao — arXiv:2312.00752), used under CC BY 4.0. The authors do not endorse this demo. See demo/ATTRIBUTION.md.

The Noto Korean fonts in backend/render/fonts/ are under the SIL Open Font License 1.1.

About

Compare RAG pipelines stage by stage — branch at any node, run the variants on the same document, and see which combination actually answers better.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages