Skip to content

Repository files navigation

RAG + Context Management MCP Server

mcp-rag-memory

Persistent long-term memory & knowledge retrieval for AI assistants. Exposes RAG (vector storage + semantic search) and context management (memories) as MCP tools, so opencode can store and recall context across sessions.

👤 Identity: The Coder — "I'm The Coder, Selamat datang dan Semoga perjalanan mu menyenangkan"

📖 Project documentation site: https://adyoi.github.io/mcp-rag-memory/ (see docs/)

What it does

  • RAG — ingest documents (text / files / whole directories), chunk + embed locally, store in SQLite, and semantically retrieve relevant context.
  • Hybrid search — FTS5 BM25 keyword hits fused with vector similarity (reciprocal-rank fusion). Tunable SEARCH_MODE=hybrid|vector|keyword.
  • Context Management — remember facts/decisions/preferences, recall them later (score-decayed by staleness), rate by importance, consolidate duplicates, filter by type/tag.
  • Zero external APIs by default — local hashing-embedder (1024-dim), built-in node:sqlite. Works fully offline. Optional EMBEDDING_PROVIDER=transformers for a higher-quality ONNX model.
  • Safe ingestion — content-hash deduplication, file size guard (RAG_MAX_FILE_MB) and an opt-in path allowlist (RAG_ALLOWED_DIRS).
  • MCP server — runs on stdio, 19 tools.

Quick start

Requires Node.js ≥ 22.12 (for node:sqlite read-only handles and busy_timeout).

npm install
npm test                 # 160 checks: unit + MCP round-trip via SDK client
npm run lint             # oxlint
npm run typecheck        # src + scripts + .opencode/plugin
npm run build            # compile to dist/
npm run cli -- docs      # CLI playground

Run from npm (npx)

Install or run directly from the npm package — no tsx, no source checkout:

npx mcp-rag-memory

Point your MCP client at npx mcp-rag-memory (the published bin), or at the local dev entry: node --import tsx src/mcp/rag-server.ts.

Works with any MCP client

This is a standard MCP server (stdio transport, official @modelcontextprotocol/sdk). It speaks the MCP spec, so any agent that supports MCP can use it — not just opencode. Memory & knowledge become shared across all your tools: remember once, recall everywhere.

opencode

{
  "mcp": {
    "rag-memory": {
      "type": "local",
      "command": ["node", "--import", "tsx", "D:/Project/mcp-server/src/mcp/rag-server.ts"],
      "cwd": "D:/Project/mcp-server",
      "enabled": true,
      "environment": { "RAG_DB_DIR": "D:/Project/mcp-server/.rag-data" }
    }
  }
}

Claude Desktop

claude_desktop_config.json (in %APPDATA%\Claude):

{
  "mcpServers": {
    "rag-memory": {
      "command": "node",
      "args": ["--import", "tsx", "D:/Project/mcp-server/src/mcp/rag-server.ts"],
      "env": { "RAG_DB_DIR": "D:/Project/mcp-server/.rag-data" }
    }
  }
}

Cursor / Windsurf

Settings → MCP → add server, same command/args pattern as Claude Desktop. No cwd needed — use absolute paths for the entry file and DB.

VS Code (Copilot)

mcp.json in .vscode/:

{
  "servers": {
    "rag-memory": {
      "type": "stdio",
      "command": "node",
      "args": ["--import", "tsx", "D:/Project/mcp-server/src/mcp/rag-server.ts"],
      "env": { "RAG_DB_DIR": "D:/Project/mcp-server/.rag-data" }
    }
  }
}

Tip: once published (or installed), just point every client at npx mcp-rag-memory — no tsx, no absolute paths, no local checkout. Locally you can still use node --import tsx <abs-path>/src/mcp/rag-server.ts.

opencode configuration

Global config (~/.config/opencode/opencode.json)

The MCP server is registered here so it works in every project, not just this folder:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "rag-memory": {
      "type": "local",
      "command": ["node", "--import", "tsx", "src/mcp/rag-server.ts"],
      "cwd": "D:/Project/mcp-server",
      "enabled": true,
      "environment": {
        "RAG_DB_DIR": "D:/Project/mcp-server/.rag-data"
      }
    }
  }
}

cwd and RAG_DB_DIR are absolute paths so the server resolves tsx from this project's node_modules and stores data in the same DB regardless of which project opencode is opened from.

Re-assert this global config with one command. opencode updates have been observed to reset/wipe the global config (the MCP server, plugin, LSP and instructions disappear from new sessions). scripts/setup-opencode.* rewrite both opencode.json and opencode.jsonc (PowerShell on Windows, portable bash elsewhere; dispatcher npm run setup-opencode), merging with what is already there — user keys are never clobbered, arrays (instructions, plugin) are unioned, broken files are backed up to .bak-<ts>, and runs are idempotent. Re-run it (and restart opencode) after any opencode update:

npm run setup-opencode          # merge-in our defaults
npm run setup-opencode -- --check   # exit 0 = up to date, 1 = changes pending

Project config (opencode.json in this repo)

Keeps project-specific settings (model, LSP, permissions, instructions) — no mcp block here:

{
  "$schema": "https://opencode.ai/config.json",
  "model": "anthropic/claude-sonnet-4-6",
  "lsp": true,
  "instructions": [".opencode/instructions.md"],
  "permission": {
    "edit": "allow",
    "bash": { "git *": "allow", "*": "ask" }
  }
}

.opencode/instructions.md tells every new session to call memory_context(topic="The Coder") first, so the assistant always knows who you are without asking. The same instructions file is also wired into the global config to apply to every project.

Notifications & sound (~/.config/opencode/tui.json)

{
  "$schema": "https://opencode.ai/tui.json",
  "attention": {
    "enabled": true,
    "sound": true,
    "notifications": true,
    "volume": 0.4
  }
}

Restart opencode after any config change.

MCP tools (19)

Group Tool Purpose
System system_stats DB dir, embedding backend, doc/memory counts
RAG ingest rag_ingest_text Store text as a knowledge document (dedup-aware)
rag_ingest_file Store a file (content_type auto-detected)
rag_ingest_dir Recursively ingest source files
RAG query rag_search Hybrid (BM25 + vector) search, ranked chunks + scores
rag_retrieve Ready-to-inject context block with token count
RAG docs rag_list_documents, rag_document_stats Inventory (rag_list_documents is paginated: limit 1–1000, offset)
rag_delete_document Remove a document + chunks + FTS rows
Session sync rag_sync_session Ingest an opencode session's user inputs (or the latest in a workspace)
Memory memory_remember Save long-term memory (type/importance/tags)
memory_recall Semantic memory search (decay + recall count)
memory_context Compact context block from memories for prompts
memory_list, memory_get, memory_update, memory_forget CRUD
memory_consolidate Dedupe near-identical (cosine and word overlap) + promote hot memories
memory_stats Counts, tokens, avg importance, by type

Run in other AI assistants

The server is a standard MCP stdio server — it works with any MCP client (Claude Desktop, Cursor, Continue, VS Code, ...), not just opencode.

Via npx (after publishing).

npx mcp-rag-memory

Claude Desktop — claude_desktop_config.json:

{
  "mcpServers": {
    "rag-memory": {
      "command": "npx",
      "args": ["-y", "mcp-rag-memory"],
      "env": { "RAG_DB_DIR": "C:\\Users\\you\\rag-data" }
    }
  }
}

Cursor (.cursor/mcp.json) — continue style JSON, same shape:

{
  "mcpServers": {
    "rag-memory": {
      "command": "node",
      "args": ["D:/Project/mcp-server/dist/index.js"],
      "env": { "RAG_DB_DIR": "D:/Project/mcp-server/.rag-data" }
    }
  }
}

Continue — ~/.continue/config.json:

{
  "mcpServers": {
    "rag-memory": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/Project/mcp-server/dist/index.js"]
    }
  }
}

Configuration. Every RAG_* variable can come from three places, with the first winning: the MCP client's env block → your shell environment → a .env file in the working directory. Copy .env.example to .env and edit, or point at any file via RAG_ENV_FILE. The store defaults to .rag-data/ in the working directory.

Custom slash commands

The 19 MCP tools are called by the AI automatically — but you can also trigger them directly with custom commands. Sources (same name, project wins):

  • Per-project: .opencode/commands/
  • Global: ~/.config/opencode/commands/
Command Backing tool Use
/remember <content> [--type][--importance] memory_remember Save a memory
/recall <topic> memory_recall Search memories semantically
/context <topic> memory_context Context block for prompts
/consolidate memory_consolidate Dedupe + promote hot memories
/search <query> rag_search Semantic search documents
/retrieve <query> rag_retrieve Raw context block + tokens
/ingest <text> rag_ingest_text Store knowledge text
/docs rag_list_documents List all documents
/stats system_stats + doc/memory stats Full statistics

CLI

npm run cli -- remember "deploys every Friday" --type task --importance 0.7
npm run cli -- recall "deployment schedule"
npm run cli -- ingest-dir ./src/rag
npm run cli -- search "vector similarity"
npm run cli -- search "auth bug" --source opencode-db   # filter by source or doc-id
npm run cli -- docs --limit 20 --offset 20            # paginated inventory
npm run cli -- mem-context "database"
npm run cli -- sessions              # list recent opencode sessions
npm run cli -- sync-session ses_123  # ingest one session's inputs
npm run cli -- sync-latest           # ingest the most recent session
npm run clean-db -- --force          # wipe ALL documents + memories (destructive)

mcp-rag-memory is also a global CLI. Install once (npm i -g mcp-rag-memory) and run mcp-rag-memory-cli search "..." anywhere; the CLI reads <cwd>/.env or RAG_ENV_FILE to locate its store, so one store per project is enough.

Semua akses juga tersedia sebagai custom commands (/remember, /recall, /search, ...) dan sebagai MCP tools.

Auto-save session inputs

Every user message typed in opencode can be mirrored into the RAG store as small searchable documents (metadata source: "opencode-db"), kept separate from your curated long-term memories. Three layers:

  1. Plugin (real-time, debounced, cross-workspace). .opencode/plugin/session-logger.ts watches message events and, at most once a minute, runs the portable CLI — npx -y -p mcp-rag-memory mcp-rag-memory-cli sync-latest --dir <workspace> — so the session you are typing in lands in the store. Because it invokes the published package (not this repo's scripts), it works in any project that adds the plugin and a .env pointing at its store. Disable with RAG_AUTOSYNC=0.
  2. CLI on demand.
    npx -y -p mcp-rag-memory mcp-rag-memory-cli sync-latest --dir "D:/Project/my-app"  # latest session in a workspace
    npx -y -p mcp-rag-memory mcp-rag-memory-cli sync-session ses_abc123                # a specific session
    npx -y -p mcp-rag-memory mcp-rag-memory-cli sync-logs                              # .session-logs/*.jsonl
  3. MCP tool. rag_sync_session exposes the same logic over MCP: { "session_id": "..." } or { "directory": "..." } for the latest.

Honest scope note: the plugin only ships with this repo out of the box, but it is a plain node script — copy it into any workspace (and give that project an .env) and auto-save just works, since it shells out to the published CLI bin (mcp-rag-memory-cli, inside the mcp-rag-memory package) via npx. Without a plugin, auto-save is "manual": run layer 2 or call layer 3 whenever you want a session ingested.

How it works: sessions are read from opencode's own DB (~/.local/share/opencode/opencode.db + opencode-local.db, override with OPENCODE_DB / OPENCODE_DB_LOCAL), each user message becomes one small document (title <session>-<ts>, metadata source: "opencode-db"). Long inputs are condensed to key points before storage (extractive, no LLM) so the store stays lean; short inputs are saved whole. Raw transcripts stay in opencode's DB; long-term memories are never mixed in. Re-runs are idempotent thanks to SHA-256 content-hash dedup (metadata.condensed marks reduced docs), so plugging this into any scheduler is safe. A schema guard raises a clear error (instead of mysterious SQL failures) if your opencode DB layout changes to an unsupported shape.

The plugin needs an opencode restart to take effect.

Performance & limits

  • Search is brute-force cosine / FTS5 over every chunk: O(N) per query (exact, deterministic — good for a personal store of thousands of chunks). ANN / HNSW indexing is planned scope (v3) for 100k+ chunks.
  • Keyword leg uses the FTS5 trigram tokenizer, so CJK text and Indonesian-style substring/inflection matching work without extra config. Consequence: keyword terms of ≤ 2 characters are skipped (trigram needs 3); the vector leg still covers them.
  • Vector embeddings are local hash-based (no external APIs, ~0 cost, offline). They are tuned for short-phrase similarity, not full-document semantics.

Architecture

src/
├── db/database.ts          node:sqlite (documents, chunks, memories, chunks_fts, meta)
├── rag/
│   ├── embedder.ts         local hashing embedder (1024-d, FNV-1a, n-grams)
│   ├── embeddings.ts       provider layer: local | transformers (optional dep)
│   ├── chunker.ts          paragraph/code-aware chunking with overlap
│   ├── vector-search.ts    vector cache + FTS5 BM25 + RRF hybrid scoring
│   └── pipeline.ts         async ingest/search/retrieve, dedup, guards, doc mgmt
├── session/transcript.ts   read opencode sessions (global+local DB), ingest inputs
├── memory/memory.ts        remember / recall / consolidate + decay + prune
├── mcp/rag-server.ts       MCP server (19 tools, stdio, npm bin)
├── cli.ts                  CLI playground (incl. sync-session / sync-latest / sync-logs)
└── test/test-all.ts        full test suite (unit + MCP round-trip)

Data lives in .rag-data/rag.sqlite (git-ignored).

Configuration (env vars)

Variable Default Purpose
RAG_DB_DIR .rag-data Where the SQLite store lives
SEARCH_MODE hybrid hybrid | vector | keyword
EMBEDDING_PROVIDER local local (zero-dep) | transformers (needs optional @huggingface/transformers)
EMBEDDING_MODEL Xenova/all-MiniLM-L6-v2 Transformers model name
RAG_MAX_FILE_MB 10 Reject files larger than this
RAG_ALLOWED_DIRS (unset = anywhere) Semicolon/pipe/comma-separated allowed ingest roots
RAG_MEMORY_HALF_LIFE_DAYS 14 Recall score decay half-life
RAG_PRUNE 0 Set 1 to allow consolidate to delete non-essential memories
RAG_PRUNE_IMPORTANCE 0.2 Delete memories below this importance
RAG_PRUNE_AGE_DAYS 90 ...and older than this (never-recalled only)
RAG_AUTOSYNC 1 Set 0 to disable the session-logger plugin's auto-sync
OPENCODE_DB / OPENCODE_DB_LOCAL ~/.local/share/opencode/*.db Where session transcripts are read from (re-read on every call, so it can be set late)
RAG_SESSION_MAX_MSGS 500 Cap on user messages read from one session
RAG_SESSION_CONDENSE 1 Set 0 to store session inputs verbatim
RAG_SESSION_CONDENSE_MIN_CHARS 120 Inputs at/below this length are saved whole
RAG_SESSION_CONDENSE_RATIO 0.35 Fraction of long-input length to keep as key points
RAG_ENV_FILE <cwd>/.env Env file to load at boot

A .env file can only set this server's own variables (RAG_*, OPENCODE_*, EMBEDDING_*, SEARCH_MODE); anything else in it is ignored with a warning on stderr, and a real environment variable always wins over the file. The loader runs from dist/index.js and dist/cli.js only — if you invoke src/mcp/rag-server.ts directly (TSX/npx tsx), import src/env.js yourself.

Testing

npm test spins up the real MCP server over stdio using the SDK client and exercises every tool end-to-end against a scratch DB (.test-data). npm run lint checks the codebase with oxlint (run in CI too).

GitHub Pages site

The docs/ directory is a self-contained static site of this project (single index.html, no build step). To publish it on GitHub Pages:

  1. Push this repo to GitHub.
  2. Repo Settings → Pages → Build and deployment → Source: Deploy from a branch.
  3. Choose branch main and folder /docs.
  4. Your site is live at https://adyoi.github.io/mcp-rag-memory/.

Troubleshooting

ConfigInvalidError: missing key "command" for typescript / javascript / json

The lsp block in opencode.json has the wrong shape. Each language key must have a command array:

"lsp": {
  "typescript": {
    "command": ["typescript-language-server", "--stdio"]
  }
}

Fields like language_id or extensions alone are not allowed — the schema enforces additionalProperties: false and requires command.

Fix: either write the command array, or (if you just want built-in LSP) replace the whole block with "lsp": true.


MCP tools not showing up in opencode

  1. Wrong config filename — opencode reads only opencode.json, opencode.jsonc, or .opencode/opencode.json. A file named opencode.jsonx is silently ignored; no error, no tools.
  2. File not in the right place — global MCP config lives at ~/.config/opencode/opencode.json, not in the project folder.
  3. Missing cwd for global use — when command uses a relative path (src/mcp/rag-server.ts), opencode resolves it against the workspace directory, not the project. Add "cwd" to point at the project that contains node_modules/tsx:
"rag-memory": {
  "type": "local",
  "command": ["node", "--import", "tsx", "src/mcp/rag-server.ts"],
  "cwd": "D:/Project/mcp-server"
}
  1. Missing env var — RAG_DB_DIR must be set (absolute path for global config) or the DB defaults to .rag-data relative to cwd.

Sound / notification not playing when response finishes

The attention feature is off by default. Create ~/.config/opencode/tui.json:

{
  "$schema": "https://opencode.ai/tui.json",
  "attention": {
    "enabled": true,
    "sound": true,
    "notifications": true,
    "volume": 0.4
  }
}

Known issue (#40445): sound silently fails when opencode runs under the Node runtime instead of Bun — the audio library depends on Bun FFI. Symptoms: no sound despite correct config. Not a configuration error.


Server slow to start on first load (~3 s)

tsx compiles TypeScript on first invocation. This is normal and only happens on cold start; subsequent tool calls within the same session are instant.


Opencode feels sluggish / UI lag after adding plugins

Plugins that run synchronous I/O or write to stderr on every tool call can slow down the TUI. This repo ships exactly one plugin (.opencode/plugin/session-logger.ts) and it is debounced to at most one npx … sync-latest per 60s, spawned detached with stdio: "ignore" and shell: false. If it still costs too much, set RAG_AUTOSYNC=0 (keeps the plugin, disables the auto-sync) or drop the plugin entry from opencode.json — the MCP server and its 19 tools keep working, you just have to call rag_sync_session yourself.


ajv-cli fails with strict mode: unknown keyword: allowComments

The opencode schema uses custom JSON-Schema extensions (allowComments, allowTrailingCommas). ajv-cli rejects these by default. Validate config manually instead:

node -e "JSON.parse(require('fs').readFileSync('opencode.json','utf8')); console.log('OK')"

About

RAG + context management MCP server: persistent memory, vector search, zero APIs

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages