Skip to content

Latest commit

 

History

History
425 lines (326 loc) · 18.8 KB

File metadata and controls

425 lines (326 loc) · 18.8 KB

Contributing to VelesDB

First off, thank you for considering contributing to VelesDB! It's people like you that make VelesDB such a great tool.

Table of Contents

Code of Conduct

This project and everyone participating in it is governed by our Code of Conduct. By participating, you are expected to uphold this code.

How Can I Contribute?

Reporting Bugs

Before creating bug reports, please check the existing issues to avoid duplicates. When you create a bug report, include as many details as possible:

  • Use a clear and descriptive title
  • Describe the exact steps to reproduce the problem
  • Provide specific examples (code snippets, configuration files)
  • Describe the behavior you observed and what you expected
  • Include your environment details (OS, Rust version, VelesDB version)

Suggesting Enhancements

Enhancement suggestions are tracked as GitHub issues. When creating an enhancement suggestion:

  • Use a clear and descriptive title
  • Provide a detailed description of the proposed enhancement
  • Explain why this enhancement would be useful
  • List any alternatives you've considered

Your First Code Contribution

Unsure where to begin? Look for issues labeled:

  • good first issue - Simple issues perfect for newcomers
  • help wanted - Issues where we need community help
  • documentation - Documentation improvements

Pull Requests

This project follows Git Flow. Feature and fix branches must target develop, not main.

Branch prefix Target Examples
feature/*, feat/* develop feature/amazing-feature
fix/*, bugfix/* develop fix/crash-on-empty-index
refactor/*, chore/*, ci/*, docs/*, style/*, perf/*, test/*, build/* develop docs/update-api-guide
release/*, hotfix/*, support/* main release/1.2.0
  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Make your changes
  4. Run tests (cargo test)
  5. Run lints (cargo clippy)
  6. Format code (cargo fmt)
  7. Commit your changes (git commit -m 'Add amazing feature')
  8. Push to the branch (git push origin feature/amazing-feature)
  9. Open a Pull Request targeting develop (not main)

Architecture Quick Start

New to the codebase? Start with these documents (in order):

  1. Project Structure — Workspace layout, crate responsibilities (~15 min)
  2. Architecture — 3-layer design (Client, API, Core), data flow (~30 min)
  3. Concurrency Model — Lock ordering, shard strategy, deadlock prevention (~20 min)
  4. Storage Format — On-disk layout, WAL, mmap (~15 min)
  5. Soundness — Unsafe code audit with invariant proofs (reference)

Workspace Crate Map

Crate Purpose
velesdb-core Core engine: HNSW, SIMD, VelesQL, collections, storage
velesdb-server Axum REST API server (54 endpoints, 61 operations; OpenAPI optional)
velesdb-cli Interactive REPL for VelesQL
velesdb-python PyO3 bindings with NumPy support
velesdb-wasm Browser-side vector search (no persistence feature)
velesdb-mobile iOS/Android bindings via UniFFI
velesdb-migrate Schema/data migration tooling
velesdb-memory MCP agent-memory server. Three tool families, not one: durable memory (remember/recall/why/relate/feedback…), the deterministic context compiler (compile_context, compile_transcript, explain_compilation, retrieve_context_source, context_savings, suggest_budget), and cross-session working-context resumption (save_working_context, load_working_context, list_working_contexts)
velesdb-node Node.js binding of the memory wedge (napi-rs; @wiscale/velesdb-memory-node)
tauri-plugin-velesdb Tauri desktop integration
velesdb-rustdoc-guard Test-only, never published: flags rustdoc link syntax left in the descriptions velesdb-memory (MCP schemas) and velesdb-server (OpenAPI) publish

Quality Gates

All new/modified code must satisfy these limits (enforced by Codacy and CI):

Metric Limit Enforcement
Cyclomatic complexity <= 8 per function Codacy
Function NLOC <= 50 lines Codacy
File NLOC <= 500 lines Code review
Code duplication < 2% jscpd
Unsafe blocks Must have // SAFETY: comment CI (verify_unsafe_safety_template.py)
TODO format Must carry an issue tag: [EPIC-XXX/US-YYY], (PREFIX-NNN), #123, or #issue CI (check-todo-annotations.py)
.unwrap() Forbidden in production code CI (check_prod_unwraps.py)
Wrapper adapters An impl Trait for Box<T>/Arc<T>/Rc<T>, or for a type DynTrait = Box<dyn Trait> alias, must forward every method the trait declares — rustc only demands the required ones, so a dropped default-bodied method compiles and silently runs the default CI (check-trait-forwarding.py)
Recall@10 >= 0.95 (if search path modified) CI + local validation

Every guard script this repository runs is declared in scripts/guards.json: which workflow and job invokes it, whether its failure blocks a merge, and whether it runs strict or advisory. Adding a scripts/check-*.py (or verify-/gate-) without an entry there fails scripts/tests/test_ci_gate_reachability.py — the registry is checked against the workflows and against the filesystem, so it cannot quietly shrink.

Some of those guards need a Python package — check-figure-sources.py parses Markdown with a real parser rather than by regex, and the CI-wiring suites read the workflows with PyYAML. Every one of them, with its pin, is declared in scripts/requirements-guards.txt:

python3 -m venv .venv-guards
.venv-guards/bin/pip install -r scripts/requirements-guards.txt

Then run the guards with .venv-guards/bin/python. A virtualenv rather than a plain pip install --user, because a system Python that follows PEP 668 — Homebrew's, Debian's, and the CI runners' — refuses to install into itself; ci.yml builds a venv for the same reason. If your Python is not managed that way, python3 -m pip install -r scripts/requirements-guards.txt is enough.

Without them a guard exits 2, which means it could not run — not that it passed, and not that it refused. The workflows install from that same file, and scripts/check_guard_python_pins.py refuses a workflow that pins those packages inline again: the pins used to be repeated across ci.yml and gate-contracts.yml, where nothing kept them in step.

Concurrency Rules

  • Use parking_lot::RwLock / Mutex (never std::sync — no poisoning, no .unwrap() on locks)
  • Follow lock ordering documented in CONCURRENCY_MODEL.md
  • Tests MUST run single-threaded: --test-threads=1 (file system isolation)

Pre-Push Validation

Run this sequence before every push (CI runs on every PR — running it locally first avoids red pipelines and wasted CI minutes):

# 1. Format
cargo fmt --all

# 2. Lint (strict — mirrors CI; node and python are linted separately, see AGENTS.md)
cargo clippy --workspace --all-targets --features persistence,gpu,update-check \
  --exclude velesdb-python --exclude velesdb-node -- -D warnings -D clippy::pedantic

# 3. Tests
cargo test -p velesdb-core --features persistence -- --test-threads=1

# 4. (If search path modified) Recall gate
cargo test -p velesdb-core --features persistence test_recall -- --test-threads=1

# 5. Feature gate check
cargo check --no-default-features
cargo check -p velesdb-wasm --no-default-features --target wasm32-unknown-unknown

# 6. (Optional) Codacy CLI — run from the repo root
codacy-cli analyze
# Windows: prefix with `wsl -- bash -c "cd $(wslpath -a .) && ..."` if the CLI runs under WSL.

Or use the local CI script: .\scripts\local-ci.ps1 (Full) or .\scripts\local-ci.ps1 -Quick (fmt + clippy only).

Git hooks are provided in .githooks/ — activate them with the setup script for your platform, ./scripts/setup-hooks.sh (Linux/macOS) or .\scripts\setup-hooks.ps1 (Windows). Both set core.hooksPath; the shell one also restores the executable bit, which a fresh clone can lose and which git ignores silently — an un-executable hook does not fail, it simply never runs.

The bare equivalent is git config core.hooksPath .githooks.

Four hooks then apply: commit-msg rejects AI-attributed authors and AI attribution trailers, pre-commit validates the change, pre-push runs the full local gate — only on direct pushes to develop/main; feature-branch pushes are waved through and CI is the gate — and post-merge warns when a merge moved the repo ahead of your installed skills.

Note (SSH pushes): the pre-push hook runs the full validation (~15 min) while git already holds the connection to GitHub open; if that SSH connection dies in the meantime, git push exits 141 (SIGPIPE) with no error message right after the validation passes. Pushing over HTTPS is immune (the ref advertisement and the pack upload are independent requests):

git config remote.origin.pushurl https://github.com/cyberlife-coder/VelesDB.git
git config credential.helper '!gh auth git-credential'

Development Setup

Prerequisites

  • Rust 1.90+ (stable) — the workspace MSRV, declared in Cargo.toml (rust-version) and pinned in rust-toolchain.toml so local matches CI; scripts/tests/test_msrv_single_source.py fails CI Success when the two disagree or when a member crate declares its own rust-version. Two things force it: avx512vpopcntdq target_feature, stabilized in 1.89 (see crates/velesdb-core/src/simd_native/x86_avx512.rs), and roaring 0.11.4, which declares rust-version = 1.90.0 — on 1.89 the workspace only builds when that dependency is already cached.
  • Docker (optional, for integration tests)

Building from Source

# Clone the repository (replace cyberlife-coder with your fork if contributing via PR)
git clone https://github.com/cyberlife-coder/VelesDB.git
cd VelesDB

# Build the project
cargo build --workspace

# No GTK/WebKit on the machine? The two Tauri crates are the only ones that
# need them (CI installs libgtk-3-dev/libwebkit2gtk-4.1-dev); everything else
# builds without system packages:
cargo check --workspace --exclude tauri-plugin-velesdb --exclude tauri-rag-app

# Run tests (single-threaded — required for file system isolation)
cargo test --workspace --features persistence,gpu,update-check \
  --exclude velesdb-python -- --test-threads=1

# Lint (strict — mirrors CI; node and python are linted separately, see AGENTS.md)
cargo clippy --workspace --all-targets --features persistence,gpu,update-check \
  --exclude velesdb-python --exclude velesdb-node -- -D warnings -D clippy::pedantic

# Run the server locally
cargo run --bin velesdb-server -- --data-dir ./data

Running Benchmarks

cargo bench -p velesdb-core --features internal-bench -- --noplot

Agent Skills — installing, checking, re-syncing

This repository is the source of truth for three agent skills: skills/velesdb-context-optimizer, skills/velesdb-learning-loop and crates/velesdb-memory/skill/velesdb-memory. An agent does not load them from here — it loads a copy installed under ~/.claude/skills.

That third copy is invisible to CI, which cannot read your home directory, and it drifted: measured on 2026-08-02, the installed velesdb-memory skill was 67 lines behind and stated the wrong argument order for the Node binding's recallFusedDated. Nothing anywhere reported a fault; an agent simply read stale instructions.

# Install (or re-sync) the managed skills into ~/.claude/skills
python3 scripts/sync-skills.py --install

# Report drift; exits non-zero when an installed copy differs
python3 scripts/sync-skills.py --check

# Same, but an ABSENT managed skill also fails (what the post-merge hook runs)
python3 scripts/sync-skills.py --check --strict

# Regenerate the copies bundled into the npm package after editing a source
python3 scripts/sync-skills.py --bundle

Each managed skill is reported as one of three states, never two:

state meaning --check --check --strict
in step the agent reads the right thing green green
drifted the agent reads the wrong thing exit 1 exit 1
absent the agent reads nothing reported, exit 0 exit 1
  • Only the managed skills are touched. Any other skill installed in that directory is left exactly as it is — the tool works from an explicit pair list, never a scan.
  • LOCAL.md is yours. A managed skill may hold a machine-local layer under that name: never committed, never a copy of the shipped SKILL.md. It is preserved across an --install and never reported as drift. Nothing else extra is forgiven — a stale file from an older version is still unexpected.
  • The npm copies are generated. crates/velesdb-node/skills/ ships inside the package; --bundle rewrites it from the same registry. Two byte-identity guards stay red until you run it.
  • --strict widens what ABSENCE costs, and nothing else. Plain --check forgives it because a contributor who never installed these must not have their work refused over a machine-local state; it still says so, since silence would leave you unable to tell "installed and correct" from "not there at all".
  • The copy is atomic. The new tree is built beside the target and moved into place by rename, so an agent reading a SKILL.md during an install sees the whole old version or the whole new one — never half of each.
  • .githooks/post-merge runs --check --strict and prints a notice when a merge has just moved the repository ahead of your install. It deliberately does not block: --check compares against the working tree, so gating a commit would force you to install unmerged work into your global skills. Hooks are also bypassable, which is why the command stays runnable by hand — and why scripts/tests/test_sync_skills.py exercises the mechanism in CI.
  • A versioned hook is not a protection until git is pointed at it. Run ./scripts/setup-hooks.sh (or .\scripts\setup-hooks.ps1) once per clone; it sets core.hooksPath and restores the executable bit a fresh checkout can drop. Verify with git rev-parse --git-path hooks, which prints the path git actually uses. A test now holds both activation scripts against the contents of .githooks/, so a hook added without being described there turns red — that check found setup-hooks.ps1 had never mentioned commit-msg, the hook enforcing the no-AI-attribution rule.

Set CLAUDE_SKILLS_DIR to point both commands at a different directory (the test suite uses it to run without touching a real install).

Pull Request Process

  1. Ensure all tests pass - Run cargo test --workspace --features persistence,gpu,update-check --exclude velesdb-python -- --test-threads=1 before submitting
  2. Update documentation - If you're adding new features, update the relevant docs
  3. Follow the style guidelines - Run cargo fmt and cargo clippy
  4. Write meaningful commit messages - Follow conventional commits format
  5. Keep PRs focused - One feature or fix per PR
  6. Be responsive - Address review feedback promptly

Commit Message Format

We follow the Conventional Commits specification:

<type>(<scope>): <description>

[optional body]

[optional footer]

Types:

  • feat: A new feature
  • fix: A bug fix
  • docs: Documentation changes
  • style: Code style changes (formatting, etc.)
  • refactor: Code refactoring
  • perf: Performance improvements
  • test: Adding or updating tests
  • chore: Maintenance tasks

Examples:

feat(search): add hybrid search support
fix(storage): resolve mmap alignment issue on ARM
docs(readme): update quick start guide

Style Guidelines

Rust Code Style

  • Follow the Rust API Guidelines
  • Use rustfmt for formatting (default configuration)
  • Use clippy for linting (fix all warnings)
  • Write documentation for public APIs
  • Keep functions under 50 lines when possible
  • Prefer composition over inheritance

Documentation Style

  • Use clear, concise language
  • Include code examples where appropriate
  • Keep README focused on getting started
  • Put detailed docs in the /docs folder

Contributor License Agreement (CLA)

By submitting a pull request, you agree to the VelesDB CLA: you grant cyberlife-coder a perpetual, worldwide, non-exclusive, royalty-free license to use, modify, and distribute your contribution under the terms of the VelesDB Core License 1.0 or any future license chosen by the project.

This is required to allow VelesDB to offer commercial licensing alongside the source-available version. All past contributions are covered retroactively.

Recognition

Contributors are recognized in:

  • CONTRIBUTORS.md — full list with PR links
  • Release notes for each version they contributed to
  • Our Discord community

Release Process

CI spans 21 workflow files; the merge gate is the CI Success summary job, whose needs: list in ci.yml is the authoritative inventory (mapped guard-by-guard in scripts/guards.json). The three you interact with most:

Workflow Purpose
ci.yml Tests, lint, security, and the CI Success gate
release.yml Full publish (binaries, crates.io, PyPI, npm)
bench-regression.yml Benchmarks

Publishing a release

# 1. Bump every manifest in lock-step
python3 scripts/bump_version.py <X.Y.Z>
python3 scripts/check-version-sync.py   # every policed manifest must align
python scripts/check-promise-contract.py # every claim must pass
cargo update --workspace                 # refresh Cargo.lock

# 2. Open release/<vX.Y.Z> -> main, wait for ALL CI green on the merge commit
# 3. Tag from main (NEVER push the tag before CI Success on main HEAD)
git checkout main && git pull origin main
git tag -a v<X.Y.Z> -m "v<X.Y.Z> -- <one-line summary>"
git push origin main --tags

The release.yml workflow automatically publishes to:

  • GitHub Releases (binaries)
  • crates.io
  • PyPI
  • npm

📖 Full guide: docs/contributing/RELEASE.md


Thank you for contributing to VelesDB! 🦀