Skip to content

Fix/general improvements - #17

Merged
Abas-Tim merged 2 commits into
mainfrom
fix/general-improvements
Aug 28, 2026
Merged

Abas-Tim merged 2 commits into
mainfrom
fix/general-improvements

Conversation

@Abas-Tim

Copy link
Copy Markdown
Owner

Summary

Pre-release hardening pass: fixes three destructive/incorrect behaviors,
removes dead code, splits oversized modules, adds lint enforcement, and
expands the machine-readable CLI surface. Two commits:

  • fix: harden index lifecycle, embedder (6fa0c3c)
  • feat: split storage and extractors, git-aware reindex, json outputs (339c1d0)

Critical fixes

  • Read paths no longer destroy embeddings. Opening an index whose stored
    vector dimension differs from embedding.dimension previously dropped the
    whole vec_units table on any open (search, status, MCP). It now raises a
    clear error; rebuilds only happen via explicit migrate opens
    (urag index, init --full, embed --reindex, MCP index_now).
  • Config round-trip preserved http_timeout (was silently reset to 30s).
  • Embedder hardening: new local model is verified before the old cache is
    purged; MCP embedder cache key now includes dimension; fallback warnings
    go to stderr instead of degrading silently.
  • Indexer crash-safety: refs_pending/extractor_version upgrade flags
    are written only after successful re-extraction; a process-wide write lock
    serializes watcher vs manual runs; index_paths honors max_file_bytes.
  • Watcher debounce race fixed (stale timer could swallow a newer flush).
  • Eval harness: --reresolve now re-derives gold for reference questions
    (stale ids survived index rebuilds, producing false regressions).

Architecture

  • db.py split: call-graph/reference queries -> db_edges.py, search and
    navigation -> db_search.py (mixins). from urag.db import Database
    unchanged.
  • native_ext.py split into per-language modules (go_ext, rust_ext,
    java_ext, c_ext, csharp_ext) + shared native_common.py;
    native_ext remains a compatibility re-export.
  • Eval orchestration moved out of the CLI into eval.run_eval (~290 lines).
  • C# extractor reuses a cached parser (grammar was recompiled per method);
    Python extractor defers tree-sitter grammar compilation to first use.
  • MCP tool naming contract documented (server names are short; harnesses
    prefix them, e.g. urag_search). No renames — renaming would double-prefix
    in harnesses like opencode.

Improvements

  • Git-aware incremental indexing: files clean vs HEAD with matching
    size+mtime skip re-hashing (modified/untracked still content-verified;
    non-git repos keep exact hashing).
  • --json for read, status, doctor (doctor exits non-zero on failure).
  • Definition queries no longer treat common English words as exact symbol
    identifiers (precision boost, fewer false positives).
  • 5 missing SQLite indexes; % wildcard escapes; evidence off-by-one;
    char-vs-byte chunk attribution; budget floors/config ceilings; top_k=0.
  • pathspec GitIgnoreSpec (removes ~3k deprecation warnings in tests).
  • Ruff config + CI lint step; license = {text = "MIT"}; dead code removed
    (~1k net lines deleted).

Verification

  • Tests: 158 passed / 1 skipped (rg not installed on Windows CI-host; skip is
    pre-existing). Ruff clean.
  • Benchmarks vs pre-branch (same questions, re-resolved gold):
    • self repo: hybrid recall 0.560 -> 0.583, mrr 0.447 -> 0.464,
      tokens/run 727 -> 701; callers/transitive tokens 952 -> 940;
      auto/lexical/references unchanged (0.947 / 0.947 / 0.928)
    • fixture suite: quality byte-identical; fresh index build 7.3s -> 6.3s
    • no regressions
  • Real-project smoke test on a C#-heavy repo (index migration, search,
    callers through interface fields, deadcode, eval).
  • uv build wheel + sdist verified; wheel ships all new modules.

- dimension mismatch no longer drops all embeddings on read opens; rebuilds only via explicit migrate paths (index, init --full, embed --reindex, MCP index_now/init_project)
- config round-trip keeps http_timeout; read-only commands no longer create .urag/urag.toml
- embedder: verify new model before purging old cache; MCP cache key includes dimension; fallback warnings to stderr
- indexer: upgrade flags written after successful re-extraction; process-wide write lock; index_paths honors max_file_bytes
- watcher: stale-timer flush race fixed
- db: missing indexes added, % wildcard escapes, evidence off-by-one, orphan cleanup on migrate opens only
- retrieve: definition-symbol regex bug, per-call regex compilation, top_k=0, budgets respect config ceiling
- eval: orchestration moved to eval.run_eval; closures bind loop vars; chunk byte offsets; reresolve re-derives reference gold
- extractors: C# parser caching, lazy tree-sitter, dead code removed
- packaging: ruff config + CI lint step, license table, dev deps, changelog
- db: call-graph/reference queries moved to db_edges, search/navigation to db_search (mixins); public imports unchanged
- extractors: native languages split into go/rust/java/c/csharp modules with native_common helpers; native_ext is a compatibility re-export
- indexer: git repos skip re-hashing files that are clean vs HEAD with matching size+mtime; non-git keeps exact hashing
- cli: --json for read, status, doctor; doctor exits non-zero on failure
- retrieve: definition queries ignore common English words as exact identifiers
- mcp: document harness-level tool prefixing (urag_search etc.); skill doc updated
- tests: git fast-path, --json outputs, stopword helper; 158 passed
@Abas-Tim
Abas-Tim merged commit 8c0047d into main Aug 28, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant