Skip to content

feat(api): revalidate the collection cache against disk on read - #90

Merged
simoncodes-ca merged 1 commit into
developfrom
feature/cache-revalidation
Sep 5, 2026
Merged

simoncodes-ca merged 1 commit into
developfrom
feature/cache-revalidation

Conversation

@simoncodes-ca

@simoncodes-ca simoncodes-ca commented Sep 5, 2026 •

Copy link
Copy Markdown
Owner

Closes #86

Problem

The app serves a collection from an in-memory tree that it patches as the UI mutates
resources. CLI commands write straight to resource_entries.json / tracker_meta.json and
never tell a running app, so the cache goes stale — and a browser refresh does not fix it,
because the page re-reads the same server-side cache. Only a restart does. Users reasonably
read that as data loss.

Why not the two approaches in the issue

Both have a WSL failure mode, and both miss changes that do not come from the CLI.

Approach WSL failure mode
CLI → POST /invalidate CLI in WSL, app on Windows (or vice versa) → localhost is not shared. WSL2→Windows host needs the gateway IP from /etc/resolv.conf, or Win11 22H2+ mirrored networking. Silent no-op exactly when it matters.
fs.watch / chokidar in the API Windows-side writes on /mnt/c (drvfs/9p) never raise inotify events in WSL — microsoft/WSL#4739. The watcher looks healthy and fires nothing. Same on several network and container mounts.

Neither sees a git checkout, a branch switch, or a hand edit.

What this does instead: revalidate on read

Before serving the tree, the cache status, or a search, the API takes a stat-only
fingerprint of the collection's translations folder and compares it with the one recorded at
index time. A mismatch drops the cache, so the request answers 202 Accepted and the next
one returns fresh data. The existing not-ready → async reindex → UI retry protocol does the
rest, so the Tracker UI needed no change.

  • computeTreeFingerprint() (libs/core/src/lib/resource/tree-fingerprint.ts) walks the
    folder with readdir + stat only, returning {fileCount, folderCount, totalSize, maxMtimeMs}. Counts and size sit alongside mtime because mtime granularity is as coarse
    as one to two seconds on drvfs and FAT, so a same-second edit can leave mtime untouched.
  • CollectionCacheService.revalidate() is throttled to 2000 ms, configurable via
    LINGO_TRACKER_REVALIDATE_INTERVAL_MS.
  • The API's own writes re-take the baseline on the next tick, coalesced, so a bulk endpoint's
    per-resource cache patches cost one scan rather than one per resource — and the app never
    reads its own write as somebody else's.

No new dependency, no watcher, no port resolution, no daemon. Behaves identically on macOS,
Linux, Windows, WSL, containers and network shares.

Verified end to end

Built API run against a scratch project on port 3931:

Action Result
add-resource via the CLI new key served on the next fetch, no restart
POST /resources via the API no false invalidation — invalidation and index counts unchanged
Hand edit of resource_entries.json detected, edited value served

Suites green: core 1227, api 200, tracker 493, plus cli and domain. pnpm lint and
pnpm typecheck clean.

The app serves a collection from an in-memory tree it patches as the UI mutates
resources. CLI commands write straight to disk, so the cache went stale and a
browser refresh did not help — it re-read the same cache. Only a restart did.

Detect the change on read instead of being told about it. Before serving the
tree, cache status or a search, the API takes a stat-only fingerprint of the
translations folder (file and folder counts, total size, newest mtime) and
compares it with the one recorded at index time. A mismatch drops the cache, so
the request answers 202 and the next one returns fresh data.

Watching the folder was the obvious alternative and is not portable: inotify
never fires for Windows-side writes on a WSL /mnt/c mount, and the same holds for
several network and container mounts. A CLI-side HTTP invalidate is not portable
either — localhost is not shared across the WSL boundary — and would miss every
change that does not come from the CLI, such as a git checkout or a hand edit.

The scan is throttled to once every 2000 ms (LINGO_TRACKER_REVALIDATE_INTERVAL_MS),
and the API's own writes re-take the fingerprint on the next tick, coalesced, so
bulk endpoints do not trigger a redundant re-index.

Closes #86

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RifJr7xyhZkaTWn4ssV1b4
@simoncodes-ca
simoncodes-ca merged commit 7ff3594 into develop Sep 5, 2026
1 check passed
@simoncodes-ca
simoncodes-ca deleted the feature/cache-revalidation branch September 5, 2026 05:29
@lingo-tracker-release

Copy link
Copy Markdown

🎉 This PR is included in version 0.18.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Invalidate the running app's index after a CLI resource mutation

2 participants