Skip to content

feat: pass litelink's read-cache settings through to every reader - #93

Merged
nhobin219 merged 2 commits into
mainfrom
read-cache
Oct 4, 2026
Merged

nhobin219 merged 2 commits into
mainfrom
read-cache

Conversation

@nhobin219

@nhobin219 nhobin219 commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Closes #77.

Following the issue's comment, cache_key is the caller's choice. Rather than keying each stream's cache by its id, streamcast passes litelink's four settings through and lets the caller decide.

What

  • Four keywords: memory_cache, disk_cache, cache_key and disk_cache_volume_limit, with litelink's defaults (memory on, disk off).
  • Every reader of the published tables takes them:
    • Stream.snapshot, scan and sql;
    • Stream.live, for its base, every rebase, and the catch-up it opens;
    • connect, for a catch-up.
  • Inside, the settings travel as one frozen ReadCache. It's part of the key the shared DuckDB databases are pooled under, so readers asking for different caching get separate databases (the issue's first "to settle" point). A local read caches only in memory, so disk settings don't split local databases.

A litelink bug found on the way: a disk-cached reader goes stale

The issue asked to check that cache_httpfs never caches version-hint.text. It does. I filed it as litelink#141. Rewrite the hint, then read it again:

same process new process, same cache
memory cache only (default) fresh —
disk_cache=True stale stale
+ exclusion regex for the hint stale fresh
  • In the same process, cache_httpfs's file-handle cache keeps the hint's handle for up to an hour, and it ignores the exclusion.
  • After a restart, its data cache serves the old hint.
  • The effect: a stale hint pins a reader to the first snapshot it saw. With litelink's documented iceberg_scan(<table path>) read, a disk-cached connection kept reading 250 rows after 50 more were published.

Fixed upstream in litelink 0.9 with litelink.current_metadata, which reads the hint outside DuckDB's caches. Table.open now resolves through it, so a reader always names the table's current metadata.json, and every file that names is written once and caches safely. This PR therefore requires litelink >=0.9.0,<0.10.

On 0.9, iceberg_scan on a table path through a disk-cached connection fails an ETag check instead of reading old rows. But reading the hint itself with read_text through that connection still returns the old one in the same process. I noted that on litelink#141; streamcast doesn't read it that way any more.

Disk bound

disk_cache_volume_limit keeps litelink's 0.8, and it's the caller's to change. That answers the issue's third "to settle" point the same way as the key.

Benchmark

Run against rustfs on localhost, since bench-snapshot only reads local tables: a full Stream.scan of 300,000 rows.

cold warm, same process cold after restart
disk cache off 0.51 s 0.20 s 0.49 s
disk cache on 0.93 s 0.13 s 0.55 s
  • On localhost the disk cache gains nothing. S3 there reads about as fast as local disk, and the first read pays for loading cache_httpfs and writing blocks.
  • Its case is reads that cross a network to object storage, which I couldn't measure from here.
  • That supports off by default and the caller's choice, and API.md says so. It isn't evidence about real S3.

Tests

  • test_a_disk_cached_reader_sees_every_publish (S3): with the disk cache on, a table reopened after a new publish sees it, and the cache directory has data in it.
  • test_different_caching_never_shares_a_database: different settings get different databases; local disk settings don't split.
  • test_the_settings_reach_every_read_of_a_stream: snapshot, the live base and a rebase all pass the caller's exact settings to the connection.
  • test_a_catch_up_reads_with_the_callers_cache_settings: connect(catch_up=True, disk_cache=True, cache_key=…) reads with them.
  • Falsified: each of these breaks fails a test:
    • reading the hint through DuckDB again;
    • keying databases by credentials only;
    • a rebase dropping the settings;
    • a catch-up dropping the settings.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi

nhobin219 and others added 2 commits October 4, 2026 07:39
memory_cache, disk_cache, cache_key and disk_cache_volume_limit, with
litelink's defaults, on Stream.snapshot, scan, sql, live (its base, each
rebase and its catch-up) and connect for a catch-up. The key is the
caller's: nothing is keyed automatically. Readers with different
settings read through different DuckDB databases, which now share by
credentials and cache settings.

version-hint.text is read around DuckDB. With the disk cache on,
cache_httpfs served an old hint for ever (litelink#141): in the same
process through its file-handle cache, and after a restart through its
data cache. Every file the hint names is written once and caches safely.

Closes #77.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi
litelink 0.9 fixes litelink#141 upstream with current_metadata, which
reads a published table's version-hint outside DuckDB's caches. Table.open
now resolves through it, so the hint has one owner, and the uncached read
helper moves back to _snapshot, unchanged from main. litelink is
>=0.9.0,<0.10.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi
@nhobin219
nhobin219 merged commit 23becb8 into main Oct 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Use litelink 0.7's read cache for snapshot, live and catch-up

1 participant