feat: pass litelink's read-cache settings through to every reader - #93
Merged
Merged
Conversation
memory_cache, disk_cache, cache_key and disk_cache_volume_limit, with litelink's defaults, on Stream.snapshot, scan, sql, live (its base, each rebase and its catch-up) and connect for a catch-up. The key is the caller's: nothing is keyed automatically. Readers with different settings read through different DuckDB databases, which now share by credentials and cache settings. version-hint.text is read around DuckDB. With the disk cache on, cache_httpfs served an old hint for ever (litelink#141): in the same process through its file-handle cache, and after a restart through its data cache. Every file the hint names is written once and caches safely. Closes #77. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi
litelink 0.9 fixes litelink#141 upstream with current_metadata, which reads a published table's version-hint outside DuckDB's caches. Table.open now resolves through it, so the hint has one owner, and the uncached read helper moves back to _snapshot, unchanged from main. litelink is >=0.9.0,<0.10. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi
This was referenced Oct 4, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #77.
Following the issue's comment,
cache_keyis the caller's choice. Rather than keying each stream's cache by its id, streamcast passes litelink's four settings through and lets the caller decide.What
memory_cache,disk_cache,cache_keyanddisk_cache_volume_limit, with litelink's defaults (memory on, disk off).Stream.snapshot,scanandsql;Stream.live, for its base, every rebase, and the catch-up it opens;connect, for a catch-up.ReadCache. It's part of the key the shared DuckDB databases are pooled under, so readers asking for different caching get separate databases (the issue's first "to settle" point). A local read caches only in memory, so disk settings don't split local databases.A litelink bug found on the way: a disk-cached reader goes stale
The issue asked to check that
cache_httpfsnever cachesversion-hint.text. It does. I filed it as litelink#141. Rewrite the hint, then read it again:disk_cache=Truecache_httpfs's file-handle cache keeps the hint's handle for up to an hour, and it ignores the exclusion.iceberg_scan(<table path>)read, a disk-cached connection kept reading 250 rows after 50 more were published.Fixed upstream in litelink 0.9 with
litelink.current_metadata, which reads the hint outside DuckDB's caches.Table.opennow resolves through it, so a reader always names the table's currentmetadata.json, and every file that names is written once and caches safely. This PR therefore requires litelink>=0.9.0,<0.10.On 0.9,
iceberg_scanon a table path through a disk-cached connection fails an ETag check instead of reading old rows. But reading the hint itself withread_textthrough that connection still returns the old one in the same process. I noted that on litelink#141; streamcast doesn't read it that way any more.Disk bound
disk_cache_volume_limitkeeps litelink's 0.8, and it's the caller's to change. That answers the issue's third "to settle" point the same way as the key.Benchmark
Run against rustfs on localhost, since
bench-snapshotonly reads local tables: a fullStream.scanof 300,000 rows.cache_httpfsand writing blocks.Tests
test_a_disk_cached_reader_sees_every_publish(S3): with the disk cache on, a table reopened after a new publish sees it, and the cache directory has data in it.test_different_caching_never_shares_a_database: different settings get different databases; local disk settings don't split.test_the_settings_reach_every_read_of_a_stream:snapshot, thelivebase and a rebase all pass the caller's exact settings to the connection.test_a_catch_up_reads_with_the_callers_cache_settings:connect(catch_up=True, disk_cache=True, cache_key=…)reads with them.🤖 Generated with Claude Code
https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi