Skip to content

[Feature](lance) Add index prewarm to reduce query-time random S3 IO #68692

Description

@Gabriel39

Goal

Add a synchronous Doris SQL statement that calls Lance SDK index prewarm on the BEs used by subsequent queries. Reduce query-time random S3 reads for index contents by loading the index into the same BE-local shared cache used by normal queries.

Parent tracker: #66340. Completion on branch-4.1 follows the parent tracker convention; this is not a commitment to a published release date.

Scope

The statement waits for all target BE prewarm calls to complete and returns success or an explicit error directly to the client.

The implementation prewarms one named logical index, including all of its active segments in a pinned dataset snapshot. It reuses the SDK's existing prewarm implementation and Doris's shared Lance session.

SQL interface

-- Synchronous; defaults to the current execution context.
WARM UP INDEX vector_idx ON lake_catalog.demo.items;

-- Optional: prewarm a different compute group before routing queries there.
WARM UP INDEX vector_idx ON lake_catalog.demo.items
WITH COMPUTE GROUP search_group;

This syntax is the interface to implement.

  • Compute-group selection is optional. By default, resolve the current session's compute group. In deployments without compute groups, use the eligible BEs selected by existing query/resource-tag policies.
  • Warm all eligible target BEs in that execution context, so subsequent normal scheduling can use the cache. Do not include unrelated compute groups.
  • Resolve and fix the target BE set and dataset snapshot at statement start. Workers must not independently resolve latest.
  • Resolve the logical index name in that snapshot. The SDK handles its active physical segments; users do not need to provide UUIDs.
  • Validate table access, compute-group usage, and the existing administrative permission selected for this resource-intensive operation.
  • A nonexistent or unsupported index, inaccessible storage, or failed target BE causes an explicit statement error. Include failed BE/index context in the error and avoid credentials or sensitive storage options.
  • Return success only when all target calls have completed successfully. A compact result can include table, index, pinned dataset version, target BE count, and elapsed time.
  • The statement uses ordinary SQL/RPC timeout and connection-error behavior. A timeout does not roll back already cached entries or guarantee immediate interruption of in-flight SDK IO.
  • A user can rerun the statement after failure. Successful cache fills on other BEs need not be rolled back. Do not silently switch to a different snapshot during execution.

Implementation

Synchronous SQL command
  -> FE validates access and resolves snapshot, index, and target BEs
  -> FE sends bounded synchronous prewarm RPCs and waits for completion
  -> BE opens the fixed snapshot through LanceSessionManager
  -> lance-c calls dataset.prewarm_index(index_name)
  -> later queries reuse the same BE-local session/index cache
  -> FE returns success or an explicit error

Lance-C: a thin binding

The upstream binding provides the following public entry point:

int32_t lance_dataset_prewarm_index(
    const LanceDataset* dataset,
    const char* index_name);

The call blocks until the SDK operation finishes and uses the existing error-reporting convention. It validates null arguments and string encoding and preserves dataset/session ownership for the call lifetime.

Lance Rust already provides prewarm by logical index name, including multiple segments. The C binding delegates to that existing method. Integration tests must verify compatibility and cache reuse.

Dependency baseline: lance-c 98468344bc9d56aa7ed4a4192e844afa62b34ba1 includes the public C/C++ prewarm binding from lance-c #94 and pins Lance 68c12dfd7efe02f90ad2d7f3a239a7eb64884e57. The Doris dependency update is tracked by Doris #68698.

Doris FE/BE

  • Add SQL parsing, analysis, and a synchronous command.
  • Reuse catalog access and snapshot resolution, including the normal temporary-credential path where applicable. Keep credentials scoped to the normal execution/access path and exclude them from logs and user-visible results.
  • Add a small BE RPC handler that opens the pinned dataset through the existing LanceSessionManager and invokes the binding. Return status directly.
  • Bound RPC fan-out and concurrent prewarm work per BE using existing execution/resource controls. The synchronous user interface does not require serial execution across every BE.
  • Reuse the SDK's bounded windowed IO. Do not implement per-partition RPCs or issue synthetic ANN queries to warm the cache.
  • Use normal mixed-version capability/error handling so an old BE cannot silently report success without doing the operation.

Cache scope and expected benefit

The existing LanceSessionManager owns one shared session per BE process. The prewarm handler must use it rather than constructing a temporary private session whose cache disappears after the statement.

Cache Role
Lance index cache Main destination: reusable index structures and partition state, shared across queries on the same BE
Lance metadata cache Metadata loaded while opening the dataset/index; governed by normal SDK cache keys
Existing Foyer data-file cache A separate data-file cache, not persistent index prewarm

Current Doris configuration exposes lance_index_cache_size_bytes and lance_metadata_cache_size_bytes, with current defaults of 10 GiB and 1 GiB per BE. These are process-scoped settings in the existing manager. Verify values against the delivery branch.

A successful statement means all requested SDK calls completed. It does not guarantee permanent or complete residency: entries remain evictable, and a working set larger than the cache cannot stay fully warm. Index files' encoded size is not necessarily their decoded memory footprint. Document capacity planning and retain bounded IO/decode concurrency.

For N target BEs and a working set of about S bytes, full replication needs approximately S on each BE and N * S aggregate capacity. It can also multiply prewarm traffic. Cache on another BE cannot substitute for missing entries on the BE executing a query.

The pinned SDK IVF prewarm implementation already groups adjacent partitions into bounded read windows. Reuse that path to move cold index reads ahead of foreground queries and reduce small remote reads during subsequent queries.

Prewarm still consumes S3 IO itself. It may increase total traffic for rarely queried indexes. It does not automatically warm row-fetch/output columns, data-file filter columns, unindexed rows, or all version-dependent metadata. It does not eliminate distance-computation/Top-K CPU work.

Cache effectiveness and invalidation

Event Behavior
Query/connection ends or a dataset handle closes BE shared cache remains available
FE restarts Cache in surviving BEs remains available
BE restarts, is replaced, or recreates the session In-memory index cache is cold; rerun prewarm if needed
Other queries/indexes cause capacity pressure Entries can be evicted and subsequent reads can reach S3 again
Queries use another compute group or new BEs Those BEs were not covered by the completed statement
Append retains old index segment identities Existing segments can be reused; new segments and unindexed data are not covered
Index rebuild/merge/replacement creates new identities New segments need warming; old entries are unused or usable by matching historical snapshots until eviction
Index/table is dropped New queries stop using that index; no immediate physical cache flush is required
Snapshot changes Reuse only matching immutable index contents; preserve version-dependent visibility and deletion semantics
Storage identity/access-isolation context changes Revalidate cache identity; index name alone cannot identify cached contents

Lance scopes index cache entries by dataset URI and index UUID, with additional remapping identity where needed. Reuse the normal query access path and check that prewarm cannot bypass authorization or alias different storage namespaces. Do not flush all immutable segments merely because the dataset version changes.

There is no new cache TTL, pinning guarantee, persistent index cache, automatic refresh, or topology-driven rewarm in this first implementation.

Tests and acceptance

  • FE tests: parsing, optional compute-group resolution, permissions, pinned snapshot/index resolution, and target selection.
  • Lance-C/BE tests: null/invalid inputs, missing indexes, multi-segment indexes, shared-session reuse across separately opened dataset handles, proper resource release, and error propagation.
  • End-to-end regression: synchronous completion, correct results before/after prewarm, multiple BEs, and an explicit error when any target fails.
  • Cache lifecycle tests: shared-session reuse, capacity eviction, BE restart, index replacement, and snapshot visibility/deletes.
  • Establish and publish the supported index-type/format and catalog-access matrix. Do not claim untested scalar/FTS or vector format combinations are covered.
  • In a controlled environment with enough cache and no eviction, subsequent queries should perform approximately zero remote index-content reads. Measure metadata and row-fetch IO separately. Test multiple vectors and nprobes values rather than replaying only one query.
  • Measure prewarm duration/traffic, subsequent S3 GET/Range count and bytes, cold/warm latency, peak memory, and interference with concurrent foreground queries. Exact new per-operation metrics are not a prerequisite; use existing trustworthy metrics or isolated instrumented tests. Do not treat a concurrent global cache-stat delta as task-local IO.
  • Add bilingual Doris website documentation for syntax, default execution context, synchronous behavior, errors, cache lifecycle, capacity planning, and benefit boundaries.

Roadmap

Stages are ordered by implementation dependency. Completion requires code, tests, documentation, and verification evidence.

M1 — Upstream Lance-C binding

Upstream implementation: lance-c #94 is merged. The synchronous C/C++ APIs, argument/error handling, multi-segment and shared-session tests, and usage documentation are available upstream. Doris dependency delivery is tracked by Doris #68698; M1 remains pending until that PR is merged into branch-4.1.

  • Add the synchronous prewarm-by-index-name C API over the existing Rust SDK method.
  • Add argument/error, multi-segment, shared-session reuse, and resource-lifetime tests plus a usage example.
  • Merge the Lance-C binding upstream.
  • Merge the Doris dependency update without a Doris-only prewarm patch.

The existing Foyer data-file cache patch is retained until lance-c #73 merges; it contains no downstream prewarm implementation.

M2 — Synchronous Doris SQL integration

  • Implement parsing/analysis and normal permission checks.
  • Resolve a fixed snapshot and eligible target BEs; implement bounded RPC fan-out and synchronous waiting.
  • Invoke the C API through LanceSessionManager and return explicit success/error.
  • Add FE/BE unit tests and external regression coverage.

M3 — Verification, documentation, and delivery

  • Verify the supported index/catalog matrix, cache lifecycle, capacity behavior, and failed-target handling.
  • Benchmark remote index-read reduction and foreground-query impact with reproducible cold/warm conditions.
  • Publish bilingual user documentation and merge delivery into branch-4.1.
  • Link upstream, Doris, documentation PRs and verification evidence before closing this issue.

Code references

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/catalogIssues or PRs related to catalog managementkind/featureCategorizes issue or PR as related to a new feature.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions