Skip to content

Commit d0c1913

Browse files
smypmsaclaude
andauthored
query: build models once per release, read snapshots in place (#7)
* query: run a build's queries in one worker A build forked one worker per query, and every worker loaded every snapshot and rebuilt every model before running its one SELECT. A project with a few models and a dozen queries paid for the same joins a dozen times, and each repetition counted against that query's 60 s deadline. build now sends all its queries to a single worker as a batch. The worker loads the snapshots, shuts external access and builds the models once, then runs each query in turn and reports one line per query. The sandbox is unchanged: every query is still admitted only as a single SELECT, held to its own row limit, and runs on the closed side of the external-access door. The deadline is per step rather than per batch: loading and models get 60 s, and each query gets its own 60 s after that. A timeout or failure now names the model or query it belongs to; "query deadline exceeded" used to say neither. runQuery is a batch of one, so query, test and describe behave as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * query: read snapshots in place, pass the memory limit through The worker copied every column of every snapshot into its in-memory database before running anything. Past the 1 GB memory cap that copy spills to disk, and every model then runs at disk speed. Snapshots are now views read in place, so a model scans only the columns it uses. File access is allowed for exactly the declared snapshot paths, set before external access is shut; DuckDB refuses to widen the list or reopen access afterwards, and a file beside a snapshot stays out of reach. CHAINPLOT_QUERY_MEMORY_LIMIT never reached the worker: its environment is stripped to PATH, HOME and LANG, so the documented override did nothing. The parent now reads and validates it and passes it in the request; a value that is not a size is refused by name. On a project with about 4M event rows behind five models and 19 queries, build went from 345 s to 12 s at the default 1 GB, with identical results. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
1 parent cbbd28b commit d0c1913

6 files changed

Lines changed: 432 additions & 153 deletions

File tree

‎docs/capabilities.md‎

Lines changed: 6 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -75,7 +75,7 @@ visible in review rather than only in production.
7575
| Chains per project | 1 | `schemas/project.schema.json` (`maxItems`) |
7676
| Contract addresses | 20 | `schemas/project.schema.json` (`maxItems`) |
7777
| Blocks per approved run | 100_000, `policy.block_budget` to change. A block count, not a duration: 100k blocks is about 14 days on Ethereum (12 s blocks), 2.3 days on Base (2 s), 7 hours on Arbitrum One (0.25 s) — set it for the chain you index | `src/plan/generate.ts` (`DEFAULT_BLOCK_BUDGET`) |
78-
| Query deadline | 60 s | `src/query/runQuery.ts` (`DEADLINE_MS`, SIGKILL) |
78+
| Query deadline | 60 s per query, and 60 s for loading snapshots and building models, which a build does once for all its queries | `src/query/runQuery.ts` (`DEADLINE_MS`, SIGKILL) |
7979
| DuckDB memory | 1 GiB, spills to a temp dir | `src/query/workerMain.ts` (`MEMORY_LIMIT`) |
8080
| Returned rows | 10_000, `policy.row_limit` to change | `src/project/limits.ts`, enforced in `src/query/workerMain.ts` (the reader stops at the limit) |
8181
| RPC job wall clock | 30 min, resumable | `src/ingest/rindexer/runBounded.ts` |
@@ -86,8 +86,11 @@ visible in review rather than only in production.
8686
| `fork` per-request timeout | 30 s | `src/fork/fetchGuard.ts` (`FORK_LIMITS`) |
8787
| `fork` redirect hops | 0 | `src/fork/fetchGuard.ts` (refused outright) |
8888

89-
`CHAINPLOT_QUERY_MEMORY_LIMIT` overrides the memory figure; the query still
90-
spills to disk rather than failing when it goes over.
89+
`CHAINPLOT_QUERY_MEMORY_LIMIT` overrides the memory figure, as a size such as
90+
`2GB` or `1536MB`; anything else is refused by name. The query still spills to
91+
disk rather than failing when it goes over. The default stays small because a
92+
forked recipe is untrusted; a project with millions of rows behind its models
93+
builds far faster with 2–3 GB, where the memory is there to give.
9194

9295
The row limit bounds the *download*, not the rendering. The table is
9396
virtualised — 5,000 rows put 26 in the DOM — so a wide result no longer

‎docs/security.md‎

Lines changed: 17 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -16,24 +16,32 @@ than the file format.
1616
## Containment
1717

1818
Query execution happens in a forked child process (`src/query/workerMain.ts`),
19-
never in the CLI process:
19+
never in the CLI process. `build` runs all of a release's queries in one such
20+
process: snapshots load and models build once, then each query runs in turn
21+
against the same session.
2022

2123
| Control | Where |
2224
|---|---|
2325
| Separate process, env stripped to `PATH`/`HOME`/`LANG` | `src/query/runQuery.ts` |
2426
| In-memory DuckDB; no database file on disk | `workerMain.ts` |
27+
| File reads allowed for exactly the declared snapshots (`allowed_paths`), not their directories | `workerMain.ts` |
2528
| Extension autoinstall and autoload disabled | `workerMain.ts` |
2629
| `enable_external_access=false` **before any project SQL runs** | `workerMain.ts` |
2730
| Single-SELECT admission control, via DuckDB's parser | `src/query/sqlGuard.ts` |
28-
| 60 s deadline, SIGKILL on expiry | `runQuery.ts` |
31+
| 60 s deadline for loading and models, then 60 s per query; SIGKILL on expiry, naming the query | `runQuery.ts` |
2932
| Row limit enforced by stopping the reader, not by truncating after | `workerMain.ts` |
3033

31-
Ordering matters and is the part that was wrong before 2026-09-15. Snapshots
32-
are read first, because `read_parquet` needs filesystem access. External
33-
access is then disabled, and only after that are models materialized and the
34-
query run. Models are project-supplied SQL like any other, so they must land
34+
Ordering matters and is the part that was wrong before 2026-09-15. The
35+
snapshot files are allowlisted first, by exact path, and external access is
36+
then disabled; DuckDB refuses both to widen that list and to re-enable access
37+
afterwards. Each snapshot is a view read in place, so a model scans only the
38+
columns it uses rather than a copy of every column held in memory. Only after
39+
that are models materialized and the queries run. Models are project-supplied SQL like any other, so they must land
3540
on the closed side of that door; DuckDB does not allow external access to be
36-
re-enabled within a session.
41+
re-enabled within a session, so the door stays shut for every query in the
42+
batch, not only the first. Sharing the session gives one query nothing over
43+
another: each is still admitted only as a single SELECT, which cannot change
44+
the session the next one runs in, and all of them come from the same recipe.
3745

3846
### Admission control
3947

@@ -70,7 +78,8 @@ so a value cannot carry markup or a scheme into an `href`.
7078
to the bucket can serve a consistent, hostile release. **Fork only from
7179
buckets you would trust with the data itself.**
7280
- **Denial of service by a hostile recipe.** Bounded, not eliminated: a forked
73-
query gets 60 s, a row limit, and a 1 GiB memory cap that spills to a temp
81+
query gets 60 s, as does loading the snapshots and building its models, plus a
82+
row limit and a 1 GiB memory cap that spills to a temp
7483
directory rather than failing. A release can still make your build slow, and
7584
can still fill that temp directory.
7685
- **Secrets you place inside the recipe directories.** The source bundle is an

‎src/publish/writeRelease.ts‎

Lines changed: 23 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ import type { CommandError } from "../cli/envelope.js";
66
import { loadProject } from "../project/load.js";
77
import { validateProject } from "../project/validate.js";
88
import { topoSortModels } from "../project/modelGraph.js";
9-
import { runQuery } from "../query/runQuery.js";
9+
import { runQueries } from "../query/runQuery.js";
1010
import { isComplete, requiredEnd } from "../ingest/coverage.js";
1111
import { readCoverageFile, segmentsFor } from "../ingest/coverageStore.js";
1212
import { lastProvenCompleteBlock } from "../ingest/coverage.js";
@@ -178,7 +178,7 @@ export async function buildRelease(
178178
);
179179
}
180180

181-
// Models materialize in dependency order for every query.
181+
// Models materialize in dependency order, once per build.
182182
const modelOrder = topoSortModels(models);
183183
const modelSql: { id: string; sql: string }[] = modelOrder.map((id) => {
184184
const model = models.find((m) => m.id === id)!;
@@ -198,29 +198,36 @@ export async function buildRelease(
198198
const files: string[] = [];
199199

200200
try {
201-
// Queries (with models materialized first).
202-
for (const query of queries) {
201+
// Queries: one isolated session for the whole release. Snapshots load and
202+
// models build once, then every query runs against them in turn.
203+
//
204+
// Every dataset is in scope, not just the declared one, so a model over
205+
// dataset A can feed a query on dataset B and a query can join across
206+
// datasets. `query.dataset` still names the provenance.
207+
const planned = queries.map((query) => {
203208
const dataset = datasetById.get(query.dataset);
204209
if (!dataset) {
205210
throw error("validation", `unknown dataset: ${query.dataset}`, {
206211
resource_id: query.id,
207212
pointer: "/queries",
208213
});
209214
}
210-
const sqlPath = path.resolve(projectDir, query.file);
211-
const data = await runQuery({
212-
sql: fs.readFileSync(sqlPath, "utf8"),
213-
// Every dataset is in scope, not just the declared one. Models are
214-
// materialized into each query's session, so loading one table meant a
215-
// model over dataset A failed every query on dataset B — which made
216-
// models unusable in any multi-dataset project. It also lets a query
217-
// join across datasets. `query.dataset` still names the provenance.
218-
tables: allTables,
215+
const sql = fs.readFileSync(path.resolve(projectDir, query.file), "utf8");
216+
return { query, dataset, sql };
217+
});
218+
const results = await runQueries({
219+
tables: allTables,
220+
models: modelSql,
221+
queries: planned.map(({ query, sql }) => ({
222+
id: query.id,
223+
sql,
219224
rawAmountColumns: rawAmountNames(query.raw_amount_columns),
220225
rowLimit: rowLimitFor(project),
221-
models: modelSql,
222-
});
226+
})),
227+
});
223228

229+
for (const [i, { query, dataset, sql }] of planned.entries()) {
230+
const data = results[i]!;
224231
const rel = path.join("results", `${query.id}.json`);
225232
writeJson(path.join(staging, rel), {
226233
schema_version: 1,
@@ -231,7 +238,7 @@ export async function buildRelease(
231238
columns: decorateColumns(data.columns, query.raw_amount_columns),
232239
rows: data.rows,
233240
snapshot: dataset.snapshot,
234-
query_digest: sha256(fs.readFileSync(sqlPath, "utf8")),
241+
query_digest: sha256(sql),
235242
raw_amount_columns: rawAmountNames(query.raw_amount_columns),
236243
});
237244
files.push(rel);

0 commit comments

Comments
 (0)