Skip to content

Commit e0b42a8

Browse files
authored
perf(webapp): avoid unindexed fileId scan in get-background-worker-by-version (#4245)
## What The `GET /api/v1/projects/:projectRef/background-workers/:envSlug/:version` endpoint loaded each file's tasks through the nested `files.tasks` relation. Prisma resolves that as a separate query: ```sql SELECT id, slug, "fileId" FROM "BackgroundWorkerTask" WHERE "fileId" IN (...) ``` `BackgroundWorkerTask.fileId` is not indexed — the FK constraint exists, but Postgres does not auto-create an index for foreign keys — so on a large table this can only run as a sequential scan, which gets progressively slower as the table grows and was observed taking minutes per call in production. The loader already loads every task for the worker via `tasks: true`, which uses the indexed `workerId` relation, and those rows already include `fileId`. This PR groups task slugs by `fileId` in memory from that already-loaded data and drops the `files.tasks` include entirely. ## Behavior change (latent bug fix) The response shape is unchanged, but there is a semantic correction for **source files reused across worker versions** (files are de-duplicated by `@@unique([projectId, contentHash])`, so one file row can be linked to many workers). - **Before:** `file.tasks` came from the `BackgroundWorkerFile.tasks` relation, i.e. *every* `BackgroundWorkerTask` with that `fileId` — across all workers sharing the file. So a worker's manifest could list tasks it doesn't actually have. - **After:** `file.tasks` is grouped from the queried worker's own tasks, so it reflects only that worker version's tasks. Verified on a local DB: 460 files are referenced by tasks from more than one worker; of 6819 (worker, file) pairs, 6 differ — all one file where the old union leaked a task slug (`cancellation-test`) into worker versions that never had it. The new per-worker behavior is the correct one for a worker-version manifest. (Thanks to the automated review for flagging this.) ## Analysis Captured the exact SQL before/after by instrumenting Prisma against real data (a worker with 62 files): - **Before:** 5 statements, including the `WHERE "fileId" IN (...)` scan. - **After:** 4 statements; the `fileId` query is gone and the other four are identical. EXPLAIN of the two access paths: ``` Before WHERE "fileId" IN (...) Seq Scan on "BackgroundWorkerTask" Filter: ("fileId" = ANY (...)) -- reads the whole table, scales with table size After WHERE "workerId" IN (...) Index Scan using "BackgroundWorkerTask_workerId_slug_key" Index Cond: ("workerId" = ...) -- bounded by matching rows, scale-independent ``` No new index is required: the `workerId` access path is already covered by the existing `BackgroundWorkerTask_workerId_slug_key` unique index. ## Testing - `pnpm run typecheck --filter webapp` passes. - Query capture + EXPLAIN performed against a local database seeded with real worker/file/task data.
1 parent 703a6dc commit e0b42a8

2 files changed

Lines changed: 21 additions & 10 deletions

File tree

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,6 @@
1+
---
2+
area: webapp
3+
type: improvement
4+
---
5+
6+
Speed up retrieving a background worker by version. The endpoint no longer runs a slow lookup that scanned the full task table for large deployments; it now reuses data it already loads, so the response is the same but returns much faster.

apps/webapp/app/routes/api.v1.projects.$projectRef.background-workers.$envSlug.$version.ts

Lines changed: 15 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -45,22 +45,27 @@ export async function loader({ params, request }: LoaderFunctionArgs) {
4545
},
4646
include: {
4747
tasks: true,
48-
files: {
49-
include: {
50-
tasks: {
51-
select: {
52-
slug: true,
53-
},
54-
},
55-
},
56-
},
48+
files: true,
5749
},
5850
});
5951

6052
if (!backgroundWorker) {
6153
return json({ error: "Background worker not found" }, { status: 404 });
6254
}
6355

56+
// Group task slugs by fileId from the already-loaded tasks (which are fetched
57+
// via the indexed workerId relation) instead of loading files.tasks, which
58+
// queries BackgroundWorkerTask by the unindexed fileId column.
59+
const taskSlugsByFileId = new Map<string, Set<string>>();
60+
for (const task of backgroundWorker.tasks) {
61+
if (!task.fileId) {
62+
continue;
63+
}
64+
const slugs = taskSlugsByFileId.get(task.fileId) ?? new Set<string>();
65+
slugs.add(task.slug);
66+
taskSlugsByFileId.set(task.fileId, slugs);
67+
}
68+
6469
return json({
6570
id: backgroundWorker.friendlyId,
6671
version: backgroundWorker.version,
@@ -82,7 +87,7 @@ export async function loader({ params, request }: LoaderFunctionArgs) {
8287
filePath: file.filePath,
8388
contentHash: file.contentHash,
8489
contents: decompressContent(file.contents),
85-
tasks: Array.from(new Set(file.tasks.map((task) => task.slug))),
90+
tasks: Array.from(taskSlugsByFileId.get(file.id) ?? []),
8691
})),
8792
});
8893
} catch (error) {

0 commit comments

Comments
 (0)