feat: add 10-minute granularity - #107
Merged
Merged
Conversation
Roll up pairs of local five-minute buckets into half-open 10m windows with the same additive coverage, union and MAAD semantics as 30m/1h. Extend CHECK constraints, verify rollup parity, compare and export, and bump the product schema, table and pipeline contract versions so 10m products never mix with older databases. Export now validates against the shared product schema instead of a stale copy.
Accept 10m/10min in API params, schemas and coverage zero-fill, add the 10 min picker option, and let adaptive range selection choose 10min for ranges up to eight days (the 5min point budget over twice the range). Drizzle stores granularity as a TS-only text enum, so no SQL migration is generated; the local SQLite CHECK constraints now admit 10m.
Remove build_rollups, whose 10m windows used UTC div_euclid instead of the local-time aggregate bounds the pipeline uses, together with its tests, the unused getBucketStartQuery and the stale StatsTable::schema_version copy of the table versions. Hour-chart clicks now drill into a seven-day 10min window, matching adaptive drag selection, and the adaptive-granularity docstring states the real cutoffs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note
🤖 Claude Opus 5.5 on behalf of Oliver
ELI5
Adds a 10-minute zoom level next to the existing 5-minute, 30-minute, hourly and daily views.
Why
30 minutes is too coarse for multi-day investigations and 5 minutes is too noisy. This is layer 1 of the MAAD stack, which ends in one fresh reprocess.
What changes
10mgranularity: half-open, local-time aligned, exactly two 5m children, and the same additive rollup, coverage and zero-fill rules as 30m/1h. Every stats table (including MAAD andbucket_coverage) gets 10m rows.Granularity, CHECK constraints, rollup scheduling,verifyrollup parity,exportandcompareunderstand10m. Contract and table versions bump, so this needs a fresh product database.build_rollups,getBucketStartQuery,StatsTable::schema_version).Cost (accepted)
A normal day goes from 361 to 505 stored buckets per scope (288 × 5m + 144 × 10m + 48 × 30m + 24 × 1h + 1 × 1d), which is +40%. Each extra bucket needs its own address unions and MAAD run. Pipeline runs were about 30–40% slower in wall-clock time on the same day and machine. That cost comes with the feature: review found no duplicated work. The owner accepted it.
How to review
Setup: copy
datasets.json.exampletodatasets.json, pointroot_pathat any nfcapd tree, and build the nfdump fork (docs/user/setup-pipeline.md)../scripts/netflow-db.sh pipeline --nfdump target/nfdump/libexec/nfdump --dataset example --start-date <day> --end-date <day> --database-path /tmp/10m.sqlite. Expected:SELECT granularity, COUNT(DISTINCT bucket_start) FROM bucket_coverage GROUP BY 1reports 144 for10mon a normal day../scripts/netflow-db.sh verify /tmp/10m.sqlite --require-rollup-parity. Expected: OK.bun run dev:web, open the dataset, choose 10 min, then click an hour on an hourly chart. Expected: the drilldown opens at 10 min.The top layer of the stack covers the real-day end-to-end check.
Edge cases and decisions
ten_minute_bounds_pair_local_five_minute_buckets_across_dst_transitions). This PR does not change the convention.compareagainst an older reference database accepts candidate-only10mrows.Verification
bun run format,bun run lint,bun run typecheck,bun run test:web,bun run test:db(includingten_minute_rollups.rsand the DST test).Made by Claude Opus 5.5 (with Opus 5.5 subagents) via Claude Code.