Skip to content

feat: add 10-minute granularity - #107

Merged
flamboh merged 3 commits into
mainfrom
maad/01-10m-buckets
Sep 28, 2026
Merged

flamboh merged 3 commits into
mainfrom
maad/01-10m-buckets

Conversation

@flamboh

@flamboh flamboh commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Note

🤖 Claude Opus 5.5 on behalf of Oliver

ELI5

Adds a 10-minute zoom level next to the existing 5-minute, 30-minute, hourly and daily views.

Why

30 minutes is too coarse for multi-day investigations and 5 minutes is too noisy. This is layer 1 of the MAAD stack, which ends in one fresh reprocess.

What changes

  • New 10m granularity: half-open, local-time aligned, exactly two 5m children, and the same additive rollup, coverage and zero-fill rules as 30m/1h. Every stats table (including MAAD and bucket_coverage) gets 10m rows.
  • Rust Granularity, CHECK constraints, rollup scheduling, verify rollup parity, export and compare understand 10m. Contract and table versions bump, so this needs a fresh product database.
  • Web: granularity parsing, pickers, auto-selection (ranges up to ~8 days default to 10 min), and hour-click drilldowns now open 10 min so they match drag selection.
  • Removed dead helpers (build_rollups, getBucketStartQuery, StatsTable::schema_version).

Cost (accepted)

A normal day goes from 361 to 505 stored buckets per scope (288 × 5m + 144 × 10m + 48 × 30m + 24 × 1h + 1 × 1d), which is +40%. Each extra bucket needs its own address unions and MAAD run. Pipeline runs were about 30–40% slower in wall-clock time on the same day and machine. That cost comes with the feature: review found no duplicated work. The owner accepted it.

How to review

Setup: copy datasets.json.example to datasets.json, point root_path at any nfcapd tree, and build the nfdump fork (docs/user/setup-pipeline.md).

  1. ./scripts/netflow-db.sh pipeline --nfdump target/nfdump/libexec/nfdump --dataset example --start-date <day> --end-date <day> --database-path /tmp/10m.sqlite. Expected: SELECT granularity, COUNT(DISTINCT bucket_start) FROM bucket_coverage GROUP BY 1 reports 144 for 10m on a normal day.
  2. ./scripts/netflow-db.sh verify /tmp/10m.sqlite --require-rollup-parity. Expected: OK.
  3. bun run dev:web, open the dataset, choose 10 min, then click an hour on an hourly chart. Expected: the drilldown opens at 10 min.

The top layer of the stack covers the real-day end-to-end check.

Edge cases and decisions

  • DST follows the existing civil-time convention. The 5-minute iterator walks local wall-clock times, so the repeated fall-back hour is not revisited. Both transition days therefore keep the existing 5m counts: 276 on spring-forward and 288 on fall-back. 10m pairs those children, which gives 138 10m buckets on spring-forward and 144 on fall-back. Tests pin this (ten_minute_bounds_pair_local_five_minute_buckets_across_dst_transitions). This PR does not change the convention.
  • compare against an older reference database accepts candidate-only 10m rows.

Verification

  • Automated (CI and local): bun run format, bun run lint, bun run typecheck, bun run test:web, bun run test:db (including ten_minute_rollups.rs and the DST test).
  • Manual: steps 1–3 above on any local capture day.

Made by Claude Opus 5.5 (with Opus 5.5 subagents) via Claude Code.

Roll up pairs of local five-minute buckets into half-open 10m windows
with the same additive coverage, union and MAAD semantics as 30m/1h.
Extend CHECK constraints, verify rollup parity, compare and export, and
bump the product schema, table and pipeline contract versions so 10m
products never mix with older databases. Export now validates against
the shared product schema instead of a stale copy.
Accept 10m/10min in API params, schemas and coverage zero-fill, add the
10 min picker option, and let adaptive range selection choose 10min for
ranges up to eight days (the 5min point budget over twice the range).
Drizzle stores granularity as a TS-only text enum, so no SQL migration
is generated; the local SQLite CHECK constraints now admit 10m.
Remove build_rollups, whose 10m windows used UTC div_euclid instead of the local-time aggregate bounds the pipeline uses, together with its tests, the unused getBucketStartQuery and the stale StatsTable::schema_version copy of the table versions. Hour-chart clicks now drill into a seven-day 10min window, matching adaptive drag selection, and the adaptive-granularity docstring states the real cutoffs.
@flamboh
flamboh added this pull request to stack #113 September 25, 2026 10:02
@flamboh flamboh changed the title maad/01 10m buckets feat: add 10-minute granularity Sep 25, 2026
@flamboh
flamboh merged commit e0d1be9 into main Sep 28, 2026
3 checks passed
@flamboh
flamboh deleted the maad/01-10m-buckets branch September 28, 2026 19:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant