Conversation
flamboh
added this pull request to stack #113
September 25, 2026 10:02
Add compute_weighted for (address, weight) pairs of either family, generic over MaadAddress like compute. Duplicate addresses sum their weights; non-finite or non-positive weights are rejected. Prefix validity and path pruning keep using distinct-address counts, while moments, D2 and D1 entropy use summed weights. Weighted results carry structure and dimensions only. The unweighted path keeps its integer power table and prefix-count layout, and its JSON output is byte-identical to the previous release on 20 real windows. Extend validate_maad.py with --weighted (ADDR,MEASURE CSV, oracle --csv --meas-col 1) and add a matching netflow-db maad --weighted flag; both combine with --ipv6.
Walk prefix levels top-down holding only a parent and a child level, so MAAD memory is linear in the address count instead of address count times prefix levels. compute_measures evaluates the distinct-address result and any number of weight columns over one sort and one walk, because validity and path pruning depend only on the distinct addresses. Weighted structure uses an exact per-q power table for small integral masses. Unweighted JSON is byte-identical to the previous layer on the conformance fixtures and 20 real windows, and weighted JSON is byte-identical to the previous estimator on 42 packet, byte and IPv6 inputs.
…lues across threads Weighted structure maps each moment's masses to indices into the sorted distinct masses, so every q computes one powf per distinct mass. Structure rows for different q values are independent and now run in parallel on the caller's rayon pool. Output stays byte-identical: unweighted on the fixtures and 20 real windows, weighted on 42 packet, byte and IPv6 inputs.
…ference compute_weighted now evaluates only its weighted result instead of building and discarding the distinct-address result, so the zero-weight fallback in the pipeline computes the unweighted measure once. Entries sort stably by address only, so repeated addresses sum their weights in input order for every column and compute_measures matches compute_weighted exactly. A weighted ordered-map reference, shared with the unweighted one, checks IPv4 with duplicates and IPv6.
flamboh
force-pushed
the
maad/05-weighted-estimator
branch
from
September 27, 2026 05:17
804d55d to
b7aff01
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note
🤖 Claude Opus 5.5 on behalf of Oliver
ELI5
Teaches the MAAD analysis to weight each address by how much traffic it sent (packets or bytes), instead of counting every address once.
Why
Address-count MAAD misses volume anomalies. In our study, one bogus flow claiming 465 G packets was invisible to address-count MAAD but drove packet-weighted D1 to 0. A source-address spike showed the opposite pattern. The two views complement each other.
Flows to exercise
cargo build --release -p atlantis-netflow-db), an upstream MAAD checkout, and anaddress,packetsCSV. Runpython3 scripts/local/validate_maad.py --weighted --rust target/release/netflow-db --haskell /path/to/MAAD name=addr_packets.csv. Expected: every structure row and D0/D1/D2 match within the validator's default 1e-10 tolerance.0, a negative weight, orNaNand runnetflow-db maad --weighted. Expected: the command exits non-zero and names the offending address.netflow-db maadwithout--weightedproduces byte-identical JSON to the base branch.Decisions and edge cases
--csv --meas-col. Distinct-address counts decide prefix validity and pruning. Summed weights drive moments, D2, and D1 entropy. Weights on duplicate addresses are summed.10.0.0.0and10.0.0.128with equal weights give D2 = 1/17 whether the weights are 1, 1e160, or 1e-200. Before this fix, 1e160 gave NaN and 1e-200 gave 0. Ordinary inputs keep the raw path, so unit weights still reproduce the unweighted result exactly.Follow-up
#112 covers the integration work: handling zero-packet flows, storing weighted results, and showing them in the dashboard. This PR only adds the estimator and the CLI/validator surface.
Verification
cargo fmt --all --check,cargo clippy --workspace --all-targets --all-features --locked -- -D warnings, andcargo test -p atlantis-netflow-db maad. The tests cover:mainon fixtures and 20 real windows. Weighted output matched the oracle to about 1e-15.Made by Claude Opus 5.5 (with Opus 5.5 subagents) via Claude Code.