Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 26 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,8 +178,22 @@ stream plus the latest published read-model. MCP-only readers pull the semantic
index and compact analysis projection. With `--athena-workgroup`, their trace
tools query time/stream-pruned raw event rows directly and do not download
`trace.json` or mirror raw chunks.
A bounded stream registry and per-stream local key cursors avoid relisting
historical chunks on each local read. The TUI builds unpublished event deltas in the background;
A bounded stream registry, per-day object watermarks, and local key cursors
avoid relisting unchanged historical chunks. A reader scans pre-watermark
history once, then lists only days whose published high-water key advanced.
New uploads use each event's UTC day, while
`event-partitions/<stream>.json` records the event-time range of every physical
day and `event-partitions/<stream>/track.<day>.json` records each immutable
object's range. Athena combines projected stream/day partitions with its hidden
`$path` column, so a wide legacy capture day can prune unrelated objects.
Once a stream index exists, readers include its legacy days and objects missing
from the object index conservatively. A stream with no partition metadata falls
back to physical days overlapping the requested window; initialize or backfill
the metadata before relying on delayed historical event-time lookups. No raw
object migration is required. A one-time metadata-only range scan is recommended
before broad historical MCP use; it writes only these small indexes and does
not copy, rewrite, or delete JSONL. The TUI builds
unpublished event deltas in the background;
`synty build` does the same explicitly, while `search` warns if raw events are
newer than the published index. One tokened machine scrapes GitHub for everyone.

Expand Down Expand Up @@ -238,12 +252,17 @@ MCP pulls the published semantic index and compact analysis projection before
serving and refreshes them on a background thread; it never mirrors the raw
event lake. In Athena mode it deliberately omits the legacy `trace.json` blob.
Each trace request is a read-only `SELECT`, partition-pruned by stream and day,
limited to seven days, 50,000 events, 64 MiB of returned envelopes, a 50-second
query timeout, and the workgroup's 20 GiB scan cutoff. Limit hits fail closed
and ask the caller for a narrower time/machine/source/operation filter.
limited to seven days, 50,000 events, 64 MiB of returned envelopes, one shared
45-second Athena budget across all queries needed by the request, 1,000 raw
object paths per query, and the workgroup's 20 GiB scan cutoff. Success, error,
timeout, and limit outcomes use the standard metrics block. Limit hits fail
closed and ask the caller for a narrower time/machine/source/operation filter
or a metadata-only object-range backfill.
`/health` reports transport liveness and `/ready` waits for the semantic index
and analysis projection, plus both dispatchers (`trace.json` is required only
in local-projection mode). Analysis tools are
and analysis projection, a compatible remote read-model format, plus both
dispatchers (`trace.json` is required only in local-projection mode). Status
and both health endpoints report the newest indexed raw-event timestamp, the
published-model timestamp/format, and the last bucket metadata check. Analysis tools are
serialized on a one-slot dispatcher so concurrent first loads cannot multiply
memory. HTTP work is bounded by separate semantic and analysis queues, a
120-second response deadline, 32 in-flight requests, 64 live connections,
Expand Down
1 change: 0 additions & 1 deletion deploy/aws/mcp-reader.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,6 @@ Parameters:
GlueTable:
Type: String
Default: raw_events

Resources:
ReaderRole:
Type: AWS::IAM::Role
Expand Down
56 changes: 39 additions & 17 deletions docs/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ flowchart LR
B --> P["Published next-plaid read model"]
E --> B
P -->|"download + mmap"| MCP["Read-only MCP"]
S3 -->|"external table; no migration"| G["Glue Data Catalog"]
S3 -->|"external table; no migration"| G["Glue raw_events"]
G --> A["Bounded Athena SELECT"]
A -->|"bounded raw envelopes"| F["Rust trace fold"]
F --> MCP
Expand All @@ -62,11 +62,13 @@ jobs. The MCP pod does not need a trace projection job or `trace.json` on its
volume. Local CLI/TUI operation stays self-contained and can use the local
projection offline.

The raw-table overlay is the zero-copy bootstrap, not a columnar rewrite:
Athena still scans the selected JSONL object bytes. If daily raw volume reaches
the workgroup cutoff, a follow-up can write a separate immutable Parquet
projection partitioned for session lookup while leaving the authoritative raw
prefix untouched. This path deliberately creates no compaction job or migration.
The raw-table overlay is a zero-copy query path, not a columnar rewrite:
Athena still scans the selected JSONL object bytes. A compact per-stream
partition-range index maps event-time windows onto physical days, including
legacy capture-day chunks; per-day object-range indexes then prune immutable
files through Athena's hidden `$path` column. Per-day object high-water keys let
local readers list only partitions that changed after a one-time compatibility
scan. Neither metadata path rewrites the authoritative objects.

## Engine

Expand Down Expand Up @@ -217,12 +219,21 @@ runs on CI or a server without a developer machine.
artifacts in the background; it does not mirror the raw event lake. With
`--athena-workgroup`, trace calls issue only bounded `SELECT` statements over
the existing S3 event chunks and fold the returned rows in Rust. The backend
discovers injected stream partitions from `event-streams/`, caps the time
window at seven days, rows at 50,000, returned bytes at 64 MiB, query time at
50 seconds, and relies on a workgroup scan cutoff as the final cost guard.
discovers injected stream partitions from `event-streams/`, resolves
event-time windows through `event-partitions/<stream>.json`, caps the time
window at seven days, rows at 50,000, returned bytes at 64 MiB, and all
Athena work for one request at a shared 45-second budget. Queries include at
most 1,000 exact raw-object `$path` values. A workgroup scan cutoff is the
final cost guard. Indexed streams include declared legacy days and physically
present objects missing from per-object metadata conservatively; a completely
unindexed stream falls back to the requested physical-day span.
`/health` remains a liveness check, while `/ready` requires the semantic
index, compact analysis projection, and both dispatchers; local-projection
mode additionally requires `trace.json`. Analysis calls use a
index, compact analysis projection, a compatible bucket read-model format,
and both dispatchers; local-projection mode additionally requires
`trace.json`. Health and status expose bucket raw/model freshness; the
unauthenticated HTTP endpoints return only a generic metadata-error indicator,
while detailed provider diagnostics stay in server logs and protected status.
Analysis calls use a
serialized one-slot dispatcher so concurrent first loads cannot multiply
memory or block semantic search. Each dispatcher has a bounded queue; HTTP
clients have a 120-second response deadline and a per-client
Expand Down Expand Up @@ -327,12 +338,19 @@ bucket and drops into the viewer, so a paste goes from nothing to tracking.
## Storage layout (bucket)

```text
events/<stream>/chunks/<track-day>/<range-hash>.jsonl
events/<stream>/chunks/track.<event-day>/<range-hash>.jsonl
immutable append deltas (source of truth);
stream = edge-<machine>-<source>, so many
trackers' files coexist without collision
event-streams/<stream> immutable bounded discovery registry;
readers continue each stream by key cursor
event-streams/<stream> immutable bounded discovery registry
event-partitions/<stream>.json physical day → complete event-time range
plus day → highest immutable object key;
legacy/unindexed days remain unconditional
query candidates (metadata-only; mutable)
event-partitions/<stream>/track.<day>.json
immutable object → event-time range;
Athena uses exact $path pruning;
readers continue each day by key cursor
members/<machine>/activation.json immutable init access marker (no session data)
embeddings/<hash[..2]>/<hash>.emb content-addressed f16 vectors (write-once)
summaries/<kh[..2]>/<kh>-<ihash>.json per-(unit, input-hash) LLM summaries
Expand Down Expand Up @@ -361,9 +379,13 @@ fleet model is **no designated builder**: every tracker pushes events; whoever
opens a local viewer pulls all raw streams and the published read-model, then
contributes a build. MCP-only readers pull semantic/analysis artifacts without
bucket write access or a raw-history mirror; an S3 reader can query trace rows
through the read-only Glue/Athena overlay. Local commands can inspect and
rebuild from every machine, while semantic results cover the latest published
build and warn when newer raw chunks are pending.
through the read-only Glue/Athena overlay. The partition-range index makes
delayed event-time windows queryable without moving JSONL: new writers publish
by event UTC day, while days created by old writers remain conservative until
an optional metadata-only range backfill proves their exact day and object
coverage. Local
commands can inspect and rebuild from every machine, while semantic results
cover the latest published build and warn when newer raw chunks are pending.
Write-once stores are the collaboration primitive: a viewer encodes and
summarizes only what no other machine has (pending lists shuffle per machine,
so concurrent viewers split the work). The lease only prevents duplicate index
Expand Down
Loading
Loading