Repository navigation
chore: remove in-tree language bindings - #9
Closed
gabewillen wants to merge 264 commits into
Closed
gabewillen wants to merge 264 commits into
gabewillen wants to merge 264 commits into
Conversation
- record planum.cpp audio and segmentation scaffold commits - keep parent repo aligned with the primary runtime submodule state
- record the normalized planum contract boundary and sink seam summary - update roadmap, requirements, and state for completed plans 01, 03, and 05
- route finalized planum perception events through ProcessTextAt - add white-box bridge tests and private-source test wiring
- add a deterministic chat-area smoke target for synthetic planum events - build the private bridge in examples without changing voice session contracts
- record summary and updated execution state for phase 01 plan 06 - advance the root repo to the nested planum.cpp task commits
- New research bench tools/accumulator_label_context_bench.py that replays the scaled multimodal sequence through a faithful LiveAccumulator shadow (mu_acc incremental mean + c_t EWMA + recent window) instead of guessing static episodes and aggregating labels over a bipartite graph. - Same label DB, same 18 signals, same 18 probes as the static-graph bench for direct comparison. - Result: target label top-5 hit rate remains 0.167 (identical to static graph). Even the true live blended context cortext actually uses for retrieval and anchoring cannot rescue raw label readout when the 256-d label vectors entangle modality/generic terms with specific identities. - Paper update in 9_experimental.qmd with commands, artifacts, table, and architectural interpretation reinforcing that cross-modal reference ultimately requires either stricter label-bank filtering or a model pass over the stored payload after the accumulator has delivered the neighborhood. This experiment directly addresses the 'guessing at episodes vs. live accumulator' distinction and narrows the solution space for the label promotion problem.
- Updated implementation notes to clarify that durable ingress inputs now preserve an ordered working-memory trace independent of the `source_id` string, which is used solely for provenance and grouping. - Replaced `chat/user` and `chat/assistant` identifiers with opaque identifiers (`Gabe` and `Julie`) across various benchmark and evaluation files to enhance clarity and maintainability. - Adjusted related code to ensure consistent handling of source identifiers in processing functions and documentation, emphasizing the importance of ordering in prompt hydration and memory management. - Improved comments and documentation to reflect these changes and their implications for memory processing and retrieval.
- watch_julie_probe_stream_judge.py: defer the cortext token-savings
fail-fast until normal RAG is under real history-budget pressure
(same condition the quality gate already used). Against a 60-token
raw history at the first probe of a short window, a 50% savings
floor is unsatisfiable for any system; this killed video_window and
audio_window at milestone 1 while Cortext was winning on quality.
A token_savings_gate_deferred check records the deferral.
- judge_julie_live_run.py: the judge_prompt_fits_context_window check
estimated tokens at chars/4, but real Gemma tokenization of judge
prompts measured <= 2.54 chars/token (83,145-char prompt >= 32,767
prompt-eval tokens via Ollama). The check passed while the prompt
overflowed a 32k window, so Ollama truncated and the model returned
a bare '{' deterministically on every retry, killing the final judge
in two consecutive release runs. Add a calibrated chars/2.5
estimator used only for context-fit accounting; packet token metrics
keep the chars/4 scale shared with the benchmark.
- run_julie_release_windows.py: custom --window entries hardcoded
probe_stride/warmup_events/min_probe_rows_after_benchmark, silently
discarding the CLI flags (an 80-message window ran with a 200-event
warmup and produced 0 probes). Custom windows now inherit the flags;
validation moved after defaults are applied.
WMBaseCapacity moves from round(lerp(5,3,S) + lerp(-1,1,F)) (range [2,6], 4 at neutral knobs) to round(lerp(8,6,S) + lerp(-1,1,F)) (range [5,9], 7 at neutral). Knob directions are unchanged; only the capacity center shifts. Motivation: in the 20260609 Julie release window, Cortext matched traditional chat-RAG on judged relevance but trailed on sufficiency (2.4 vs 3.6 mean) with ~4 working-memory slots producing ~200-token packets. Capacity is the direct lever on packet completeness. To be validated by a paired rerun of the frozen early_text_image window under the same judge protocol.
Rearm consolidation drift and bound Natural active work
Release Cortext v1.2.3
Language bindings should resolve shared libraries and model chunks from cortext release assets instead of re-vendoring ~160 MiB per repo.
Track under-100 MiB model parts (no LFS) and bake them into the shared library so consumers need no download or CORTEXT_AIST_MODEL_PATH. At load the library assembles and checksum-verifies the full GGUF into the process cache. Opt out with -DCORTEXT_EMBED_AIST_MODEL=OFF or -Dembed-aist-model=false. Release packaging still publishes natives (and optional model trees) for binding installers on GitHub Releases.
Checkout the requested release tag for dispatch builds, create missing releases as drafts, validate chunk size, skip packing the destination tarball, keep top-level optimize aligned with reused natives, and verify cached model digests before sharding. Document that CORTEXT_ASSETS_DIR is a binding install layout, not a core env var. CI smoke builds opt out of model embed; ubuntu-aist keeps embed on.
Document CORTEXT_LIBRARY_PATH + reassemble for CORTEXT_AIST_MODEL_PATH instead of a non-existent CORTEXT_ASSETS_DIR. Skip packing the .sha256 sibling when --tarball lands under --output.
Embed-off CI builds still compile aist_embedded_model.cpp; leave the cache/sha/materialize helpers out of that TU so -Werror=unused-function does not fail ubuntu-native and ubuntu-sanitizers.
- Link shell32 on Windows Zig builds for SHGetFolderPathA - Move aist_embedded_model.hpp under src/ (not installed public API) - Emit .note.GNU-stack for ELF embedded blob assembly - Unique temp paths + tolerant replace for concurrent materialize - Python/JS default model bootstrap no longer forces HF download that would shadow embedded assemble-at-load natives
…age models. - prepare_embedded_aist: COFF .rdata for Windows, ELF .rodata, Mach-O const; hide blob labels (.hidden / .private_extern); GNU-stack note ELF-only - ReplaceFile only discards tmp when dest matches expected digest/size - build.zig.zon packages models/ for default embed-on Zig consumers - MSVC defaults CORTEXT_FETCH_AIST_MODEL=ON when embed is unavailable
Publish shared release assets for language bindings
Ship shared release assets and default AIST model embed from main after PR #8.
Language packages live in standalone repos (cortext.py/ts/go/dart/wasm). Keep optional N-API glue at ffi/node/addon.cpp for engine-side addon builds. Update docs, release packaging, and the web demo import path accordingly.
Contributor
Author
|
Restacked onto rewritten
|
Contributor
Author
|
Closed after |
1 of 2 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
bindings/language packages (Python, Go, JS, Dart, WASM wrappers). Language packages live in standalone repos:cortext.py,cortext.ts,cortext.go,cortext.dart,cortext.wasm.ffi/node/addon.cppso the engine can still producecortext.nodefor consumers.build_python_package.py/build_javascript_package.py; release assets still built here for shared natives/models.examples/web/cortext-wasm.js.Test plan
bindings/ffi/node/addon.cppFollow-up (out of this PR's merge, same request)
Repo rename
augmem/cortext→augmem/cortext.cpp(local + GitHub) after or with this change.