diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 7ba69c2..aa2c562 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -11,7 +11,7 @@ jobs: strategy: fail-fast: false matrix: - python-version: ["3.13"] + python-version: ["3.10", "3.11", "3.12", "3.13"] steps: - uses: actions/checkout@v4 diff --git a/README.md b/README.md index 8ffd3ba..ee47ad9 100644 --- a/README.md +++ b/README.md @@ -1,203 +1,236 @@ -# ACK — Agent Context Kernel +
-**Intelligent memory middleware for AI coding agents.** ACK is a transparent -proxy that sits between you and any terminal-based AI agent (Aider, Claude -Code, etc.). It intercepts the agent's output in real time, **prunes the -noisy, repetitive, high-token blobs** (stack traces, build-error walls, log -floods) down to their actionable signal, and **archives the full, untouched -output** to a local, full-text-searchable database. +# ACK: Agent Context Kernel -The agent — and your context window — only sees the summary. The full log is -always one `ack search` away. +**Intelligent memory middleware for terminal AI coding agents.** + +ACK is a transparent proxy between you and any terminal AI agent (Aider, Claude Code, and others). It collapses the noisy, high-token output that pollutes the context window (stack traces, build-error walls, log floods) into its actionable signal, and archives the full untouched stream to a local searchable database. + +Your agent sees the summary. The full log is one `ack search` away. [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE) -![Python](https://img.shields.io/badge/python-3.13%2B-blue.svg) +![Python](https://img.shields.io/badge/python-3.10%2B-blue.svg) ![Platform](https://img.shields.io/badge/platform-POSIX-lightgrey.svg) +![Tests](https://img.shields.io/badge/tests-97%20passing-brightgreen.svg) + +
+ +![ACK compressing a log flood and a traceback in real time](./assets/context_kernel.gif) + + +--- + +## Results + +Measured on a realistic corpus of tracebacks, Rust/GCC builds, and log floods. No LLM in the loop, reproducible, and run in CI on every push. + +| Suite | Token compression | Signature fidelity | +| --- | :---: | :---: | +| Python tracebacks | **93%** | **100%** | +| Rust + GCC builds | **91%** | **100%** | +| Log floods | **99%** | **100%** | +| Mixed realistic | **80%** | **100%** | +| **Global** | **~95%** | **100%** | + +**62,772 of 66,158 corpus tokens removed with zero loss of error signatures.** Every exception class, error code, and `file:line` survives compression. + +```bash +poetry run python scripts/run_benchmarks.py +``` --- -## The problem +## Why it matters + +A single failed test can dump a 60-frame traceback. A broken build repeats the same `error[E0308]` once per crate. A runaway logger floods 400 identical lines. Those tokens do two costly things: -Long AI coding sessions rot. A single failed test can dump a 60-frame -traceback; a broken build can repeat the same `error[E0308]` once per crate; a -runaway logger can flood 400 identical lines. Every one of those tokens: +1. **They drown the signal** the model reasons over. The real information (one exception, two user frames, a unique error code) is tiny. +2. **They invalidate the prompt cache.** Inject a 2,000-token wall and the cached prefix shifts, so the next turn is recomputed from scratch: slower and more expensive. -- **pollutes the context window** the model reasons over, and -- **invalidates the prompt cache**, so the next turn is slower and costlier. +ACK keeps the signal, drops the noise, and never loses the original. The pruning is regex/structural, deterministic, and free. There is no second LLM. -The signal in that wall of text is tiny — one exception type, a couple of user -frames, a unique error code. ACK keeps the signal and drops the noise, with no -new LLM in the loop. +--- ## How it works -ACK spawns your agent inside a real **pseudo-terminal (PTY)**, so interactive -prompts, colors, and cursor movement behave exactly as if you'd run the agent -directly. It multiplexes I/O in a single `select()` loop: +ACK spawns your agent inside a real pseudo-terminal (`openpty` + `fork` + `setsid`), so prompts, colors, and cursor movement behave exactly as if you launched it directly. A single `select()` loop multiplexes your keyboard, the agent, the screen, and the database. + +![Arch diagram](./assets/context-kernel.drawio.png) + + +--- + +Every flushed buffer is handled in order: + +1. **Small output passes through verbatim** (below `--prune-threshold`, default 30 lines). +2. **Interactive prompts are never pruned.** A `[Y/n]`, `(yes/no)`, or `Continue?` tail means the agent is waiting, so you always see it. +3. **Large blobs run through ordered pruners.** First match wins: a hit injects a compact summary, a miss passes through. +4. **Everything is persisted.** The summary is a view, never a deletion. + +When a pruner fires you see exactly what happened: ``` - ┌─────────────────────────── ACK ───────────────────────────┐ - you ──stdin──┤ │ - │ keystrokes ─────────────────────────────────► master_fd │──► agent - │ (PTY) │ (in PTY) - you ◄─stdout─┤ agent output ─► buffer ─► pruners ─► stdout │◄── agent - │ │ │ - │ └─► full raw text ─► SQLite + FTS5 │ - └────────────────────────────────────────────────────────────┘ +[ACK] Compressed 63 lines → 4 lines (full log stored in DB) +── Traceback ── + at /app/src/views/checkout.py:201 in post + at /app/src/models/cart.py:88 in checkout + ↳ KeyError: 'card_token' ``` -- Output below a configurable line threshold passes through **verbatim**. -- Interactive prompts (`[Y/n]`, `Continue?`, `Enter key:`) are **never** pruned — - the agent is waiting and you need to see the question. -- Larger blobs are routed through ordered **Pruners** (first match wins). A - match injects a compact summary into the live stream; a miss passes through. -- Every chunk — pruned or not — is persisted to SQLite so nothing is lost. +--- ## Install -ACK is POSIX-only (it uses `pty`, `fork`, and `termios`) and requires -**Python 3.13+**. +POSIX only (uses `pty`, `fork`, `termios`), Python 3.10+. macOS and Linux native; on Windows use WSL. ```bash -# with Poetry (recommended for development) -git clone https://github.com/vaibhavtripathi/context-kernel -cd context-kernel -poetry install +git clone https://github.com/Vaibhavtripathi7/context_kernel +cd context_kernel +poetry install # or: pip install . poetry run ack --help - -# or with pip -pip install . -ack --help ``` -`tree-sitter` is an optional dependency used by `ack toc` for accurate parsing; -if it isn't available, ACK falls back to a regex parser automatically. +`tree-sitter` powers `ack toc`; without it, ACK falls back to a regex parser automatically. + +--- ## Quick start ```bash -# Wrap any agent — ACK is transparent +# Wrap any agent. Everything after `--` is the agent command. ack run -- aider --model gpt-4o ack run -- claude --dangerously-skip-permissions -# Tune when pruning kicks in (output lines before pruners activate) -ack run --prune-threshold 50 -- aider +# Search the full archive of every session (FTS5 syntax, BM25 ranked) +ack search "ImportError OR ModuleNotFoundError" -# Search the full archived output of every session (FTS5 syntax) -ack search "ImportError" -ack search "ModuleNotFoundError OR FileNotFoundError" -ack search "context window" --session abc123 - -# Print a symbol table-of-contents for a file instead of dumping the whole thing +# Symbol map of a file instead of dumping the whole thing ack toc context_kernel/core/orchestrator.py -# List recent sessions +# Recent sessions ack sessions ``` -### Try it in 10 seconds (no agent required) +**Try it in 10 seconds, no agent required.** A bundled script emits a 200-line flood plus a deep traceback: ```bash -# A bundled script that emits a 200-line log flood + a real traceback. -# Watch ACK collapse it live, then search the full archived output. ack run -- python examples/crashing_agent.py ack search "KeyError" ``` +The flood collapses to a one-line frequency table, the traceback to its two user frames plus the exception, and the original is still searchable. When the agent exits, ACK prints exactly what it saved: + +``` +[ACK] Session summary + Chunks intercepted : 8 + Chunks pruned : 4 + Tokens saved : 3,438 (~97% of pruned output) + Est. cost saved : $ 0.01 (at $3/M input tokens) + Elapsed : 6s +``` + +--- + ## What gets compressed -The built-in `ShellPruner` handles the three biggest context-window offenders: +The built-in `ShellPruner` targets the three biggest offenders. The original bytes on disk are never modified. | Input | Strategy | Result | | --- | --- | --- | -| **Python tracebacks** | Keep the exception line + 3 deepest *user* frames; drop stdlib/venv noise. Chained exceptions each summarized. | `── Traceback ──` with the frames that matter | -| **Rust / GCC / Clang errors** | Deduplicate by error code / diagnostic line; surface counts + unique list. | `[Rust build: 32 errors → 8 unique]` | -| **Repetitive log floods** | Detect when one line dominates (≥60%); replace with a frequency table. | `× 400 WARNING:root:retrying...` | +| **Python tracebacks** | Keep the exception + 3 deepest user frames; drop stdlib/venv noise. Chained exceptions each summarized. | `── Traceback ──` with only the frames that matter. | +| **Rust / GCC / Clang errors** | Deduplicate by error code or diagnostic line; surface counts and the unique set. | `[Rust build: 32 errors → 8 unique]` | +| **Log floods** | Detect when one line dominates (≥ 60%) and replace it with a frequency table. | `× 400 WARNING:root:retrying…` | -The original bytes are never modified on disk — only the *live stream* the -agent sees is compressed. +--- -## Benchmark results +## Use cases -ACK ships a reproducible A/B benchmark (no LLM calls, free to run) that -measures token reduction and **signature fidelity** — whether every actionable -error signature (exception class, error code, file:line) survives compression. +- **Long refactors that hit failing tests.** The agent sees the exception and the 3 user frames it needs, not the 60-frame async wall. +- **Large Rust / C++ builds.** One type error re-emitted per crate becomes "8 unique errors," not 200 lines. +- **Noisy services.** Repetitive retry/heartbeat spam collapses to a frequency table so the rare line stays visible. +- **A searchable audit log.** `ack search` finds output across all sessions, weeks later, even after it scrolled off screen. -```bash -poetry run python scripts/run_benchmarks.py -``` +--- -Representative output across the bundled corpus (tracebacks, Rust/GCC builds, -log floods, mixed realistic output): +## Command reference -| Suite | Compression | Signature fidelity | +**`ack run -- `** + +| Option | Default | Description | | --- | --- | --- | -| Python Traceback | 93% | 100% | -| Rust + GCC Builds | 91% | 100% | -| Log Flood | 99% | 100% | -| Mixed Realistic | 80% | 100% | -| **Global** | **~95%** | **100%** | +| `--db PATH` | `~/.local/share/ack/kernel.db` | SQLite database path. | +| `--prune-threshold N` | `30` | Output lines buffered before pruners activate. | +| `--no-annotate` | off | Suppress the `[ACK]` banners. | +| `--tui` | off | Experimental Textual stats dashboard. | + +**`ack search ""`** takes an FTS5 expression (`AND`/`OR`/`NOT`, prefix, phrase) and ranks results by BM25. Options: `--session`, `--limit` (default 20), `--db`. -≈ 62,000 of 66,000 corpus tokens removed with zero loss of error signatures. +**`ack toc `** prints a symbol table-of-contents (Python today). + +**`ack sessions`** lists recent sessions. Options: `--limit`, `--db`. + +--- ## Architecture +No second LLM, no pipes (PTY only), no heavy database. + | Module | Responsibility | | --- | --- | -| `core/orchestrator.py` | PTY spawn (`openpty` + `fork` + `setsid`), `select()` I/O loop, buffer management, stream injection, terminal raw/cooked mode, SIGWINCH forwarding. | -| `memory/storage.py` | SQLite in WAL mode + FTS5 content-table (trigger-synced, zero row duplication), BM25 search, per-session stats. | -| `memory/pager.py` | tree-sitter (with regex fallback) symbol mapper — gives an agent a file's table-of-contents so it can page in just one function. | -| `pruners/base.py` | `BasePruner` ABC: `matches()` fast gate + `compress()`. | -| `pruners/shell_pruner.py` | The built-in traceback / build-error / log-flood pruner. | -| `cli.py` | `click` CLI (`run`, `search`, `toc`, `sessions`) + experimental Textual stats dashboard. | +| `core/orchestrator.py` | PTY spawn, `select()` loop, buffering, stream injection, raw/cooked mode, `SIGWINCH` forwarding. | +| `memory/storage.py` | SQLite (WAL) + FTS5 content table kept in sync by triggers (text stored once), BM25 search, per-session stats. | +| `memory/pager.py` | tree-sitter symbol mapper with regex fallback. | +| `pruners/base.py` | `BasePruner` ABC: cheap `matches()` gate + `compress()`. | +| `pruners/shell_pruner.py` | Built-in traceback / build-error / log-flood pruner. | +| `cli.py` | `click` CLI plus the experimental Textual dashboard. | -The full session archive lives at `~/.local/share/ack/kernel.db` by default -(override with `--db`). +WAL mode lets the orchestrator write while readers query without blocking. The archive lives at `~/.local/share/ack/kernel.db` (override with `--db`). + +--- ## Writing your own pruner ```python -from typing import Optional from context_kernel.pruners.base import BasePruner, PrunerMetadata + class MyPruner(BasePruner): - metadata = PrunerMetadata(name="my_pruner", description="...") + metadata = PrunerMetadata(name="my_pruner", description="Collapses my tool's output") def matches(self, text: str) -> bool: - # Cheap gate — called on every flush. - return "my-pattern" in text + return "my-pattern" in text # cheap gate, runs on every flush - def compress(self, text: str) -> Optional[str]: - # Return a summary string, or None to pass through verbatim. - return summarize(text) + def compress(self, text: str) -> str | None: + return summarize(text) # return a summary, or None to pass through ``` -Register it on the `Orchestrator` (`pruners=[MyPruner(), ShellPruner()]`) or at -runtime via `Orchestrator.add_pruner()`. Pruners run in order; the first to -return a non-`None` summary wins. +Register with `pruners=[MyPruner(), ShellPruner()]` on the `Orchestrator`. Pruners run in order; first non-`None` wins, so list specific pruners before general ones. + +--- ## Development ```bash poetry install -poetry run pytest # full suite (unit + integration) +poetry run pytest # 97 tests (unit + integration) poetry run pytest -m "not integration" # fast unit tests only poetry run python scripts/run_benchmarks.py -poetry run ruff check context_kernel -poetry run mypy context_kernel +poetry run ruff check context_kernel scripts +poetry run mypy context_kernel # strict ``` -## Limitations & roadmap - -- **POSIX only** — relies on `pty`/`fork`/`termios`. No Windows support (WSL works). -- **`ack toc` is Python-only** today; the tree-sitter integration is structured - to add more languages. -- **`--tui` dashboard is experimental**: Textual and the PTY interceptor both - want to own the terminal, so in `--tui` mode the agent runs non-interactively - and output is captured but not mirrored live. Headless mode (`ack run`) is the - recommended path. A future split-pane (tmux/Pilot) approach can lift this. -- **Heuristic pruning**: pruners use regex/structural heuristics, not an LLM, by - design — they're fast, deterministic, and free. +CI runs ruff, mypy `--strict`, the full suite, and the benchmark on every push and PR. + +--- + +## Limitations + +- **POSIX only** (uses `pty`/`fork`/`termios`); use WSL on Windows. +- **`ack toc` is Python-only** today; tree-sitter is structured to add languages. +- **`--tui` is experimental:** Textual and the PTY interceptor both want the terminal, so the agent runs non-interactively in that mode. Plain `ack run` is the recommended interactive path. +- **Heuristic by design.** Pruners are fast, deterministic, and free; a new output format needs a new pruner (easy to write, see above). + +--- ## License diff --git a/assets/context-kernel.drawio.png b/assets/context-kernel.drawio.png new file mode 100644 index 0000000..d5433df Binary files /dev/null and b/assets/context-kernel.drawio.png differ diff --git a/assets/context_kernel.gif b/assets/context_kernel.gif new file mode 100644 index 0000000..91d6ea1 Binary files /dev/null and b/assets/context_kernel.gif differ diff --git a/context_kernel/cli.py b/context_kernel/cli.py index 025e3d1..c108bf5 100644 --- a/context_kernel/cli.py +++ b/context_kernel/cli.py @@ -6,6 +6,7 @@ """ from __future__ import annotations +import re import sys import threading import time @@ -23,6 +24,55 @@ from .memory.storage import StorageEngine from .pruners.shell_pruner import ShellPruner +_USD_PER_MILLION_INPUT_TOKENS = 3.0 + +_ANSI_OSC = re.compile(r"\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)") +_ANSI_CSI = re.compile(r"\x1b\[[0-9;?]*[ -/]*[@-~]") +_ANSI_OTHER = re.compile(r"\x1b[@-Z\\-_]") +_CTRL_CHARS = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]") + + +def _sanitize(text: str) -> str: + """Strip escape sequences and control chars from stored output before echo.""" + text = _ANSI_OSC.sub("", text) + text = _ANSI_CSI.sub("", text) + text = _ANSI_OTHER.sub("", text) + return _CTRL_CHARS.sub("", text) + + +def _print_session_summary( + storage: StorageEngine, + session_id: str, + stats: OrchestratorStats, +) -> None: + """Print a one-glance summary of what ACK saved this session (to stderr).""" + db_stats = storage.stats(session_id) + chunks = db_stats["total_entries"] + pruned = db_stats["pruned_entries"] + raw_pruned_tokens = db_stats["tokens_saved"] + saved = stats.tokens_saved + + if chunks == 0: + return + + elapsed = max(1, int(time.monotonic() - stats.session_start)) + pct = (saved / raw_pruned_tokens * 100) if raw_pruned_tokens else 0.0 + cost = saved / 1_000_000 * _USD_PER_MILLION_INPUT_TOKENS + rate = f"{_USD_PER_MILLION_INPUT_TOKENS:g}" + + lines = [ + click.style("[ACK] Session summary", fg="cyan", bold=True), + f" Chunks intercepted : {chunks:>8,}", + f" Chunks pruned : {pruned:>8,}", + f" Tokens saved : {saved:>8,} (~{pct:.0f}% of pruned output)", + f" Est. cost saved : ${cost:>7.2f} (at ${rate}/M input tokens)", + f" Elapsed : {elapsed:>7}s", + ] + try: + click.echo("\n" + "\n".join(lines), err=True) + except (BrokenPipeError, OSError): + pass + class StatsPanel(Static): stats: reactive[OrchestratorStats] = reactive(OrchestratorStats()) @@ -209,10 +259,14 @@ def cmd_run( ) if tui: - exit_code = AckDashboard(orchestrator=orch).run() - sys.exit(exit_code or 0) + exit_code = AckDashboard(orchestrator=orch).run() or 0 else: - sys.exit(orch.run()) + exit_code = orch.run() + if not no_annotate: + _print_session_summary(storage, session.session_id, orch.stats) + + storage.close() + sys.exit(exit_code) @main.command(name="search") @@ -259,7 +313,7 @@ def cmd_search( pruned = " [pruned]" if row["was_pruned"] else "" click.echo(click.style(f"[{ts}] session={sid_short}… type={etype}{pruned}", fg="cyan")) - content: str = row["compressed_summary"] or row["raw_content"] + content = _sanitize(row["compressed_summary"] or row["raw_content"]) preview = content[:300].strip() if len(content) > 300: preview += "\n …" @@ -308,7 +362,7 @@ def cmd_sessions(limit: int, db: Path | None) -> None: click.echo("─" * 90) for row in rows: sid = row["session_id"] - cmd = row["agent_command"][:40] + cmd = _sanitize(row["agent_command"])[:40] ts = time.strftime("%Y-%m-%d %H:%M:%S", time.localtime(row["started_at"])) ended = " ✓" if row["ended_at"] else " …" click.echo(f"{sid} {ts} {cmd}{ended}") diff --git a/context_kernel/core/orchestrator.py b/context_kernel/core/orchestrator.py index 8984755..1d47a09 100644 --- a/context_kernel/core/orchestrator.py +++ b/context_kernel/core/orchestrator.py @@ -10,11 +10,13 @@ import fcntl import os import pty +import re import select import signal import struct import sys import termios +import threading import time import tty from collections.abc import Callable @@ -26,10 +28,10 @@ _DIM = "\033[2m" _CYAN = "\033[36m" -_BOLD = "\033[1m" _RESET = "\033[0m" -_SELECT_TIMEOUT = 0.08 +_ANSI_ESC = re.compile(r"\x1b\[[0-9;]*[mGKHFJA-Z]") +_TRACEBACK_MARKER = "Traceback (most recent call last):" @dataclass @@ -38,6 +40,7 @@ class OrchestratorConfig: buffer_flush_timeout: float = 0.15 read_chunk_bytes: int = 8192 annotate_injections: bool = True + max_buffer_bytes: int = 262144 @dataclass @@ -86,7 +89,7 @@ def __init__( self.command = command self.session_id = session_id self.storage = storage - self.pruners:list[BasePruner] = pruners or [] + self.pruners: list[BasePruner] = pruners or [] self.config = config or OrchestratorConfig() self.stats_callback = stats_callback self.text_callback: Callable[[str], None] | None = None @@ -102,9 +105,6 @@ def __init__( def stats(self) -> OrchestratorStats: return self._stats - def add_pruner(self, pruner: BasePruner) -> None: - self.pruners.append(pruner) - def run(self) -> int: """Spawn the agent, block until it exits, and return its exit code. @@ -114,7 +114,7 @@ def run(self) -> int: self._child_pid = child_pid self._enter_raw_mode() - self._install_sigwinch_handler() + self._install_signal_handlers() try: return self._io_loop(child_pid) @@ -153,27 +153,34 @@ def _spawn_in_pty(self) -> tuple[int, int]: os.close(slave_fd) os.close(master_fd) - os.execvp(self.command[0], self.command) + try: + os.execvp(self.command[0], self.command) + except OSError as exc: + os.write(2, f"ack: cannot run {self.command[0]!r}: {exc.strerror}\n".encode()) os._exit(127) os.close(slave_fd) return master_fd, child_pid def _io_loop(self, child_pid: int) -> int: + """Pump I/O between the user and the child until the child exits. + + While nothing is buffered the loop blocks in select(); it only polls on + the flush timeout while it is still holding data to emit, so it stays at + ~0% CPU when idle. stdin is dropped from the watch set once it reaches + EOF so a closed/piped input never spins the loop. + """ assert self._master_fd is not None master_fd = self._master_fd stdin_fd = sys.stdin.fileno() stdout_fd = sys.stdout.fileno() exit_code = 0 + watched = [master_fd, stdin_fd] while True: + timeout = self.config.buffer_flush_timeout if self._buffer else None try: - rlist, _, _ = select.select( - [master_fd, stdin_fd], - [], - [], - _SELECT_TIMEOUT, - ) + rlist, _, _ = select.select(watched, [], [], timeout) except InterruptedError: continue except (ValueError, OSError): @@ -201,13 +208,13 @@ def _io_loop(self, child_pid: int) -> int: os.write(master_fd, keys) except OSError: pass + else: + watched = [master_fd] - elapsed_since_data = time.monotonic() - self._last_data_monotonic - if ( - self._buffer - and elapsed_since_data >= self.config.buffer_flush_timeout - ): - self._flush_buffer(stdout_fd, force=True) + if self._buffer: + elapsed_since_data = time.monotonic() - self._last_data_monotonic + if elapsed_since_data >= self.config.buffer_flush_timeout: + self._flush_buffer(stdout_fd, force=True) if self._buffer: self._flush_buffer(stdout_fd, force=True) @@ -229,9 +236,10 @@ def _accumulate(self, chunk: bytes) -> None: def _flush_buffer(self, stdout_fd: int, *, force: bool = False) -> None: """Emit the buffered output, pruning it first if it qualifies. - force controls only *whether to emit now* (silence timeout / EOF), - never *whether to prune*: output below pruning_threshold_lines and - interactive prompts always pass through verbatim. + force controls only whether to emit now (silence timeout / EOF), never + whether to prune: output below pruning_threshold_lines and interactive + prompts always pass through verbatim. An unfinished traceback is held + (up to max_buffer_bytes) so it prunes as one unit across PTY reads. """ if not self._buffer: return @@ -251,6 +259,14 @@ def _flush_buffer(self, stdout_fd: int, *, force: bool = False) -> None: self._buffer.append(raw) return + if ( + not force + and len(raw) < self.config.max_buffer_bytes + and self._is_incomplete_traceback(text) + ): + self._buffer.append(raw) + return + if below_threshold or self._text_is_prompt(text): self._emit(stdout_fd, raw) self._persist(text, pruned=False) @@ -260,19 +276,18 @@ def _flush_buffer(self, stdout_fd: int, *, force: bool = False) -> None: summary: str | None = None for pruner in self.pruners: - if pruner.matches(text): - result = pruner.compress(text) - if result is not None: - summary = result - saved = max(0, (len(text) - len(summary)) // 4) - self._stats.tokens_saved += saved - self._stats.total_pruner_hits += 1 - break + result = pruner.compress(text) + if result is not None: + summary = result + saved = max(0, (len(text) - len(summary)) // 4) + self._stats.tokens_saved += saved + self._stats.total_pruner_hits += 1 + break if summary is not None: self._persist(text, pruned=True, summary=summary) injection = self._format_injection(summary, line_count) - injected = injection.encode("utf-8") + injected = self._terminal_newlines(injection).encode("utf-8") self._stats.total_bytes_injected += len(injected) self._emit(stdout_fd, injected) _display = injection @@ -287,6 +302,17 @@ def _flush_buffer(self, stdout_fd: int, *, force: bool = False) -> None: if self.stats_callback is not None: self.stats_callback(self._stats) + def _terminal_newlines(self, text: str) -> str: + """Convert text ACK generates itself to CRLF while the terminal is raw. + + Raw mode disables the terminal's NL->CRLF output mapping, so a bare + ``\\n`` would leave the cursor in the same column (the "staircase" + effect). Child passthrough already carries CRLF from its own PTY. + """ + if self._saved_tty is None: + return text + return text.replace("\r\n", "\n").replace("\n", "\r\n") + def _emit(self, fd: int, data: bytes) -> None: offset = 0 while offset < len(data): @@ -321,6 +347,24 @@ def _format_injection(self, summary: str, original_lines: int) -> str: ) return banner + body + def _is_incomplete_traceback(self, text: str) -> bool: + """True if text holds a traceback still streaming its frames. + + A finished traceback ends in a non-indented exception line; while frames + are still arriving the last non-blank line is an indented frame line, or + the header itself. + """ + if _TRACEBACK_MARKER not in text: + return False + clean = _ANSI_ESC.sub("", text) + nonblank = [ln for ln in clean.splitlines() if ln.strip()] + if not nonblank: + return False + last = nonblank[-1] + if _TRACEBACK_MARKER in last: + return True + return last[:1].isspace() + def _tail_is_prompt(self, chunk: bytes) -> bool: tail = chunk[-200:] return any(p in tail for p in _PROMPT_BYTES) @@ -345,13 +389,27 @@ def _restore_terminal(self) -> None: termios.tcsetattr(sys.stdin.fileno(), termios.TCSAFLUSH, self._saved_tty) self._saved_tty = None - def _install_sigwinch_handler(self) -> None: - def _handler(signum: int, frame: object) -> None: # noqa: ARG001 - if self._master_fd is not None and sys.stdout.isatty(): - rows, cols = self._get_terminal_size() - self._set_winsize(self._master_fd, rows, cols) + def _install_signal_handlers(self) -> None: + if threading.current_thread() is not threading.main_thread(): + return + signal.signal(signal.SIGWINCH, self._on_sigwinch) + signal.signal(signal.SIGTERM, self._on_terminate) + signal.signal(signal.SIGHUP, self._on_terminate) + + def _on_sigwinch(self, signum: int, frame: object) -> None: # noqa: ARG002 + if self._master_fd is not None and sys.stdout.isatty(): + rows, cols = self._get_terminal_size() + self._set_winsize(self._master_fd, rows, cols) - signal.signal(signal.SIGWINCH, _handler) + def _on_terminate(self, signum: int, frame: object) -> None: # noqa: ARG002 + """Restore the terminal and forward the signal to the child before exit.""" + self._restore_terminal() + if self._child_pid is not None: + try: + os.killpg(self._child_pid, signum) + except OSError: + pass + os._exit(128 + signum) @staticmethod def _get_terminal_size() -> tuple[int, int]: diff --git a/context_kernel/memory/pager.py b/context_kernel/memory/pager.py index e06a95b..7f85be7 100644 --- a/context_kernel/memory/pager.py +++ b/context_kernel/memory/pager.py @@ -93,24 +93,6 @@ def map_file(self, path: Path) -> FileSymbolMap: self._cache[cache_key] = fmap return fmap - def page_symbol(self, path: Path, symbol_name: str) -> str | None: - """Return the source text of one named symbol, or None if not found.""" - fmap = self.map_file(path) - entry = next((s for s in fmap.symbols if s.name == symbol_name), None) - if entry is None: - return None - - source_lines = path.read_text(encoding="utf-8", errors="replace").splitlines() - return "\n".join(source_lines[entry.start_line - 1 : entry.end_line]) - - def toc(self, path: Path) -> str: - return self.map_file(path).to_toc() - - def invalidate(self, path: Path) -> None: - stale = [k for k in self._cache if k.startswith(str(path.resolve()))] - for k in stale: - del self._cache[k] - def _parse_with_tree_sitter( self, path: Path, diff --git a/context_kernel/pruners/shell_pruner.py b/context_kernel/pruners/shell_pruner.py index fb25e98..9e83fc6 100644 --- a/context_kernel/pruners/shell_pruner.py +++ b/context_kernel/pruners/shell_pruner.py @@ -132,7 +132,7 @@ def _compress_python_tracebacks(self, text: str) -> str: (i for i, ln in enumerate(lines) if _PY_TRACEBACK_HDR.search(ln)), 0, ) - preamble = "\n".join(lines[:first_tb]).strip() + preamble = self._condense_preamble(lines[:first_tb]) result: list[str] = [] if preamble: @@ -140,6 +140,20 @@ def _compress_python_tracebacks(self, text: str) -> str: result.extend(parts) return "\n\n".join(result) + def _condense_preamble(self, preamble_lines: list[str]) -> str: + """Summarise output that precedes a traceback in the same buffer. + + The common "logs, then a crash" pattern means a log flood often shares a + buffer with the traceback. Collapse a repetitive preamble to a frequency + table instead of dumping it verbatim; keep short, varied preambles as-is. + """ + text = "\n".join(preamble_lines).strip() + if not text: + return "" + if self._is_highly_repetitive(text): + return self._compress_repetitive(text) + return text + def _compress_rust_errors(self, text: str) -> str: lines = text.splitlines() error_lines = [ln for ln in lines if _RUST_ERROR_HDR.match(ln)] diff --git a/examples/crashing_agent.py b/examples/crashing_agent.py index 9aec7db..768831e 100644 --- a/examples/crashing_agent.py +++ b/examples/crashing_agent.py @@ -1,24 +1,60 @@ #!/usr/bin/env python3 """Demo "agent" that emits the noisy output ACK is meant to compress. -Run it through ACK to watch the pruners fire: +Run it through ACK to watch both pruners fire: ack run -- python examples/crashing_agent.py -It prints a 200-line log flood (repetition pruner) and then raises a real -exception (traceback pruner). Run it directly to see the uncompressed output. +It simulates an agent doing a build-and-test run: a few status lines pass through +untouched, a 200-line log flood collapses to a frequency table, and a deep +test-failure traceback (two app frames buried under framework/stdlib noise) +collapses to the exception plus the user frames. The short pauses between phases +keep it readable (and make a nice screen recording); run it directly, without +ACK, to see the full uncompressed output for comparison. """ +import sys +import time -def level_three() -> None: - raise KeyError("target key not found") +def say(message: str, pause: float = 0.9) -> None: + """Print a status line the way an agent would, then pause briefly.""" + print(message) + sys.stdout.flush() + time.sleep(pause) -def level_two() -> None: - for _ in range(200): +def emit_log_flood(n_lines: int = 200) -> None: + for _ in range(n_lines): print("DEBUG: allocating buffer space...") - level_three() + sys.stdout.flush() + + +def emit_buried_traceback(n_noise_frames: int = 40) -> None: + """Relay a realistic traceback: 2 app frames under a wall of framework noise.""" + frames = [ + "Traceback (most recent call last):", + ' File "/app/src/views/checkout.py", line 201, in post', + " order = cart.checkout(user_id=user.pk, payment=payload)", + ' File "/app/src/models/cart.py", line 88, in checkout', + " charge_result = gateway.charge(amount, card_token)", + ] + for i in range(n_noise_frames): + frames.append( + ' File "/home/user/.venv/lib/python3.11/site-packages/django/core/' + f'handlers/base.py", line {200 + i}, in _get_response' + ) + frames.append(" response = wrapped_callback(request, *args, **kwargs)") + frames.append("KeyError: 'card_token'") + print("\n".join(frames)) + sys.stdout.flush() if __name__ == "__main__": - level_two() + say("[agent] Building project and running the test suite...", 1.2) + say("[agent] Compiling dependencies (verbose output follows)...", 1.0) + emit_log_flood() + time.sleep(1.4) + say("[agent] Compile done. Running pytest...", 1.2) + emit_buried_traceback() + time.sleep(1.0) + say("[agent] Test run failed (see compressed traceback above).", 0.6) diff --git a/poetry.lock b/poetry.lock index 599c63f..9ce8a3c 100644 --- a/poetry.lock +++ b/poetry.lock @@ -28,6 +28,25 @@ files = [ ] markers = {main = "platform_system == \"Windows\"", dev = "sys_platform == \"win32\""} +[[package]] +name = "exceptiongroup" +version = "1.3.1" +description = "Backport of PEP 654 (exception groups)" +optional = false +python-versions = ">=3.7" +groups = ["dev"] +markers = "python_version == \"3.10\"" +files = [ + {file = "exceptiongroup-1.3.1-py3-none-any.whl", hash = "sha256:a7a39a3bd276781e98394987d3a5701d0c4edffb633bb7a5144577f82c773598"}, + {file = "exceptiongroup-1.3.1.tar.gz", hash = "sha256:8b412432c6055b0b7d14c310000ae93352ed6754f70fa8f7c34141f91c4e3219"}, +] + +[package.dependencies] +typing-extensions = {version = ">=4.6.0", markers = "python_version < \"3.13\""} + +[package.extras] +test = ["pytest (>=6)"] + [[package]] name = "iniconfig" version = "2.3.0" @@ -278,6 +297,7 @@ files = [ librt = {version = ">=0.8.0", markers = "platform_python_implementation != \"PyPy\""} mypy_extensions = ">=1.0.0" pathspec = ">=1.0.0" +tomli = {version = ">=1.1.0", markers = "python_version < \"3.11\""} typing_extensions = [ {version = ">=4.6.0", markers = "python_version < \"3.15\""}, {version = ">=4.14.0", markers = "python_version >= \"3.15\""}, @@ -377,10 +397,12 @@ files = [ [package.dependencies] colorama = {version = ">=0.4", markers = "sys_platform == \"win32\""} +exceptiongroup = {version = ">=1", markers = "python_version < \"3.11\""} iniconfig = ">=1" packaging = ">=20" pluggy = ">=1.5,<2" pygments = ">=2.7.2" +tomli = {version = ">=1", markers = "python_version < \"3.11\""} [package.extras] dev = ["argcomplete", "attrs (>=19.2)", "hypothesis (>=3.56)", "mock", "requests", "setuptools", "xmlschema"] @@ -470,6 +492,64 @@ typing-extensions = ">=4.4.0,<5.0.0" [package.extras] syntax = ["tree-sitter (>=0.20.1,<0.21.0)", "tree-sitter-languages (==1.10.2)"] +[[package]] +name = "tomli" +version = "2.4.1" +description = "A lil' TOML parser" +optional = false +python-versions = ">=3.8" +groups = ["dev"] +markers = "python_version == \"3.10\"" +files = [ + {file = "tomli-2.4.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:f8f0fc26ec2cc2b965b7a3b87cd19c5c6b8c5e5f436b984e85f486d652285c30"}, + {file = "tomli-2.4.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:4ab97e64ccda8756376892c53a72bd1f964e519c77236368527f758fbc36a53a"}, + {file = "tomli-2.4.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:96481a5786729fd470164b47cdb3e0e58062a496f455ee41b4403be77cb5a076"}, + {file = "tomli-2.4.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:5a881ab208c0baf688221f8cecc5401bd291d67e38a1ac884d6736cbcd8247e9"}, + {file = "tomli-2.4.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:47149d5bd38761ac8be13a84864bf0b7b70bc051806bc3669ab1cbc56216b23c"}, + {file = "tomli-2.4.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:ec9bfaf3ad2df51ace80688143a6a4ebc09a248f6ff781a9945e51937008fcbc"}, + {file = "tomli-2.4.1-cp311-cp311-win32.whl", hash = "sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049"}, + {file = "tomli-2.4.1-cp311-cp311-win_amd64.whl", hash = "sha256:5ee18d9ebdb417e384b58fe414e8d6af9f4e7a0ae761519fb50f721de398dd4e"}, + {file = "tomli-2.4.1-cp311-cp311-win_arm64.whl", hash = "sha256:c2541745709bad0264b7d4705ad453b76ccd191e64aa6f0fc66b69a293a45ece"}, + {file = "tomli-2.4.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:c742f741d58a28940ce01d58f0ab2ea3ced8b12402f162f4d534dfe18ba1cd6a"}, + {file = "tomli-2.4.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:7f86fd587c4ed9dd76f318225e7d9b29cfc5a9d43de44e5754db8d1128487085"}, + {file = "tomli-2.4.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9"}, + {file = "tomli-2.4.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:136443dbd7e1dee43c68ac2694fde36b2849865fa258d39bf822c10e8068eac5"}, + {file = "tomli-2.4.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5e262d41726bc187e69af7825504c933b6794dc3fbd5945e41a79bb14c31f585"}, + {file = "tomli-2.4.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:5cb41aa38891e073ee49d55fbc7839cfdb2bc0e600add13874d048c94aadddd1"}, + {file = "tomli-2.4.1-cp312-cp312-win32.whl", hash = "sha256:da25dc3563bff5965356133435b757a795a17b17d01dbc0f42fb32447ddfd917"}, + {file = "tomli-2.4.1-cp312-cp312-win_amd64.whl", hash = "sha256:52c8ef851d9a240f11a88c003eacb03c31fc1c9c4ec64a99a0f922b93874fda9"}, + {file = "tomli-2.4.1-cp312-cp312-win_arm64.whl", hash = "sha256:f758f1b9299d059cc3f6546ae2af89670cb1c4d48ea29c3cacc4fe7de3058257"}, + {file = "tomli-2.4.1-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:36d2bd2ad5fb9eaddba5226aa02c8ec3fa4f192631e347b3ed28186d43be6b54"}, + {file = "tomli-2.4.1-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:eb0dc4e38e6a1fd579e5d50369aa2e10acfc9cace504579b2faabb478e76941a"}, + {file = "tomli-2.4.1-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c7f2c7f2b9ca6bdeef8f0fa897f8e05085923eb091721675170254cbc5b02897"}, + {file = "tomli-2.4.1-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f3c6818a1a86dd6dca7ddcaaf76947d5ba31aecc28cb1b67009a5877c9a64f3f"}, + {file = "tomli-2.4.1-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:d312ef37c91508b0ab2cee7da26ec0b3ed2f03ce12bd87a588d771ae15dcf82d"}, + {file = "tomli-2.4.1-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:51529d40e3ca50046d7606fa99ce3956a617f9b36380da3b7f0dd3dd28e68cb5"}, + {file = "tomli-2.4.1-cp313-cp313-win32.whl", hash = "sha256:2190f2e9dd7508d2a90ded5ed369255980a1bcdd58e52f7fe24b8162bf9fedbd"}, + {file = "tomli-2.4.1-cp313-cp313-win_amd64.whl", hash = "sha256:8d65a2fbf9d2f8352685bc1364177ee3923d6baf5e7f43ea4959d7d8bc326a36"}, + {file = "tomli-2.4.1-cp313-cp313-win_arm64.whl", hash = "sha256:4b605484e43cdc43f0954ddae319fb75f04cc10dd80d830540060ee7cd0243cd"}, + {file = "tomli-2.4.1-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:fd0409a3653af6c147209d267a0e4243f0ae46b011aa978b1080359fddc9b6cf"}, + {file = "tomli-2.4.1-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:a120733b01c45e9a0c34aeef92bf0cf1d56cfe81ed9d47d562f9ed591a9828ac"}, + {file = "tomli-2.4.1-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:559db847dc486944896521f68d8190be1c9e719fced785720d2216fe7022b662"}, + {file = "tomli-2.4.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853"}, + {file = "tomli-2.4.1-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:7f94b27a62cfad8496c8d2513e1a222dd446f095fca8987fceef261225538a15"}, + {file = "tomli-2.4.1-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:ede3e6487c5ef5d28634ba3f31f989030ad6af71edfb0055cbbd14189ff240ba"}, + {file = "tomli-2.4.1-cp314-cp314-win32.whl", hash = "sha256:3d48a93ee1c9b79c04bb38772ee1b64dcf18ff43085896ea460ca8dec96f35f6"}, + {file = "tomli-2.4.1-cp314-cp314-win_amd64.whl", hash = "sha256:88dceee75c2c63af144e456745e10101eb67361050196b0b6af5d717254dddf7"}, + {file = "tomli-2.4.1-cp314-cp314-win_arm64.whl", hash = "sha256:b8c198f8c1805dc42708689ed6864951fd2494f924149d3e4bce7710f8eb5232"}, + {file = "tomli-2.4.1-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:d4d8fe59808a54658fcc0160ecfb1b30f9089906c50b23bcb4c69eddc19ec2b4"}, + {file = "tomli-2.4.1-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:7008df2e7655c495dd12d2a4ad038ff878d4ca4b81fccaf82b714e07eae4402c"}, + {file = "tomli-2.4.1-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1d8591993e228b0c930c4bb0db464bdad97b3289fb981255d6c9a41aedc84b2d"}, + {file = "tomli-2.4.1-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:734e20b57ba95624ecf1841e72b53f6e186355e216e5412de414e3c51e5e3c41"}, + {file = "tomli-2.4.1-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:8a650c2dbafa08d42e51ba0b62740dae4ecb9338eefa093aa5c78ceb546fcd5c"}, + {file = "tomli-2.4.1-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:504aa796fe0569bb43171066009ead363de03675276d2d121ac1a4572397870f"}, + {file = "tomli-2.4.1-cp314-cp314t-win32.whl", hash = "sha256:b1d22e6e9387bf4739fbe23bfa80e93f6b0373a7f1b96c6227c32bef95a4d7a8"}, + {file = "tomli-2.4.1-cp314-cp314t-win_amd64.whl", hash = "sha256:2c1c351919aca02858f740c6d33adea0c5deea37f9ecca1cc1ef9e884a619d26"}, + {file = "tomli-2.4.1-cp314-cp314t-win_arm64.whl", hash = "sha256:eab21f45c7f66c13f2a9e0e1535309cee140182a9cdae1e041d02e47291e8396"}, + {file = "tomli-2.4.1-py3-none-any.whl", hash = "sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe"}, + {file = "tomli-2.4.1.tar.gz", hash = "sha256:7c7e1a961a0b2f2472c1ac5b69affa0ae1132c39adcb67aba98568702b9cc23f"}, +] + [[package]] name = "tree-sitter" version = "0.23.2" @@ -575,5 +655,5 @@ test = ["coverage", "pytest", "pytest-cov"] [metadata] lock-version = "2.1" -python-versions = "^3.13" -content-hash = "add8f71304ebc906ff8c6c8dd3010ccfa1a990b2d8171039ce3dcea7643b06d0" +python-versions = "^3.10" +content-hash = "6d78740037a10637aa02c99d5b667f619b59b21d874ee4f0c4c2a017ee8943aa" diff --git a/pyproject.toml b/pyproject.toml index 48b5063..2dbcc0f 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -5,7 +5,7 @@ description = "ACK: Agent Context Kernel — intelligent memory middleware for A authors = ["Vaibhav Tripathi "] license = "MIT" readme = "README.md" -repository = "https://github.com/vaibhavtripathi/context-kernel" +repository = "https://github.com/Vaibhavtripathi7/context_kernel" keywords = ["llm", "ai-agents", "context-window", "pty", "developer-tools"] classifiers = [ "Development Status :: 4 - Beta", @@ -13,13 +13,16 @@ classifiers = [ "Intended Audience :: Developers", "License :: OSI Approved :: MIT License", "Operating System :: POSIX", + "Programming Language :: Python :: 3.10", + "Programming Language :: Python :: 3.11", + "Programming Language :: Python :: 3.12", "Programming Language :: Python :: 3.13", "Topic :: Software Development :: Build Tools", ] packages = [{ include = "context_kernel" }] [tool.poetry.dependencies] -python = "^3.13" +python = "^3.10" textual = "^0.65.0" click = "^8.1" # tree-sitter 0.23 ships the new one-step Language(grammar.language()) API. @@ -42,7 +45,7 @@ ack = "context_kernel.cli:main" # ── mypy ──────────────────────────────────────────────────────────────────── [tool.mypy] -python_version = "3.13" +python_version = "3.10" strict = true warn_return_any = true warn_unused_configs = true @@ -61,7 +64,7 @@ ignore_missing_imports = true # ── ruff ──────────────────────────────────────────────────────────────────── [tool.ruff] line-length = 100 -target-version = "py313" +target-version = "py310" [tool.ruff.lint] select = ["E", "F", "I", "N", "UP", "ANN", "S", "B", "C4", "SIM"] diff --git a/tests/test_cli.py b/tests/test_cli.py new file mode 100644 index 0000000..5304f02 --- /dev/null +++ b/tests/test_cli.py @@ -0,0 +1,31 @@ +"""CLI helper tests: terminal-escape sanitisation of stored output on display.""" +from __future__ import annotations + +from context_kernel.cli import _sanitize + + +class TestSanitize: + """Stored agent output carries raw escape sequences; displaying them must + never be able to reprogram the user's terminal.""" + + def test_strips_color_csi(self) -> None: + assert _sanitize("\x1b[31mred\x1b[0m text") == "red text" + + def test_strips_truecolor_csi(self) -> None: + assert _sanitize("\x1b[38;2;255;0;0mx\x1b[0m") == "x" + + def test_strips_mouse_and_app_mode_toggles(self) -> None: + evil = "\x1b[?1000h\x1b[?1006hclick\x1b[?2004h" + assert _sanitize(evil) == "click" + + def test_strips_osc_title_sequence(self) -> None: + assert _sanitize("\x1b]0;malicious title\x07ok") == "ok" + + def test_strips_lone_control_chars(self) -> None: + assert _sanitize("a\x00b\x07c\x7f") == "abc" + + def test_preserves_newlines_and_tabs(self) -> None: + assert _sanitize("line1\nline2\tend") == "line1\nline2\tend" + + def test_plain_text_unchanged(self) -> None: + assert _sanitize("KeyError: 'card_token'") == "KeyError: 'card_token'" diff --git a/tests/test_orchestrator.py b/tests/test_orchestrator.py index bec36c3..2e36625 100644 --- a/tests/test_orchestrator.py +++ b/tests/test_orchestrator.py @@ -309,6 +309,68 @@ def compress(self, text: str) -> Optional[str]: assert orch.stats.tokens_saved == (len(original_text) - len(summary_text)) // 4 +class TestIncompleteTracebackBuffering: + """ + A traceback can span several PTY reads. ACK must hold an unfinished one + (still streaming frames, no exception line yet) so it prunes as a single + unit instead of a broken half. + """ + + @staticmethod + def _frames(n: int) -> bytes: + return b"".join( + b' File "/app/module_%d.py", line %d, in fn\n do_something()\n' % (i, i) + for i in range(n) + ) + + def test_incomplete_traceback_is_held( + self, storage: StorageEngine, pipe_pair: tuple[int, int] + ) -> None: + r_fd, w_fd = pipe_pair + orch = _make_orchestrator(storage, threshold=3) + orch._buffer = [b"Traceback (most recent call last):\n" + self._frames(10)] + + orch._flush_buffer(w_fd, force=False) # type: ignore[attr-defined] + + data = _read_pipe(r_fd, timeout=0.1) + assert data == b"", "An unfinished traceback must not be emitted yet." + assert orch._buffer, "The partial traceback must remain buffered." + + def test_complete_traceback_is_flushed( + self, storage: StorageEngine, pipe_pair: tuple[int, int] + ) -> None: + r_fd, w_fd = pipe_pair + from context_kernel.pruners.shell_pruner import ShellPruner + + orch = _make_orchestrator(storage, pruners=[ShellPruner()], threshold=3) + full = ( + b"Traceback (most recent call last):\n" + + self._frames(10) + + b"ValueError: boom\n" + ) + orch._buffer = [full] + + orch._flush_buffer(w_fd, force=False) # type: ignore[attr-defined] + + data = _read_pipe(r_fd) + assert b"ValueError" in data, "A finished traceback must be emitted." + assert not orch._buffer + + def test_incomplete_traceback_flushed_when_forced( + self, storage: StorageEngine, pipe_pair: tuple[int, int] + ) -> None: + """The silence timeout / EOF (force=True) must flush even a partial.""" + r_fd, w_fd = pipe_pair + orch = _make_orchestrator(storage, threshold=3) + orch._buffer = [b"Traceback (most recent call last):\n" + self._frames(10)] + + orch._flush_buffer(w_fd, force=True) # type: ignore[attr-defined] + + data = _read_pipe(r_fd) + assert b"Traceback" in data + assert not orch._buffer + + class TestStatsTracking: """Verify the OrchestratorStats counters accumulate correctly."""