Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 50 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,12 +5,12 @@

ACK is a transparent proxy between you and any terminal AI agent (Aider, Claude Code, and others). It collapses the noisy, high-token output that pollutes the context window (stack traces, build-error walls, log floods) into its actionable signal, and archives the full untouched stream to a local searchable database.

Your agent sees the summary. The full log is one `ack search` away.
Your agent sees the signal. The full log is one `ack recall` away — for you *and* the agent itself.

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
![Python](https://img.shields.io/badge/python-3.10%2B-blue.svg)
![Platform](https://img.shields.io/badge/platform-POSIX-lightgrey.svg)
![Tests](https://img.shields.io/badge/tests-97%20passing-brightgreen.svg)
![Tests](https://img.shields.io/badge/tests-107%20passing-brightgreen.svg)

---

Expand Down Expand Up @@ -69,16 +69,27 @@ Every flushed buffer is handled in order:
3. **Large blobs run through ordered pruners.** First match wins: a hit injects a compact summary, a miss passes through.
4. **Everything is persisted.** The summary is a view, never a deletion.

When a pruner fires you see exactly what happened:
When a pruner fires you see exactly what happened — and a stable handle to page the original back:

```
[ACK] Compressed 63 lines → 4 lines (full log stored in DB)
[ACK] Compressed 63 lines → 4 lines (recall: ack #42)
── Traceback ──
at /app/src/views/checkout.py:201 in post
at /app/src/models/cart.py:88 in checkout
↳ KeyError: 'card_token'
```

### Recall: the archive is readable, not just written

Pruning that you can't undo is just lossy compression. ACK stamps every pruned injection with a handle (`ack #42`) and exposes `ack recall` as a plain shell command — so when the summary isn't enough, the exact original bytes come back on demand, by handle or by search:

```bash
ack recall 42 # page back the full log behind that banner
ack recall "card_token" # or find it by content, scoped to this session
```

Because it's an ordinary command, the **agent** can run it too: it drops the 60-frame wall, keeps working from the summary, and pulls the verbatim detail back only if it actually needs it. A [benchmark](scripts/run_recall_benchmark.py) shows this recovers every dropped detail at ~80% fewer tokens than never pruning at all.

---

## Install
Expand Down Expand Up @@ -106,6 +117,10 @@ ack run -- claude --dangerously-skip-permissions
# Search the full archive of every session (FTS5 syntax, BM25 ranked)
ack search "ImportError OR ModuleNotFoundError"

# Page a pruned log back — by its `ack #N` handle, or by content
ack recall 42
ack recall "ImportError"

# Symbol map of a file instead of dumping the whole thing
ack toc context_kernel/core/orchestrator.py

Expand All @@ -117,10 +132,10 @@ ack sessions

```bash
ack run -- python examples/crashing_agent.py
ack search "KeyError"
ack recall "KeyError" # page the full traceback back by content
```

The flood collapses to a one-line frequency table, the traceback to its two user frames plus the exception, and the original is still searchable. When the agent exits, ACK prints exactly what it saved:
The flood collapses to a one-line frequency table, the traceback to its two user frames plus the exception, and the original is still recoverable verbatim. When the agent exits, ACK prints exactly what it saved:

```
[ACK] Session summary
Expand Down Expand Up @@ -167,6 +182,8 @@ The built-in `ShellPruner` targets the three biggest offenders. The original byt

**`ack search "<query>"`** takes an FTS5 expression (`AND`/`OR`/`NOT`, prefix, phrase) and ranks results by BM25. Options: `--session`, `--limit` (default 20), `--db`.

**`ack recall <id|query>`** pages a stored log back. A numeric argument is an `ack #N` handle (exact lookup); anything else is a search, scoped to the most recent session by default. Options: `--session`, `--all` (search every session), `--limit` (default 1), `--raw` (skip escape-sequence sanitisation), `--db`.

**`ack toc <file>`** prints a symbol table-of-contents (Python today).

**`ack sessions`** lists recent sessions. Options: `--limit`, `--db`.
Expand Down Expand Up @@ -214,9 +231,10 @@ Register with `pruners=[MyPruner(), ShellPruner()]` on the `Orchestrator`. Prune

```bash
poetry install
poetry run pytest # 97 tests (unit + integration)
poetry run pytest # 107 tests (unit + integration)
poetry run pytest -m "not integration" # fast unit tests only
poetry run python scripts/run_benchmarks.py
poetry run python scripts/run_benchmarks.py # L1 compression / fidelity
poetry run python scripts/run_recall_benchmark.py # L2 needle recovery
poetry run ruff check context_kernel scripts
poetry run mypy context_kernel # strict
```
Expand All @@ -225,6 +243,30 @@ CI runs ruff, mypy `--strict`, the full suite, and the benchmark on every push a

---

## Roadmap

ACK is one piece of a larger idea: treat the context window as a scarce resource
to be managed, and keep deterministic work out of the model's way. The pruner you
see today is the first of three layers.

- **L1 — output reduction** *(shipped).* The pruners. Collapse deterministic
noise — tracebacks, build-error walls, log floods — to its signal before it
ever reaches the model.
- **L2 — memory paging** *(in progress).* The archive plus `ack recall`. Pruned
detail is recoverable on demand, so compression is never a one-way loss.
Shipped: stable `ack #N` handles and recall by id or content. Next:
- **Proactive dedup** — content-hash repeated output so the same error isn't re-paged across turns.
- **Pager narrowing** — recall just the errored function, not the whole log.
- **A context-health signal** — detect repetition and re-run-the-same-command loops as a deterministic paging trigger, instead of a fixed token threshold.
- **L3 — execution offload** *(exploring).* Run well-specified, deterministic
sub-tasks outside the model entirely and hand back only the result.

Layers compound: L1 shrinks what enters the window, L2 makes that shrink safe to
undo, L3 keeps whole tasks out of the window to begin with. Issues and PRs
against any layer are welcome — see [CONTRIBUTING](CONTRIBUTING.md).

---

## Limitations

- **POSIX only** (uses `pty`/`fork`/`termios`); use WSL on Windows.
Expand Down
75 changes: 75 additions & 0 deletions context_kernel/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
import sys
import threading
import time
from collections.abc import Callable
from pathlib import Path

import click
Expand Down Expand Up @@ -320,6 +321,80 @@ def cmd_search(
click.echo(f" {preview}\n")


@main.command(name="recall")
@click.argument("target")
@click.option(
"--session",
"session_id",
default=None,
metavar="SESSION_ID",
help="Session to search when TARGET is a query (default: most recent).",
)
@click.option(
"--all",
"all_sessions",
is_flag=True,
default=False,
help="Search every session, not just the most recent.",
)
@click.option(
"--limit",
default=1,
show_default=True,
type=int,
help="Max entries to return when TARGET is a query.",
)
@click.option(
"--raw",
is_flag=True,
default=False,
help="Print exact stored bytes without stripping escape sequences.",
)
@click.option(
"--db",
default=None,
type=click.Path(path_type=Path),
help="Path to the ACK SQLite database.",
)
def cmd_recall(
target: str,
session_id: str | None,
all_sessions: bool,
limit: int,
raw: bool,
db: Path | None,
) -> None:
"""Page a stored log back into view by handle or query.

TARGET is either a numeric recall handle (the "ack #N" shown on a pruned
banner -> `ack recall N`) or an FTS5 query, in which case the best match
from the most recent session is returned. Use --all to widen the search.
"""
render: Callable[[str], str] = (lambda t: t) if raw else _sanitize
storage = StorageEngine(db_path=db) if db else StorageEngine()
with storage:
if target.isdigit():
row = storage.get_entry(int(target))
rows = [row] if row is not None else []
else:
scope = session_id
if scope is None and not all_sessions:
recent = storage.list_sessions(limit=1)
scope = recent[0]["session_id"] if recent else None
rows = storage.search(target, session_id=scope, limit=limit)

if not rows:
click.echo(f"No entry for: {target!r}", err=True)
sys.exit(1)

for row in rows:
sid_short = row["session_id"][:8]
ts = time.strftime("%Y-%m-%d %H:%M:%S", time.localtime(row["timestamp"]))
pruned = " [pruned]" if row["was_pruned"] else ""
click.echo(click.style(f"[ack #{row['id']}] session={sid_short}… {ts}{pruned}", fg="cyan"))
click.echo(render(row["raw_content"]))


@main.command(name="toc")
@click.argument("file", type=click.Path(exists=True, path_type=Path))
def cmd_toc(file: Path) -> None:
Expand Down
19 changes: 12 additions & 7 deletions context_kernel/core/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -285,8 +285,8 @@ def _flush_buffer(self, stdout_fd: int, *, force: bool = False) -> None:
break

if summary is not None:
self._persist(text, pruned=True, summary=summary)
injection = self._format_injection(summary, line_count)
entry_id = self._persist(text, pruned=True, summary=summary)
injection = self._format_injection(summary, line_count, entry_id)
injected = self._terminal_newlines(injection).encode("utf-8")
self._stats.total_bytes_injected += len(injected)
self._emit(stdout_fd, injected)
Expand Down Expand Up @@ -321,7 +321,9 @@ def _emit(self, fd: int, data: bytes) -> None:
except OSError:
break

def _persist(self, text: str, *, pruned: bool, summary: str = "") -> None:
def _persist(self, text: str, *, pruned: bool, summary: str = "") -> int | None:
"""Store one entry and return its row id (the recall handle), or None
if the write failed — a failed archive must never abort the session."""
entry = LogEntry(
session_id=self.session_id,
raw_content=text,
Expand All @@ -331,19 +333,22 @@ def _persist(self, text: str, *, pruned: bool, summary: str = "") -> None:
was_pruned=pruned,
)
try:
self.storage.insert_entry(entry)
return self.storage.insert_entry(entry)
except Exception: # noqa: BLE001
pass
return None

def _format_injection(self, summary: str, original_lines: int) -> str:
def _format_injection(
self, summary: str, original_lines: int, entry_id: int | None = None
) -> str:
body = f"{_CYAN}{summary}{_RESET}\n"
if not self.config.annotate_injections:
return body

summary_lines = len(summary.splitlines())
recall = f"recall: ack #{entry_id}" if entry_id is not None else "full log stored in DB"
banner = (
f"{_DIM}[ACK] Compressed {original_lines} lines → "
f"{summary_lines} lines (full log stored in DB){_RESET}\n"
f"{summary_lines} lines ({recall}){_RESET}\n"
)
return banner + body

Expand Down
12 changes: 12 additions & 0 deletions context_kernel/memory/storage.py
Original file line number Diff line number Diff line change
Expand Up @@ -295,6 +295,18 @@ def search(
(query, limit),
).fetchall()

def get_entry(self, entry_id: int) -> sqlite3.Row | None:
"""Fetch a single log entry by its primary-key id, or None if absent.

This is the direct-lookup path behind `ack recall <id>`: the recall
handle stamped on a pruned injection is exactly this id.
"""
row: sqlite3.Row | None = self._db.execute(
"SELECT * FROM log_entries WHERE id = ?",
(entry_id,),
).fetchone()
return row

def get_recent_entries(self, session_id: str, limit: int = 50) -> list[sqlite3.Row]:
return self._db.execute(
"""
Expand Down
Loading
Loading