From 7d393c4b8b07cc1f429a6ea13e8416c871e1d49f Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Mon, 22 Jun 2026 19:38:27 -0400 Subject: [PATCH 01/20] =?UTF-8?q?docs:=20update=20Development=20Workflow?= =?UTF-8?q?=20=E2=80=94=20iterate=20fast=20within=20a=20feature,=20gate=20?= =?UTF-8?q?merges?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - No more stop-per-task; commit task-by-task continuously within a feature instead - Stop only at feature completion, with a concise bulleted (not verbose) summary - Claude now creates feature branches and merges them into the phase branch autonomously, and solo-handles PR creation + CI fixes for the phase -> develop -> main flow - Hard rule, no exceptions: every merge and every PR creation still requires an explicit go-ahead first --- .claude/CLAUDE.md | 23 +++++++++++++++-------- 1 file changed, 15 insertions(+), 8 deletions(-) diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index d45b899..ed8a8c5 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -67,16 +67,23 @@ Follow this exactly when starting any phase or feature work: existing convention (git branch tree, feature checklist, endpoints, schema, Definition of Done). List every feature branch the phase needs. 2. **One feature at a time.** Never start a second feature before the current one is fully closed out. + Create each feature branch off the phase branch automatically — no need to ask. 3. **Within a feature, one task at a time, TDD style** (`superpowers:test-driven-development`): - write the failing test → write the code → run it → confirm it passes. -4. **Stop after each task.** Give a short, clear, non-verbose explanation of what was done, then - prompt for a commit. Don't move to the next task until told to. -5. **Update the phase plan checklist after every single task** is completed — check it off in + write the failing test → write the code → run it → confirm it passes → commit immediately. + Keep iterating task to task without stopping to ask — quick, continuous commits. +4. **Update the phase plan checklist after every single task** is completed — check it off in `Phase Plans/Phase_X_Name.md` before moving on, not just at the end of the feature. -6. **Stop after the feature is fully done.** Give a bulleted summary of everything the feature did, - then prompt for a commit covering the whole feature. - -Never commit or push automatically — always stop and ask. +5. **Stop only once the feature is fully done.** Give a concise, bulleted, one-line-per-point + summary of what was implemented (architecture/approach) — not verbose, no long prose. +6. **Merge the feature branch into the phase branch autonomously** — but always pause and ask + ("merge about to happen, proceed?") before that merge actually executes. Same rule applies to + merging the phase branch into `develop`, opening the `develop` → `main` PR, and the final merge + to `main`: handle CI fixes and the mechanics solo, but never execute a merge or open a PR + without an explicit go-ahead first. + +Never commit code changes automatically without asking — except the task-by-task commits inside +an in-progress feature, which proceed without stopping. Merges and PR creation always require +an explicit go-ahead first, no exceptions. --- From f0aacb1df443b4fe84b2c17362b96c31d8b577c2 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Mon, 22 Jun 2026 19:45:24 -0400 Subject: [PATCH 02/20] docs: add subagent-spawning rule to Development Workflow MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Only spawn subagents for genuinely independent work (parallel research, isolated implementation pieces, context-heavy lookups) — not sequential feature-building work, which is most of what this project's workflow looks like. Always announce before spawning one, never silently. --- .claude/CLAUDE.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index ed8a8c5..5b28a0d 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -85,6 +85,12 @@ Never commit code changes automatically without asking — except the task-by-ta an in-progress feature, which proceed without stopping. Merges and PR creation always require an explicit go-ahead first, no exceptions. +**Subagents:** only spawn one when the task is genuinely independent (parallel research/exploration, +isolated implementation pieces with no shared state, or context-heavy lookups worth keeping out of +the main thread) — not for sequential work where each step needs the last step's result, which is +most of what building a feature looks like. Before spawning any subagent, say so and what it's for +first; don't spawn silently. + --- ## Communication Style From d520b9ef1044633f1391f1ffaded95536064eb8c Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Mon, 22 Jun 2026 22:42:34 -0400 Subject: [PATCH 03/20] docs: explanations must cover architectural reasoning, not just what was built MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every explanation of completed work now includes: key architectural choices made, why each was made, and the deciding factor whenever a new library/framework/service is introduced. Same concise one-liner bullet standard as the rest of Communication Style — reasoning in one tight line, not a paragraph. --- .claude/CLAUDE.md | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index 5b28a0d..6ed385e 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -97,3 +97,10 @@ first; don't spawn silently. Be concise. No filler, no over-explaining, no restating the obvious. Same standard applies to any doc, plan, or explanation written for this project. + +**Every explanation includes the "why," not just the "what."** When explaining work done +(a feature summary, a bug fix, anything), also cover: the key architectural choices made, +the reason each one was made, and — whenever a library/framework/service was introduced — +why that one specifically and what the deciding factor was over the alternatives. Keep this +at the same concise, one-line-per-point bullet standard — reasoning in one tight line, not a +paragraph. From 867b946c27c6c4e6cad76f0925f2923b957e0aa2 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:04:40 -0400 Subject: [PATCH 04/20] docs: closing task summaries always written in caveman style Keeps deliverables (code, docs, commits, app content) in normal prose per existing rules, but standardizes the final wrap-up line so task completion is always signaled the same terse way. --- .claude/CLAUDE.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index 6ed385e..1b36897 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -104,3 +104,8 @@ the reason each one was made, and — whenever a library/framework/service was i why that one specifically and what the deciding factor was over the alternatives. Keep this at the same concise, one-line-per-point bullet standard — reasoning in one tight line, not a paragraph. + +**Final task summary always in caveman.** After completing any task, the closing one-liner +("done, here's what changed") must be written in caveman style, regardless of whether caveman +mode applied to the task body itself. The deliverable (code, docs, commits, application content) +stays in normal prose per the rules above — only the wrap-up line switches to caveman. From 86cbe387ed8412510e987d710cc15557bc6b09bf Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:05:12 -0400 Subject: [PATCH 05/20] =?UTF-8?q?docs:=20add=20Phase=205=20plan=20?= =?UTF-8?q?=E2=80=94=20Argus=20Brain=20skeleton?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Stands up Argus as living infra: outcome ledger, supervisor graph, tool registry, chat API + UI. Architected so later phases (Cashflow, Goal Planning, Card Routing, Credit) plug in as new tools without rearchitecting this layer. --- Phase Plans/Phase_5_ArgusBrain.md | 142 ++++++++++++++++++++++++++++++ 1 file changed, 142 insertions(+) create mode 100644 Phase Plans/Phase_5_ArgusBrain.md diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md new file mode 100644 index 0000000..35f982f --- /dev/null +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -0,0 +1,142 @@ +# ArgusAI — Phase 5: Argus Brain (Skeleton) + +> Weeks 11–12. Goal: stand up Argus as living infrastructure — alive with whatever tools exist today (categorization, bills, subscriptions from Phase 3.5), built to grow as later phases add engines without rearchitecting. Most important phase in the remaining roadmap; everything after this is Argus gaining new senses. + +--- + +## What This Phase Covers + +| Layer | Goal | +|---|---| +| Database | `ai_predictions` outcome ledger table | +| Backend | LangGraph supervisor graph, tool registry, outcome-resolution Celery task, `POST /argus/chat` SSE endpoint with RAG | +| Frontend | Argus chat page, Cmd+K side panel (context-aware, persists across navigation) | +| Infra | New Celery task for periodic outcome resolution | + +**Note:** `backend/agents/` already exists from Phase 3.5 (`graph.py`, `analyst_agent.py`, `enrichment_agent.py`, `memory_agent.py`) — that's a background pipeline that runs after transaction sync. Phase 5's supervisor graph is a separate, new graph for live interactive chat — different purpose, same package. + +--- + +## Git Branch Structure + +``` +develop +└── phase/5-argus-brain + ├── feature/outcome-ledger + ├── feature/supervisor-graph + ├── feature/argus-chat-api + ├── feature/argus-chat-ui + └── feature/argus-side-panel +``` + +--- + +## Execution Checklist + +### `feature/outcome-ledger` ⬜ +*Every prediction Argus makes gets logged and later graded* + +**Backend:** +- [ ] Migration — `ai_predictions` table: `id`, `user_id`, `prediction_type`, `prediction_payload JSONB`, `predicted_at`, `resolves_at`, `actual_outcome JSONB`, `was_accurate BOOLEAN`, RLS policy +- [ ] `backend/tasks/resolve_predictions.py` — Celery task, runs periodically, checks unresolved predictions (`resolves_at <= now()`) against actual transaction/balance data, fills `actual_outcome` + `was_accurate` +- [ ] Tests for the resolution logic (pure function, mocked Supabase) +- [ ] Merge → `develop` + +--- + +### `feature/supervisor-graph` ⬜ +*Argus's routing brain — starts small, grows without rearchitecting* + +**Backend:** +- [ ] `backend/agents/tools.py` — tool registry pattern; new engine tools register themselves here as later phases ship +- [ ] Register Phase 3.5's existing capabilities (categorization, bills, subscriptions) as the first tools +- [ ] `backend/agents/supervisor.py` — LangGraph supervisor graph routing queries to specialist nodes based on available tools +- [ ] Tests for tool registration and routing logic +- [ ] Merge → `develop` + +--- + +### `feature/argus-chat-api` ⬜ +*The endpoint the frontend actually talks to* + +**Backend:** +- [ ] `backend/routers/argus.py` — `POST /argus/chat` SSE streaming endpoint +- [ ] RAG retrieval wired in: hot transactions (pgvector, last 90 days) + distilled monthly summaries (`ai_insights`) + profile (`user_financial_profiles`) + outcome ledger (relevant past predictions + accuracy for this user) +- [ ] Every response that makes a prediction/recommendation logs to `ai_predictions` +- [ ] Tests for the endpoint (mocked RAG + supervisor) +- [ ] Merge → `develop` + +--- + +### `feature/argus-chat-ui` ⬜ +*Where users actually talk to Argus* + +**Frontend:** +- [ ] `app/(app)/argus/page.tsx` — chat page, SSE streaming consumption +- [ ] Responses render as charts/tables/verdict cards — no paragraph-only responses (per the Specificity rule in `Argus Details/product-detail.md`) +- [ ] Merge → `develop` + +--- + +### `feature/argus-side-panel` ⬜ +*Argus everywhere, not just its own page* + +**Frontend:** +- [ ] Side panel component — slides in from right (380px) +- [ ] Cmd+K trigger from any screen +- [ ] Context-aware per current screen (knows what page the user is on) +- [ ] Conversation persists across navigation (not reset per route change) +- [ ] Merge → `develop` + +--- + +### Phase 5 Close +- [ ] Merge `phase/5-argus-brain` → `develop` +- [ ] Open PR `develop` → `main`, wait for CI, merge +- [ ] Delete all feature branches + `phase/5-argus-brain` +- [ ] Mark Phase 5 as ✅ Complete in `Argus Details/product-plan.md` + +--- + +## New Backend Endpoints + +| Method | Path | Description | +|---|---|---| +| `POST` | `/argus/chat` | SSE streaming chat endpoint, RAG-grounded, routes through supervisor graph | + +--- + +## Database Tables Used + +```sql +ai_predictions ( + id UUID PRIMARY KEY, + user_id UUID REFERENCES users, + prediction_type TEXT, + prediction_payload JSONB, + predicted_at TIMESTAMPTZ, + resolves_at TIMESTAMPTZ, + actual_outcome JSONB, + was_accurate BOOLEAN +) +``` + +--- + +## Definition of Done + +- [ ] Outcome ledger live, predictions logged and later resolved +- [ ] Supervisor graph routes to whatever tools exist today (categorization, bills, subscriptions) +- [ ] `/argus/chat` answers questions grounded in real data, never invents a number +- [ ] Chat page + Cmd+K side panel both live, side panel works from any screen +- [ ] CI green on `main` + +--- + +## Critical Files + +| File | Why it can't be skipped | +|---|---| +| `backend/agents/tools.py` | The registry pattern every later phase's engine (Cashflow, Goal Planning, Card Routing, Credit) plugs into — get this wrong and every later phase needs rearchitecting | +| `backend/migrations/0XX_ai_predictions.sql` | Self-improvement loop depends on this existing before Argus ever answers its first question | +| `backend/routers/argus.py` | The actual product — every other Phase 5 file exists to support this endpoint | From 6816d588f61c095b846e2f09af70cb85e88dfa93 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:06:57 -0400 Subject: [PATCH 06/20] feat(argus): add outcome ledger table and prediction resolution task ai_predictions table logs every prediction Argus makes; resolve_due_predictions Celery task grades resolved predictions against actual account balances. Evaluation logic kept as a pure function (_evaluate_prediction) so accuracy checks are testable without hitting Supabase, matching the existing detect_bills/detect_subscriptions pattern in this codebase. --- backend/migrations/015_ai_predictions.sql | 23 +++++++++ backend/tasks/resolve_predictions.py | 61 +++++++++++++++++++++++ backend/tests/test_resolve_predictions.py | 42 ++++++++++++++++ 3 files changed, 126 insertions(+) create mode 100644 backend/migrations/015_ai_predictions.sql create mode 100644 backend/tasks/resolve_predictions.py create mode 100644 backend/tests/test_resolve_predictions.py diff --git a/backend/migrations/015_ai_predictions.sql b/backend/migrations/015_ai_predictions.sql new file mode 100644 index 0000000..f546836 --- /dev/null +++ b/backend/migrations/015_ai_predictions.sql @@ -0,0 +1,23 @@ +-- backend/migrations/015_ai_predictions.sql + +CREATE TABLE IF NOT EXISTS ai_predictions ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + user_id UUID NOT NULL REFERENCES users ON DELETE CASCADE, + prediction_type TEXT NOT NULL, + prediction_payload JSONB NOT NULL, + predicted_at TIMESTAMPTZ NOT NULL DEFAULT now(), + resolves_at TIMESTAMPTZ NOT NULL, + actual_outcome JSONB, + was_accurate BOOLEAN +); + +CREATE INDEX IF NOT EXISTS ai_predictions_resolution_idx + ON ai_predictions (resolves_at) + WHERE actual_outcome IS NULL; + +ALTER TABLE ai_predictions ENABLE ROW LEVEL SECURITY; + +CREATE POLICY "ai_predictions_user_policy" + ON ai_predictions + FOR ALL + USING (user_id = auth.uid()); diff --git a/backend/tasks/resolve_predictions.py b/backend/tasks/resolve_predictions.py new file mode 100644 index 0000000..aa18c9c --- /dev/null +++ b/backend/tasks/resolve_predictions.py @@ -0,0 +1,61 @@ +from datetime import UTC, datetime + +from celery_app import celery +from db.client import get_supabase + + +def _evaluate_prediction( + prediction_type: str, payload: dict, actual_balance: float | None +) -> bool | None: + if actual_balance is None: + return None + + if prediction_type == "balance_below_threshold": + return actual_balance < payload["threshold"] + + return None + + +@celery.task(name="tasks.resolve_predictions.resolve_due_predictions") +def resolve_due_predictions() -> dict: + supabase = get_supabase() + + due = ( + supabase.table("ai_predictions") + .select("id, prediction_type, prediction_payload") + .lte("resolves_at", datetime.now(UTC).isoformat()) + .is_("actual_outcome", "null") + .execute() + ).data or [] + + resolved = 0 + for prediction in due: + payload = prediction["prediction_payload"] + account_id = payload.get("account_id") + actual_balance = None + + if account_id: + account = ( + supabase.table("accounts") + .select("balance") + .eq("id", account_id) + .execute() + ).data + if account: + actual_balance = account[0]["balance"] + + was_accurate = _evaluate_prediction( + prediction["prediction_type"], payload, actual_balance + ) + if was_accurate is None: + continue + + supabase.table("ai_predictions").update( + { + "actual_outcome": {"balance": actual_balance}, + "was_accurate": was_accurate, + } + ).eq("id", prediction["id"]).execute() + resolved += 1 + + return {"checked": len(due), "resolved": resolved} diff --git a/backend/tests/test_resolve_predictions.py b/backend/tests/test_resolve_predictions.py new file mode 100644 index 0000000..66e775e --- /dev/null +++ b/backend/tests/test_resolve_predictions.py @@ -0,0 +1,42 @@ +import os + +os.environ.setdefault("JWT_SECRET", "test-secret-key-for-unit-tests-only") +os.environ.setdefault("SUPABASE_URL", "https://placeholder.supabase.co") +os.environ.setdefault("SUPABASE_SERVICE_ROLE_KEY", "placeholder-service-role-key") +os.environ.setdefault("REDIS_URL", "redis://localhost:6379/0") +os.environ.setdefault("OPENAI_API_KEY", "sk-placeholder") +os.environ.setdefault("PLAID_CLIENT_ID", "placeholder") +os.environ.setdefault("PLAID_SECRET", "placeholder") +os.environ.setdefault("PLAID_ENV", "sandbox") +os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) + +from tasks.resolve_predictions import _evaluate_prediction # noqa: E402 + + +def test_balance_below_threshold_accurate_when_actual_is_lower(): + payload = {"account_id": "acct-1", "threshold": 0} + result = _evaluate_prediction("balance_below_threshold", payload, actual_balance=-12.50) + assert result is True + + +def test_balance_below_threshold_inaccurate_when_actual_is_higher(): + payload = {"account_id": "acct-1", "threshold": 0} + result = _evaluate_prediction("balance_below_threshold", payload, actual_balance=45.00) + assert result is False + + +def test_balance_below_threshold_boundary_is_not_below(): + payload = {"account_id": "acct-1", "threshold": 100} + result = _evaluate_prediction("balance_below_threshold", payload, actual_balance=100) + assert result is False + + +def test_unknown_prediction_type_returns_none(): + result = _evaluate_prediction("unknown_type", {}, actual_balance=10) + assert result is None + + +def test_missing_actual_balance_returns_none(): + payload = {"account_id": "acct-1", "threshold": 0} + result = _evaluate_prediction("balance_below_threshold", payload, actual_balance=None) + assert result is None From 585270edcd5a80f20b94bbc3fc466a031b725094 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:07:08 -0400 Subject: [PATCH 07/20] docs: check off outcome-ledger backend tasks in Phase 5 plan --- Phase Plans/Phase_5_ArgusBrain.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md index 35f982f..be2d15b 100644 --- a/Phase Plans/Phase_5_ArgusBrain.md +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -37,9 +37,9 @@ develop *Every prediction Argus makes gets logged and later graded* **Backend:** -- [ ] Migration — `ai_predictions` table: `id`, `user_id`, `prediction_type`, `prediction_payload JSONB`, `predicted_at`, `resolves_at`, `actual_outcome JSONB`, `was_accurate BOOLEAN`, RLS policy -- [ ] `backend/tasks/resolve_predictions.py` — Celery task, runs periodically, checks unresolved predictions (`resolves_at <= now()`) against actual transaction/balance data, fills `actual_outcome` + `was_accurate` -- [ ] Tests for the resolution logic (pure function, mocked Supabase) +- [x] Migration — `ai_predictions` table: `id`, `user_id`, `prediction_type`, `prediction_payload JSONB`, `predicted_at`, `resolves_at`, `actual_outcome JSONB`, `was_accurate BOOLEAN`, RLS policy +- [x] `backend/tasks/resolve_predictions.py` — Celery task, runs periodically, checks unresolved predictions (`resolves_at <= now()`) against actual transaction/balance data, fills `actual_outcome` + `was_accurate` +- [x] Tests for the resolution logic (pure function, mocked Supabase) - [ ] Merge → `develop` --- From b5375008260a9a09e7f666bf5c63be70dd96fb3b Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:12:26 -0400 Subject: [PATCH 08/20] feat(argus): add tool registry and supervisor routing graph MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit agents/tools.py is the registry pattern every later engine (Cashflow, Goal Planning, Card Routing, Credit) plugs into via @register_tool — decided on a decorator + dict registry over a class hierarchy so new phases add a function, not a subclass. Registers Phase 3.5's existing capabilities (bills, subscriptions, spending) as the first tools. agents/supervisor.py routes a chat query to whichever registered tools are relevant via keyword matching, falls back to all tools when nothing matches. Kept routing as a pure, deterministic function rather than an LLM tool-call loop for this skeleton phase — testable without mocking Anthropic, and the LangGraph node is a thin wrapper around it so an LLM-driven router can replace the matching logic later without touching the graph structure. --- backend/agents/supervisor.py | 47 ++++++++++++++++ backend/agents/tools.py | 89 +++++++++++++++++++++++++++++++ backend/tests/test_agent_tools.py | 45 ++++++++++++++++ backend/tests/test_supervisor.py | 51 ++++++++++++++++++ 4 files changed, 232 insertions(+) create mode 100644 backend/agents/supervisor.py create mode 100644 backend/agents/tools.py create mode 100644 backend/tests/test_agent_tools.py create mode 100644 backend/tests/test_supervisor.py diff --git a/backend/agents/supervisor.py b/backend/agents/supervisor.py new file mode 100644 index 0000000..937f295 --- /dev/null +++ b/backend/agents/supervisor.py @@ -0,0 +1,47 @@ +from typing import TypedDict + +from langgraph.graph import END, StateGraph + +from agents.tools import call_tool, get_registered_tools + + +class SupervisorState(TypedDict): + user_id: str + query: str + selected_tools: list[str] + tool_results: dict + + +def _select_tools_for_query(query: str, tools: dict[str, dict]) -> list[str]: + query_lower = query.lower() + matched = [ + name + for name, tool in tools.items() + if any(keyword in query_lower for keyword in tool["keywords"]) + ] + return matched or list(tools.keys()) + + +def route_and_execute_node(state: SupervisorState) -> dict: + tools = get_registered_tools() + selected = _select_tools_for_query(state["query"], tools) + + results: dict = {} + for name in selected: + results.update(call_tool(name, state["user_id"])) + + return {"selected_tools": selected, "tool_results": results} + + +def build_supervisor_graph(): + builder = StateGraph(SupervisorState) + + builder.add_node("route_and_execute", route_and_execute_node) + + builder.set_entry_point("route_and_execute") + builder.add_edge("route_and_execute", END) + + return builder.compile() + + +supervisor_graph = build_supervisor_graph() diff --git a/backend/agents/tools.py b/backend/agents/tools.py new file mode 100644 index 0000000..72a3416 --- /dev/null +++ b/backend/agents/tools.py @@ -0,0 +1,89 @@ +from collections.abc import Callable + +from db.client import get_supabase + +ToolFn = Callable[[str], dict] + +_TOOL_REGISTRY: dict[str, dict] = {} + + +def register_tool(name: str, description: str, keywords: list[str]): + def decorator(fn: ToolFn) -> ToolFn: + _TOOL_REGISTRY[name] = { + "name": name, + "description": description, + "keywords": keywords, + "fn": fn, + } + return fn + + return decorator + + +def get_registered_tools() -> dict[str, dict]: + return dict(_TOOL_REGISTRY) + + +def call_tool(name: str, user_id: str) -> dict: + tool = _TOOL_REGISTRY.get(name) + if tool is None: + raise KeyError(f"Unknown tool: {name}") + return tool["fn"](user_id) + + +@register_tool( + "get_bills", + "Returns the user's upcoming and recurring bills", + keywords=["bill", "bills", "due", "payment"], +) +def _get_bills_tool(user_id: str) -> dict: + supabase = get_supabase() + result = ( + supabase.table("bills") + .select("merchant, avg_amount, recurrence_pattern, next_due_date") + .eq("user_id", user_id) + .execute() + ) + return {"bills": result.data or []} + + +@register_tool( + "get_subscriptions", + "Returns the user's active subscriptions", + keywords=["subscription", "subscriptions", "sub", "recurring charge"], +) +def _get_subscriptions_tool(user_id: str) -> dict: + supabase = get_supabase() + result = ( + supabase.table("subscriptions") + .select("merchant, avg_amount, price_change_pct, billing_cycle") + .eq("user_id", user_id) + .eq("is_active", True) + .execute() + ) + return {"subscriptions": result.data or []} + + +@register_tool( + "get_spending_by_category", + "Returns the user's spending broken down by transaction category", + keywords=["spend", "spent", "spending", "category", "categories"], +) +def _get_spending_by_category_tool(user_id: str) -> dict: + supabase = get_supabase() + accounts = ( + supabase.table("accounts").select("id").eq("user_id", user_id).execute() + ).data or [] + account_ids = [a["id"] for a in accounts] + if not account_ids: + return {"transactions": []} + + result = ( + supabase.table("transactions") + .select("category, amount, timestamp") + .in_("account_id", account_ids) + .order("timestamp", desc=True) + .limit(500) + .execute() + ) + return {"transactions": result.data or []} diff --git a/backend/tests/test_agent_tools.py b/backend/tests/test_agent_tools.py new file mode 100644 index 0000000..e5ad9b0 --- /dev/null +++ b/backend/tests/test_agent_tools.py @@ -0,0 +1,45 @@ +import os +from unittest.mock import MagicMock, patch + +os.environ.setdefault("JWT_SECRET", "test-secret-key-for-unit-tests-only") +os.environ.setdefault("SUPABASE_URL", "https://placeholder.supabase.co") +os.environ.setdefault("SUPABASE_SERVICE_ROLE_KEY", "placeholder-service-role-key") +os.environ.setdefault("REDIS_URL", "redis://localhost:6379/0") +os.environ.setdefault("OPENAI_API_KEY", "sk-placeholder") +os.environ.setdefault("PLAID_CLIENT_ID", "placeholder") +os.environ.setdefault("PLAID_SECRET", "placeholder") +os.environ.setdefault("PLAID_ENV", "sandbox") +os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) + +from agents.tools import call_tool, get_registered_tools # noqa: E402 + + +def test_phase_3_5_tools_are_registered(): + tools = get_registered_tools() + assert {"get_bills", "get_subscriptions", "get_spending_by_category"} <= tools.keys() + + +def test_registered_tool_has_description_and_keywords(): + tools = get_registered_tools() + bills_tool = tools["get_bills"] + assert bills_tool["description"] + assert "bill" in bills_tool["keywords"] + + +def test_call_tool_unknown_name_raises(): + try: + call_tool("not_a_real_tool", "user-123") + assert False, "expected KeyError" + except KeyError: + pass + + +def test_call_tool_get_bills_queries_supabase(): + mock_supabase = MagicMock() + mock_query = mock_supabase.table.return_value.select.return_value.eq.return_value + mock_query.execute.return_value.data = [{"merchant": "Netflix", "avg_amount": 15.99}] + + with patch("agents.tools.get_supabase", return_value=mock_supabase): + result = call_tool("get_bills", "user-123") + + assert result == {"bills": [{"merchant": "Netflix", "avg_amount": 15.99}]} diff --git a/backend/tests/test_supervisor.py b/backend/tests/test_supervisor.py new file mode 100644 index 0000000..728f998 --- /dev/null +++ b/backend/tests/test_supervisor.py @@ -0,0 +1,51 @@ +import os +from unittest.mock import patch + +os.environ.setdefault("JWT_SECRET", "test-secret-key-for-unit-tests-only") +os.environ.setdefault("SUPABASE_URL", "https://placeholder.supabase.co") +os.environ.setdefault("SUPABASE_SERVICE_ROLE_KEY", "placeholder-service-role-key") +os.environ.setdefault("REDIS_URL", "redis://localhost:6379/0") +os.environ.setdefault("OPENAI_API_KEY", "sk-placeholder") +os.environ.setdefault("PLAID_CLIENT_ID", "placeholder") +os.environ.setdefault("PLAID_SECRET", "placeholder") +os.environ.setdefault("PLAID_ENV", "sandbox") +os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) + +from agents.supervisor import _select_tools_for_query, route_and_execute_node # noqa: E402 + +_FAKE_TOOLS = { + "get_bills": {"keywords": ["bill", "bills"]}, + "get_subscriptions": {"keywords": ["subscription", "sub"]}, + "get_spending_by_category": {"keywords": ["spend", "category"]}, +} + + +def test_select_tools_matches_bills_keyword(): + assert _select_tools_for_query("when is my next bill due", _FAKE_TOOLS) == ["get_bills"] + + +def test_select_tools_matches_subscriptions_keyword(): + result = _select_tools_for_query("cancel this subscription", _FAKE_TOOLS) + assert result == ["get_subscriptions"] + + +def test_select_tools_matches_multiple_tools(): + result = _select_tools_for_query("how much did I spend on subscriptions", _FAKE_TOOLS) + assert set(result) == {"get_subscriptions", "get_spending_by_category"} + + +def test_select_tools_falls_back_to_all_when_no_match(): + result = _select_tools_for_query("what should I do today", _FAKE_TOOLS) + assert set(result) == set(_FAKE_TOOLS.keys()) + + +def test_route_and_execute_node_calls_selected_tools(): + state = {"user_id": "user-123", "query": "what bills do I have"} + + with patch("agents.supervisor.get_registered_tools", return_value=_FAKE_TOOLS), \ + patch("agents.supervisor.call_tool", return_value={"bills": []}) as mock_call: + result = route_and_execute_node(state) + + mock_call.assert_called_once_with("get_bills", "user-123") + assert result["selected_tools"] == ["get_bills"] + assert result["tool_results"] == {"bills": []} From a4a18927e421496d2402811f7fc9199d4bdf0855 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:12:41 -0400 Subject: [PATCH 09/20] docs: check off supervisor-graph backend tasks in Phase 5 plan --- Phase Plans/Phase_5_ArgusBrain.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md index be2d15b..60b703b 100644 --- a/Phase Plans/Phase_5_ArgusBrain.md +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -48,10 +48,10 @@ develop *Argus's routing brain — starts small, grows without rearchitecting* **Backend:** -- [ ] `backend/agents/tools.py` — tool registry pattern; new engine tools register themselves here as later phases ship -- [ ] Register Phase 3.5's existing capabilities (categorization, bills, subscriptions) as the first tools -- [ ] `backend/agents/supervisor.py` — LangGraph supervisor graph routing queries to specialist nodes based on available tools -- [ ] Tests for tool registration and routing logic +- [x] `backend/agents/tools.py` — tool registry pattern; new engine tools register themselves here as later phases ship +- [x] Register Phase 3.5's existing capabilities (categorization, bills, subscriptions) as the first tools +- [x] `backend/agents/supervisor.py` — LangGraph supervisor graph routing queries to specialist nodes based on available tools +- [x] Tests for tool registration and routing logic - [ ] Merge → `develop` --- From ffc87711c735ff98c8189c11c5ace7d891404bec Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:15:19 -0400 Subject: [PATCH 10/20] feat(argus): add POST /argus/chat SSE endpoint with RAG + outcome logging MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit agents/chat.py wires RAG context from three sources: supervisor_graph (live data via registered tools), user_financial_profiles (static profile), and the existing _retrieve_relevant_insights embedding search from the Phase 3.5 pipeline (distilled monthly summaries) — reused instead of duplicated since it already does pgvector retrieval against ai_insights. Also pulls the user's own outcome ledger (past predictions + accuracy) into context so Argus can calibrate against its own track record per the self-improving design in product-detail.md. Predictions are captured via a fenced ```prediction JSON block the system prompt asks Claude to emit after any verifiable claim — parsed by a pure, testable function (_extract_prediction_block) rather than trying to classify free text. Logged to ai_predictions, then stripped from the text shown to the user. Endpoint streams via SSE (StreamingResponse + text/event-stream) since the product spec requires live token-by-token chat, not request/response. --- backend/agents/chat.py | 109 +++++++++++++++++++++++++++++ backend/main.py | 3 +- backend/routers/argus.py | 60 ++++++++++++++++ backend/tests/test_argus_chat.py | 106 ++++++++++++++++++++++++++++ backend/tests/test_argus_router.py | 64 +++++++++++++++++ 5 files changed, 341 insertions(+), 1 deletion(-) create mode 100644 backend/agents/chat.py create mode 100644 backend/routers/argus.py create mode 100644 backend/tests/test_argus_chat.py create mode 100644 backend/tests/test_argus_router.py diff --git a/backend/agents/chat.py b/backend/agents/chat.py new file mode 100644 index 0000000..1c8861e --- /dev/null +++ b/backend/agents/chat.py @@ -0,0 +1,109 @@ +import json +import re +import uuid +from datetime import UTC, datetime + +from agents.supervisor import supervisor_graph +from db.client import get_supabase +from tasks.run_intelligence_pipeline import _retrieve_relevant_insights + +_PREDICTION_BLOCK_RE = re.compile(r"```prediction\s*(\{.*?\})\s*```", re.DOTALL) + +_CHAT_SYSTEM_PROMPT = """You are Argus, ArgusAI's financial brain. You speak with merchant +names, dollar amounts, and timeframes — never vague generalities. If you cannot back a +statement with a specific merchant, amount, and timeframe, stay silent on that point. + +You never invent a number — only reason over the data given to you. + +When you make a verifiable prediction (e.g. "you'll overdraft on the 27th"), emit a fenced +block immediately after your answer in this exact form so it can be logged and graded later: + +```prediction +{"prediction_type": "balance_below_threshold", "payload": {"account_id": "", +"threshold": }, "resolves_at": ""} +``` + +Omit the block entirely if you made no verifiable prediction.""" + + +def _strip_prediction_block(response_text: str) -> str: + return _PREDICTION_BLOCK_RE.sub("", response_text).strip() + + +def _extract_prediction_block(response_text: str) -> dict | None: + match = _PREDICTION_BLOCK_RE.search(response_text) + if not match: + return None + try: + parsed = json.loads(match.group(1)) + except (json.JSONDecodeError, ValueError): + return None + if not isinstance(parsed, dict): + return None + if not {"prediction_type", "payload", "resolves_at"} <= parsed.keys(): + return None + return parsed + + +def _log_prediction(user_id: str, prediction: dict) -> None: + supabase = get_supabase() + supabase.table("ai_predictions").insert({ + "id": str(uuid.uuid4()), + "user_id": user_id, + "prediction_type": prediction["prediction_type"], + "prediction_payload": prediction["payload"], + "predicted_at": datetime.now(UTC).isoformat(), + "resolves_at": prediction["resolves_at"], + }).execute() + + +def _retrieve_chat_context(user_id: str, query: str) -> dict: + supervisor_result = supervisor_graph.invoke({ + "user_id": user_id, + "query": query, + "selected_tools": [], + "tool_results": {}, + }) + + supabase = get_supabase() + + profile_result = ( + supabase.table("user_financial_profiles") + .select("profile") + .eq("user_id", user_id) + .execute() + ) + profile = profile_result.data[0]["profile"] if profile_result.data else {} + + predictions_result = ( + supabase.table("ai_predictions") + .select("prediction_type, prediction_payload, was_accurate") + .eq("user_id", user_id) + .order("predicted_at", desc=True) + .limit(10) + .execute() + ) + past_predictions = [ + p for p in (predictions_result.data or []) if p.get("was_accurate") is not None + ] + + distilled_insights = _retrieve_relevant_insights(user_id, query) + + return { + "tool_results": supervisor_result["tool_results"], + "profile": profile, + "past_predictions": past_predictions, + "distilled_insights": distilled_insights, + } + + +def _build_chat_brief(query: str, context: dict) -> str: + return ( + f"USER QUESTION:\n{query}\n\n" + f"USER PROFILE:\n{json.dumps(context['profile'], indent=2)}\n\n" + f"RELEVANT LIVE DATA:\n{json.dumps(context['tool_results'], indent=2)}\n\n" + f"RELEVANT PAST INSIGHTS:\n{json.dumps(context['distilled_insights'], indent=2)}\n\n" + f"YOUR TRACK RECORD WITH THIS USER " + f"(past predictions and whether they came true):\n" + f"{json.dumps(context['past_predictions'], indent=2)}" + ) diff --git a/backend/main.py b/backend/main.py index 76e078e..fa5eb69 100644 --- a/backend/main.py +++ b/backend/main.py @@ -4,7 +4,7 @@ from fastapi import FastAPI from fastapi.middleware.cors import CORSMiddleware -from routers import auth, bills, insights, onboarding, plaid, subscriptions, transactions +from routers import argus, auth, bills, insights, onboarding, plaid, subscriptions, transactions load_dotenv() @@ -35,6 +35,7 @@ app.include_router(subscriptions.router) app.include_router(insights.router) app.include_router(onboarding.router) +app.include_router(argus.router) @app.get("/health", tags=["system"]) diff --git a/backend/routers/argus.py b/backend/routers/argus.py new file mode 100644 index 0000000..c04c713 --- /dev/null +++ b/backend/routers/argus.py @@ -0,0 +1,60 @@ +import json +import os +from collections.abc import AsyncGenerator + +import anthropic +from fastapi import APIRouter, Depends +from fastapi.responses import StreamingResponse +from pydantic import BaseModel + +from agents.chat import ( + _CHAT_SYSTEM_PROMPT, + _build_chat_brief, + _extract_prediction_block, + _log_prediction, + _retrieve_chat_context, + _strip_prediction_block, +) +from middleware.auth import get_current_user + +router = APIRouter(prefix="/argus", tags=["argus"]) + + +class ChatRequest(BaseModel): + query: str + + +async def _stream_chat_response(user_id: str, query: str) -> AsyncGenerator[str, None]: + context = _retrieve_chat_context(user_id, query) + brief = _build_chat_brief(query, context) + + client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"]) + full_text = "" + + with client.messages.stream( + model="claude-sonnet-4-6", + max_tokens=2048, + system=[{ + "type": "text", + "text": _CHAT_SYSTEM_PROMPT, + "cache_control": {"type": "ephemeral"}, + }], + messages=[{"role": "user", "content": brief}], + ) as stream: + for chunk in stream.text_stream: + full_text += chunk + yield f"data: {json.dumps({'text': chunk})}\n\n" + + prediction = _extract_prediction_block(full_text) + if prediction is not None: + _log_prediction(user_id, prediction) + + yield f"data: {json.dumps({'done': True, 'text': _strip_prediction_block(full_text)})}\n\n" + + +@router.post("/chat") +async def chat(payload: ChatRequest, user_id: str = Depends(get_current_user)): + return StreamingResponse( + _stream_chat_response(user_id, payload.query), + media_type="text/event-stream", + ) diff --git a/backend/tests/test_argus_chat.py b/backend/tests/test_argus_chat.py new file mode 100644 index 0000000..416e8f4 --- /dev/null +++ b/backend/tests/test_argus_chat.py @@ -0,0 +1,106 @@ +import os +from unittest.mock import MagicMock, patch + +os.environ.setdefault("JWT_SECRET", "test-secret-key-for-unit-tests-only") +os.environ.setdefault("SUPABASE_URL", "https://placeholder.supabase.co") +os.environ.setdefault("SUPABASE_SERVICE_ROLE_KEY", "placeholder-service-role-key") +os.environ.setdefault("REDIS_URL", "redis://localhost:6379/0") +os.environ.setdefault("OPENAI_API_KEY", "sk-placeholder") +os.environ.setdefault("ANTHROPIC_API_KEY", "sk-ant-placeholder") +os.environ.setdefault("PLAID_CLIENT_ID", "placeholder") +os.environ.setdefault("PLAID_SECRET", "placeholder") +os.environ.setdefault("PLAID_ENV", "sandbox") +os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) + +from agents.chat import ( # noqa: E402 + _extract_prediction_block, + _log_prediction, + _retrieve_chat_context, + _strip_prediction_block, +) + +_VALID_BLOCK = ( + "You'll overdraft on the 27th.\n\n" + "```prediction\n" + '{"prediction_type": "balance_below_threshold", ' + '"payload": {"account_id": "acct-1", "threshold": 0}, ' + '"resolves_at": "2026-06-27T00:00:00Z"}\n' + "```" +) + + +def test_extract_prediction_block_parses_valid_block(): + result = _extract_prediction_block(_VALID_BLOCK) + assert result["prediction_type"] == "balance_below_threshold" + assert result["payload"]["threshold"] == 0 + assert result["resolves_at"] == "2026-06-27T00:00:00Z" + + +def test_extract_prediction_block_returns_none_when_absent(): + assert _extract_prediction_block("Just a plain answer, no prediction here.") is None + + +def test_extract_prediction_block_returns_none_when_malformed_json(): + text = "```prediction\nnot json\n```" + assert _extract_prediction_block(text) is None + + +def test_extract_prediction_block_returns_none_when_missing_fields(): + text = '```prediction\n{"prediction_type": "x"}\n```' + assert _extract_prediction_block(text) is None + + +def test_strip_prediction_block_removes_block_keeps_answer(): + result = _strip_prediction_block(_VALID_BLOCK) + assert "```prediction" not in result + assert result.startswith("You'll overdraft on the 27th.") + + +def test_strip_prediction_block_no_op_when_no_block(): + text = "Plain answer." + assert _strip_prediction_block(text) == text + + +def test_log_prediction_inserts_row(): + mock_supabase = MagicMock() + prediction = { + "prediction_type": "balance_below_threshold", + "payload": {"account_id": "acct-1", "threshold": 0}, + "resolves_at": "2026-06-27T00:00:00Z", + } + + with patch("agents.chat.get_supabase", return_value=mock_supabase): + _log_prediction("user-123", prediction) + + inserted = mock_supabase.table.return_value.insert.call_args[0][0] + assert inserted["user_id"] == "user-123" + assert inserted["prediction_type"] == "balance_below_threshold" + assert inserted["prediction_payload"] == prediction["payload"] + assert inserted["resolves_at"] == "2026-06-27T00:00:00Z" + + +def test_retrieve_chat_context_combines_sources(): + mock_supervisor_result = {"tool_results": {"bills": []}} + mock_supabase = MagicMock() + profile_chain = mock_supabase.table.return_value.select.return_value.eq.return_value + profile_chain.execute.return_value.data = [{"profile": {"income": 5000}}] + + predictions_chain = ( + mock_supabase.table.return_value.select.return_value.eq.return_value + .order.return_value.limit.return_value + ) + predictions_chain.execute.return_value.data = [ + {"prediction_type": "balance_below_threshold", "was_accurate": True}, + {"prediction_type": "balance_below_threshold", "was_accurate": None}, + ] + + with patch("agents.chat.supervisor_graph") as mock_graph, \ + patch("agents.chat.get_supabase", return_value=mock_supabase), \ + patch("agents.chat._retrieve_relevant_insights", return_value=[{"summary": "x"}]): + mock_graph.invoke.return_value = mock_supervisor_result + result = _retrieve_chat_context("user-123", "what bills do I have") + + assert result["tool_results"] == {"bills": []} + assert result["distilled_insights"] == [{"summary": "x"}] + assert len(result["past_predictions"]) == 1 + assert result["past_predictions"][0]["was_accurate"] is True diff --git a/backend/tests/test_argus_router.py b/backend/tests/test_argus_router.py new file mode 100644 index 0000000..85f2d56 --- /dev/null +++ b/backend/tests/test_argus_router.py @@ -0,0 +1,64 @@ +from contextlib import contextmanager +from unittest.mock import MagicMock, patch + +from fastapi.testclient import TestClient + +from main import app +from middleware.auth import get_current_user + +app.dependency_overrides[get_current_user] = lambda: "test-user-id" +client = TestClient(app) + + +@contextmanager +def _fake_stream(chunks: list[str]): + mock_stream = MagicMock() + mock_stream.text_stream = iter(chunks) + yield mock_stream + + +def test_chat_streams_sse_response(): + mock_anthropic_client = MagicMock() + mock_anthropic_client.messages.stream.side_effect = lambda **_: _fake_stream( + ["You'll ", "overdraft on the 27th."] + ) + + with patch("routers.argus._retrieve_chat_context", return_value={ + "tool_results": {}, "profile": {}, "past_predictions": [], "distilled_insights": [], + }), patch("routers.argus._build_chat_brief", return_value="brief"), \ + patch("routers.argus.anthropic.Anthropic", return_value=mock_anthropic_client), \ + patch("routers.argus._extract_prediction_block", return_value=None): + resp = client.post("/argus/chat", json={"query": "when will I overdraft"}) + + assert resp.status_code == 200 + assert resp.headers["content-type"].startswith("text/event-stream") + assert "You'll " in resp.text + assert "overdraft on the 27th." in resp.text + + +def test_chat_logs_prediction_when_present(): + mock_anthropic_client = MagicMock() + mock_anthropic_client.messages.stream.side_effect = lambda **_: _fake_stream(["answer"]) + prediction = { + "prediction_type": "balance_below_threshold", + "payload": {"account_id": "acct-1", "threshold": 0}, + "resolves_at": "2026-06-27T00:00:00Z", + } + + with patch("routers.argus._retrieve_chat_context", return_value={ + "tool_results": {}, "profile": {}, "past_predictions": [], "distilled_insights": [], + }), patch("routers.argus._build_chat_brief", return_value="brief"), \ + patch("routers.argus.anthropic.Anthropic", return_value=mock_anthropic_client), \ + patch("routers.argus._extract_prediction_block", return_value=prediction), \ + patch("routers.argus._log_prediction") as mock_log: + resp = client.post("/argus/chat", json={"query": "when will I overdraft"}) + + assert resp.status_code == 200 + mock_log.assert_called_once_with("test-user-id", prediction) + + +def test_chat_requires_auth(): + app.dependency_overrides.pop(get_current_user, None) + resp = client.post("/argus/chat", json={"query": "hi"}) + assert resp.status_code == 401 + app.dependency_overrides[get_current_user] = lambda: "test-user-id" From 05047349c9fec6a622dad1d254590135be26f79a Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:15:38 -0400 Subject: [PATCH 11/20] docs: check off argus-chat-api backend tasks in Phase 5 plan --- Phase Plans/Phase_5_ArgusBrain.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md index 60b703b..1ece757 100644 --- a/Phase Plans/Phase_5_ArgusBrain.md +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -60,10 +60,10 @@ develop *The endpoint the frontend actually talks to* **Backend:** -- [ ] `backend/routers/argus.py` — `POST /argus/chat` SSE streaming endpoint -- [ ] RAG retrieval wired in: hot transactions (pgvector, last 90 days) + distilled monthly summaries (`ai_insights`) + profile (`user_financial_profiles`) + outcome ledger (relevant past predictions + accuracy for this user) -- [ ] Every response that makes a prediction/recommendation logs to `ai_predictions` -- [ ] Tests for the endpoint (mocked RAG + supervisor) +- [x] `backend/routers/argus.py` — `POST /argus/chat` SSE streaming endpoint +- [x] RAG retrieval wired in: hot transactions (pgvector, last 90 days) + distilled monthly summaries (`ai_insights`) + profile (`user_financial_profiles`) + outcome ledger (relevant past predictions + accuracy for this user) +- [x] Every response that makes a prediction/recommendation logs to `ai_predictions` +- [x] Tests for the endpoint (mocked RAG + supervisor) - [ ] Merge → `develop` --- From e96d2c27a32a4dfbc8be4230717b0eb9aa3c5a5f Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 14:18:40 -0400 Subject: [PATCH 12/20] feat(argus): add Argus chat page with SSE streaming consumption MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Added streamChat() to lib/api.ts as a separate path from the existing apiFetch helper — SSE needs raw ReadableStream/TextDecoder parsing of `data: {...}` frames, not JSON.parse on a single response body, so it couldn't reuse the existing JSON-only client. Chat page wraps useSearchParams() in Suspense (Next 16 build requirement for any client component reading search params) so /argus?q=... deep links from AskArgusBar prerender correctly. Wired the existing dashboard AskArgusBar — previously routed to /intelligence as a placeholder — to this new route now that the dedicated endpoint exists. Added Argus to the app sidebar nav, mirroring the existing nav-item pattern. No frontend test runner exists in this repo (eslint-only, no jest/ vitest configured for app/) — verified via next build (prerenders clean) and eslint (no new violations) instead, matching how every other page in app/(app)/ is currently validated. --- frontend/app/(app)/argus/page.tsx | 135 ++++++++++++++++++ .../dashboard/_components/AskArgusBar.tsx | 3 +- frontend/app/(app)/layout.tsx | 10 ++ frontend/lib/api.ts | 47 ++++++ 4 files changed, 193 insertions(+), 2 deletions(-) create mode 100644 frontend/app/(app)/argus/page.tsx diff --git a/frontend/app/(app)/argus/page.tsx b/frontend/app/(app)/argus/page.tsx new file mode 100644 index 0000000..6a817bb --- /dev/null +++ b/frontend/app/(app)/argus/page.tsx @@ -0,0 +1,135 @@ +"use client"; + +export const dynamic = "force-dynamic"; + +import { Suspense, useEffect, useRef, useState } from "react"; +import { useSearchParams } from "next/navigation"; +import { Sparkles, ArrowUp } from "lucide-react"; + +import { streamChat } from "@/lib/api"; + +type ChatMessage = { + id: string; + role: "user" | "argus"; + text: string; +}; + +export default function ArgusChatPage() { + return ( + + + + ); +} + +function ArgusChat() { + const searchParams = useSearchParams(); + const [messages, setMessages] = useState([]); + const [input, setInput] = useState(""); + const [streaming, setStreaming] = useState(false); + const bottomRef = useRef(null); + const initialQuery = useRef(searchParams.get("q")); + + useEffect(() => { + bottomRef.current?.scrollIntoView({ behavior: "smooth" }); + }, [messages]); + + useEffect(() => { + if (initialQuery.current) { + const q = initialQuery.current; + initialQuery.current = null; + send(q); + } + // eslint-disable-next-line react-hooks/exhaustive-deps + }, []); + + function send(query: string) { + const trimmed = query.trim(); + if (!trimmed || streaming) return; + + const userMessage: ChatMessage = { id: crypto.randomUUID(), role: "user", text: trimmed }; + const argusId = crypto.randomUUID(); + setMessages((prev) => [...prev, userMessage, { id: argusId, role: "argus", text: "" }]); + setInput(""); + setStreaming(true); + + streamChat( + trimmed, + (chunk) => { + setMessages((prev) => + prev.map((m) => (m.id === argusId ? { ...m, text: m.text + chunk } : m)) + ); + }, + (fullText) => { + setMessages((prev) => + prev.map((m) => (m.id === argusId ? { ...m, text: fullText } : m)) + ); + setStreaming(false); + } + ).catch((err) => { + setMessages((prev) => + prev.map((m) => + m.id === argusId ? { ...m, text: `Couldn't reach Argus: ${err.message}` } : m + ) + ); + setStreaming(false); + }); + } + + return ( +
+
+

+ + Argus +

+

+ Ask about your spending, bills, or what's coming. +

+
+ +
+ {messages.length === 0 ? ( +
+

No messages yet.

+

Ask Argus anything about your money.

+
+ ) : ( + messages.map((m) => ( +
+ {m.text || (streaming && m.role === "argus" ? "…" : "")} +
+ )) + )} +
+
+ +
+ setInput(e.target.value)} + onKeyDown={(e) => { + if (e.key === "Enter") send(input); + }} + placeholder="Ask Argus..." + className="flex-1 bg-transparent outline-none text-sm text-white placeholder:text-gray-600" + /> + +
+
+ ); +} diff --git a/frontend/app/(app)/dashboard/_components/AskArgusBar.tsx b/frontend/app/(app)/dashboard/_components/AskArgusBar.tsx index 8bbec23..d36fc9e 100644 --- a/frontend/app/(app)/dashboard/_components/AskArgusBar.tsx +++ b/frontend/app/(app)/dashboard/_components/AskArgusBar.tsx @@ -36,9 +36,8 @@ export function AskArgusBar() { }, [askIdx]); function submit() { - // Routes to /intelligence pending a dedicated AI Copilot chat route (see project plan: "Copilot + Simulations" phase) const q = askInput.trim(); - router.push(q ? `/intelligence?q=${encodeURIComponent(q)}` : "/intelligence"); + router.push(q ? `/argus?q=${encodeURIComponent(q)}` : "/argus"); } return ( diff --git a/frontend/app/(app)/layout.tsx b/frontend/app/(app)/layout.tsx index 5f23d8f..32ad7d4 100644 --- a/frontend/app/(app)/layout.tsx +++ b/frontend/app/(app)/layout.tsx @@ -62,6 +62,16 @@ const NAV = [ ), }, + { + href: "/argus", + title: "Argus", + icon: ( + + + + + ), + }, ]; const BOTTOM = [ diff --git a/frontend/lib/api.ts b/frontend/lib/api.ts index 0a515e9..84e109d 100644 --- a/frontend/lib/api.ts +++ b/frontend/lib/api.ts @@ -33,3 +33,50 @@ export const api = { body: body !== undefined ? JSON.stringify(body) : undefined, }), }; + +export async function streamChat( + query: string, + onChunk: (text: string) => void, + onDone: (fullText: string) => void, +): Promise { + const token = await getToken(); + const res = await fetch(`${API_URL}/argus/chat`, { + method: "POST", + headers: { + "Content-Type": "application/json", + Authorization: `Bearer ${token}`, + }, + body: JSON.stringify({ query }), + }); + + if (!res.ok || !res.body) { + const err = await res.json().catch(() => ({ detail: res.statusText })); + throw new Error((err as { detail?: string }).detail ?? res.statusText); + } + + const reader = res.body.getReader(); + const decoder = new TextDecoder(); + let buffer = ""; + + while (true) { + const { done, value } = await reader.read(); + if (done) break; + + buffer += decoder.decode(value, { stream: true }); + const lines = buffer.split("\n\n"); + buffer = lines.pop() ?? ""; + + for (const line of lines) { + if (!line.startsWith("data: ")) continue; + const payload = JSON.parse(line.slice("data: ".length)) as { + text?: string; + done?: boolean; + }; + if (payload.done) { + onDone(payload.text ?? ""); + } else if (payload.text) { + onChunk(payload.text); + } + } + } +} From 9f2c4ff370a83eb23f1a42dd2a0c0265ee754484 Mon Sep 17 00:00:00 2001 From: Evin Bento Date: Tue, 23 Jun 2026 15:27:15 -0400 Subject: [PATCH 13/20] feat(argus): structure chat responses as cards instead of prose MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per the Specificity rule in product-detail.md, Argus answers should render as verdict cards, tables, or charts — never paragraphs. Added a second fenced-block convention (\`\`\`argus-card) alongside the existing prediction block: the system prompt now requires every answer to be one or more typed JSON cards (verdict/table/chart), parsed by _extract_card_blocks and stripped from the displayed text the same way predictions already were. SSE 'done' payload now carries a `cards` array the frontend renders directly, instead of dumping structured data into prose the UI would have to re-parse. --- backend/agents/chat.py | 52 +++++++++++++++++++++++++-- backend/routers/argus.py | 7 +++- backend/tests/test_argus_chat.py | 58 ++++++++++++++++++++++++++++++ backend/tests/test_argus_router.py | 17 +++++++++ 4 files changed, 130 insertions(+), 4 deletions(-) diff --git a/backend/agents/chat.py b/backend/agents/chat.py index 1c8861e..84b7633 100644 --- a/backend/agents/chat.py +++ b/backend/agents/chat.py @@ -8,6 +8,8 @@ from tasks.run_intelligence_pipeline import _retrieve_relevant_insights _PREDICTION_BLOCK_RE = re.compile(r"```prediction\s*(\{.*?\})\s*```", re.DOTALL) +_CARD_BLOCK_RE = re.compile(r"```argus-card\s*(\{.*?\})\s*```", re.DOTALL) +_CARD_TYPES = {"verdict", "table", "chart"} _CHAT_SYSTEM_PROMPT = """You are Argus, ArgusAI's financial brain. You speak with merchant names, dollar amounts, and timeframes — never vague generalities. If you cannot back a @@ -15,21 +17,65 @@ You never invent a number — only reason over the data given to you. -When you make a verifiable prediction (e.g. "you'll overdraft on the 27th"), emit a fenced -block immediately after your answer in this exact form so it can be logged and graded later: +Never answer in plain paragraphs. Every answer is delivered as one or more fenced +```argus-card blocks of structured JSON — the UI renders these as cards, not prose. +Each block has this shape: + +```argus-card +{"type": "verdict", "data": {"label": "", "value": "", +"detail": ""}} +``` + +```argus-card +{"type": "table", "data": {"title": "", "columns": ["<col1>", "<col2>"], +"rows": [["<cell>", "<cell>"]]}} +``` + +```argus-card +{"type": "chart", "data": {"title": "<title>", "points": [{"label": "<x>", "value": <y>}]}} +``` + +Use "verdict" for a single answer (e.g. affordability, a yes/no, a dollar figure). +Use "table" for lists (bills, subscriptions, merchants). Use "chart" for trends over time. +Emit only the card blocks needed to answer — no surrounding text. + +When you make a verifiable prediction (e.g. "you'll overdraft on the 27th"), also emit a +fenced block immediately after your cards in this exact form so it can be logged and +graded later: ```prediction {"prediction_type": "balance_below_threshold", "payload": {"account_id": "<id>", "threshold": <number>}, "resolves_at": "<ISO 8601 datetime>"} ``` -Omit the block entirely if you made no verifiable prediction.""" +Omit the prediction block entirely if you made no verifiable prediction.""" def _strip_prediction_block(response_text: str) -> str: return _PREDICTION_BLOCK_RE.sub("", response_text).strip() +def _extract_card_blocks(response_text: str) -> list[dict]: + cards = [] + for match in _CARD_BLOCK_RE.finditer(response_text): + try: + parsed = json.loads(match.group(1)) + except (json.JSONDecodeError, ValueError): + continue + if not isinstance(parsed, dict): + continue + if parsed.get("type") not in _CARD_TYPES: + continue + if "data" not in parsed: + continue + cards.append(parsed) + return cards + + +def _strip_card_blocks(response_text: str) -> str: + return _CARD_BLOCK_RE.sub("", response_text).strip() + + def _extract_prediction_block(response_text: str) -> dict | None: match = _PREDICTION_BLOCK_RE.search(response_text) if not match: diff --git a/backend/routers/argus.py b/backend/routers/argus.py index c04c713..de4126b 100644 --- a/backend/routers/argus.py +++ b/backend/routers/argus.py @@ -10,9 +10,11 @@ from agents.chat import ( _CHAT_SYSTEM_PROMPT, _build_chat_brief, + _extract_card_blocks, _extract_prediction_block, _log_prediction, _retrieve_chat_context, + _strip_card_blocks, _strip_prediction_block, ) from middleware.auth import get_current_user @@ -49,7 +51,10 @@ async def _stream_chat_response(user_id: str, query: str) -> AsyncGenerator[str, if prediction is not None: _log_prediction(user_id, prediction) - yield f"data: {json.dumps({'done': True, 'text': _strip_prediction_block(full_text)})}\n\n" + cards = _extract_card_blocks(full_text) + remaining_text = _strip_card_blocks(_strip_prediction_block(full_text)) + + yield f"data: {json.dumps({'done': True, 'text': remaining_text, 'cards': cards})}\n\n" @router.post("/chat") diff --git a/backend/tests/test_argus_chat.py b/backend/tests/test_argus_chat.py index 416e8f4..bcfe0f1 100644 --- a/backend/tests/test_argus_chat.py +++ b/backend/tests/test_argus_chat.py @@ -13,9 +13,11 @@ os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) from agents.chat import ( # noqa: E402 + _extract_card_blocks, _extract_prediction_block, _log_prediction, _retrieve_chat_context, + _strip_card_blocks, _strip_prediction_block, ) @@ -104,3 +106,59 @@ def test_retrieve_chat_context_combines_sources(): assert result["distilled_insights"] == [{"summary": "x"}] assert len(result["past_predictions"]) == 1 assert result["past_predictions"][0]["was_accurate"] is True + + +_VERDICT_CARD = ( + "```argus-card\n" + '{"type": "verdict", "data": {"label": "Safe to spend", "value": "$214",' + ' "detail": "After Netflix $15.99 due Jul 1"}}\n' + "```" +) +_TABLE_CARD = ( + "```argus-card\n" + '{"type": "table", "data": {"title": "Bills", "columns": ["Merchant", "Amount"],' + ' "rows": [["Netflix", "$15.99"]]}}\n' + "```" +) + + +def test_extract_card_blocks_parses_single_card(): + result = _extract_card_blocks(_VERDICT_CARD) + assert len(result) == 1 + assert result[0]["type"] == "verdict" + assert result[0]["data"]["value"] == "$214" + + +def test_extract_card_blocks_parses_multiple_cards(): + text = f"{_VERDICT_CARD}\n\n{_TABLE_CARD}" + result = _extract_card_blocks(text) + assert [c["type"] for c in result] == ["verdict", "table"] + + +def test_extract_card_blocks_skips_unknown_type(): + text = '```argus-card\n{"type": "paragraph", "data": {}}\n```' + assert _extract_card_blocks(text) == [] + + +def test_extract_card_blocks_skips_malformed_json(): + text = "```argus-card\nnot json\n```" + assert _extract_card_blocks(text) == [] + + +def test_extract_card_blocks_skips_missing_data(): + text = '```argus-card\n{"type": "verdict"}\n```' + assert _extract_card_blocks(text) == [] + + +def test_extract_card_blocks_returns_empty_when_no_blocks(): + assert _extract_card_blocks("plain text") == [] + + +def test_strip_card_blocks_removes_cards(): + result = _strip_card_blocks(_VERDICT_CARD) + assert "argus-card" not in result + assert result == "" + + +def test_strip_card_blocks_no_op_when_no_cards(): + assert _strip_card_blocks("plain text") == "plain text" diff --git a/backend/tests/test_argus_router.py b/backend/tests/test_argus_router.py index 85f2d56..c68c037 100644 --- a/backend/tests/test_argus_router.py +++ b/backend/tests/test_argus_router.py @@ -36,6 +36,23 @@ def test_chat_streams_sse_response(): assert "overdraft on the 27th." in resp.text +def test_chat_returns_cards_in_done_payload(): + mock_anthropic_client = MagicMock() + mock_anthropic_client.messages.stream.side_effect = lambda **_: _fake_stream(["text"]) + cards = [{"type": "verdict", "data": {"label": "Safe to spend", "value": "$214"}}] + + with patch("routers.argus._retrieve_chat_context", return_value={ + "tool_results": {}, "profile": {}, "past_predictions": [], "distilled_insights": [], + }), patch("routers.argus._build_chat_brief", return_value="brief"), \ + patch("routers.argus.anthropic.Anthropic", return_value=mock_anthropic_client), \ + patch("routers.argus._extract_prediction_block", return_value=None), \ + patch("routers.argus._extract_card_blocks", return_value=cards): + resp = client.post("/argus/chat", json={"query": "can I afford this"}) + + assert resp.status_code == 200 + assert '"cards": [{"type": "verdict"' in resp.text + + def test_chat_logs_prediction_when_present(): mock_anthropic_client = MagicMock() mock_anthropic_client.messages.stream.side_effect = lambda **_: _fake_stream(["answer"]) From 939797f4794a7d04d939d371d25b6438a4fbf6ed Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:28:51 -0400 Subject: [PATCH 14/20] feat(argus): render chat responses as verdict/table/chart cards MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Cards.tsx adds the three card types matching the backend's argus-card contract. Used recharts for the chart card — already a project dependency, unused elsewhere yet, so no new dependency added. Chat page no longer streams raw text into the bubble: card JSON arrives fenced mid-stream and would render as garbage if shown live, so chunks now just keep a "thinking" indicator up until the done event delivers parsed cards. Falls back to a plain text bubble only for client-side errors (e.g. network failure) where there's no card to show. --- .../app/(app)/argus/_components/Cards.tsx | 94 +++++++++++++++++++ frontend/app/(app)/argus/page.tsx | 73 +++++++++----- frontend/lib/api.ts | 5 +- 3 files changed, 149 insertions(+), 23 deletions(-) create mode 100644 frontend/app/(app)/argus/_components/Cards.tsx diff --git a/frontend/app/(app)/argus/_components/Cards.tsx b/frontend/app/(app)/argus/_components/Cards.tsx new file mode 100644 index 0000000..50af469 --- /dev/null +++ b/frontend/app/(app)/argus/_components/Cards.tsx @@ -0,0 +1,94 @@ +"use client"; + +import { Bar, BarChart, ResponsiveContainer, Tooltip, XAxis, YAxis } from "recharts"; + +export type VerdictCardData = { label: string; value: string; detail?: string }; +export type TableCardData = { title: string; columns: string[]; rows: string[][] }; +export type ChartCardData = { title: string; points: { label: string; value: number }[] }; + +export type ArgusCardModel = + | { type: "verdict"; data: VerdictCardData } + | { type: "table"; data: TableCardData } + | { type: "chart"; data: ChartCardData }; + +export function ArgusCard({ card }: { card: ArgusCardModel }) { + if (card.type === "verdict") return <VerdictCard data={card.data} />; + if (card.type === "table") return <TableCard data={card.data} />; + return <ChartCard data={card.data} />; +} + +function VerdictCard({ data }: { data: VerdictCardData }) { + return ( + <div className="bg-gray-900 rounded-2xl border border-gray-800 px-5 py-4"> + <p className="text-xs text-gray-500 uppercase tracking-wider mb-1">{data.label}</p> + <p className="text-2xl font-bold text-white">{data.value}</p> + {data.detail && <p className="text-xs text-gray-500 mt-1.5">{data.detail}</p>} + </div> + ); +} + +function TableCard({ data }: { data: TableCardData }) { + return ( + <div className="bg-gray-900 rounded-2xl border border-gray-800 overflow-hidden"> + <p className="text-xs text-gray-500 uppercase tracking-wider px-5 pt-4 pb-2"> + {data.title} + </p> + <table className="w-full text-sm"> + <thead> + <tr className="border-t border-gray-800"> + {data.columns.map((col) => ( + <th + key={col} + className="text-left text-xs font-medium text-gray-500 px-5 py-2" + > + {col} + </th> + ))} + </tr> + </thead> + <tbody> + {data.rows.map((row, ri) => ( + <tr key={ri} className="border-t border-gray-800/60"> + {row.map((cell, ci) => ( + <td key={ci} className="px-5 py-2 text-gray-200"> + {cell} + </td> + ))} + </tr> + ))} + </tbody> + </table> + </div> + ); +} + +function ChartCard({ data }: { data: ChartCardData }) { + return ( + <div className="bg-gray-900 rounded-2xl border border-gray-800 px-5 py-4"> + <p className="text-xs text-gray-500 uppercase tracking-wider mb-3">{data.title}</p> + <div style={{ width: "100%", height: 160 }}> + <ResponsiveContainer> + <BarChart data={data.points}> + <XAxis + dataKey="label" + stroke="#6b7280" + fontSize={11} + tickLine={false} + axisLine={false} + /> + <YAxis stroke="#6b7280" fontSize={11} tickLine={false} axisLine={false} /> + <Tooltip + contentStyle={{ + background: "#111827", + border: "1px solid #1f2937", + borderRadius: 8, + fontSize: 12, + }} + /> + <Bar dataKey="value" fill="rgb(255, 103, 0)" radius={[4, 4, 0, 0]} /> + </BarChart> + </ResponsiveContainer> + </div> + </div> + ); +} diff --git a/frontend/app/(app)/argus/page.tsx b/frontend/app/(app)/argus/page.tsx index 6a817bb..14a1d65 100644 --- a/frontend/app/(app)/argus/page.tsx +++ b/frontend/app/(app)/argus/page.tsx @@ -8,10 +8,14 @@ import { Sparkles, ArrowUp } from "lucide-react"; import { streamChat } from "@/lib/api"; +import { ArgusCard, type ArgusCardModel } from "./_components/Cards"; + type ChatMessage = { id: string; role: "user" | "argus"; text: string; + cards: ArgusCardModel[]; + pending: boolean; }; export default function ArgusChatPage() { @@ -47,29 +51,44 @@ function ArgusChat() { const trimmed = query.trim(); if (!trimmed || streaming) return; - const userMessage: ChatMessage = { id: crypto.randomUUID(), role: "user", text: trimmed }; + const userMessage: ChatMessage = { + id: crypto.randomUUID(), + role: "user", + text: trimmed, + cards: [], + pending: false, + }; const argusId = crypto.randomUUID(); - setMessages((prev) => [...prev, userMessage, { id: argusId, role: "argus", text: "" }]); + setMessages((prev) => [ + ...prev, + userMessage, + { id: argusId, role: "argus", text: "", cards: [], pending: true }, + ]); setInput(""); setStreaming(true); streamChat( trimmed, - (chunk) => { - setMessages((prev) => - prev.map((m) => (m.id === argusId ? { ...m, text: m.text + chunk } : m)) - ); + () => { + // Card JSON streams in fenced blocks mid-flight — not shown raw; + // the "thinking" indicator covers it until the parsed cards land. }, - (fullText) => { + (fullText, cards) => { setMessages((prev) => - prev.map((m) => (m.id === argusId ? { ...m, text: fullText } : m)) + prev.map((m) => + m.id === argusId + ? { ...m, text: fullText, cards: cards as ArgusCardModel[], pending: false } + : m + ) ); setStreaming(false); } ).catch((err) => { setMessages((prev) => prev.map((m) => - m.id === argusId ? { ...m, text: `Couldn't reach Argus: ${err.message}` } : m + m.id === argusId + ? { ...m, text: `Couldn't reach Argus: ${err.message}`, pending: false } + : m ) ); setStreaming(false); @@ -95,18 +114,30 @@ function ArgusChat() { <p className="text-xs text-gray-600 mt-1">Ask Argus anything about your money.</p> </div> ) : ( - messages.map((m) => ( - <div - key={m.id} - className={`rounded-2xl border px-5 py-3.5 text-sm leading-relaxed whitespace-pre-wrap ${ - m.role === "user" - ? "bg-indigo-900/30 border-indigo-800/40 text-indigo-100 ml-auto max-w-[80%]" - : "bg-gray-900 border-gray-800 text-gray-200 max-w-[85%]" - }`} - > - {m.text || (streaming && m.role === "argus" ? "…" : "")} - </div> - )) + messages.map((m) => + m.role === "user" ? ( + <div + key={m.id} + className="rounded-2xl border px-5 py-3.5 text-sm leading-relaxed whitespace-pre-wrap bg-indigo-900/30 border-indigo-800/40 text-indigo-100 ml-auto max-w-[80%]" + > + {m.text} + </div> + ) : ( + <div key={m.id} className="max-w-[85%] space-y-3"> + {m.pending ? ( + <div className="bg-gray-900 rounded-2xl border border-gray-800 px-5 py-3.5 text-sm text-gray-500"> + Argus is thinking… + </div> + ) : m.cards.length > 0 ? ( + m.cards.map((card, ci) => <ArgusCard key={ci} card={card} />) + ) : ( + <div className="bg-gray-900 rounded-2xl border border-gray-800 px-5 py-3.5 text-sm leading-relaxed whitespace-pre-wrap text-gray-200"> + {m.text} + </div> + )} + </div> + ) + ) )} <div ref={bottomRef} /> </div> diff --git a/frontend/lib/api.ts b/frontend/lib/api.ts index 84e109d..183a38f 100644 --- a/frontend/lib/api.ts +++ b/frontend/lib/api.ts @@ -37,7 +37,7 @@ export const api = { export async function streamChat( query: string, onChunk: (text: string) => void, - onDone: (fullText: string) => void, + onDone: (fullText: string, cards: unknown[]) => void, ): Promise<void> { const token = await getToken(); const res = await fetch(`${API_URL}/argus/chat`, { @@ -71,9 +71,10 @@ export async function streamChat( const payload = JSON.parse(line.slice("data: ".length)) as { text?: string; done?: boolean; + cards?: unknown[]; }; if (payload.done) { - onDone(payload.text ?? ""); + onDone(payload.text ?? "", payload.cards ?? []); } else if (payload.text) { onChunk(payload.text); } From 53a64b4c0ba219d8e06d0e17e2116c3f762a206a Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:29:23 -0400 Subject: [PATCH 15/20] docs: check off argus-chat-ui frontend tasks in Phase 5 plan --- Phase Plans/Phase_5_ArgusBrain.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md index 1ece757..af7ca05 100644 --- a/Phase Plans/Phase_5_ArgusBrain.md +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -72,8 +72,8 @@ develop *Where users actually talk to Argus* **Frontend:** -- [ ] `app/(app)/argus/page.tsx` — chat page, SSE streaming consumption -- [ ] Responses render as charts/tables/verdict cards — no paragraph-only responses (per the Specificity rule in `Argus Details/product-detail.md`) +- [x] `app/(app)/argus/page.tsx` — chat page, SSE streaming consumption +- [x] Responses render as charts/tables/verdict cards — no paragraph-only responses (per the Specificity rule in `Argus Details/product-detail.md`) - [ ] Merge → `develop` --- From 2445bcb61c85d593f9506284f1d7c5c30a2eed79 Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:31:20 -0400 Subject: [PATCH 16/20] feat(argus): chat endpoint accepts current-screen context MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds optional `page` field to ChatRequest, threaded through to _build_chat_brief as a CURRENT SCREEN line. Needed for the side panel, which is context-aware per the product spec — Argus should know what screen triggered the question (e.g. asked from /bills vs /subscriptions) without the frontend having to fold that into the query text itself. --- backend/agents/chat.py | 4 +++- backend/routers/argus.py | 9 ++++++--- backend/tests/test_argus_chat.py | 24 ++++++++++++++++++++++++ backend/tests/test_argus_router.py | 18 ++++++++++++++++++ 4 files changed, 51 insertions(+), 4 deletions(-) diff --git a/backend/agents/chat.py b/backend/agents/chat.py index 84b7633..aa83a76 100644 --- a/backend/agents/chat.py +++ b/backend/agents/chat.py @@ -143,8 +143,10 @@ def _retrieve_chat_context(user_id: str, query: str) -> dict: } -def _build_chat_brief(query: str, context: dict) -> str: +def _build_chat_brief(query: str, context: dict, page: str | None = None) -> str: + screen_line = f"CURRENT SCREEN: {page}\n\n" if page else "" return ( + f"{screen_line}" f"USER QUESTION:\n{query}\n\n" f"USER PROFILE:\n{json.dumps(context['profile'], indent=2)}\n\n" f"RELEVANT LIVE DATA:\n{json.dumps(context['tool_results'], indent=2)}\n\n" diff --git a/backend/routers/argus.py b/backend/routers/argus.py index de4126b..47474d7 100644 --- a/backend/routers/argus.py +++ b/backend/routers/argus.py @@ -24,11 +24,14 @@ class ChatRequest(BaseModel): query: str + page: str | None = None -async def _stream_chat_response(user_id: str, query: str) -> AsyncGenerator[str, None]: +async def _stream_chat_response( + user_id: str, query: str, page: str | None = None +) -> AsyncGenerator[str, None]: context = _retrieve_chat_context(user_id, query) - brief = _build_chat_brief(query, context) + brief = _build_chat_brief(query, context, page=page) client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"]) full_text = "" @@ -60,6 +63,6 @@ async def _stream_chat_response(user_id: str, query: str) -> AsyncGenerator[str, @router.post("/chat") async def chat(payload: ChatRequest, user_id: str = Depends(get_current_user)): return StreamingResponse( - _stream_chat_response(user_id, payload.query), + _stream_chat_response(user_id, payload.query, page=payload.page), media_type="text/event-stream", ) diff --git a/backend/tests/test_argus_chat.py b/backend/tests/test_argus_chat.py index bcfe0f1..34dd8aa 100644 --- a/backend/tests/test_argus_chat.py +++ b/backend/tests/test_argus_chat.py @@ -13,6 +13,7 @@ os.environ.setdefault("PLAID_TOKEN_ENCRYPTION_KEY", "a" * 64) from agents.chat import ( # noqa: E402 + _build_chat_brief, _extract_card_blocks, _extract_prediction_block, _log_prediction, @@ -162,3 +163,26 @@ def test_strip_card_blocks_removes_cards(): def test_strip_card_blocks_no_op_when_no_cards(): assert _strip_card_blocks("plain text") == "plain text" + + +def test_build_chat_brief_includes_current_screen_when_given(): + context = { + "profile": {}, + "tool_results": {}, + "distilled_insights": [], + "past_predictions": [], + } + brief = _build_chat_brief("what's this", context, page="/bills") + assert "CURRENT SCREEN" in brief + assert "/bills" in brief + + +def test_build_chat_brief_omits_current_screen_when_absent(): + context = { + "profile": {}, + "tool_results": {}, + "distilled_insights": [], + "past_predictions": [], + } + brief = _build_chat_brief("what's this", context) + assert "CURRENT SCREEN" not in brief diff --git a/backend/tests/test_argus_router.py b/backend/tests/test_argus_router.py index c68c037..161e021 100644 --- a/backend/tests/test_argus_router.py +++ b/backend/tests/test_argus_router.py @@ -74,6 +74,24 @@ def test_chat_logs_prediction_when_present(): mock_log.assert_called_once_with("test-user-id", prediction) +def test_chat_passes_page_to_brief(): + mock_anthropic_client = MagicMock() + mock_anthropic_client.messages.stream.side_effect = lambda **_: _fake_stream(["text"]) + + with patch("routers.argus._retrieve_chat_context", return_value={ + "tool_results": {}, "profile": {}, "past_predictions": [], "distilled_insights": [], + }), patch("routers.argus._build_chat_brief", return_value="brief") as mock_brief, \ + patch("routers.argus.anthropic.Anthropic", return_value=mock_anthropic_client), \ + patch("routers.argus._extract_prediction_block", return_value=None): + resp = client.post( + "/argus/chat", json={"query": "what bills do I have", "page": "/bills"} + ) + + assert resp.status_code == 200 + _, kwargs = mock_brief.call_args + assert kwargs["page"] == "/bills" + + def test_chat_requires_auth(): app.dependency_overrides.pop(get_current_user, None) resp = client.post("/argus/chat", json={"query": "hi"}) From 23a9cdc3785f64dde9f5e7bb2bacdba6718ea771 Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:32:44 -0400 Subject: [PATCH 17/20] feat(argus): add Cmd+K side panel, context-aware and nav-persistent MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ArgusSidePanel mounts once in app/(app)/layout.tsx — the shared layout that wraps every (app) route and doesn't remount on navigation in the App Router — so its message state survives moving between screens instead of resetting per route. Visibility is just a CSS transform (right: 14px vs -396px) rather than a conditional mount, which is what makes the persistence work: an unmounted component loses state, a hidden one doesn't. Cmd+K / Ctrl+K toggles it from any screen; Escape closes it. Passes the current pathname as `page` into streamChat so Argus knows what screen the question came from, using the page-context field added to the backend in the prior commit. --- .../app/(app)/_components/ArgusSidePanel.tsx | 164 ++++++++++++++++++ frontend/app/(app)/layout.tsx | 21 +++ frontend/lib/api.ts | 3 +- 3 files changed, 187 insertions(+), 1 deletion(-) create mode 100644 frontend/app/(app)/_components/ArgusSidePanel.tsx diff --git a/frontend/app/(app)/_components/ArgusSidePanel.tsx b/frontend/app/(app)/_components/ArgusSidePanel.tsx new file mode 100644 index 0000000..b6ddc2f --- /dev/null +++ b/frontend/app/(app)/_components/ArgusSidePanel.tsx @@ -0,0 +1,164 @@ +"use client"; + +import { useEffect, useRef, useState } from "react"; +import { Sparkles, ArrowUp, X } from "lucide-react"; + +import { streamChat } from "@/lib/api"; +import { ArgusCard, type ArgusCardModel } from "../argus/_components/Cards"; + +type ChatMessage = { + id: string; + role: "user" | "argus"; + text: string; + cards: ArgusCardModel[]; + pending: boolean; +}; + +export function ArgusSidePanel({ + open, + onClose, + page, +}: { + open: boolean; + onClose: () => void; + page: string; +}) { + const [messages, setMessages] = useState<ChatMessage[]>([]); + const [input, setInput] = useState(""); + const [streaming, setStreaming] = useState(false); + const bottomRef = useRef<HTMLDivElement>(null); + + useEffect(() => { + if (open) bottomRef.current?.scrollIntoView({ behavior: "smooth" }); + }, [messages, open]); + + function send() { + const trimmed = input.trim(); + if (!trimmed || streaming) return; + + const userMessage: ChatMessage = { + id: crypto.randomUUID(), + role: "user", + text: trimmed, + cards: [], + pending: false, + }; + const argusId = crypto.randomUUID(); + setMessages((prev) => [ + ...prev, + userMessage, + { id: argusId, role: "argus", text: "", cards: [], pending: true }, + ]); + setInput(""); + setStreaming(true); + + streamChat( + trimmed, + () => {}, + (fullText, cards) => { + setMessages((prev) => + prev.map((m) => + m.id === argusId + ? { ...m, text: fullText, cards: cards as ArgusCardModel[], pending: false } + : m + ) + ); + setStreaming(false); + }, + page + ).catch((err) => { + setMessages((prev) => + prev.map((m) => + m.id === argusId + ? { ...m, text: `Couldn't reach Argus: ${err.message}`, pending: false } + : m + ) + ); + setStreaming(false); + }); + } + + return ( + <aside + style={{ + position: "fixed", + top: 14, + right: open ? 14 : -396, + bottom: 14, + width: 380, + background: "var(--surface-1)", + border: "1px solid var(--surface-3)", + borderRadius: "var(--r-xl)", + boxShadow: "0 12px 40px rgba(0,0,0,0.45)", + display: "flex", + flexDirection: "column", + zIndex: 100, + transition: "right 220ms cubic-bezier(0.23,1,0.32,1)", + }} + > + <div className="flex items-center justify-between px-5 py-4 border-b border-gray-800"> + <div className="flex items-center gap-2"> + <Sparkles size={16} color="rgb(255, 103, 0)" /> + <span className="text-sm font-semibold text-white">Argus</span> + </div> + <button onClick={onClose} className="text-gray-500 hover:text-white"> + <X size={16} /> + </button> + </div> + + <div className="flex-1 overflow-y-auto px-4 py-4 space-y-3"> + {messages.length === 0 ? ( + <p className="text-xs text-gray-600 text-center mt-8"> + Ask Argus anything — it knows you're on {page}. + </p> + ) : ( + messages.map((m) => + m.role === "user" ? ( + <div + key={m.id} + className="rounded-xl border px-4 py-2.5 text-sm leading-relaxed whitespace-pre-wrap bg-indigo-900/30 border-indigo-800/40 text-indigo-100 ml-auto max-w-[85%]" + > + {m.text} + </div> + ) : ( + <div key={m.id} className="max-w-[90%] space-y-2"> + {m.pending ? ( + <div className="bg-gray-900 rounded-xl border border-gray-800 px-4 py-2.5 text-xs text-gray-500"> + Argus is thinking… + </div> + ) : m.cards.length > 0 ? ( + m.cards.map((card, ci) => <ArgusCard key={ci} card={card} />) + ) : ( + <div className="bg-gray-900 rounded-xl border border-gray-800 px-4 py-2.5 text-sm leading-relaxed whitespace-pre-wrap text-gray-200"> + {m.text} + </div> + )} + </div> + ) + ) + )} + <div ref={bottomRef} /> + </div> + + <div className="flex items-center gap-2 border-t border-gray-800 px-3 py-3"> + <input + type="text" + value={input} + onChange={(e) => setInput(e.target.value)} + onKeyDown={(e) => { + if (e.key === "Enter") send(); + }} + placeholder="Ask Argus..." + className="flex-1 bg-transparent outline-none text-sm text-white placeholder:text-gray-600 px-2" + /> + <button + onClick={send} + disabled={streaming || !input.trim()} + className="rounded-full w-7 h-7 flex items-center justify-center bg-indigo-600 disabled:opacity-40 transition-opacity flex-shrink-0" + > + <ArrowUp size={13} color="#fff" /> + </button> + </div> + </aside> + ); +} diff --git a/frontend/app/(app)/layout.tsx b/frontend/app/(app)/layout.tsx index 32ad7d4..0d1afc8 100644 --- a/frontend/app/(app)/layout.tsx +++ b/frontend/app/(app)/layout.tsx @@ -6,6 +6,8 @@ import Link from "next/link"; import { motion } from "framer-motion"; import { createClient } from "@/lib/supabase/client"; +import { ArgusSidePanel } from "./_components/ArgusSidePanel"; + const NAV = [ { href: "/dashboard", @@ -152,6 +154,7 @@ export default function AppLayout({ children }: { children: React.ReactNode }) { const pathname = usePathname(); const router = useRouter(); const [theme, setTheme] = useState<"dark" | "light">("dark"); + const [argusPanelOpen, setArgusPanelOpen] = useState(false); useEffect(() => { const stored = (localStorage.getItem("argus-theme") as "dark" | "light") ?? "dark"; @@ -159,6 +162,18 @@ export default function AppLayout({ children }: { children: React.ReactNode }) { document.body.classList.toggle("light", stored === "light"); }, []); + useEffect(() => { + function handleKeydown(e: KeyboardEvent) { + if ((e.metaKey || e.ctrlKey) && e.key.toLowerCase() === "k") { + e.preventDefault(); + setArgusPanelOpen((prev) => !prev); + } + if (e.key === "Escape") setArgusPanelOpen(false); + } + window.addEventListener("keydown", handleKeydown); + return () => window.removeEventListener("keydown", handleKeydown); + }, []); + function toggleTheme(mode: "dark" | "light") { setTheme(mode); localStorage.setItem("argus-theme", mode); @@ -253,6 +268,12 @@ export default function AppLayout({ children }: { children: React.ReactNode }) { <main style={{ flex: 1, minWidth: 0, display: "flex", flexDirection: "column", overflowY: "auto" }}> {children} </main> + + <ArgusSidePanel + open={argusPanelOpen} + onClose={() => setArgusPanelOpen(false)} + page={pathname} + /> </div> ); } diff --git a/frontend/lib/api.ts b/frontend/lib/api.ts index 183a38f..927f14b 100644 --- a/frontend/lib/api.ts +++ b/frontend/lib/api.ts @@ -38,6 +38,7 @@ export async function streamChat( query: string, onChunk: (text: string) => void, onDone: (fullText: string, cards: unknown[]) => void, + page?: string, ): Promise<void> { const token = await getToken(); const res = await fetch(`${API_URL}/argus/chat`, { @@ -46,7 +47,7 @@ export async function streamChat( "Content-Type": "application/json", Authorization: `Bearer ${token}`, }, - body: JSON.stringify({ query }), + body: JSON.stringify({ query, page }), }); if (!res.ok || !res.body) { From ae6baffa5dee5cc392043859ff3d4fd8eaa90fd4 Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:32:55 -0400 Subject: [PATCH 18/20] docs: check off argus-side-panel frontend tasks in Phase 5 plan --- Phase Plans/Phase_5_ArgusBrain.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/Phase Plans/Phase_5_ArgusBrain.md b/Phase Plans/Phase_5_ArgusBrain.md index af7ca05..72d4495 100644 --- a/Phase Plans/Phase_5_ArgusBrain.md +++ b/Phase Plans/Phase_5_ArgusBrain.md @@ -82,10 +82,10 @@ develop *Argus everywhere, not just its own page* **Frontend:** -- [ ] Side panel component — slides in from right (380px) -- [ ] Cmd+K trigger from any screen -- [ ] Context-aware per current screen (knows what page the user is on) -- [ ] Conversation persists across navigation (not reset per route change) +- [x] Side panel component — slides in from right (380px) +- [x] Cmd+K trigger from any screen +- [x] Context-aware per current screen (knows what page the user is on) +- [x] Conversation persists across navigation (not reset per route change) - [ ] Merge → `develop` --- From ac7b587e9bc735022101b5f75acd464eb8ae3872 Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 15:52:57 -0400 Subject: [PATCH 19/20] chore: update CLAUDE.md communication style + add tooling config - Bullets and phase summaries now use caveman style, jargon-stripped - .agents/ and .continue/ skill configs tracked - skills-lock.json added Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --- .agents/skills/cavecrew/README.md | 41 +++ .agents/skills/cavecrew/SKILL.md | 82 +++++ .agents/skills/caveman-commit/README.md | 44 +++ .agents/skills/caveman-commit/SKILL.md | 65 ++++ .agents/skills/caveman-compress/README.md | 163 +++++++++ .agents/skills/caveman-compress/SECURITY.md | 31 ++ .agents/skills/caveman-compress/SKILL.md | 111 ++++++ .../caveman-compress/scripts/__init__.py | 9 + .../caveman-compress/scripts/__main__.py | 3 + .../caveman-compress/scripts/benchmark.py | 80 ++++ .../skills/caveman-compress/scripts/cli.py | 85 +++++ .../caveman-compress/scripts/compress.py | 342 ++++++++++++++++++ .../skills/caveman-compress/scripts/detect.py | 121 +++++++ .../caveman-compress/scripts/validate.py | 213 +++++++++++ .agents/skills/caveman-help/README.md | 38 ++ .agents/skills/caveman-help/SKILL.md | 63 ++++ .agents/skills/caveman-review/README.md | 33 ++ .agents/skills/caveman-review/SKILL.md | 55 +++ .agents/skills/caveman-stats/README.md | 30 ++ .agents/skills/caveman-stats/SKILL.md | 10 + .agents/skills/caveman/README.md | 48 +++ .agents/skills/caveman/SKILL.md | 78 ++++ .claude/CLAUDE.md | 4 + .continue/skills/cavecrew/README.md | 41 +++ .continue/skills/cavecrew/SKILL.md | 82 +++++ .continue/skills/caveman-commit/README.md | 44 +++ .continue/skills/caveman-commit/SKILL.md | 65 ++++ .continue/skills/caveman-compress/README.md | 163 +++++++++ .continue/skills/caveman-compress/SECURITY.md | 31 ++ .continue/skills/caveman-compress/SKILL.md | 111 ++++++ .../caveman-compress/scripts/__init__.py | 9 + .../caveman-compress/scripts/__main__.py | 3 + .../caveman-compress/scripts/benchmark.py | 80 ++++ .../skills/caveman-compress/scripts/cli.py | 85 +++++ .../caveman-compress/scripts/compress.py | 342 ++++++++++++++++++ .../skills/caveman-compress/scripts/detect.py | 121 +++++++ .../caveman-compress/scripts/validate.py | 213 +++++++++++ .continue/skills/caveman-help/README.md | 38 ++ .continue/skills/caveman-help/SKILL.md | 63 ++++ .continue/skills/caveman-review/README.md | 33 ++ .continue/skills/caveman-review/SKILL.md | 55 +++ .continue/skills/caveman-stats/README.md | 30 ++ .continue/skills/caveman-stats/SKILL.md | 10 + .continue/skills/caveman/README.md | 48 +++ .continue/skills/caveman/SKILL.md | 78 ++++ skills-lock.json | 47 +++ 46 files changed, 3541 insertions(+) create mode 100644 .agents/skills/cavecrew/README.md create mode 100644 .agents/skills/cavecrew/SKILL.md create mode 100644 .agents/skills/caveman-commit/README.md create mode 100644 .agents/skills/caveman-commit/SKILL.md create mode 100644 .agents/skills/caveman-compress/README.md create mode 100644 .agents/skills/caveman-compress/SECURITY.md create mode 100644 .agents/skills/caveman-compress/SKILL.md create mode 100644 .agents/skills/caveman-compress/scripts/__init__.py create mode 100644 .agents/skills/caveman-compress/scripts/__main__.py create mode 100644 .agents/skills/caveman-compress/scripts/benchmark.py create mode 100644 .agents/skills/caveman-compress/scripts/cli.py create mode 100644 .agents/skills/caveman-compress/scripts/compress.py create mode 100644 .agents/skills/caveman-compress/scripts/detect.py create mode 100644 .agents/skills/caveman-compress/scripts/validate.py create mode 100644 .agents/skills/caveman-help/README.md create mode 100644 .agents/skills/caveman-help/SKILL.md create mode 100644 .agents/skills/caveman-review/README.md create mode 100644 .agents/skills/caveman-review/SKILL.md create mode 100644 .agents/skills/caveman-stats/README.md create mode 100644 .agents/skills/caveman-stats/SKILL.md create mode 100644 .agents/skills/caveman/README.md create mode 100644 .agents/skills/caveman/SKILL.md create mode 100644 .continue/skills/cavecrew/README.md create mode 100644 .continue/skills/cavecrew/SKILL.md create mode 100644 .continue/skills/caveman-commit/README.md create mode 100644 .continue/skills/caveman-commit/SKILL.md create mode 100644 .continue/skills/caveman-compress/README.md create mode 100644 .continue/skills/caveman-compress/SECURITY.md create mode 100644 .continue/skills/caveman-compress/SKILL.md create mode 100644 .continue/skills/caveman-compress/scripts/__init__.py create mode 100644 .continue/skills/caveman-compress/scripts/__main__.py create mode 100644 .continue/skills/caveman-compress/scripts/benchmark.py create mode 100644 .continue/skills/caveman-compress/scripts/cli.py create mode 100644 .continue/skills/caveman-compress/scripts/compress.py create mode 100644 .continue/skills/caveman-compress/scripts/detect.py create mode 100644 .continue/skills/caveman-compress/scripts/validate.py create mode 100644 .continue/skills/caveman-help/README.md create mode 100644 .continue/skills/caveman-help/SKILL.md create mode 100644 .continue/skills/caveman-review/README.md create mode 100644 .continue/skills/caveman-review/SKILL.md create mode 100644 .continue/skills/caveman-stats/README.md create mode 100644 .continue/skills/caveman-stats/SKILL.md create mode 100644 .continue/skills/caveman/README.md create mode 100644 .continue/skills/caveman/SKILL.md create mode 100644 skills-lock.json diff --git a/.agents/skills/cavecrew/README.md b/.agents/skills/cavecrew/README.md new file mode 100644 index 0000000..e722952 --- /dev/null +++ b/.agents/skills/cavecrew/README.md @@ -0,0 +1,41 @@ +# cavecrew + +Decision guide. When to delegate to caveman subagents instead of doing the work inline. + +## What it does + +Tells the main thread when to spawn a caveman-style subagent versus the vanilla equivalent. The win: subagent tool-results inject back into main context verbatim, and caveman output is roughly 1/3 the size of vanilla prose. Across 20 delegations in one session, that is the difference between context exhaustion and finishing the task. + +Three subagents: + +| Subagent | Job | Use when | +|----------|-----|----------| +| `cavecrew-investigator` | Locate code (read-only) | "Where is X defined / what calls Y / list uses of Z" | +| `cavecrew-builder` | Surgical edit, 1-2 files | Scope is obvious, ≤2 files. Refuses 3+ file scope. | +| `cavecrew-reviewer` | Diff/file review | One-line findings with severity emoji | + +Use vanilla `Explore` or `Code Reviewer` when you want prose, architecture commentary, or rationale. Use main thread directly for one-line answers and 3+ file refactors. + +This skill is a decision guide, not a slash command. It activates when the conversation mentions delegation. + +## How to invoke + +Triggers on phrases like "delegate to subagent", "use cavecrew", "spawn investigator", "save context", "compressed agent output". + +## Example chaining + +Locate → fix → verify (most common): + +1. `cavecrew-investigator` returns site list (`path:line — symbol — note`) +2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder` +3. `cavecrew-reviewer` audits the resulting diff + +Parallel scout: spawn 2-3 `cavecrew-investigator` calls in one message with different angles (defs, callers, tests). Aggregate in main. + +## See also + +- [`SKILL.md`](./SKILL.md) — full decision matrix and output contracts +- [`agents/cavecrew-investigator.md`](../../agents/cavecrew-investigator.md) +- [`agents/cavecrew-builder.md`](../../agents/cavecrew-builder.md) +- [`agents/cavecrew-reviewer.md`](../../agents/cavecrew-reviewer.md) +- [Caveman README](../../README.md) — repo overview diff --git a/.agents/skills/cavecrew/SKILL.md b/.agents/skills/cavecrew/SKILL.md new file mode 100644 index 0000000..efa413f --- /dev/null +++ b/.agents/skills/cavecrew/SKILL.md @@ -0,0 +1,82 @@ +--- +name: cavecrew +description: > + Decision guide for delegating to caveman-style subagents. Tells the main + thread WHEN to spawn `cavecrew-investigator` (locate code), `cavecrew-builder` + (1-2 file edit), or `cavecrew-reviewer` (diff review) instead of doing the + work inline or using vanilla `Explore`. Subagent output is caveman-compressed + so the tool-result injected back into main context is ~60% smaller — main + context lasts longer across long sessions. + Trigger: "delegate to subagent", "use cavecrew", "spawn investigator/builder/reviewer", + "save context", "compressed agent output". +--- + +Cavecrew = three subagent presets that emit caveman output. Same job as Anthropic defaults (`Explore`, edit-style agents, reviewer); difference is the tool-result they return is compressed, so main context shrinks per delegation. + +## When to use cavecrew vs alternatives + +| Task | Use | +|---|---| +| "Where is X defined / what calls Y / list uses of Z" | `cavecrew-investigator` | +| Same but you also want suggestions/architecture commentary | `Explore` (vanilla) | +| Surgical edit, ≤2 files, scope obvious | `cavecrew-builder` | +| New feature / 3+ files / cross-cutting refactor | Main thread or `feature-dev:code-architect` | +| Review diff, branch, or file for bugs | `cavecrew-reviewer` | +| Deep code review with rationale + alternatives | `Code Reviewer` (vanilla) | +| One-line answer you already know | Main thread, no subagent | + +Rule of thumb: **if you'd want the subagent's output in 1/3 the tokens, pick cavecrew. If you'd want prose, pick vanilla.** + +## Why this exists (the real win) + +Subagent tool results get injected into main context verbatim. A vanilla `Explore` that returns 2k tokens of prose costs 2k tokens of main-context budget every time. The same finding from `cavecrew-investigator` returns ~700 tokens. Across 20 delegations in one session that's the difference between context exhaustion and finishing the task. + +## Output contracts + +What main thread can rely on per agent: + +**`cavecrew-investigator`** +``` +<Header>: +- path:line — `symbol` — short note +totals: <counts>. +``` +Or `No match.` Always file-path-first, line-number-attached, backticked symbols. Safe to grep with `path:\d+`. + +**`cavecrew-builder`** +``` +<path:line-range> — <change ≤10 words>. +verified: <re-read OK | mismatch @ path:line>. +``` +Or one of: `too-big.` / `needs-confirm.` / `ambiguous.` / `regressed.` (terminal first token). + +**`cavecrew-reviewer`** +``` +path:line: <emoji> <severity>: <problem>. <fix>. +totals: N🔴 N🟡 N🔵 N❓ +``` +Or `No issues.` Findings sorted file → line ascending. + +## Chaining patterns + +**Locate → fix → verify** (most common): +1. `cavecrew-investigator` returns site list. +2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder`. +3. `cavecrew-reviewer` audits the diff. + +**Parallel scout** (when investigation is broad): +Spawn 2-3 `cavecrew-investigator` calls in one message (different angles: defs vs callers vs tests). Aggregate in main thread. + +**Single-shot edit** (when site is already known): +Skip investigator. Hand exact path:line to `cavecrew-builder` directly. + +## What NOT to do + +- Don't use `cavecrew-builder` when you don't already know the file. Spawn investigator first or main thread will eat tokens passing context. +- Don't chain `cavecrew-investigator → cavecrew-builder` for a 5-file refactor. Builder will return `too-big.` and you'll have wasted a turn. +- Don't ask `cavecrew-reviewer` for "general feedback" — it returns findings only, no architecture opinions. Use `Code Reviewer` for that. +- Don't expect prose. Cavecrew output is structured, sometimes terse to the point of cryptic. If a human will read it directly, paraphrase. + +## Auto-clarity (inherited) + +Subagents drop caveman → normal English for security warnings, irreversible-action confirmations, and any output where fragment ambiguity could be misread. Resume caveman after. diff --git a/.agents/skills/caveman-commit/README.md b/.agents/skills/caveman-commit/README.md new file mode 100644 index 0000000..d5aee01 --- /dev/null +++ b/.agents/skills/caveman-commit/README.md @@ -0,0 +1,44 @@ +# caveman-commit + +Terse Conventional Commits. Why over what. + +## What it does + +Generates commit messages in Conventional Commits format. Subject ≤50 chars, hard cap 72. Imperative mood. Body only when the *why* is non-obvious or there are breaking changes. No AI attribution, no "this commit does X", no emoji unless the project uses them. Body always required for breaking changes, security fixes, data migrations, and reverts — future debuggers need the context. + +Outputs only the message. Does not stage, commit, or amend. + +## How to invoke + +``` +/caveman-commit +``` + +Also triggers on phrases like "write a commit", "commit message", "generate commit". + +## Example output + +Diff: new endpoint for user profile. + +``` +feat(api): add GET /users/:id/profile + +Mobile client needs profile data without the full user payload +to reduce LTE bandwidth on cold-launch screens. + +Closes #128 +``` + +Diff: breaking API rename. + +``` +feat(api)!: rename /v1/orders to /v1/checkout + +BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout +before 2026-06-01. Old route returns 410 after that date. +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview diff --git a/.agents/skills/caveman-commit/SKILL.md b/.agents/skills/caveman-commit/SKILL.md new file mode 100644 index 0000000..b9999e3 --- /dev/null +++ b/.agents/skills/caveman-commit/SKILL.md @@ -0,0 +1,65 @@ +--- +name: caveman-commit +description: > + Ultra-compressed commit message generator. Cuts noise from commit messages while preserving + intent and reasoning. Conventional Commits format. Subject ≤50 chars, body only when "why" + isn't obvious. Use when user says "write a commit", "commit message", "generate commit", + "/commit", or invokes /caveman-commit. Auto-triggers when staging changes. +--- + +Write commit messages terse and exact. Conventional Commits format. No fluff. Why over what. + +## Rules + +**Subject line:** +- `<type>(<scope>): <imperative summary>` — `<scope>` optional +- Types: `feat`, `fix`, `refactor`, `perf`, `docs`, `test`, `chore`, `build`, `ci`, `style`, `revert` +- Imperative mood: "add", "fix", "remove" — not "added", "adds", "adding" +- ≤50 chars when possible, hard cap 72 +- No trailing period +- Match project convention for capitalization after the colon + +**Body (only if needed):** +- Skip entirely when subject is self-explanatory +- Add body only for: non-obvious *why*, breaking changes, migration notes, linked issues +- Wrap at 72 chars +- Bullets `-` not `*` +- Reference issues/PRs at end: `Closes #42`, `Refs #17` + +**What NEVER goes in:** +- "This commit does X", "I", "we", "now", "currently" — the diff says what +- "As requested by..." — use Co-authored-by trailer +- "Generated with Claude Code" or any AI attribution — unless the user's own rule requires an `Assisted-by`/AI-attribution trailer, then add it as a trailer +- Emoji (unless project convention requires) +- Restating the file name when scope already says it + +## Examples + +Diff: new endpoint for user profile with body explaining the why +- ❌ "feat: add a new endpoint to get user profile information from the database" +- ✅ + ``` + feat(api): add GET /users/:id/profile + + Mobile client needs profile data without the full user payload + to reduce LTE bandwidth on cold-launch screens. + + Closes #128 + ``` + +Diff: breaking API change +- ✅ + ``` + feat(api)!: rename /v1/orders to /v1/checkout + + BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout + before 2026-06-01. Old route returns 410 after that date. + ``` + +## Auto-Clarity + +Always include body for: breaking changes, security fixes, data migrations, anything reverting a prior commit. Never compress these into subject-only — future debuggers need the context. + +## Boundaries + +Only generates the commit message. Does not run `git commit`, does not stage files, does not amend. Output the message as a code block ready to paste. "stop caveman-commit" or "normal mode": revert to verbose commit style. diff --git a/.agents/skills/caveman-compress/README.md b/.agents/skills/caveman-compress/README.md new file mode 100644 index 0000000..3ef6922 --- /dev/null +++ b/.agents/skills/caveman-compress/README.md @@ -0,0 +1,163 @@ +<p align="center"> + <img src="https://em-content.zobj.net/source/apple/391/rock_1faa8.png" width="80" /> +</p> + +<h1 align="center">caveman-compress</h1> + +<p align="center"> + <strong>shrink memory file. save token every session.</strong> +</p> + +--- + +A Claude Code skill that compresses your project memory files (`CLAUDE.md`, todos, preferences) into caveman format — so every session loads fewer tokens automatically. + +Claude read `CLAUDE.md` on every session start. If file big, cost big. Caveman make file small. Cost go down forever. + +## What It Do + +``` +/caveman-compress CLAUDE.md +``` + +``` +CLAUDE.md ← compressed (Claude reads this — fewer tokens every session) +CLAUDE.original.md ← human-readable backup (you edit this) +``` + +Original never lost. You can read and edit `.original.md`. Run skill again to re-compress after edits. + +## Benchmarks + +Real results on real project files: + +| File | Original | Compressed | Saved | +|------|----------:|----------:|------:| +| `claude-md-preferences.md` | 706 | 285 | **59.6%** | +| `project-notes.md` | 1145 | 535 | **53.3%** | +| `claude-md-project.md` | 1122 | 636 | **43.3%** | +| `todo-list.md` | 627 | 388 | **38.1%** | +| `mixed-with-code.md` | 888 | 560 | **36.9%** | +| **Average** | **898** | **481** | **46%** | + +All validations passed ✅ — headings, code blocks, URLs, file paths preserved exactly. + +## Before / After + +<table> +<tr> +<td width="50%"> + +### 📄 Original (706 tokens) + +> "I strongly prefer TypeScript with strict mode enabled for all new code. Please don't use `any` type unless there's genuinely no way around it, and if you do, leave a comment explaining the reasoning. I find that taking the time to properly type things catches a lot of bugs before they ever make it to runtime." + +</td> +<td width="50%"> + +### <img src="../../docs/assets/dancing-rock.svg" width="20" height="20" alt="rock"/> Caveman (285 tokens) + +> "Prefer TypeScript strict mode always. No `any` unless unavoidable — comment why if used. Proper types catch bugs early." + +</td> +</tr> +</table> + +**Same instructions. 60% fewer tokens. Every. Single. Session.** + +## Security + +`caveman-compress` is flagged as Snyk High Risk due to subprocess and file I/O patterns detected by static analysis. This is a false positive — see [SECURITY.md](./SECURITY.md) for a full explanation of what the skill does and does not do. + +## Install + +Compress is built in with the `caveman` plugin. Install `caveman` once, then use `/caveman-compress`. + +If you need local files, the compress skill lives at: + +```bash +caveman-compress/ +``` + +**Requires:** Python 3.10+ + +## Usage + +``` +/caveman-compress <filepath> +``` + +Examples: +``` +/caveman-compress CLAUDE.md +/caveman-compress docs/preferences.md +/caveman-compress todos.md +``` + +### What files work + +| Type | Compress? | +|------|-----------| +| `.md`, `.txt`, `.rst`, `.typ`, `.typst`, `.tex` | ✅ Yes | +| Extensionless natural language | ✅ Yes | +| `.py`, `.js`, `.ts`, `.json`, `.yaml` | ❌ Skip (code/config) | +| `*.original.md` | ❌ Skip (backup files) | + +## How It Work + +``` +/caveman-compress CLAUDE.md + ↓ +detect file type (no tokens) + ↓ +Claude compresses (tokens — one call) + ↓ +validate output (no tokens) + checks: headings, code blocks, URLs, file paths, bullets + ↓ +if errors: Claude fixes cherry-picked issues only (tokens — targeted fix) + does NOT recompress — only patches broken parts + ↓ +retry up to 2 times + ↓ +write compressed → CLAUDE.md +write original → CLAUDE.original.md +``` + +Only two things use tokens: initial compression + targeted fix if validation fails. Everything else is local Python. + +## What Is Preserved + +Caveman compress natural language. It never touch: + +- Code blocks (` ``` ` fenced or indented) +- Inline code (`` `backtick content` ``) +- URLs and links +- File paths (`/src/components/...`) +- Commands (`npm install`, `git commit`) +- Technical terms, library names, API names +- Headings (exact text preserved) +- Tables (structure preserved, cell text compressed) +- Dates, version numbers, numeric values + +## Why This Matter + +`CLAUDE.md` loads on **every session start**. A 1000-token project memory file costs tokens every single time you open a project. Over 100 sessions that's 100,000 tokens of overhead — just for context you already wrote. + +Caveman cut that by ~46% on average. Same instructions. Same accuracy. Less waste. + +``` +┌────────────────────────────────────────────┐ +│ TOKEN SAVINGS PER FILE █████ 46% │ +│ SESSIONS THAT BENEFIT ██████████ 100% │ +│ INFORMATION PRESERVED ██████████ 100% │ +│ SETUP TIME █ 1x │ +└────────────────────────────────────────────┘ +``` + +## Part of Caveman + +This skill is part of the [caveman](https://github.com/JuliusBrussee/caveman) toolkit — making Claude use fewer tokens without losing accuracy. + +- **caveman** — make Claude *speak* like caveman (cuts response tokens ~65%) +- **caveman-compress** — make Claude *read* less (cuts context tokens ~46%) diff --git a/.agents/skills/caveman-compress/SECURITY.md b/.agents/skills/caveman-compress/SECURITY.md new file mode 100644 index 0000000..693108c --- /dev/null +++ b/.agents/skills/caveman-compress/SECURITY.md @@ -0,0 +1,31 @@ +# Security + +## Snyk High Risk Rating + +`caveman-compress` receives a Snyk High Risk rating due to static analysis heuristics. This document explains what the skill does and does not do. + +### What triggers the rating + +1. **subprocess usage**: The skill calls the `claude` CLI via `subprocess.run()` as a fallback when `ANTHROPIC_API_KEY` is not set. The subprocess call uses a fixed argument list — no shell interpolation occurs. User file content is passed via stdin, not as a shell argument. + +2. **File read/write**: The skill reads the file the user explicitly points it at, compresses it, and writes the result back to the same path. A `.original.md` backup is saved alongside it. No files outside the user-specified path are read or written. + +### What the skill does NOT do + +- Does not execute user file content as code +- Does not make network requests except to Anthropic's API (via SDK or CLI) +- Does not access files outside the path the user provides +- Does not use shell=True or string interpolation in subprocess calls +- Does not collect or transmit any data beyond the file being compressed + +### Auth behavior + +If `ANTHROPIC_API_KEY` is set, the skill uses the Anthropic Python SDK directly (no subprocess). If not set, it falls back to the `claude` CLI, which uses the user's existing Claude desktop authentication. + +### File size limit + +Files larger than 500KB are rejected before any API call is made. + +### Reporting a vulnerability + +If you believe you've found a genuine security issue, please open a GitHub issue with the label `security`. diff --git a/.agents/skills/caveman-compress/SKILL.md b/.agents/skills/caveman-compress/SKILL.md new file mode 100644 index 0000000..00ce454 --- /dev/null +++ b/.agents/skills/caveman-compress/SKILL.md @@ -0,0 +1,111 @@ +--- +name: caveman-compress +description: > + Compress natural language memory files (CLAUDE.md, todos, preferences) into caveman format + to save input tokens. Preserves all technical substance, code, URLs, and structure. + Compressed version overwrites the original file. Human-readable backup saved as FILE.original.md. + Trigger: /caveman-compress FILEPATH or "compress memory file" +--- + +# Caveman Compress + +## Purpose + +Compress natural language files (CLAUDE.md, todos, preferences) into caveman-speak to reduce input tokens. Compressed version overwrites original. Human-readable backup saved as `<filename>.original.md`. + +## Trigger + +`/caveman-compress <filepath>` or when user asks to compress a memory file. + +## Process + +1. The compression scripts live in `scripts/` (adjacent to this SKILL.md). If the path is not immediately available, search for `scripts/__main__.py` next to this SKILL.md. + +2. From the directory containing this SKILL.md, run: + +python3 -m scripts <absolute_filepath> + +3. The CLI will: +- detect file type (no tokens) +- call Claude to compress +- validate output (no tokens) +- if errors: cherry-pick fix with Claude (targeted fixes only, no recompression) +- retry up to 2 times +- if still failing after 2 retries: report error to user, leave original file untouched + +4. Return result to user + +## Compression Rules + +### Remove +- Articles: a, an, the +- Filler: just, really, basically, actually, simply, essentially, generally +- Pleasantries: "sure", "certainly", "of course", "happy to", "I'd recommend" +- Hedging: "it might be worth", "you could consider", "it would be good to" +- Redundant phrasing: "in order to" → "to", "make sure to" → "ensure", "the reason is because" → "because" +- Connective fluff: "however", "furthermore", "additionally", "in addition" + +### Preserve EXACTLY (never modify) +- Code blocks (fenced ``` and indented) +- Inline code (`backtick content`) +- URLs and links (full URLs, markdown links) +- File paths (`/src/components/...`, `./config.yaml`) +- Commands (`npm install`, `git commit`, `docker build`) +- Technical terms (library names, API names, protocols, algorithms) +- Proper nouns (project names, people, companies) +- Dates, version numbers, numeric values +- Environment variables (`$HOME`, `NODE_ENV`) + +### Preserve Structure +- All markdown headings (keep exact heading text, compress body below) +- Bullet point hierarchy (keep nesting level) +- Numbered lists (keep numbering) +- Tables (compress cell text, keep structure) +- Frontmatter/YAML headers in markdown files + +### Compress +- Use short synonyms: "big" not "extensive", "fix" not "implement a solution for", "use" not "utilize" +- Fragments OK: "Run tests before commit" not "You should always run tests before committing" +- Drop "you should", "make sure to", "remember to" — just state the action +- Merge redundant bullets that say the same thing differently +- Keep one example where multiple examples show the same pattern + +CRITICAL RULE: +Anything inside ``` ... ``` must be copied EXACTLY. +Do not: +- remove comments +- remove spacing +- reorder lines +- shorten commands +- simplify anything + +Inline code (`...`) must be preserved EXACTLY. +Do not modify anything inside backticks. + +If file contains code blocks: +- Treat code blocks as read-only regions +- Only compress text outside them +- Do not merge sections around code + +## Pattern + +Original: +> You should always make sure to run the test suite before pushing any changes to the main branch. This is important because it helps catch bugs early and prevents broken builds from being deployed to production. + +Compressed: +> Run tests before push to main. Catch bugs early, prevent broken prod deploys. + +Original: +> The application uses a microservices architecture with the following components. The API gateway handles all incoming requests and routes them to the appropriate service. The authentication service is responsible for managing user sessions and JWT tokens. + +Compressed: +> Microservices architecture. API gateway route all requests to services. Auth service manage user sessions + JWT tokens. + +## Boundaries + +- ONLY compress natural language files (.md, .txt, .typ, .typst, .tex, extensionless) +- NEVER modify: .py, .js, .ts, .json, .yaml, .yml, .toml, .env, .lock, .css, .html, .xml, .sql, .sh +- If file has mixed content (prose + code), compress ONLY the prose sections +- If unsure whether something is code or prose, leave it unchanged +- Original file is backed up as FILE.original.md before overwriting +- Never compress FILE.original.md (skip it) diff --git a/.agents/skills/caveman-compress/scripts/__init__.py b/.agents/skills/caveman-compress/scripts/__init__.py new file mode 100644 index 0000000..16b8c53 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/__init__.py @@ -0,0 +1,9 @@ +"""Caveman compress scripts. + +This package provides tools to compress natural language markdown files +into caveman format to save input tokens. +""" + +__all__ = ["cli", "compress", "detect", "validate"] + +__version__ = "1.0.0" diff --git a/.agents/skills/caveman-compress/scripts/__main__.py b/.agents/skills/caveman-compress/scripts/__main__.py new file mode 100644 index 0000000..4e28416 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/__main__.py @@ -0,0 +1,3 @@ +from .cli import main + +main() diff --git a/.agents/skills/caveman-compress/scripts/benchmark.py b/.agents/skills/caveman-compress/scripts/benchmark.py new file mode 100644 index 0000000..f9e2ee0 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/benchmark.py @@ -0,0 +1,80 @@ +#!/usr/bin/env python3 +from pathlib import Path +import sys + +# Support both direct execution and module import +try: + from .validate import validate +except ImportError: + sys.path.insert(0, str(Path(__file__).parent)) + from validate import validate + +try: + import tiktoken + _enc = tiktoken.get_encoding("o200k_base") +except ImportError: + _enc = None + + +def count_tokens(text): + if _enc is None: + return len(text.split()) # fallback: word count + return len(_enc.encode(text)) + + +def benchmark_pair(orig_path: Path, comp_path: Path): + orig_text = orig_path.read_text() + comp_text = comp_path.read_text() + + orig_tokens = count_tokens(orig_text) + comp_tokens = count_tokens(comp_text) + saved = 100 * (orig_tokens - comp_tokens) / orig_tokens if orig_tokens > 0 else 0.0 + result = validate(orig_path, comp_path) + + return (comp_path.name, orig_tokens, comp_tokens, saved, result.is_valid) + + +def print_table(rows): + print("\n| File | Original | Compressed | Saved % | Valid |") + print("|------|----------|------------|---------|-------|") + for r in rows: + print(f"| {r[0]} | {r[1]} | {r[2]} | {r[3]:.1f}% | {'✅' if r[4] else '❌'} |") + + +def main(): + # Direct file pair: python3 benchmark.py original.md compressed.md + if len(sys.argv) == 3: + orig = Path(sys.argv[1]).resolve() + comp = Path(sys.argv[2]).resolve() + if not orig.exists(): + print(f"❌ Not found: {orig}") + sys.exit(1) + if not comp.exists(): + print(f"❌ Not found: {comp}") + sys.exit(1) + print_table([benchmark_pair(orig, comp)]) + return + + # Glob mode: repo_root/tests/caveman-compress/ + # __file__ lives at <repo_root>/skills/caveman-compress/scripts/benchmark.py + # Walk up four dirs: scripts → caveman-compress → skills → repo_root. + tests_dir = Path(__file__).resolve().parents[3] / "tests" / "caveman-compress" + if not tests_dir.exists(): + print(f"❌ Tests dir not found: {tests_dir}") + sys.exit(1) + + rows = [] + for orig in sorted(tests_dir.glob("*.original.md")): + comp = orig.with_name(orig.stem.removesuffix(".original") + ".md") + if comp.exists(): + rows.append(benchmark_pair(orig, comp)) + + if not rows: + print("No compressed file pairs found.") + return + + print_table(rows) + + +if __name__ == "__main__": + main() diff --git a/.agents/skills/caveman-compress/scripts/cli.py b/.agents/skills/caveman-compress/scripts/cli.py new file mode 100644 index 0000000..75ea8a6 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/cli.py @@ -0,0 +1,85 @@ +#!/usr/bin/env python3 +""" +Caveman Compress CLI + +Usage: + caveman <filepath> +""" + +import sys + +# Force UTF-8 on stdout/stderr before any code can print. Windows consoles +# default to cp1252 and crash on the ❌ glyphs in error/validation branches, +# masking the real error and leaving the user with a half-compressed file. +for _stream in (sys.stdout, sys.stderr): + reconfigure = getattr(_stream, "reconfigure", None) + if callable(reconfigure): + try: + reconfigure(encoding="utf-8", errors="replace") + except Exception: + pass + +from pathlib import Path + +from .compress import backup_dir_for, compress_file +from .detect import detect_file_type, should_compress + + +def print_usage(): + print("Usage: caveman <filepath>") + + +def main(): + if len(sys.argv) != 2: + print_usage() + sys.exit(1) + + filepath = Path(sys.argv[1]) + + # Check file exists + if not filepath.exists(): + print(f"❌ File not found: {filepath}") + sys.exit(1) + + if not filepath.is_file(): + print(f"❌ Not a file: {filepath}") + sys.exit(1) + + filepath = filepath.resolve() + + # Detect file type + file_type = detect_file_type(filepath) + + print(f"Detected: {file_type}") + + # Check if compressible + if not should_compress(filepath): + print("Skipping: file is not natural language (code/config)") + sys.exit(0) + + print("Starting caveman compression...\n") + + try: + success = compress_file(filepath) + + if success: + print("\nCompression completed successfully") + backup_path = backup_dir_for(filepath) / (filepath.stem + ".original.md") + print(f"Compressed: {filepath}") + print(f"Original: {backup_path}") + sys.exit(0) + else: + print("\n❌ Compression failed after retries") + sys.exit(2) + + except KeyboardInterrupt: + print("\nInterrupted by user") + sys.exit(130) + + except Exception as e: + print(f"\n❌ Error: {e}") + sys.exit(1) + + +if __name__ == "__main__": + main() diff --git a/.agents/skills/caveman-compress/scripts/compress.py b/.agents/skills/caveman-compress/scripts/compress.py new file mode 100644 index 0000000..e93d934 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/compress.py @@ -0,0 +1,342 @@ +#!/usr/bin/env python3 +""" +Caveman Memory Compression Orchestrator + +Usage: + python scripts/compress.py <filepath> +""" + +import os +import re +import shutil +import subprocess +import sys +from pathlib import Path +from typing import List + +OUTER_FENCE_REGEX = re.compile( + r"\A\s*(`{3,}|~{3,})[^\n]*\n(.*)\n\1\s*\Z", re.DOTALL +) + +# YAML frontmatter: starts at file start with --- on its own line, ends with --- on its own line. +# Captures the entire block (including delimiters and trailing newline) and the body after. +FRONTMATTER_REGEX = re.compile( + r"\A(---\r?\n.*?\r?\n---\r?\n)(.*)", re.DOTALL +) + + +def split_frontmatter(text: str): + """Split YAML frontmatter from body. Returns (frontmatter, body). + + Memory files (and many other markdown docs) start with a YAML frontmatter + block delimited by `---` lines. The compression LLM has a habit of stripping + or rewriting these despite preserve-structure rules in the prompt — so we + surgically remove the frontmatter before compression and prepend it back + verbatim to the output. Files without frontmatter pass through unchanged. + """ + m = FRONTMATTER_REGEX.match(text) + if m: + return m.group(1), m.group(2) + return "", text + +# Filenames and paths that almost certainly hold secrets or PII. Compressing +# them ships raw bytes to the Anthropic API — a third-party data boundary that +# developers on sensitive codebases cannot cross. detect.py already skips .env +# by extension, but credentials.md / secrets.txt / ~/.aws/credentials would +# slip through the natural-language filter. This is a hard refuse before read. +SENSITIVE_BASENAME_REGEX = re.compile( + r"(?ix)^(" + r"\.env(\..+)?" + r"|\.netrc" + r"|credentials(\..+)?" + r"|secrets?(\..+)?" + r"|passwords?(\..+)?" + r"|id_(rsa|dsa|ecdsa|ed25519)(\.pub)?" + r"|authorized_keys" + r"|known_hosts" + r"|.*\.(pem|key|p12|pfx|crt|cer|jks|keystore|asc|gpg)" + r")$" +) + +SENSITIVE_PATH_COMPONENTS = frozenset({".ssh", ".aws", ".gnupg", ".kube", ".docker"}) + +SENSITIVE_NAME_TOKENS = ( + "secret", "credential", "password", "passwd", + "apikey", "accesskey", "token", "privatekey", +) + + +def backup_dir_for(filepath: Path) -> Path: + """Resolve the out-of-tree backup directory for a given source file. + + Backups must live OUTSIDE the source directory so skill auto-loaders + (Claude Code rules/, opencode instructions/, etc.) stop re-ingesting the + `.original.md` copies as live files. Base dir is platform-aware: + - Windows: %LOCALAPPDATA%\\caveman-compress\\backups + - else: $XDG_DATA_HOME/caveman-compress/backups if set, + else ~/.local/share/caveman-compress/backups + + The source file's parent-dir name is mirrored under the base to reduce + cross-project collisions (e.g. two `task.md` files in different repos). + """ + if os.name == "nt" or sys.platform == "win32": + local_appdata = os.environ.get("LOCALAPPDATA") + base = Path(local_appdata) if local_appdata else Path.home() / "AppData" / "Local" + base = base / "caveman-compress" / "backups" + else: + xdg = os.environ.get("XDG_DATA_HOME") + base = Path(xdg) if xdg else Path.home() / ".local" / "share" + base = base / "caveman-compress" / "backups" + return base / filepath.parent.name + + +def is_sensitive_path(filepath: Path) -> bool: + """Heuristic denylist for files that must never be shipped to a third-party API.""" + name = filepath.name + if SENSITIVE_BASENAME_REGEX.match(name): + return True + lowered_parts = {p.lower() for p in filepath.parts} + if lowered_parts & SENSITIVE_PATH_COMPONENTS: + return True + # Normalize separators so "api-key" and "api_key" both match "apikey". + lower = re.sub(r"[_\-\s.]", "", name.lower()) + return any(tok in lower for tok in SENSITIVE_NAME_TOKENS) + + +def strip_llm_wrapper(text: str) -> str: + """Strip outer ```markdown ... ``` fence when it wraps the entire output.""" + m = OUTER_FENCE_REGEX.match(text) + if m: + return m.group(2) + return text + +from .detect import should_compress +from .validate import validate + +MAX_RETRIES = 2 + + +# ---------- Claude Calls ---------- + + +def call_claude(prompt: str) -> str: + """Send a prompt to Claude. + + Prefers the Anthropic SDK when ANTHROPIC_API_KEY is set; otherwise falls + back to the ``claude --print`` CLI (which handles desktop auth). + + On Windows the CLI subprocess decoding defaults to the system codepage + (cp1251 / cp1252) and crashes on UTF-8 output — see issue #152. Pinning + ``encoding="utf-8"`` with ``errors="replace"`` matches the CLI's actual + native I/O and prevents the UnicodeDecodeError before validation can + report. Windows users with non-ASCII content can also set + ``ANTHROPIC_API_KEY`` to route through the SDK and skip the subprocess. + """ + api_key = os.environ.get("ANTHROPIC_API_KEY") + if api_key: + try: + import anthropic + + client = anthropic.Anthropic(api_key=api_key) + msg = client.messages.create( + model=os.environ.get("CAVEMAN_MODEL", "claude-sonnet-4-5"), + max_tokens=8192, + messages=[{"role": "user", "content": prompt}], + ) + return strip_llm_wrapper(msg.content[0].text.strip()) + except ImportError: + pass # anthropic not installed, fall back to CLI + # Fallback: use claude CLI (handles desktop auth). + # Resolve binary via shutil.which so Windows .cmd/.bat shims (e.g. + # %APPDATA%\npm\claude.CMD) work without shell=True. On POSIX, + # shutil.which returns the same absolute path as the implicit lookup, + # so this is a no-op there. Falls back to bare "claude" if not found + # on PATH so subprocess raises a clear FileNotFoundError. + claude_bin = shutil.which("claude") or "claude" + try: + result = subprocess.run( + [claude_bin, "--print"], + input=prompt, + text=True, + capture_output=True, + check=True, + encoding="utf-8", + errors="replace", + ) + return strip_llm_wrapper(result.stdout.strip()) + except subprocess.CalledProcessError as e: + raise RuntimeError(f"Claude call failed:\n{e.stderr}") + + +def build_compress_prompt(original: str) -> str: + return f""" +Compress this markdown into caveman format. + +STRICT RULES: +- Do NOT modify anything inside ``` code blocks +- Do NOT modify anything inside inline backticks +- Preserve ALL URLs exactly +- Preserve ALL headings exactly +- Preserve file paths and commands +- Return ONLY the compressed markdown body — do NOT wrap the entire output in a ```markdown fence or any other fence. Inner code blocks from the original stay as-is; do not add a new outer fence around the whole file. + +Only compress natural language. + +TEXT: +{original} +""" + + +def build_fix_prompt(original: str, compressed: str, errors: List[str]) -> str: + errors_str = "\n".join(f"- {e}" for e in errors) + return f"""You are fixing a caveman-compressed markdown file. Specific validation errors were found. + +CRITICAL RULES: +- DO NOT recompress or rephrase the file +- ONLY fix the listed errors — leave everything else exactly as-is +- The ORIGINAL is provided as reference only (to restore missing content) +- Preserve caveman style in all untouched sections + +ERRORS TO FIX: +{errors_str} + +HOW TO FIX: +- Missing URL: find it in ORIGINAL, restore it exactly where it belongs in COMPRESSED +- Code block mismatch: find the exact code block in ORIGINAL, restore it in COMPRESSED +- Heading mismatch: restore the exact heading text from ORIGINAL into COMPRESSED +- Do not touch any section not mentioned in the errors + +ORIGINAL (reference only): +{original} + +COMPRESSED (fix this): +{compressed} + +Return ONLY the fixed compressed file. No explanation. +""" + + +# ---------- Core Logic ---------- + + +def compress_file(filepath: Path) -> bool: + # Resolve and validate path + filepath = filepath.resolve() + MAX_FILE_SIZE = 500_000 # 500KB + if not filepath.exists(): + raise FileNotFoundError(f"File not found: {filepath}") + if filepath.stat().st_size > MAX_FILE_SIZE: + raise ValueError(f"File too large to compress safely (max 500KB): {filepath}") + + # Refuse files that look like they contain secrets or PII. Compressing ships + # the raw bytes to the Anthropic API — a third-party boundary — so we fail + # loudly rather than silently exfiltrate credentials or keys. Override is + # intentional: the user must rename the file if the heuristic is wrong. + if is_sensitive_path(filepath): + raise ValueError( + f"Refusing to compress {filepath}: filename looks sensitive " + "(credentials, keys, secrets, or known private paths). " + "Compression sends file contents to the Anthropic API. " + "Rename the file if this is a false positive." + ) + + print(f"Processing: {filepath}") + + if not should_compress(filepath): + print("Skipping (not natural language)") + return False + + original_text = filepath.read_text(errors="ignore") + # Store backup outside the source directory so skill auto-loaders don't + # re-ingest the `.original.md` copy as a live file. Mirror the source's + # parent-dir name + stem under a platform-aware base to reduce collisions. + backup_dir = backup_dir_for(filepath) + backup_dir.mkdir(parents=True, exist_ok=True) + backup_path = backup_dir / (filepath.stem + ".original.md") + + if not original_text.strip(): + print("❌ Refusing to compress: file is empty or whitespace-only.") + return False + + # Check if backup already exists to prevent accidental overwriting + if backup_path.exists(): + print(f"⚠️ Backup file already exists: {backup_path}") + print("The original backup may contain important content.") + print("Aborting to prevent data loss. Please remove or rename the backup file if you want to proceed.") + return False + + # Split YAML frontmatter off before compression. Claude tends to strip or + # rewrite frontmatter despite preserve-structure rules; we keep it verbatim + # by removing it from the input and re-prepending it to the output. + frontmatter, body = split_frontmatter(original_text) + if frontmatter: + print(f"Detected YAML frontmatter ({len(frontmatter)} chars) — preserving verbatim") + + if not body.strip(): + print("❌ Refusing to compress: body is empty after frontmatter removal.") + return False + + # Step 1: Compress (body only, frontmatter excluded) + print("Compressing with Claude...") + compressed_body = call_claude(build_compress_prompt(body)) + + if compressed_body is None or not compressed_body.strip(): + print("❌ Compression aborted: Claude returned an empty response.") + print(" Original file is untouched (no backup created).") + return False + + # Compare the BODY (not the whole file) — frontmatter is preserved verbatim + # and would never change, so identity must be judged on the compressible part. + if compressed_body.strip() == body.strip(): + print("❌ Compression aborted: output is identical to input.") + print(" Likely causes: Claude refused, returned the prompt verbatim, or the file is") + print(" already in caveman form. Original file is untouched (no backup created).") + return False + + # Reassemble: frontmatter (verbatim) + compressed body + compressed = frontmatter + compressed_body + + # Save original as backup, then verify the backup readback before + # touching the input file. If the filesystem dropped bytes (encoding, + # antivirus, disk full), unlink the bad backup and abort instead of + # leaving the user with a corrupt backup + compressed primary. + backup_path.write_text(original_text) + backup_readback = backup_path.read_text(errors="ignore") + if backup_readback != original_text: + print(f"❌ Backup write verification failed: {backup_path}") + print(" In-memory original differs from on-disk backup. Aborting before touching the input file.") + try: + backup_path.unlink() + except OSError: + pass + return False + filepath.write_text(compressed) + + # Step 2: Validate + Retry + for attempt in range(MAX_RETRIES): + print(f"\nValidation attempt {attempt + 1}") + + result = validate(backup_path, filepath) + + if result.is_valid: + print("Validation passed") + break + + print("❌ Validation failed:") + for err in result.errors: + print(f" - {err}") + + if attempt == MAX_RETRIES - 1: + # Restore original on failure + filepath.write_text(original_text) + backup_path.unlink(missing_ok=True) + print("❌ Failed after retries — original restored") + return False + + print("Fixing with Claude...") + compressed = call_claude( + build_fix_prompt(original_text, compressed, result.errors) + ) + filepath.write_text(compressed) + + return True diff --git a/.agents/skills/caveman-compress/scripts/detect.py b/.agents/skills/caveman-compress/scripts/detect.py new file mode 100644 index 0000000..8d5f6d7 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/detect.py @@ -0,0 +1,121 @@ +#!/usr/bin/env python3 +"""Detect whether a file is natural language (compressible) or code/config (skip).""" + +import json +import re +from pathlib import Path + +# Extensions that are natural language and compressible +COMPRESSIBLE_EXTENSIONS = {".md", ".txt", ".markdown", ".rst", ".typ", ".typst", ".tex"} + +# Extensions that are code/config and should be skipped +SKIP_EXTENSIONS = { + ".py", ".js", ".ts", ".tsx", ".jsx", ".json", ".yaml", ".yml", + ".toml", ".env", ".lock", ".css", ".scss", ".html", ".xml", + ".sql", ".sh", ".bash", ".zsh", ".go", ".rs", ".java", ".c", + ".cpp", ".h", ".hpp", ".rb", ".php", ".swift", ".kt", ".lua", + ".dockerfile", ".makefile", ".csv", ".ini", ".cfg", +} + +# Patterns that indicate a line is code +CODE_PATTERNS = [ + re.compile(r"^\s*(import |from .+ import |require\(|const |let |var )"), + re.compile(r"^\s*(def |class |function |async function |export )"), + re.compile(r"^\s*(if\s*\(|for\s*\(|while\s*\(|switch\s*\(|try\s*\{)"), + re.compile(r"^\s*[\}\]\);]+\s*$"), # closing braces/brackets + re.compile(r"^\s*@\w+"), # decorators/annotations + re.compile(r'^\s*"[^"]+"\s*:\s*'), # JSON-like key-value + re.compile(r"^\s*\w+\s*=\s*[{\[\(\"']"), # assignment with literal +] + + +def _is_code_line(line: str) -> bool: + """Check if a line looks like code.""" + return any(p.match(line) for p in CODE_PATTERNS) + + +def _is_json_content(text: str) -> bool: + """Check if content is valid JSON.""" + try: + json.loads(text) + return True + except (json.JSONDecodeError, ValueError): + return False + + +def _is_yaml_content(lines: list[str]) -> bool: + """Heuristic: check if content looks like YAML.""" + yaml_indicators = 0 + for line in lines[:30]: + stripped = line.strip() + if stripped.startswith("---"): + yaml_indicators += 1 + elif re.match(r"^\w[\w\s]*:\s", stripped): + yaml_indicators += 1 + elif stripped.startswith("- ") and ":" in stripped: + yaml_indicators += 1 + # If most non-empty lines look like YAML + non_empty = sum(1 for l in lines[:30] if l.strip()) + return non_empty > 0 and yaml_indicators / non_empty > 0.6 + + +def detect_file_type(filepath: Path) -> str: + """Classify a file as 'natural_language', 'code', 'config', or 'unknown'. + + Returns: + One of: 'natural_language', 'code', 'config', 'unknown' + """ + ext = filepath.suffix.lower() + + # Extension-based classification + if ext in COMPRESSIBLE_EXTENSIONS: + return "natural_language" + if ext in SKIP_EXTENSIONS: + return "code" if ext not in {".json", ".yaml", ".yml", ".toml", ".ini", ".cfg", ".env"} else "config" + + # Extensionless files (like CLAUDE.md, TODO) — check content + if not ext: + try: + text = filepath.read_text(errors="ignore") + except (OSError, PermissionError): + return "unknown" + + lines = text.splitlines()[:50] + + if _is_json_content(text[:10000]): + return "config" + if _is_yaml_content(lines): + return "config" + + code_lines = sum(1 for l in lines if l.strip() and _is_code_line(l)) + non_empty = sum(1 for l in lines if l.strip()) + if non_empty > 0 and code_lines / non_empty > 0.4: + return "code" + + return "natural_language" + + return "unknown" + + +def should_compress(filepath: Path) -> bool: + """Return True if the file is natural language and should be compressed.""" + if not filepath.is_file(): + return False + # Skip backup files + if filepath.name.endswith(".original.md"): + return False + return detect_file_type(filepath) == "natural_language" + + +if __name__ == "__main__": + import sys + + if len(sys.argv) < 2: + print("Usage: python detect.py <file1> [file2] ...") + sys.exit(1) + + for path_str in sys.argv[1:]: + p = Path(path_str).resolve() + file_type = detect_file_type(p) + compress = should_compress(p) + print(f" {p.name:30s} type={file_type:20s} compress={compress}") diff --git a/.agents/skills/caveman-compress/scripts/validate.py b/.agents/skills/caveman-compress/scripts/validate.py new file mode 100644 index 0000000..dc07307 --- /dev/null +++ b/.agents/skills/caveman-compress/scripts/validate.py @@ -0,0 +1,213 @@ +#!/usr/bin/env python3 +import re +from collections import Counter +from pathlib import Path + +URL_REGEX = re.compile(r"https?://[^\s)]+") +FENCE_OPEN_REGEX = re.compile(r"^(\s{0,3})(`{3,}|~{3,})(.*)$") +HEADING_REGEX = re.compile(r"^(#{1,6})\s+(.*)", re.MULTILINE) +BULLET_REGEX = re.compile(r"^\s*[-*+]\s+", re.MULTILINE) + +# crude but effective path detection +# Requires either a path prefix (./ ../ / or drive letter) or a slash/backslash within the match +PATH_REGEX = re.compile(r"(?:\./|\.\./|/|[A-Za-z]:\\)[\w\-/\\\.]+|[\w\-\.]+[/\\][\w\-/\\\.]+") + + +class ValidationResult: + def __init__(self): + self.is_valid = True + self.errors = [] + self.warnings = [] + + def add_error(self, msg): + self.is_valid = False + self.errors.append(msg) + + def add_warning(self, msg): + self.warnings.append(msg) + + +def read_file(path: Path) -> str: + return path.read_text(errors="ignore") + + +# ---------- Extractors ---------- + + +def extract_headings(text): + return [(level, title.strip()) for level, title in HEADING_REGEX.findall(text)] + + +def extract_code_blocks(text): + """Line-based fenced code block extractor. + + Handles ``` and ~~~ fences with variable length (CommonMark: closing + fence must use same char and be at least as long as opening). Supports + nested fences (e.g. an outer 4-backtick block wrapping inner 3-backtick + content). + """ + blocks = [] + lines = text.split("\n") + i = 0 + n = len(lines) + while i < n: + m = FENCE_OPEN_REGEX.match(lines[i]) + if not m: + i += 1 + continue + fence_char = m.group(2)[0] + fence_len = len(m.group(2)) + open_line = lines[i] + block_lines = [open_line] + i += 1 + closed = False + while i < n: + close_m = FENCE_OPEN_REGEX.match(lines[i]) + if ( + close_m + and close_m.group(2)[0] == fence_char + and len(close_m.group(2)) >= fence_len + and close_m.group(3).strip() == "" + ): + block_lines.append(lines[i]) + closed = True + i += 1 + break + block_lines.append(lines[i]) + i += 1 + if closed: + blocks.append("\n".join(block_lines)) + # Unclosed fences are silently skipped — they indicate malformed markdown + # and including them would cause false-positive validation failures. + return blocks + + +def extract_urls(text): + return set(URL_REGEX.findall(text)) + + +def extract_paths(text): + return set(PATH_REGEX.findall(text)) + + +def count_bullets(text): + return len(BULLET_REGEX.findall(text)) + + +def extract_inline_codes(text): + text_without_fences = re.sub(r"^```[\s\S]*?^```", "", text, flags=re.MULTILINE) + text_without_fences = re.sub(r"^~~~[\s\S]*?^~~~", "", text_without_fences, flags=re.MULTILINE) + return re.findall(r"`([^`]+)`", text_without_fences) + + +# ---------- Validators ---------- + + +def validate_headings(orig, comp, result): + h1 = extract_headings(orig) + h2 = extract_headings(comp) + + if len(h1) != len(h2): + result.add_error(f"Heading count mismatch: {len(h1)} vs {len(h2)}") + + if h1 != h2: + result.add_warning("Heading text/order changed") + + +def validate_code_blocks(orig, comp, result): + c1 = extract_code_blocks(orig) + c2 = extract_code_blocks(comp) + + if c1 != c2: + result.add_error("Code blocks not preserved exactly") + + +def validate_urls(orig, comp, result): + u1 = extract_urls(orig) + u2 = extract_urls(comp) + + if u1 != u2: + result.add_error(f"URL mismatch: lost={u1 - u2}, added={u2 - u1}") + + +def validate_paths(orig, comp, result): + p1 = extract_paths(orig) + p2 = extract_paths(comp) + + if p1 != p2: + result.add_warning(f"Path mismatch: lost={p1 - p2}, added={p2 - p1}") + + +def validate_bullets(orig, comp, result): + b1 = count_bullets(orig) + b2 = count_bullets(comp) + + if b1 == 0: + return + + diff = abs(b1 - b2) / b1 + + if diff > 0.15: + result.add_warning(f"Bullet count changed too much: {b1} -> {b2}") + + +def validate_inline_codes(orig, comp, result): + c1 = Counter(extract_inline_codes(orig)) + c2 = Counter(extract_inline_codes(comp)) + + if c1 != c2: + lost = set(c1.keys()) - set(c2.keys()) + added = set(c2.keys()) - set(c1.keys()) + for code, count in c1.items(): + if code in c2 and c2[code] < count: + lost.add(f"{code} (lost {count - c2[code]} of {count} occurrences)") + if lost: + result.add_error(f"Inline code lost: {lost}") + if added: + result.add_warning(f"Inline code added: {added}") + + +# ---------- Main ---------- + + +def validate(original_path: Path, compressed_path: Path) -> ValidationResult: + result = ValidationResult() + + orig = read_file(original_path) + comp = read_file(compressed_path) + + validate_headings(orig, comp, result) + validate_code_blocks(orig, comp, result) + validate_urls(orig, comp, result) + validate_paths(orig, comp, result) + validate_bullets(orig, comp, result) + validate_inline_codes(orig, comp, result) + + return result + + +# ---------- CLI ---------- + +if __name__ == "__main__": + import sys + + if len(sys.argv) != 3: + print("Usage: python validate.py <original> <compressed>") + sys.exit(1) + + orig = Path(sys.argv[1]).resolve() + comp = Path(sys.argv[2]).resolve() + + res = validate(orig, comp) + + print(f"\nValid: {res.is_valid}") + + if res.errors: + print("\nErrors:") + for e in res.errors: + print(f" - {e}") + + if res.warnings: + print("\nWarnings:") + for w in res.warnings: + print(f" - {w}") diff --git a/.agents/skills/caveman-help/README.md b/.agents/skills/caveman-help/README.md new file mode 100644 index 0000000..5841256 --- /dev/null +++ b/.agents/skills/caveman-help/README.md @@ -0,0 +1,38 @@ +# caveman-help + +Quick-reference card. One shot, no mode change. + +## What it does + +Prints a cheat sheet of all caveman modes, sibling skills, deactivation triggers, and how to set the default mode via env var or config file. One-shot display — does not flip the active mode, write flag files, or persist anything. Use when you forget the slash commands. + +## How to invoke + +``` +/caveman-help +``` + +Also triggers on "caveman help", "what caveman commands", "how do I use caveman". + +## Example output + +``` +Modes: + /caveman full (default) + /caveman lite lighter + /caveman ultra extreme + /caveman wenyan classical Chinese + +Skills: + /caveman-commit terse Conventional Commits + /caveman-review one-line PR comments + /caveman-stats session token savings + +Deactivate: + "stop caveman" or "normal mode" +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full reference card +- [Caveman README](../../README.md) — repo overview diff --git a/.agents/skills/caveman-help/SKILL.md b/.agents/skills/caveman-help/SKILL.md new file mode 100644 index 0000000..ce1156f --- /dev/null +++ b/.agents/skills/caveman-help/SKILL.md @@ -0,0 +1,63 @@ +--- +name: caveman-help +description: > + Quick-reference card for all caveman modes, skills, and commands. + One-shot display, not a persistent mode. Trigger: /caveman-help, + "caveman help", "what caveman commands", "how do I use caveman". +--- + +# Caveman Help + +Display this reference card when invoked. One-shot — do NOT change mode, write flag files, or persist anything. Output in caveman style. + +## Modes + +| Mode | Trigger | What change | +|------|---------|-------------| +| **Lite** | `/caveman lite` | Drop filler. Keep sentence structure. | +| **Full** | `/caveman` | Drop articles, filler, pleasantries, hedging. Fragments OK. Default. | +| **Ultra** | `/caveman ultra` | Extreme compression. Bare fragments. Tables over prose. | +| **Wenyan-Lite** | `/caveman wenyan-lite` | Classical Chinese style, light compression. | +| **Wenyan-Full** | `/caveman wenyan` | Full 文言文. Maximum classical terseness. | +| **Wenyan-Ultra** | `/caveman wenyan-ultra` | Extreme. Ancient scholar on a budget. | + +Mode stick until changed or session end. + +## Skills + +| Skill | Trigger | What it do | +|-------|---------|-----------| +| **caveman-commit** | `/caveman-commit` | Terse commit messages. Conventional Commits. ≤50 char subject. | +| **caveman-review** | `/caveman-review` | One-line PR comments: `L42: bug: user null. Add guard.` | +| **caveman-compress** | `/caveman-compress <file>` | Compress .md files to caveman prose. Saves ~46% input tokens. | +| **caveman-help** | `/caveman-help` | This card. | + +## Deactivate + +Say "stop caveman" or "normal mode". Resume anytime with `/caveman`. + +## Language + +Keep user's language by default. User write Portuguese → reply Portuguese caveman. Compress the style, not the language. Technical terms, code, commands, commit types, and exact error strings stay verbatim unless user ask for translation. + +## Configure Default Mode + +Default mode = `full`. Change it: + +**Environment variable** (highest priority): +```bash +export CAVEMAN_DEFAULT_MODE=ultra +``` + +**Config file** (`~/.config/caveman/config.json`): +```json +{ "defaultMode": "lite" } +``` + +Set `"off"` to disable auto-activation on session start. User can still activate manually with `/caveman`. + +Resolution: env var > config file > `full`. + +## More + +Full docs: https://github.com/JuliusBrussee/caveman diff --git a/.agents/skills/caveman-review/README.md b/.agents/skills/caveman-review/README.md new file mode 100644 index 0000000..acf519f --- /dev/null +++ b/.agents/skills/caveman-review/README.md @@ -0,0 +1,33 @@ +# caveman-review + +One-line PR comments. Location, problem, fix. No throat-clearing. + +## What it does + +Generates code review comments in `L<line>: <severity> <problem>. <fix>.` format. One line per finding. Severity emoji: 🔴 bug, 🟡 risk, 🔵 nit, ❓ question. Drops "I noticed that...", hedging, and restating what the diff already shows. Keeps exact line numbers, backticked symbols, and concrete fixes. + +Auto-clarity: drops terse mode for CVE-class security findings, architectural disagreements, and onboarding contexts where the author needs the *why*. Resumes terse for the rest. + +Output only — does not approve, request changes, or run linters. + +## How to invoke + +``` +/caveman-review +``` + +Also triggers on "review this PR", "code review", "review the diff". + +## Example output + +``` +L42: 🔴 bug: user can be null after .find(). Add guard before .email. +L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist. +L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3). +L107: ❓ q: why drop the cache here? Reads on next request will miss. +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview diff --git a/.agents/skills/caveman-review/SKILL.md b/.agents/skills/caveman-review/SKILL.md new file mode 100644 index 0000000..48f4adb --- /dev/null +++ b/.agents/skills/caveman-review/SKILL.md @@ -0,0 +1,55 @@ +--- +name: caveman-review +description: > + Ultra-compressed code review comments. Cuts noise from PR feedback while preserving + the actionable signal. Each comment is one line: location, problem, fix. Use when user + says "review this PR", "code review", "review the diff", "/review", or invokes + /caveman-review. Auto-triggers when reviewing pull requests. +--- + +Write code review comments terse and actionable. One line per finding. Location, problem, fix. No throat-clearing. + +## Rules + +**Format:** `L<line>: <problem>. <fix>.` — or `<file>:L<line>: ...` when reviewing multi-file diffs. + +**Severity prefix (optional, when mixed):** +- `🔴 bug:` — broken behavior, will cause incident +- `🟡 risk:` — works but fragile (race, missing null check, swallowed error) +- `🔵 nit:` — style, naming, micro-optim. Author can ignore +- `❓ q:` — genuine question, not a suggestion + +**Drop:** +- "I noticed that...", "It seems like...", "You might want to consider..." +- "This is just a suggestion but..." — use `nit:` instead +- "Great work!", "Looks good overall but..." — say it once at the top, not per comment +- Restating what the line does — the reviewer can read the diff +- Hedging ("perhaps", "maybe", "I think") — if unsure use `q:` + +**Keep:** +- Exact line numbers +- Exact symbol/function/variable names in backticks +- Concrete fix, not "consider refactoring this" +- The *why* if the fix isn't obvious from the problem statement + +## Examples + +❌ "I noticed that on line 42 you're not checking if the user object is null before accessing the email property. This could potentially cause a crash if the user is not found in the database. You might want to add a null check here." + +✅ `L42: 🔴 bug: user can be null after .find(). Add guard before .email.` + +❌ "It looks like this function is doing a lot of things and might benefit from being broken up into smaller functions for readability." + +✅ `L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist.` + +❌ "Have you considered what happens if the API returns a 429? I think we should probably handle that case." + +✅ `L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3).` + +## Auto-Clarity + +Drop terse mode for: security findings (CVE-class bugs need full explanation + reference), architectural disagreements (need rationale, not just a one-liner), and onboarding contexts where the author is new and needs the "why". In those cases write a normal paragraph, then resume terse for the rest. + +## Boundaries + +Reviews only — does not write the code fix, does not approve/request-changes, does not run linters. Output the comment(s) ready to paste into the PR. "stop caveman-review" or "normal mode": revert to verbose review style. \ No newline at end of file diff --git a/.agents/skills/caveman-stats/README.md b/.agents/skills/caveman-stats/README.md new file mode 100644 index 0000000..30b28f9 --- /dev/null +++ b/.agents/skills/caveman-stats/README.md @@ -0,0 +1,30 @@ +# caveman-stats + +Real session token receipts. No AI estimation. + +## What it does + +Reads the current Claude Code session log directly and reports actual input/output token usage plus estimated savings versus a non-caveman baseline. Numbers come from the JSONL session log on disk — the model itself does not compute or estimate them. Output is injected by the `caveman-mode-tracker` hook, which intercepts `/caveman-stats` and returns the formatted stats as a blocked-decision reason. + +Each run also writes a lifetime-savings suffix file used by the statusline badge (`⛏ 12.4k`). + +## How to invoke + +``` +/caveman-stats +``` + +## Example output + +``` +Session: 47 turns +Input: 12,304 tokens +Output: 3,891 tokens (caveman) +Baseline: 11,247 tokens (estimated without caveman) +Saved: 7,356 tokens (~65%) +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — hook contract and mechanics +- [Caveman README](../../README.md) — repo overview diff --git a/.agents/skills/caveman-stats/SKILL.md b/.agents/skills/caveman-stats/SKILL.md new file mode 100644 index 0000000..a7348ac --- /dev/null +++ b/.agents/skills/caveman-stats/SKILL.md @@ -0,0 +1,10 @@ +--- +name: caveman-stats +description: > + Show real token usage and estimated savings for the current session. + Reads directly from the Claude Code session log — no AI estimation. + Triggers on /caveman-stats. Output is injected by the mode-tracker hook; + the model itself does not compute the numbers. +--- + +This skill is delivered by `hooks/caveman-stats.js` (read by `hooks/caveman-mode-tracker.js` on `/caveman-stats`). The model does not need to do anything when this skill fires — the hook returns `decision: "block"` with the formatted stats as the reason. The user sees the numbers immediately. diff --git a/.agents/skills/caveman/README.md b/.agents/skills/caveman/README.md new file mode 100644 index 0000000..d749b83 --- /dev/null +++ b/.agents/skills/caveman/README.md @@ -0,0 +1,48 @@ +# caveman + +Talk like smart caveman. Same brain, fewer tokens. + +## What it does + +Compress every model response to caveman-style prose. Drops articles, filler, pleasantries, and hedging. Keeps every technical detail, code block, error string, and symbol exact. Cuts ~65-75% of output tokens with full accuracy preserved. Mode persists for the whole session until changed or stopped. + +Six intensity levels: + +| Level | What change | +|-------|-------------| +| `lite` | Drop filler/hedging. Sentences stay full. Professional but tight. | +| `full` | Default. Drop articles, fragments OK, short synonyms. | +| `ultra` | Bare fragments. Abbreviations (DB, auth, fn). Arrows for causality. | +| `wenyan-lite` | Classical Chinese register, light compression. | +| `wenyan-full` | Maximum 文言文. 80-90% character reduction. | +| `wenyan-ultra` | Extreme classical compression. | + +Auto-clarity rule: caveman drops to normal prose for security warnings, irreversible-action confirmations, multi-step sequences where fragment ambiguity risks misread, and when user repeats a question. Resumes after the clear part. + +## How to invoke + +``` +/caveman # full mode (default) +/caveman lite # lighter compression +/caveman ultra # extreme compression +/caveman wenyan # classical Chinese +stop caveman # back to normal prose +``` + +## Example output + +Question: "Why does my React component re-render?" + +Normal prose: +> Your component re-renders because you create a new object reference each render. Wrapping it in `useMemo` will fix the issue. + +Caveman (full): +> New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`. + +Caveman (ultra): +> Inline obj prop → new ref → re-render. `useMemo`. + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview, install, benchmarks diff --git a/.agents/skills/caveman/SKILL.md b/.agents/skills/caveman/SKILL.md new file mode 100644 index 0000000..8792c1a --- /dev/null +++ b/.agents/skills/caveman/SKILL.md @@ -0,0 +1,78 @@ +--- +name: caveman +description: > + Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman + while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, + wenyan-lite, wenyan-full, wenyan-ultra. + Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", + "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested. +--- + +Respond terse like smart caveman. All technical substance stay. Only fluff die. + +## Persistence + +ACTIVE EVERY RESPONSE. No revert after many turns. No filler drift. Still active if unsure. Off only: "stop caveman" / "normal mode". + +Default: **full**. Switch: `/caveman lite|full|ultra`. + +## Rules + +Drop: articles (a/an/the), filler (just/really/basically/actually/simply), pleasantries (sure/certainly/of course/happy to), hedging. Fragments OK. Short synonyms (big not extensive, fix not "implement a solution for"). No tool-call narration, no decorative tables/emoji, no dumping long raw error logs unless asked — quote shortest decisive line. Standard well-known tech acronyms OK (DB/API/HTTP); never invent new abbreviations reader can't decode. Technical terms exact. Code blocks unchanged. Errors quoted exact. + +Preserve user's dominant language. User write Portuguese → reply Portuguese caveman. User write Spanish → reply Spanish caveman. Compress the style, not the language. No forced English openings or status phrases. ALWAYS keep technical terms, code, API names, CLI commands, commit-type keywords (feat/fix/...), and exact error strings verbatim — unless user explicitly ask for translation. + +No self-reference. Never name or announce the style. No "caveman mode on", "me caveman think", no third-person caveman tags. Output caveman-only — never normal answer plus "Caveman:" recap. Exception: user explicitly ask what the mode is. + +Pattern: `[thing] [action] [reason]. [next step].` + +Not: "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by..." +Yes: "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:" + +## Intensity + +| Level | What change | +|-------|------------| +| **lite** | No filler/hedging. Keep articles + full sentences. Professional but tight | +| **full** | Drop articles, fragments OK, short synonyms. Classic caveman. No tool-call narration, no decorative tables/emoji, no long raw error-log dumps unless asked. Standard acronyms OK; no invented abbreviations | +| **ultra** | Abbreviate prose words (DB/auth/config/req/res/fn/impl) — prose words only, never real code symbols/function names. Strip conjunctions, arrows for causality (X → Y), one word when one word enough. Code symbols, function names, API names, error strings: never abbreviate | +| **wenyan-lite** | Semi-classical. Drop filler/hedging but keep grammar structure, classical register | +| **wenyan-full** | Maximum classical terseness. Fully 文言文. 80-90% character reduction. Classical sentence patterns, verbs precede objects, subjects often omitted, classical particles (之/乃/為/其) | +| **wenyan-ultra** | Extreme abbreviation while keeping classical Chinese feel. Maximum compression, ultra terse | + +Example — "Why React component re-render?" +- lite: "Your component re-renders because you create a new object reference each render. Wrap it in `useMemo`." +- full: "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`." +- ultra: "Inline obj prop → new ref → re-render. `useMemo`." +- wenyan-lite: "組件頻重繪,以每繪新生對象參照故。以 useMemo 包之。" +- wenyan-full: "每繪新生對象參照,故重繪;以 useMemo 包之則免。" +- wenyan-ultra: "新參照→重繪。useMemo Wrap。" + +Example — "Explain database connection pooling." +- lite: "Connection pooling reuses open connections instead of creating new ones per request. Avoids repeated handshake overhead." +- full: "Pool reuse open DB connections. No new connection per request. Skip handshake overhead." +- ultra: "Pool = reuse DB conn. Skip handshake → fast under load." +- wenyan-full: "池reuse open connection。不每req新開。skip handshake overhead。" +- wenyan-ultra: "池reuse conn。skip handshake → fast。" + +## Auto-Clarity + +Drop caveman when: +- Security warnings +- Irreversible action confirmations +- Multi-step sequences where fragment order or omitted conjunctions risk misread +- Compression itself creates technical ambiguity (e.g., `"migrate table drop column backup first"` — order unclear without articles/conjunctions) +- User asks to clarify or repeats question + +Resume caveman after clear part done. + +Example — destructive op: +> **Warning:** This will permanently delete all rows in the `users` table and cannot be undone. +> ```sql +> DROP TABLE users; +> ``` +> Caveman resume. Verify backup exist first. + +## Boundaries + +Code/commits/PRs: write normal. "stop caveman" or "normal mode": revert. Level persist until changed or session end. \ No newline at end of file diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index 1b36897..e617ac7 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -75,6 +75,10 @@ Follow this exactly when starting any phase or feature work: `Phase Plans/Phase_X_Name.md` before moving on, not just at the end of the feature. 5. **Stop only once the feature is fully done.** Give a concise, bulleted, one-line-per-point summary of what was implemented (architecture/approach) — not verbose, no long prose. + Write bullets in caveman style. Strip technical terms — explain what the user can now *do* + or what the app can now *do*, not how it works internally. At phase close, add a plain-English + paragraph (2–3 sentences max, caveman tone) covering what the phase aimed to achieve and what + was actually built. 6. **Merge the feature branch into the phase branch autonomously** — but always pause and ask ("merge about to happen, proceed?") before that merge actually executes. Same rule applies to merging the phase branch into `develop`, opening the `develop` → `main` PR, and the final merge diff --git a/.continue/skills/cavecrew/README.md b/.continue/skills/cavecrew/README.md new file mode 100644 index 0000000..e722952 --- /dev/null +++ b/.continue/skills/cavecrew/README.md @@ -0,0 +1,41 @@ +# cavecrew + +Decision guide. When to delegate to caveman subagents instead of doing the work inline. + +## What it does + +Tells the main thread when to spawn a caveman-style subagent versus the vanilla equivalent. The win: subagent tool-results inject back into main context verbatim, and caveman output is roughly 1/3 the size of vanilla prose. Across 20 delegations in one session, that is the difference between context exhaustion and finishing the task. + +Three subagents: + +| Subagent | Job | Use when | +|----------|-----|----------| +| `cavecrew-investigator` | Locate code (read-only) | "Where is X defined / what calls Y / list uses of Z" | +| `cavecrew-builder` | Surgical edit, 1-2 files | Scope is obvious, ≤2 files. Refuses 3+ file scope. | +| `cavecrew-reviewer` | Diff/file review | One-line findings with severity emoji | + +Use vanilla `Explore` or `Code Reviewer` when you want prose, architecture commentary, or rationale. Use main thread directly for one-line answers and 3+ file refactors. + +This skill is a decision guide, not a slash command. It activates when the conversation mentions delegation. + +## How to invoke + +Triggers on phrases like "delegate to subagent", "use cavecrew", "spawn investigator", "save context", "compressed agent output". + +## Example chaining + +Locate → fix → verify (most common): + +1. `cavecrew-investigator` returns site list (`path:line — symbol — note`) +2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder` +3. `cavecrew-reviewer` audits the resulting diff + +Parallel scout: spawn 2-3 `cavecrew-investigator` calls in one message with different angles (defs, callers, tests). Aggregate in main. + +## See also + +- [`SKILL.md`](./SKILL.md) — full decision matrix and output contracts +- [`agents/cavecrew-investigator.md`](../../agents/cavecrew-investigator.md) +- [`agents/cavecrew-builder.md`](../../agents/cavecrew-builder.md) +- [`agents/cavecrew-reviewer.md`](../../agents/cavecrew-reviewer.md) +- [Caveman README](../../README.md) — repo overview diff --git a/.continue/skills/cavecrew/SKILL.md b/.continue/skills/cavecrew/SKILL.md new file mode 100644 index 0000000..efa413f --- /dev/null +++ b/.continue/skills/cavecrew/SKILL.md @@ -0,0 +1,82 @@ +--- +name: cavecrew +description: > + Decision guide for delegating to caveman-style subagents. Tells the main + thread WHEN to spawn `cavecrew-investigator` (locate code), `cavecrew-builder` + (1-2 file edit), or `cavecrew-reviewer` (diff review) instead of doing the + work inline or using vanilla `Explore`. Subagent output is caveman-compressed + so the tool-result injected back into main context is ~60% smaller — main + context lasts longer across long sessions. + Trigger: "delegate to subagent", "use cavecrew", "spawn investigator/builder/reviewer", + "save context", "compressed agent output". +--- + +Cavecrew = three subagent presets that emit caveman output. Same job as Anthropic defaults (`Explore`, edit-style agents, reviewer); difference is the tool-result they return is compressed, so main context shrinks per delegation. + +## When to use cavecrew vs alternatives + +| Task | Use | +|---|---| +| "Where is X defined / what calls Y / list uses of Z" | `cavecrew-investigator` | +| Same but you also want suggestions/architecture commentary | `Explore` (vanilla) | +| Surgical edit, ≤2 files, scope obvious | `cavecrew-builder` | +| New feature / 3+ files / cross-cutting refactor | Main thread or `feature-dev:code-architect` | +| Review diff, branch, or file for bugs | `cavecrew-reviewer` | +| Deep code review with rationale + alternatives | `Code Reviewer` (vanilla) | +| One-line answer you already know | Main thread, no subagent | + +Rule of thumb: **if you'd want the subagent's output in 1/3 the tokens, pick cavecrew. If you'd want prose, pick vanilla.** + +## Why this exists (the real win) + +Subagent tool results get injected into main context verbatim. A vanilla `Explore` that returns 2k tokens of prose costs 2k tokens of main-context budget every time. The same finding from `cavecrew-investigator` returns ~700 tokens. Across 20 delegations in one session that's the difference between context exhaustion and finishing the task. + +## Output contracts + +What main thread can rely on per agent: + +**`cavecrew-investigator`** +``` +<Header>: +- path:line — `symbol` — short note +totals: <counts>. +``` +Or `No match.` Always file-path-first, line-number-attached, backticked symbols. Safe to grep with `path:\d+`. + +**`cavecrew-builder`** +``` +<path:line-range> — <change ≤10 words>. +verified: <re-read OK | mismatch @ path:line>. +``` +Or one of: `too-big.` / `needs-confirm.` / `ambiguous.` / `regressed.` (terminal first token). + +**`cavecrew-reviewer`** +``` +path:line: <emoji> <severity>: <problem>. <fix>. +totals: N🔴 N🟡 N🔵 N❓ +``` +Or `No issues.` Findings sorted file → line ascending. + +## Chaining patterns + +**Locate → fix → verify** (most common): +1. `cavecrew-investigator` returns site list. +2. Main thread picks 1-2 sites, hands paths to `cavecrew-builder`. +3. `cavecrew-reviewer` audits the diff. + +**Parallel scout** (when investigation is broad): +Spawn 2-3 `cavecrew-investigator` calls in one message (different angles: defs vs callers vs tests). Aggregate in main thread. + +**Single-shot edit** (when site is already known): +Skip investigator. Hand exact path:line to `cavecrew-builder` directly. + +## What NOT to do + +- Don't use `cavecrew-builder` when you don't already know the file. Spawn investigator first or main thread will eat tokens passing context. +- Don't chain `cavecrew-investigator → cavecrew-builder` for a 5-file refactor. Builder will return `too-big.` and you'll have wasted a turn. +- Don't ask `cavecrew-reviewer` for "general feedback" — it returns findings only, no architecture opinions. Use `Code Reviewer` for that. +- Don't expect prose. Cavecrew output is structured, sometimes terse to the point of cryptic. If a human will read it directly, paraphrase. + +## Auto-clarity (inherited) + +Subagents drop caveman → normal English for security warnings, irreversible-action confirmations, and any output where fragment ambiguity could be misread. Resume caveman after. diff --git a/.continue/skills/caveman-commit/README.md b/.continue/skills/caveman-commit/README.md new file mode 100644 index 0000000..d5aee01 --- /dev/null +++ b/.continue/skills/caveman-commit/README.md @@ -0,0 +1,44 @@ +# caveman-commit + +Terse Conventional Commits. Why over what. + +## What it does + +Generates commit messages in Conventional Commits format. Subject ≤50 chars, hard cap 72. Imperative mood. Body only when the *why* is non-obvious or there are breaking changes. No AI attribution, no "this commit does X", no emoji unless the project uses them. Body always required for breaking changes, security fixes, data migrations, and reverts — future debuggers need the context. + +Outputs only the message. Does not stage, commit, or amend. + +## How to invoke + +``` +/caveman-commit +``` + +Also triggers on phrases like "write a commit", "commit message", "generate commit". + +## Example output + +Diff: new endpoint for user profile. + +``` +feat(api): add GET /users/:id/profile + +Mobile client needs profile data without the full user payload +to reduce LTE bandwidth on cold-launch screens. + +Closes #128 +``` + +Diff: breaking API rename. + +``` +feat(api)!: rename /v1/orders to /v1/checkout + +BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout +before 2026-06-01. Old route returns 410 after that date. +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview diff --git a/.continue/skills/caveman-commit/SKILL.md b/.continue/skills/caveman-commit/SKILL.md new file mode 100644 index 0000000..b9999e3 --- /dev/null +++ b/.continue/skills/caveman-commit/SKILL.md @@ -0,0 +1,65 @@ +--- +name: caveman-commit +description: > + Ultra-compressed commit message generator. Cuts noise from commit messages while preserving + intent and reasoning. Conventional Commits format. Subject ≤50 chars, body only when "why" + isn't obvious. Use when user says "write a commit", "commit message", "generate commit", + "/commit", or invokes /caveman-commit. Auto-triggers when staging changes. +--- + +Write commit messages terse and exact. Conventional Commits format. No fluff. Why over what. + +## Rules + +**Subject line:** +- `<type>(<scope>): <imperative summary>` — `<scope>` optional +- Types: `feat`, `fix`, `refactor`, `perf`, `docs`, `test`, `chore`, `build`, `ci`, `style`, `revert` +- Imperative mood: "add", "fix", "remove" — not "added", "adds", "adding" +- ≤50 chars when possible, hard cap 72 +- No trailing period +- Match project convention for capitalization after the colon + +**Body (only if needed):** +- Skip entirely when subject is self-explanatory +- Add body only for: non-obvious *why*, breaking changes, migration notes, linked issues +- Wrap at 72 chars +- Bullets `-` not `*` +- Reference issues/PRs at end: `Closes #42`, `Refs #17` + +**What NEVER goes in:** +- "This commit does X", "I", "we", "now", "currently" — the diff says what +- "As requested by..." — use Co-authored-by trailer +- "Generated with Claude Code" or any AI attribution — unless the user's own rule requires an `Assisted-by`/AI-attribution trailer, then add it as a trailer +- Emoji (unless project convention requires) +- Restating the file name when scope already says it + +## Examples + +Diff: new endpoint for user profile with body explaining the why +- ❌ "feat: add a new endpoint to get user profile information from the database" +- ✅ + ``` + feat(api): add GET /users/:id/profile + + Mobile client needs profile data without the full user payload + to reduce LTE bandwidth on cold-launch screens. + + Closes #128 + ``` + +Diff: breaking API change +- ✅ + ``` + feat(api)!: rename /v1/orders to /v1/checkout + + BREAKING CHANGE: clients on /v1/orders must migrate to /v1/checkout + before 2026-06-01. Old route returns 410 after that date. + ``` + +## Auto-Clarity + +Always include body for: breaking changes, security fixes, data migrations, anything reverting a prior commit. Never compress these into subject-only — future debuggers need the context. + +## Boundaries + +Only generates the commit message. Does not run `git commit`, does not stage files, does not amend. Output the message as a code block ready to paste. "stop caveman-commit" or "normal mode": revert to verbose commit style. diff --git a/.continue/skills/caveman-compress/README.md b/.continue/skills/caveman-compress/README.md new file mode 100644 index 0000000..3ef6922 --- /dev/null +++ b/.continue/skills/caveman-compress/README.md @@ -0,0 +1,163 @@ +<p align="center"> + <img src="https://em-content.zobj.net/source/apple/391/rock_1faa8.png" width="80" /> +</p> + +<h1 align="center">caveman-compress</h1> + +<p align="center"> + <strong>shrink memory file. save token every session.</strong> +</p> + +--- + +A Claude Code skill that compresses your project memory files (`CLAUDE.md`, todos, preferences) into caveman format — so every session loads fewer tokens automatically. + +Claude read `CLAUDE.md` on every session start. If file big, cost big. Caveman make file small. Cost go down forever. + +## What It Do + +``` +/caveman-compress CLAUDE.md +``` + +``` +CLAUDE.md ← compressed (Claude reads this — fewer tokens every session) +CLAUDE.original.md ← human-readable backup (you edit this) +``` + +Original never lost. You can read and edit `.original.md`. Run skill again to re-compress after edits. + +## Benchmarks + +Real results on real project files: + +| File | Original | Compressed | Saved | +|------|----------:|----------:|------:| +| `claude-md-preferences.md` | 706 | 285 | **59.6%** | +| `project-notes.md` | 1145 | 535 | **53.3%** | +| `claude-md-project.md` | 1122 | 636 | **43.3%** | +| `todo-list.md` | 627 | 388 | **38.1%** | +| `mixed-with-code.md` | 888 | 560 | **36.9%** | +| **Average** | **898** | **481** | **46%** | + +All validations passed ✅ — headings, code blocks, URLs, file paths preserved exactly. + +## Before / After + +<table> +<tr> +<td width="50%"> + +### 📄 Original (706 tokens) + +> "I strongly prefer TypeScript with strict mode enabled for all new code. Please don't use `any` type unless there's genuinely no way around it, and if you do, leave a comment explaining the reasoning. I find that taking the time to properly type things catches a lot of bugs before they ever make it to runtime." + +</td> +<td width="50%"> + +### <img src="../../docs/assets/dancing-rock.svg" width="20" height="20" alt="rock"/> Caveman (285 tokens) + +> "Prefer TypeScript strict mode always. No `any` unless unavoidable — comment why if used. Proper types catch bugs early." + +</td> +</tr> +</table> + +**Same instructions. 60% fewer tokens. Every. Single. Session.** + +## Security + +`caveman-compress` is flagged as Snyk High Risk due to subprocess and file I/O patterns detected by static analysis. This is a false positive — see [SECURITY.md](./SECURITY.md) for a full explanation of what the skill does and does not do. + +## Install + +Compress is built in with the `caveman` plugin. Install `caveman` once, then use `/caveman-compress`. + +If you need local files, the compress skill lives at: + +```bash +caveman-compress/ +``` + +**Requires:** Python 3.10+ + +## Usage + +``` +/caveman-compress <filepath> +``` + +Examples: +``` +/caveman-compress CLAUDE.md +/caveman-compress docs/preferences.md +/caveman-compress todos.md +``` + +### What files work + +| Type | Compress? | +|------|-----------| +| `.md`, `.txt`, `.rst`, `.typ`, `.typst`, `.tex` | ✅ Yes | +| Extensionless natural language | ✅ Yes | +| `.py`, `.js`, `.ts`, `.json`, `.yaml` | ❌ Skip (code/config) | +| `*.original.md` | ❌ Skip (backup files) | + +## How It Work + +``` +/caveman-compress CLAUDE.md + ↓ +detect file type (no tokens) + ↓ +Claude compresses (tokens — one call) + ↓ +validate output (no tokens) + checks: headings, code blocks, URLs, file paths, bullets + ↓ +if errors: Claude fixes cherry-picked issues only (tokens — targeted fix) + does NOT recompress — only patches broken parts + ↓ +retry up to 2 times + ↓ +write compressed → CLAUDE.md +write original → CLAUDE.original.md +``` + +Only two things use tokens: initial compression + targeted fix if validation fails. Everything else is local Python. + +## What Is Preserved + +Caveman compress natural language. It never touch: + +- Code blocks (` ``` ` fenced or indented) +- Inline code (`` `backtick content` ``) +- URLs and links +- File paths (`/src/components/...`) +- Commands (`npm install`, `git commit`) +- Technical terms, library names, API names +- Headings (exact text preserved) +- Tables (structure preserved, cell text compressed) +- Dates, version numbers, numeric values + +## Why This Matter + +`CLAUDE.md` loads on **every session start**. A 1000-token project memory file costs tokens every single time you open a project. Over 100 sessions that's 100,000 tokens of overhead — just for context you already wrote. + +Caveman cut that by ~46% on average. Same instructions. Same accuracy. Less waste. + +``` +┌────────────────────────────────────────────┐ +│ TOKEN SAVINGS PER FILE █████ 46% │ +│ SESSIONS THAT BENEFIT ██████████ 100% │ +│ INFORMATION PRESERVED ██████████ 100% │ +│ SETUP TIME █ 1x │ +└────────────────────────────────────────────┘ +``` + +## Part of Caveman + +This skill is part of the [caveman](https://github.com/JuliusBrussee/caveman) toolkit — making Claude use fewer tokens without losing accuracy. + +- **caveman** — make Claude *speak* like caveman (cuts response tokens ~65%) +- **caveman-compress** — make Claude *read* less (cuts context tokens ~46%) diff --git a/.continue/skills/caveman-compress/SECURITY.md b/.continue/skills/caveman-compress/SECURITY.md new file mode 100644 index 0000000..693108c --- /dev/null +++ b/.continue/skills/caveman-compress/SECURITY.md @@ -0,0 +1,31 @@ +# Security + +## Snyk High Risk Rating + +`caveman-compress` receives a Snyk High Risk rating due to static analysis heuristics. This document explains what the skill does and does not do. + +### What triggers the rating + +1. **subprocess usage**: The skill calls the `claude` CLI via `subprocess.run()` as a fallback when `ANTHROPIC_API_KEY` is not set. The subprocess call uses a fixed argument list — no shell interpolation occurs. User file content is passed via stdin, not as a shell argument. + +2. **File read/write**: The skill reads the file the user explicitly points it at, compresses it, and writes the result back to the same path. A `.original.md` backup is saved alongside it. No files outside the user-specified path are read or written. + +### What the skill does NOT do + +- Does not execute user file content as code +- Does not make network requests except to Anthropic's API (via SDK or CLI) +- Does not access files outside the path the user provides +- Does not use shell=True or string interpolation in subprocess calls +- Does not collect or transmit any data beyond the file being compressed + +### Auth behavior + +If `ANTHROPIC_API_KEY` is set, the skill uses the Anthropic Python SDK directly (no subprocess). If not set, it falls back to the `claude` CLI, which uses the user's existing Claude desktop authentication. + +### File size limit + +Files larger than 500KB are rejected before any API call is made. + +### Reporting a vulnerability + +If you believe you've found a genuine security issue, please open a GitHub issue with the label `security`. diff --git a/.continue/skills/caveman-compress/SKILL.md b/.continue/skills/caveman-compress/SKILL.md new file mode 100644 index 0000000..00ce454 --- /dev/null +++ b/.continue/skills/caveman-compress/SKILL.md @@ -0,0 +1,111 @@ +--- +name: caveman-compress +description: > + Compress natural language memory files (CLAUDE.md, todos, preferences) into caveman format + to save input tokens. Preserves all technical substance, code, URLs, and structure. + Compressed version overwrites the original file. Human-readable backup saved as FILE.original.md. + Trigger: /caveman-compress FILEPATH or "compress memory file" +--- + +# Caveman Compress + +## Purpose + +Compress natural language files (CLAUDE.md, todos, preferences) into caveman-speak to reduce input tokens. Compressed version overwrites original. Human-readable backup saved as `<filename>.original.md`. + +## Trigger + +`/caveman-compress <filepath>` or when user asks to compress a memory file. + +## Process + +1. The compression scripts live in `scripts/` (adjacent to this SKILL.md). If the path is not immediately available, search for `scripts/__main__.py` next to this SKILL.md. + +2. From the directory containing this SKILL.md, run: + +python3 -m scripts <absolute_filepath> + +3. The CLI will: +- detect file type (no tokens) +- call Claude to compress +- validate output (no tokens) +- if errors: cherry-pick fix with Claude (targeted fixes only, no recompression) +- retry up to 2 times +- if still failing after 2 retries: report error to user, leave original file untouched + +4. Return result to user + +## Compression Rules + +### Remove +- Articles: a, an, the +- Filler: just, really, basically, actually, simply, essentially, generally +- Pleasantries: "sure", "certainly", "of course", "happy to", "I'd recommend" +- Hedging: "it might be worth", "you could consider", "it would be good to" +- Redundant phrasing: "in order to" → "to", "make sure to" → "ensure", "the reason is because" → "because" +- Connective fluff: "however", "furthermore", "additionally", "in addition" + +### Preserve EXACTLY (never modify) +- Code blocks (fenced ``` and indented) +- Inline code (`backtick content`) +- URLs and links (full URLs, markdown links) +- File paths (`/src/components/...`, `./config.yaml`) +- Commands (`npm install`, `git commit`, `docker build`) +- Technical terms (library names, API names, protocols, algorithms) +- Proper nouns (project names, people, companies) +- Dates, version numbers, numeric values +- Environment variables (`$HOME`, `NODE_ENV`) + +### Preserve Structure +- All markdown headings (keep exact heading text, compress body below) +- Bullet point hierarchy (keep nesting level) +- Numbered lists (keep numbering) +- Tables (compress cell text, keep structure) +- Frontmatter/YAML headers in markdown files + +### Compress +- Use short synonyms: "big" not "extensive", "fix" not "implement a solution for", "use" not "utilize" +- Fragments OK: "Run tests before commit" not "You should always run tests before committing" +- Drop "you should", "make sure to", "remember to" — just state the action +- Merge redundant bullets that say the same thing differently +- Keep one example where multiple examples show the same pattern + +CRITICAL RULE: +Anything inside ``` ... ``` must be copied EXACTLY. +Do not: +- remove comments +- remove spacing +- reorder lines +- shorten commands +- simplify anything + +Inline code (`...`) must be preserved EXACTLY. +Do not modify anything inside backticks. + +If file contains code blocks: +- Treat code blocks as read-only regions +- Only compress text outside them +- Do not merge sections around code + +## Pattern + +Original: +> You should always make sure to run the test suite before pushing any changes to the main branch. This is important because it helps catch bugs early and prevents broken builds from being deployed to production. + +Compressed: +> Run tests before push to main. Catch bugs early, prevent broken prod deploys. + +Original: +> The application uses a microservices architecture with the following components. The API gateway handles all incoming requests and routes them to the appropriate service. The authentication service is responsible for managing user sessions and JWT tokens. + +Compressed: +> Microservices architecture. API gateway route all requests to services. Auth service manage user sessions + JWT tokens. + +## Boundaries + +- ONLY compress natural language files (.md, .txt, .typ, .typst, .tex, extensionless) +- NEVER modify: .py, .js, .ts, .json, .yaml, .yml, .toml, .env, .lock, .css, .html, .xml, .sql, .sh +- If file has mixed content (prose + code), compress ONLY the prose sections +- If unsure whether something is code or prose, leave it unchanged +- Original file is backed up as FILE.original.md before overwriting +- Never compress FILE.original.md (skip it) diff --git a/.continue/skills/caveman-compress/scripts/__init__.py b/.continue/skills/caveman-compress/scripts/__init__.py new file mode 100644 index 0000000..16b8c53 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/__init__.py @@ -0,0 +1,9 @@ +"""Caveman compress scripts. + +This package provides tools to compress natural language markdown files +into caveman format to save input tokens. +""" + +__all__ = ["cli", "compress", "detect", "validate"] + +__version__ = "1.0.0" diff --git a/.continue/skills/caveman-compress/scripts/__main__.py b/.continue/skills/caveman-compress/scripts/__main__.py new file mode 100644 index 0000000..4e28416 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/__main__.py @@ -0,0 +1,3 @@ +from .cli import main + +main() diff --git a/.continue/skills/caveman-compress/scripts/benchmark.py b/.continue/skills/caveman-compress/scripts/benchmark.py new file mode 100644 index 0000000..f9e2ee0 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/benchmark.py @@ -0,0 +1,80 @@ +#!/usr/bin/env python3 +from pathlib import Path +import sys + +# Support both direct execution and module import +try: + from .validate import validate +except ImportError: + sys.path.insert(0, str(Path(__file__).parent)) + from validate import validate + +try: + import tiktoken + _enc = tiktoken.get_encoding("o200k_base") +except ImportError: + _enc = None + + +def count_tokens(text): + if _enc is None: + return len(text.split()) # fallback: word count + return len(_enc.encode(text)) + + +def benchmark_pair(orig_path: Path, comp_path: Path): + orig_text = orig_path.read_text() + comp_text = comp_path.read_text() + + orig_tokens = count_tokens(orig_text) + comp_tokens = count_tokens(comp_text) + saved = 100 * (orig_tokens - comp_tokens) / orig_tokens if orig_tokens > 0 else 0.0 + result = validate(orig_path, comp_path) + + return (comp_path.name, orig_tokens, comp_tokens, saved, result.is_valid) + + +def print_table(rows): + print("\n| File | Original | Compressed | Saved % | Valid |") + print("|------|----------|------------|---------|-------|") + for r in rows: + print(f"| {r[0]} | {r[1]} | {r[2]} | {r[3]:.1f}% | {'✅' if r[4] else '❌'} |") + + +def main(): + # Direct file pair: python3 benchmark.py original.md compressed.md + if len(sys.argv) == 3: + orig = Path(sys.argv[1]).resolve() + comp = Path(sys.argv[2]).resolve() + if not orig.exists(): + print(f"❌ Not found: {orig}") + sys.exit(1) + if not comp.exists(): + print(f"❌ Not found: {comp}") + sys.exit(1) + print_table([benchmark_pair(orig, comp)]) + return + + # Glob mode: repo_root/tests/caveman-compress/ + # __file__ lives at <repo_root>/skills/caveman-compress/scripts/benchmark.py + # Walk up four dirs: scripts → caveman-compress → skills → repo_root. + tests_dir = Path(__file__).resolve().parents[3] / "tests" / "caveman-compress" + if not tests_dir.exists(): + print(f"❌ Tests dir not found: {tests_dir}") + sys.exit(1) + + rows = [] + for orig in sorted(tests_dir.glob("*.original.md")): + comp = orig.with_name(orig.stem.removesuffix(".original") + ".md") + if comp.exists(): + rows.append(benchmark_pair(orig, comp)) + + if not rows: + print("No compressed file pairs found.") + return + + print_table(rows) + + +if __name__ == "__main__": + main() diff --git a/.continue/skills/caveman-compress/scripts/cli.py b/.continue/skills/caveman-compress/scripts/cli.py new file mode 100644 index 0000000..75ea8a6 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/cli.py @@ -0,0 +1,85 @@ +#!/usr/bin/env python3 +""" +Caveman Compress CLI + +Usage: + caveman <filepath> +""" + +import sys + +# Force UTF-8 on stdout/stderr before any code can print. Windows consoles +# default to cp1252 and crash on the ❌ glyphs in error/validation branches, +# masking the real error and leaving the user with a half-compressed file. +for _stream in (sys.stdout, sys.stderr): + reconfigure = getattr(_stream, "reconfigure", None) + if callable(reconfigure): + try: + reconfigure(encoding="utf-8", errors="replace") + except Exception: + pass + +from pathlib import Path + +from .compress import backup_dir_for, compress_file +from .detect import detect_file_type, should_compress + + +def print_usage(): + print("Usage: caveman <filepath>") + + +def main(): + if len(sys.argv) != 2: + print_usage() + sys.exit(1) + + filepath = Path(sys.argv[1]) + + # Check file exists + if not filepath.exists(): + print(f"❌ File not found: {filepath}") + sys.exit(1) + + if not filepath.is_file(): + print(f"❌ Not a file: {filepath}") + sys.exit(1) + + filepath = filepath.resolve() + + # Detect file type + file_type = detect_file_type(filepath) + + print(f"Detected: {file_type}") + + # Check if compressible + if not should_compress(filepath): + print("Skipping: file is not natural language (code/config)") + sys.exit(0) + + print("Starting caveman compression...\n") + + try: + success = compress_file(filepath) + + if success: + print("\nCompression completed successfully") + backup_path = backup_dir_for(filepath) / (filepath.stem + ".original.md") + print(f"Compressed: {filepath}") + print(f"Original: {backup_path}") + sys.exit(0) + else: + print("\n❌ Compression failed after retries") + sys.exit(2) + + except KeyboardInterrupt: + print("\nInterrupted by user") + sys.exit(130) + + except Exception as e: + print(f"\n❌ Error: {e}") + sys.exit(1) + + +if __name__ == "__main__": + main() diff --git a/.continue/skills/caveman-compress/scripts/compress.py b/.continue/skills/caveman-compress/scripts/compress.py new file mode 100644 index 0000000..e93d934 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/compress.py @@ -0,0 +1,342 @@ +#!/usr/bin/env python3 +""" +Caveman Memory Compression Orchestrator + +Usage: + python scripts/compress.py <filepath> +""" + +import os +import re +import shutil +import subprocess +import sys +from pathlib import Path +from typing import List + +OUTER_FENCE_REGEX = re.compile( + r"\A\s*(`{3,}|~{3,})[^\n]*\n(.*)\n\1\s*\Z", re.DOTALL +) + +# YAML frontmatter: starts at file start with --- on its own line, ends with --- on its own line. +# Captures the entire block (including delimiters and trailing newline) and the body after. +FRONTMATTER_REGEX = re.compile( + r"\A(---\r?\n.*?\r?\n---\r?\n)(.*)", re.DOTALL +) + + +def split_frontmatter(text: str): + """Split YAML frontmatter from body. Returns (frontmatter, body). + + Memory files (and many other markdown docs) start with a YAML frontmatter + block delimited by `---` lines. The compression LLM has a habit of stripping + or rewriting these despite preserve-structure rules in the prompt — so we + surgically remove the frontmatter before compression and prepend it back + verbatim to the output. Files without frontmatter pass through unchanged. + """ + m = FRONTMATTER_REGEX.match(text) + if m: + return m.group(1), m.group(2) + return "", text + +# Filenames and paths that almost certainly hold secrets or PII. Compressing +# them ships raw bytes to the Anthropic API — a third-party data boundary that +# developers on sensitive codebases cannot cross. detect.py already skips .env +# by extension, but credentials.md / secrets.txt / ~/.aws/credentials would +# slip through the natural-language filter. This is a hard refuse before read. +SENSITIVE_BASENAME_REGEX = re.compile( + r"(?ix)^(" + r"\.env(\..+)?" + r"|\.netrc" + r"|credentials(\..+)?" + r"|secrets?(\..+)?" + r"|passwords?(\..+)?" + r"|id_(rsa|dsa|ecdsa|ed25519)(\.pub)?" + r"|authorized_keys" + r"|known_hosts" + r"|.*\.(pem|key|p12|pfx|crt|cer|jks|keystore|asc|gpg)" + r")$" +) + +SENSITIVE_PATH_COMPONENTS = frozenset({".ssh", ".aws", ".gnupg", ".kube", ".docker"}) + +SENSITIVE_NAME_TOKENS = ( + "secret", "credential", "password", "passwd", + "apikey", "accesskey", "token", "privatekey", +) + + +def backup_dir_for(filepath: Path) -> Path: + """Resolve the out-of-tree backup directory for a given source file. + + Backups must live OUTSIDE the source directory so skill auto-loaders + (Claude Code rules/, opencode instructions/, etc.) stop re-ingesting the + `.original.md` copies as live files. Base dir is platform-aware: + - Windows: %LOCALAPPDATA%\\caveman-compress\\backups + - else: $XDG_DATA_HOME/caveman-compress/backups if set, + else ~/.local/share/caveman-compress/backups + + The source file's parent-dir name is mirrored under the base to reduce + cross-project collisions (e.g. two `task.md` files in different repos). + """ + if os.name == "nt" or sys.platform == "win32": + local_appdata = os.environ.get("LOCALAPPDATA") + base = Path(local_appdata) if local_appdata else Path.home() / "AppData" / "Local" + base = base / "caveman-compress" / "backups" + else: + xdg = os.environ.get("XDG_DATA_HOME") + base = Path(xdg) if xdg else Path.home() / ".local" / "share" + base = base / "caveman-compress" / "backups" + return base / filepath.parent.name + + +def is_sensitive_path(filepath: Path) -> bool: + """Heuristic denylist for files that must never be shipped to a third-party API.""" + name = filepath.name + if SENSITIVE_BASENAME_REGEX.match(name): + return True + lowered_parts = {p.lower() for p in filepath.parts} + if lowered_parts & SENSITIVE_PATH_COMPONENTS: + return True + # Normalize separators so "api-key" and "api_key" both match "apikey". + lower = re.sub(r"[_\-\s.]", "", name.lower()) + return any(tok in lower for tok in SENSITIVE_NAME_TOKENS) + + +def strip_llm_wrapper(text: str) -> str: + """Strip outer ```markdown ... ``` fence when it wraps the entire output.""" + m = OUTER_FENCE_REGEX.match(text) + if m: + return m.group(2) + return text + +from .detect import should_compress +from .validate import validate + +MAX_RETRIES = 2 + + +# ---------- Claude Calls ---------- + + +def call_claude(prompt: str) -> str: + """Send a prompt to Claude. + + Prefers the Anthropic SDK when ANTHROPIC_API_KEY is set; otherwise falls + back to the ``claude --print`` CLI (which handles desktop auth). + + On Windows the CLI subprocess decoding defaults to the system codepage + (cp1251 / cp1252) and crashes on UTF-8 output — see issue #152. Pinning + ``encoding="utf-8"`` with ``errors="replace"`` matches the CLI's actual + native I/O and prevents the UnicodeDecodeError before validation can + report. Windows users with non-ASCII content can also set + ``ANTHROPIC_API_KEY`` to route through the SDK and skip the subprocess. + """ + api_key = os.environ.get("ANTHROPIC_API_KEY") + if api_key: + try: + import anthropic + + client = anthropic.Anthropic(api_key=api_key) + msg = client.messages.create( + model=os.environ.get("CAVEMAN_MODEL", "claude-sonnet-4-5"), + max_tokens=8192, + messages=[{"role": "user", "content": prompt}], + ) + return strip_llm_wrapper(msg.content[0].text.strip()) + except ImportError: + pass # anthropic not installed, fall back to CLI + # Fallback: use claude CLI (handles desktop auth). + # Resolve binary via shutil.which so Windows .cmd/.bat shims (e.g. + # %APPDATA%\npm\claude.CMD) work without shell=True. On POSIX, + # shutil.which returns the same absolute path as the implicit lookup, + # so this is a no-op there. Falls back to bare "claude" if not found + # on PATH so subprocess raises a clear FileNotFoundError. + claude_bin = shutil.which("claude") or "claude" + try: + result = subprocess.run( + [claude_bin, "--print"], + input=prompt, + text=True, + capture_output=True, + check=True, + encoding="utf-8", + errors="replace", + ) + return strip_llm_wrapper(result.stdout.strip()) + except subprocess.CalledProcessError as e: + raise RuntimeError(f"Claude call failed:\n{e.stderr}") + + +def build_compress_prompt(original: str) -> str: + return f""" +Compress this markdown into caveman format. + +STRICT RULES: +- Do NOT modify anything inside ``` code blocks +- Do NOT modify anything inside inline backticks +- Preserve ALL URLs exactly +- Preserve ALL headings exactly +- Preserve file paths and commands +- Return ONLY the compressed markdown body — do NOT wrap the entire output in a ```markdown fence or any other fence. Inner code blocks from the original stay as-is; do not add a new outer fence around the whole file. + +Only compress natural language. + +TEXT: +{original} +""" + + +def build_fix_prompt(original: str, compressed: str, errors: List[str]) -> str: + errors_str = "\n".join(f"- {e}" for e in errors) + return f"""You are fixing a caveman-compressed markdown file. Specific validation errors were found. + +CRITICAL RULES: +- DO NOT recompress or rephrase the file +- ONLY fix the listed errors — leave everything else exactly as-is +- The ORIGINAL is provided as reference only (to restore missing content) +- Preserve caveman style in all untouched sections + +ERRORS TO FIX: +{errors_str} + +HOW TO FIX: +- Missing URL: find it in ORIGINAL, restore it exactly where it belongs in COMPRESSED +- Code block mismatch: find the exact code block in ORIGINAL, restore it in COMPRESSED +- Heading mismatch: restore the exact heading text from ORIGINAL into COMPRESSED +- Do not touch any section not mentioned in the errors + +ORIGINAL (reference only): +{original} + +COMPRESSED (fix this): +{compressed} + +Return ONLY the fixed compressed file. No explanation. +""" + + +# ---------- Core Logic ---------- + + +def compress_file(filepath: Path) -> bool: + # Resolve and validate path + filepath = filepath.resolve() + MAX_FILE_SIZE = 500_000 # 500KB + if not filepath.exists(): + raise FileNotFoundError(f"File not found: {filepath}") + if filepath.stat().st_size > MAX_FILE_SIZE: + raise ValueError(f"File too large to compress safely (max 500KB): {filepath}") + + # Refuse files that look like they contain secrets or PII. Compressing ships + # the raw bytes to the Anthropic API — a third-party boundary — so we fail + # loudly rather than silently exfiltrate credentials or keys. Override is + # intentional: the user must rename the file if the heuristic is wrong. + if is_sensitive_path(filepath): + raise ValueError( + f"Refusing to compress {filepath}: filename looks sensitive " + "(credentials, keys, secrets, or known private paths). " + "Compression sends file contents to the Anthropic API. " + "Rename the file if this is a false positive." + ) + + print(f"Processing: {filepath}") + + if not should_compress(filepath): + print("Skipping (not natural language)") + return False + + original_text = filepath.read_text(errors="ignore") + # Store backup outside the source directory so skill auto-loaders don't + # re-ingest the `.original.md` copy as a live file. Mirror the source's + # parent-dir name + stem under a platform-aware base to reduce collisions. + backup_dir = backup_dir_for(filepath) + backup_dir.mkdir(parents=True, exist_ok=True) + backup_path = backup_dir / (filepath.stem + ".original.md") + + if not original_text.strip(): + print("❌ Refusing to compress: file is empty or whitespace-only.") + return False + + # Check if backup already exists to prevent accidental overwriting + if backup_path.exists(): + print(f"⚠️ Backup file already exists: {backup_path}") + print("The original backup may contain important content.") + print("Aborting to prevent data loss. Please remove or rename the backup file if you want to proceed.") + return False + + # Split YAML frontmatter off before compression. Claude tends to strip or + # rewrite frontmatter despite preserve-structure rules; we keep it verbatim + # by removing it from the input and re-prepending it to the output. + frontmatter, body = split_frontmatter(original_text) + if frontmatter: + print(f"Detected YAML frontmatter ({len(frontmatter)} chars) — preserving verbatim") + + if not body.strip(): + print("❌ Refusing to compress: body is empty after frontmatter removal.") + return False + + # Step 1: Compress (body only, frontmatter excluded) + print("Compressing with Claude...") + compressed_body = call_claude(build_compress_prompt(body)) + + if compressed_body is None or not compressed_body.strip(): + print("❌ Compression aborted: Claude returned an empty response.") + print(" Original file is untouched (no backup created).") + return False + + # Compare the BODY (not the whole file) — frontmatter is preserved verbatim + # and would never change, so identity must be judged on the compressible part. + if compressed_body.strip() == body.strip(): + print("❌ Compression aborted: output is identical to input.") + print(" Likely causes: Claude refused, returned the prompt verbatim, or the file is") + print(" already in caveman form. Original file is untouched (no backup created).") + return False + + # Reassemble: frontmatter (verbatim) + compressed body + compressed = frontmatter + compressed_body + + # Save original as backup, then verify the backup readback before + # touching the input file. If the filesystem dropped bytes (encoding, + # antivirus, disk full), unlink the bad backup and abort instead of + # leaving the user with a corrupt backup + compressed primary. + backup_path.write_text(original_text) + backup_readback = backup_path.read_text(errors="ignore") + if backup_readback != original_text: + print(f"❌ Backup write verification failed: {backup_path}") + print(" In-memory original differs from on-disk backup. Aborting before touching the input file.") + try: + backup_path.unlink() + except OSError: + pass + return False + filepath.write_text(compressed) + + # Step 2: Validate + Retry + for attempt in range(MAX_RETRIES): + print(f"\nValidation attempt {attempt + 1}") + + result = validate(backup_path, filepath) + + if result.is_valid: + print("Validation passed") + break + + print("❌ Validation failed:") + for err in result.errors: + print(f" - {err}") + + if attempt == MAX_RETRIES - 1: + # Restore original on failure + filepath.write_text(original_text) + backup_path.unlink(missing_ok=True) + print("❌ Failed after retries — original restored") + return False + + print("Fixing with Claude...") + compressed = call_claude( + build_fix_prompt(original_text, compressed, result.errors) + ) + filepath.write_text(compressed) + + return True diff --git a/.continue/skills/caveman-compress/scripts/detect.py b/.continue/skills/caveman-compress/scripts/detect.py new file mode 100644 index 0000000..8d5f6d7 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/detect.py @@ -0,0 +1,121 @@ +#!/usr/bin/env python3 +"""Detect whether a file is natural language (compressible) or code/config (skip).""" + +import json +import re +from pathlib import Path + +# Extensions that are natural language and compressible +COMPRESSIBLE_EXTENSIONS = {".md", ".txt", ".markdown", ".rst", ".typ", ".typst", ".tex"} + +# Extensions that are code/config and should be skipped +SKIP_EXTENSIONS = { + ".py", ".js", ".ts", ".tsx", ".jsx", ".json", ".yaml", ".yml", + ".toml", ".env", ".lock", ".css", ".scss", ".html", ".xml", + ".sql", ".sh", ".bash", ".zsh", ".go", ".rs", ".java", ".c", + ".cpp", ".h", ".hpp", ".rb", ".php", ".swift", ".kt", ".lua", + ".dockerfile", ".makefile", ".csv", ".ini", ".cfg", +} + +# Patterns that indicate a line is code +CODE_PATTERNS = [ + re.compile(r"^\s*(import |from .+ import |require\(|const |let |var )"), + re.compile(r"^\s*(def |class |function |async function |export )"), + re.compile(r"^\s*(if\s*\(|for\s*\(|while\s*\(|switch\s*\(|try\s*\{)"), + re.compile(r"^\s*[\}\]\);]+\s*$"), # closing braces/brackets + re.compile(r"^\s*@\w+"), # decorators/annotations + re.compile(r'^\s*"[^"]+"\s*:\s*'), # JSON-like key-value + re.compile(r"^\s*\w+\s*=\s*[{\[\(\"']"), # assignment with literal +] + + +def _is_code_line(line: str) -> bool: + """Check if a line looks like code.""" + return any(p.match(line) for p in CODE_PATTERNS) + + +def _is_json_content(text: str) -> bool: + """Check if content is valid JSON.""" + try: + json.loads(text) + return True + except (json.JSONDecodeError, ValueError): + return False + + +def _is_yaml_content(lines: list[str]) -> bool: + """Heuristic: check if content looks like YAML.""" + yaml_indicators = 0 + for line in lines[:30]: + stripped = line.strip() + if stripped.startswith("---"): + yaml_indicators += 1 + elif re.match(r"^\w[\w\s]*:\s", stripped): + yaml_indicators += 1 + elif stripped.startswith("- ") and ":" in stripped: + yaml_indicators += 1 + # If most non-empty lines look like YAML + non_empty = sum(1 for l in lines[:30] if l.strip()) + return non_empty > 0 and yaml_indicators / non_empty > 0.6 + + +def detect_file_type(filepath: Path) -> str: + """Classify a file as 'natural_language', 'code', 'config', or 'unknown'. + + Returns: + One of: 'natural_language', 'code', 'config', 'unknown' + """ + ext = filepath.suffix.lower() + + # Extension-based classification + if ext in COMPRESSIBLE_EXTENSIONS: + return "natural_language" + if ext in SKIP_EXTENSIONS: + return "code" if ext not in {".json", ".yaml", ".yml", ".toml", ".ini", ".cfg", ".env"} else "config" + + # Extensionless files (like CLAUDE.md, TODO) — check content + if not ext: + try: + text = filepath.read_text(errors="ignore") + except (OSError, PermissionError): + return "unknown" + + lines = text.splitlines()[:50] + + if _is_json_content(text[:10000]): + return "config" + if _is_yaml_content(lines): + return "config" + + code_lines = sum(1 for l in lines if l.strip() and _is_code_line(l)) + non_empty = sum(1 for l in lines if l.strip()) + if non_empty > 0 and code_lines / non_empty > 0.4: + return "code" + + return "natural_language" + + return "unknown" + + +def should_compress(filepath: Path) -> bool: + """Return True if the file is natural language and should be compressed.""" + if not filepath.is_file(): + return False + # Skip backup files + if filepath.name.endswith(".original.md"): + return False + return detect_file_type(filepath) == "natural_language" + + +if __name__ == "__main__": + import sys + + if len(sys.argv) < 2: + print("Usage: python detect.py <file1> [file2] ...") + sys.exit(1) + + for path_str in sys.argv[1:]: + p = Path(path_str).resolve() + file_type = detect_file_type(p) + compress = should_compress(p) + print(f" {p.name:30s} type={file_type:20s} compress={compress}") diff --git a/.continue/skills/caveman-compress/scripts/validate.py b/.continue/skills/caveman-compress/scripts/validate.py new file mode 100644 index 0000000..dc07307 --- /dev/null +++ b/.continue/skills/caveman-compress/scripts/validate.py @@ -0,0 +1,213 @@ +#!/usr/bin/env python3 +import re +from collections import Counter +from pathlib import Path + +URL_REGEX = re.compile(r"https?://[^\s)]+") +FENCE_OPEN_REGEX = re.compile(r"^(\s{0,3})(`{3,}|~{3,})(.*)$") +HEADING_REGEX = re.compile(r"^(#{1,6})\s+(.*)", re.MULTILINE) +BULLET_REGEX = re.compile(r"^\s*[-*+]\s+", re.MULTILINE) + +# crude but effective path detection +# Requires either a path prefix (./ ../ / or drive letter) or a slash/backslash within the match +PATH_REGEX = re.compile(r"(?:\./|\.\./|/|[A-Za-z]:\\)[\w\-/\\\.]+|[\w\-\.]+[/\\][\w\-/\\\.]+") + + +class ValidationResult: + def __init__(self): + self.is_valid = True + self.errors = [] + self.warnings = [] + + def add_error(self, msg): + self.is_valid = False + self.errors.append(msg) + + def add_warning(self, msg): + self.warnings.append(msg) + + +def read_file(path: Path) -> str: + return path.read_text(errors="ignore") + + +# ---------- Extractors ---------- + + +def extract_headings(text): + return [(level, title.strip()) for level, title in HEADING_REGEX.findall(text)] + + +def extract_code_blocks(text): + """Line-based fenced code block extractor. + + Handles ``` and ~~~ fences with variable length (CommonMark: closing + fence must use same char and be at least as long as opening). Supports + nested fences (e.g. an outer 4-backtick block wrapping inner 3-backtick + content). + """ + blocks = [] + lines = text.split("\n") + i = 0 + n = len(lines) + while i < n: + m = FENCE_OPEN_REGEX.match(lines[i]) + if not m: + i += 1 + continue + fence_char = m.group(2)[0] + fence_len = len(m.group(2)) + open_line = lines[i] + block_lines = [open_line] + i += 1 + closed = False + while i < n: + close_m = FENCE_OPEN_REGEX.match(lines[i]) + if ( + close_m + and close_m.group(2)[0] == fence_char + and len(close_m.group(2)) >= fence_len + and close_m.group(3).strip() == "" + ): + block_lines.append(lines[i]) + closed = True + i += 1 + break + block_lines.append(lines[i]) + i += 1 + if closed: + blocks.append("\n".join(block_lines)) + # Unclosed fences are silently skipped — they indicate malformed markdown + # and including them would cause false-positive validation failures. + return blocks + + +def extract_urls(text): + return set(URL_REGEX.findall(text)) + + +def extract_paths(text): + return set(PATH_REGEX.findall(text)) + + +def count_bullets(text): + return len(BULLET_REGEX.findall(text)) + + +def extract_inline_codes(text): + text_without_fences = re.sub(r"^```[\s\S]*?^```", "", text, flags=re.MULTILINE) + text_without_fences = re.sub(r"^~~~[\s\S]*?^~~~", "", text_without_fences, flags=re.MULTILINE) + return re.findall(r"`([^`]+)`", text_without_fences) + + +# ---------- Validators ---------- + + +def validate_headings(orig, comp, result): + h1 = extract_headings(orig) + h2 = extract_headings(comp) + + if len(h1) != len(h2): + result.add_error(f"Heading count mismatch: {len(h1)} vs {len(h2)}") + + if h1 != h2: + result.add_warning("Heading text/order changed") + + +def validate_code_blocks(orig, comp, result): + c1 = extract_code_blocks(orig) + c2 = extract_code_blocks(comp) + + if c1 != c2: + result.add_error("Code blocks not preserved exactly") + + +def validate_urls(orig, comp, result): + u1 = extract_urls(orig) + u2 = extract_urls(comp) + + if u1 != u2: + result.add_error(f"URL mismatch: lost={u1 - u2}, added={u2 - u1}") + + +def validate_paths(orig, comp, result): + p1 = extract_paths(orig) + p2 = extract_paths(comp) + + if p1 != p2: + result.add_warning(f"Path mismatch: lost={p1 - p2}, added={p2 - p1}") + + +def validate_bullets(orig, comp, result): + b1 = count_bullets(orig) + b2 = count_bullets(comp) + + if b1 == 0: + return + + diff = abs(b1 - b2) / b1 + + if diff > 0.15: + result.add_warning(f"Bullet count changed too much: {b1} -> {b2}") + + +def validate_inline_codes(orig, comp, result): + c1 = Counter(extract_inline_codes(orig)) + c2 = Counter(extract_inline_codes(comp)) + + if c1 != c2: + lost = set(c1.keys()) - set(c2.keys()) + added = set(c2.keys()) - set(c1.keys()) + for code, count in c1.items(): + if code in c2 and c2[code] < count: + lost.add(f"{code} (lost {count - c2[code]} of {count} occurrences)") + if lost: + result.add_error(f"Inline code lost: {lost}") + if added: + result.add_warning(f"Inline code added: {added}") + + +# ---------- Main ---------- + + +def validate(original_path: Path, compressed_path: Path) -> ValidationResult: + result = ValidationResult() + + orig = read_file(original_path) + comp = read_file(compressed_path) + + validate_headings(orig, comp, result) + validate_code_blocks(orig, comp, result) + validate_urls(orig, comp, result) + validate_paths(orig, comp, result) + validate_bullets(orig, comp, result) + validate_inline_codes(orig, comp, result) + + return result + + +# ---------- CLI ---------- + +if __name__ == "__main__": + import sys + + if len(sys.argv) != 3: + print("Usage: python validate.py <original> <compressed>") + sys.exit(1) + + orig = Path(sys.argv[1]).resolve() + comp = Path(sys.argv[2]).resolve() + + res = validate(orig, comp) + + print(f"\nValid: {res.is_valid}") + + if res.errors: + print("\nErrors:") + for e in res.errors: + print(f" - {e}") + + if res.warnings: + print("\nWarnings:") + for w in res.warnings: + print(f" - {w}") diff --git a/.continue/skills/caveman-help/README.md b/.continue/skills/caveman-help/README.md new file mode 100644 index 0000000..5841256 --- /dev/null +++ b/.continue/skills/caveman-help/README.md @@ -0,0 +1,38 @@ +# caveman-help + +Quick-reference card. One shot, no mode change. + +## What it does + +Prints a cheat sheet of all caveman modes, sibling skills, deactivation triggers, and how to set the default mode via env var or config file. One-shot display — does not flip the active mode, write flag files, or persist anything. Use when you forget the slash commands. + +## How to invoke + +``` +/caveman-help +``` + +Also triggers on "caveman help", "what caveman commands", "how do I use caveman". + +## Example output + +``` +Modes: + /caveman full (default) + /caveman lite lighter + /caveman ultra extreme + /caveman wenyan classical Chinese + +Skills: + /caveman-commit terse Conventional Commits + /caveman-review one-line PR comments + /caveman-stats session token savings + +Deactivate: + "stop caveman" or "normal mode" +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full reference card +- [Caveman README](../../README.md) — repo overview diff --git a/.continue/skills/caveman-help/SKILL.md b/.continue/skills/caveman-help/SKILL.md new file mode 100644 index 0000000..ce1156f --- /dev/null +++ b/.continue/skills/caveman-help/SKILL.md @@ -0,0 +1,63 @@ +--- +name: caveman-help +description: > + Quick-reference card for all caveman modes, skills, and commands. + One-shot display, not a persistent mode. Trigger: /caveman-help, + "caveman help", "what caveman commands", "how do I use caveman". +--- + +# Caveman Help + +Display this reference card when invoked. One-shot — do NOT change mode, write flag files, or persist anything. Output in caveman style. + +## Modes + +| Mode | Trigger | What change | +|------|---------|-------------| +| **Lite** | `/caveman lite` | Drop filler. Keep sentence structure. | +| **Full** | `/caveman` | Drop articles, filler, pleasantries, hedging. Fragments OK. Default. | +| **Ultra** | `/caveman ultra` | Extreme compression. Bare fragments. Tables over prose. | +| **Wenyan-Lite** | `/caveman wenyan-lite` | Classical Chinese style, light compression. | +| **Wenyan-Full** | `/caveman wenyan` | Full 文言文. Maximum classical terseness. | +| **Wenyan-Ultra** | `/caveman wenyan-ultra` | Extreme. Ancient scholar on a budget. | + +Mode stick until changed or session end. + +## Skills + +| Skill | Trigger | What it do | +|-------|---------|-----------| +| **caveman-commit** | `/caveman-commit` | Terse commit messages. Conventional Commits. ≤50 char subject. | +| **caveman-review** | `/caveman-review` | One-line PR comments: `L42: bug: user null. Add guard.` | +| **caveman-compress** | `/caveman-compress <file>` | Compress .md files to caveman prose. Saves ~46% input tokens. | +| **caveman-help** | `/caveman-help` | This card. | + +## Deactivate + +Say "stop caveman" or "normal mode". Resume anytime with `/caveman`. + +## Language + +Keep user's language by default. User write Portuguese → reply Portuguese caveman. Compress the style, not the language. Technical terms, code, commands, commit types, and exact error strings stay verbatim unless user ask for translation. + +## Configure Default Mode + +Default mode = `full`. Change it: + +**Environment variable** (highest priority): +```bash +export CAVEMAN_DEFAULT_MODE=ultra +``` + +**Config file** (`~/.config/caveman/config.json`): +```json +{ "defaultMode": "lite" } +``` + +Set `"off"` to disable auto-activation on session start. User can still activate manually with `/caveman`. + +Resolution: env var > config file > `full`. + +## More + +Full docs: https://github.com/JuliusBrussee/caveman diff --git a/.continue/skills/caveman-review/README.md b/.continue/skills/caveman-review/README.md new file mode 100644 index 0000000..acf519f --- /dev/null +++ b/.continue/skills/caveman-review/README.md @@ -0,0 +1,33 @@ +# caveman-review + +One-line PR comments. Location, problem, fix. No throat-clearing. + +## What it does + +Generates code review comments in `L<line>: <severity> <problem>. <fix>.` format. One line per finding. Severity emoji: 🔴 bug, 🟡 risk, 🔵 nit, ❓ question. Drops "I noticed that...", hedging, and restating what the diff already shows. Keeps exact line numbers, backticked symbols, and concrete fixes. + +Auto-clarity: drops terse mode for CVE-class security findings, architectural disagreements, and onboarding contexts where the author needs the *why*. Resumes terse for the rest. + +Output only — does not approve, request changes, or run linters. + +## How to invoke + +``` +/caveman-review +``` + +Also triggers on "review this PR", "code review", "review the diff". + +## Example output + +``` +L42: 🔴 bug: user can be null after .find(). Add guard before .email. +L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist. +L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3). +L107: ❓ q: why drop the cache here? Reads on next request will miss. +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview diff --git a/.continue/skills/caveman-review/SKILL.md b/.continue/skills/caveman-review/SKILL.md new file mode 100644 index 0000000..48f4adb --- /dev/null +++ b/.continue/skills/caveman-review/SKILL.md @@ -0,0 +1,55 @@ +--- +name: caveman-review +description: > + Ultra-compressed code review comments. Cuts noise from PR feedback while preserving + the actionable signal. Each comment is one line: location, problem, fix. Use when user + says "review this PR", "code review", "review the diff", "/review", or invokes + /caveman-review. Auto-triggers when reviewing pull requests. +--- + +Write code review comments terse and actionable. One line per finding. Location, problem, fix. No throat-clearing. + +## Rules + +**Format:** `L<line>: <problem>. <fix>.` — or `<file>:L<line>: ...` when reviewing multi-file diffs. + +**Severity prefix (optional, when mixed):** +- `🔴 bug:` — broken behavior, will cause incident +- `🟡 risk:` — works but fragile (race, missing null check, swallowed error) +- `🔵 nit:` — style, naming, micro-optim. Author can ignore +- `❓ q:` — genuine question, not a suggestion + +**Drop:** +- "I noticed that...", "It seems like...", "You might want to consider..." +- "This is just a suggestion but..." — use `nit:` instead +- "Great work!", "Looks good overall but..." — say it once at the top, not per comment +- Restating what the line does — the reviewer can read the diff +- Hedging ("perhaps", "maybe", "I think") — if unsure use `q:` + +**Keep:** +- Exact line numbers +- Exact symbol/function/variable names in backticks +- Concrete fix, not "consider refactoring this" +- The *why* if the fix isn't obvious from the problem statement + +## Examples + +❌ "I noticed that on line 42 you're not checking if the user object is null before accessing the email property. This could potentially cause a crash if the user is not found in the database. You might want to add a null check here." + +✅ `L42: 🔴 bug: user can be null after .find(). Add guard before .email.` + +❌ "It looks like this function is doing a lot of things and might benefit from being broken up into smaller functions for readability." + +✅ `L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist.` + +❌ "Have you considered what happens if the API returns a 429? I think we should probably handle that case." + +✅ `L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3).` + +## Auto-Clarity + +Drop terse mode for: security findings (CVE-class bugs need full explanation + reference), architectural disagreements (need rationale, not just a one-liner), and onboarding contexts where the author is new and needs the "why". In those cases write a normal paragraph, then resume terse for the rest. + +## Boundaries + +Reviews only — does not write the code fix, does not approve/request-changes, does not run linters. Output the comment(s) ready to paste into the PR. "stop caveman-review" or "normal mode": revert to verbose review style. \ No newline at end of file diff --git a/.continue/skills/caveman-stats/README.md b/.continue/skills/caveman-stats/README.md new file mode 100644 index 0000000..30b28f9 --- /dev/null +++ b/.continue/skills/caveman-stats/README.md @@ -0,0 +1,30 @@ +# caveman-stats + +Real session token receipts. No AI estimation. + +## What it does + +Reads the current Claude Code session log directly and reports actual input/output token usage plus estimated savings versus a non-caveman baseline. Numbers come from the JSONL session log on disk — the model itself does not compute or estimate them. Output is injected by the `caveman-mode-tracker` hook, which intercepts `/caveman-stats` and returns the formatted stats as a blocked-decision reason. + +Each run also writes a lifetime-savings suffix file used by the statusline badge (`⛏ 12.4k`). + +## How to invoke + +``` +/caveman-stats +``` + +## Example output + +``` +Session: 47 turns +Input: 12,304 tokens +Output: 3,891 tokens (caveman) +Baseline: 11,247 tokens (estimated without caveman) +Saved: 7,356 tokens (~65%) +``` + +## See also + +- [`SKILL.md`](./SKILL.md) — hook contract and mechanics +- [Caveman README](../../README.md) — repo overview diff --git a/.continue/skills/caveman-stats/SKILL.md b/.continue/skills/caveman-stats/SKILL.md new file mode 100644 index 0000000..a7348ac --- /dev/null +++ b/.continue/skills/caveman-stats/SKILL.md @@ -0,0 +1,10 @@ +--- +name: caveman-stats +description: > + Show real token usage and estimated savings for the current session. + Reads directly from the Claude Code session log — no AI estimation. + Triggers on /caveman-stats. Output is injected by the mode-tracker hook; + the model itself does not compute the numbers. +--- + +This skill is delivered by `hooks/caveman-stats.js` (read by `hooks/caveman-mode-tracker.js` on `/caveman-stats`). The model does not need to do anything when this skill fires — the hook returns `decision: "block"` with the formatted stats as the reason. The user sees the numbers immediately. diff --git a/.continue/skills/caveman/README.md b/.continue/skills/caveman/README.md new file mode 100644 index 0000000..d749b83 --- /dev/null +++ b/.continue/skills/caveman/README.md @@ -0,0 +1,48 @@ +# caveman + +Talk like smart caveman. Same brain, fewer tokens. + +## What it does + +Compress every model response to caveman-style prose. Drops articles, filler, pleasantries, and hedging. Keeps every technical detail, code block, error string, and symbol exact. Cuts ~65-75% of output tokens with full accuracy preserved. Mode persists for the whole session until changed or stopped. + +Six intensity levels: + +| Level | What change | +|-------|-------------| +| `lite` | Drop filler/hedging. Sentences stay full. Professional but tight. | +| `full` | Default. Drop articles, fragments OK, short synonyms. | +| `ultra` | Bare fragments. Abbreviations (DB, auth, fn). Arrows for causality. | +| `wenyan-lite` | Classical Chinese register, light compression. | +| `wenyan-full` | Maximum 文言文. 80-90% character reduction. | +| `wenyan-ultra` | Extreme classical compression. | + +Auto-clarity rule: caveman drops to normal prose for security warnings, irreversible-action confirmations, multi-step sequences where fragment ambiguity risks misread, and when user repeats a question. Resumes after the clear part. + +## How to invoke + +``` +/caveman # full mode (default) +/caveman lite # lighter compression +/caveman ultra # extreme compression +/caveman wenyan # classical Chinese +stop caveman # back to normal prose +``` + +## Example output + +Question: "Why does my React component re-render?" + +Normal prose: +> Your component re-renders because you create a new object reference each render. Wrapping it in `useMemo` will fix the issue. + +Caveman (full): +> New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`. + +Caveman (ultra): +> Inline obj prop → new ref → re-render. `useMemo`. + +## See also + +- [`SKILL.md`](./SKILL.md) — full LLM-facing instructions +- [Caveman README](../../README.md) — repo overview, install, benchmarks diff --git a/.continue/skills/caveman/SKILL.md b/.continue/skills/caveman/SKILL.md new file mode 100644 index 0000000..8792c1a --- /dev/null +++ b/.continue/skills/caveman/SKILL.md @@ -0,0 +1,78 @@ +--- +name: caveman +description: > + Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman + while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, + wenyan-lite, wenyan-full, wenyan-ultra. + Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", + "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested. +--- + +Respond terse like smart caveman. All technical substance stay. Only fluff die. + +## Persistence + +ACTIVE EVERY RESPONSE. No revert after many turns. No filler drift. Still active if unsure. Off only: "stop caveman" / "normal mode". + +Default: **full**. Switch: `/caveman lite|full|ultra`. + +## Rules + +Drop: articles (a/an/the), filler (just/really/basically/actually/simply), pleasantries (sure/certainly/of course/happy to), hedging. Fragments OK. Short synonyms (big not extensive, fix not "implement a solution for"). No tool-call narration, no decorative tables/emoji, no dumping long raw error logs unless asked — quote shortest decisive line. Standard well-known tech acronyms OK (DB/API/HTTP); never invent new abbreviations reader can't decode. Technical terms exact. Code blocks unchanged. Errors quoted exact. + +Preserve user's dominant language. User write Portuguese → reply Portuguese caveman. User write Spanish → reply Spanish caveman. Compress the style, not the language. No forced English openings or status phrases. ALWAYS keep technical terms, code, API names, CLI commands, commit-type keywords (feat/fix/...), and exact error strings verbatim — unless user explicitly ask for translation. + +No self-reference. Never name or announce the style. No "caveman mode on", "me caveman think", no third-person caveman tags. Output caveman-only — never normal answer plus "Caveman:" recap. Exception: user explicitly ask what the mode is. + +Pattern: `[thing] [action] [reason]. [next step].` + +Not: "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by..." +Yes: "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:" + +## Intensity + +| Level | What change | +|-------|------------| +| **lite** | No filler/hedging. Keep articles + full sentences. Professional but tight | +| **full** | Drop articles, fragments OK, short synonyms. Classic caveman. No tool-call narration, no decorative tables/emoji, no long raw error-log dumps unless asked. Standard acronyms OK; no invented abbreviations | +| **ultra** | Abbreviate prose words (DB/auth/config/req/res/fn/impl) — prose words only, never real code symbols/function names. Strip conjunctions, arrows for causality (X → Y), one word when one word enough. Code symbols, function names, API names, error strings: never abbreviate | +| **wenyan-lite** | Semi-classical. Drop filler/hedging but keep grammar structure, classical register | +| **wenyan-full** | Maximum classical terseness. Fully 文言文. 80-90% character reduction. Classical sentence patterns, verbs precede objects, subjects often omitted, classical particles (之/乃/為/其) | +| **wenyan-ultra** | Extreme abbreviation while keeping classical Chinese feel. Maximum compression, ultra terse | + +Example — "Why React component re-render?" +- lite: "Your component re-renders because you create a new object reference each render. Wrap it in `useMemo`." +- full: "New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`." +- ultra: "Inline obj prop → new ref → re-render. `useMemo`." +- wenyan-lite: "組件頻重繪,以每繪新生對象參照故。以 useMemo 包之。" +- wenyan-full: "每繪新生對象參照,故重繪;以 useMemo 包之則免。" +- wenyan-ultra: "新參照→重繪。useMemo Wrap。" + +Example — "Explain database connection pooling." +- lite: "Connection pooling reuses open connections instead of creating new ones per request. Avoids repeated handshake overhead." +- full: "Pool reuse open DB connections. No new connection per request. Skip handshake overhead." +- ultra: "Pool = reuse DB conn. Skip handshake → fast under load." +- wenyan-full: "池reuse open connection。不每req新開。skip handshake overhead。" +- wenyan-ultra: "池reuse conn。skip handshake → fast。" + +## Auto-Clarity + +Drop caveman when: +- Security warnings +- Irreversible action confirmations +- Multi-step sequences where fragment order or omitted conjunctions risk misread +- Compression itself creates technical ambiguity (e.g., `"migrate table drop column backup first"` — order unclear without articles/conjunctions) +- User asks to clarify or repeats question + +Resume caveman after clear part done. + +Example — destructive op: +> **Warning:** This will permanently delete all rows in the `users` table and cannot be undone. +> ```sql +> DROP TABLE users; +> ``` +> Caveman resume. Verify backup exist first. + +## Boundaries + +Code/commits/PRs: write normal. "stop caveman" or "normal mode": revert. Level persist until changed or session end. \ No newline at end of file diff --git a/skills-lock.json b/skills-lock.json new file mode 100644 index 0000000..c3044cd --- /dev/null +++ b/skills-lock.json @@ -0,0 +1,47 @@ +{ + "version": 1, + "skills": { + "cavecrew": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/cavecrew/SKILL.md", + "computedHash": "505d836228d1c5e14834ff5d62aad72390c7d27f79c6aa7f9a7a55ed6606d6a2" + }, + "caveman": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman/SKILL.md", + "computedHash": "1902fa0b569912d0c05736d8d98a72097d9b82719aac88c0c1d03bb546f9176d" + }, + "caveman-commit": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman-commit/SKILL.md", + "computedHash": "790a4eeace0be35c6691faf923518ba5bd50f1f1305d1101d09dd4971be94e00" + }, + "caveman-compress": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman-compress/SKILL.md", + "computedHash": "1e9b3e2bf68b75dc0252c4328c8514bea79d4d8f7d7259da616ba7d12cad6865" + }, + "caveman-help": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman-help/SKILL.md", + "computedHash": "dd85267e76baad76995157e7b9f762dfa557cd58951ee92af0c283f48aa26537" + }, + "caveman-review": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman-review/SKILL.md", + "computedHash": "fb7214a1c5793bae6ba8b1be4329e2e6f40dbec6dd911dfb335ad29f09c316a1" + }, + "caveman-stats": { + "source": "JuliusBrussee/caveman", + "sourceType": "github", + "skillPath": "skills/caveman-stats/SKILL.md", + "computedHash": "47ce2de3d6cb39a75047b5c962e4eb3da15594e7397c94103e9a104d42626553" + } + } +} From b830621b28e48834531a0118e14abbdf9fcb3548 Mon Sep 17 00:00:00 2001 From: Evin Bento <evin.dev009@gmail.com> Date: Tue, 23 Jun 2026 16:23:11 -0400 Subject: [PATCH 20/20] docs: add Phase 6.5 UI Design Overhaul plan + update product roadmap - Phase_6.5_UIOverhaul.md: 7-branch overhaul plan for all unstyled pages - product-plan.md: Phase 6.5 entry inserted after Phase 6 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --- Argus Details/product-plan.md | 11 +++ Phase Plans/Phase_6.5_UIOverhaul.md | 111 ++++++++++++++++++++++++++++ 2 files changed, 122 insertions(+) create mode 100644 Phase Plans/Phase_6.5_UIOverhaul.md diff --git a/Argus Details/product-plan.md b/Argus Details/product-plan.md index 4e9cd87..e65b00b 100644 --- a/Argus Details/product-plan.md +++ b/Argus Details/product-plan.md @@ -140,6 +140,17 @@ LangGraph multi-agent pipeline (enrichment + analyst + memory nodes), pgvector R --- +## Phase 6.5 — UI Design Overhaul ⬜ +*After Phase 6* + +**Goal:** Bring every app page into the Argus design language. Pages were built functional-first with zero token usage. This phase makes the product look cohesive before Phase 7 ships new surfaces. + +Pages redesigned: login, signup, verify-email, transactions, bills, bills calendar, subscriptions, accounts, settings, intelligence feed, Argus chat page + response cards, Argus side panel. + +**Deliverable:** Every page uses design tokens consistently — no raw color/spacing utilities, typography scale applied everywhere, side panel visually matches the main chat page. + +--- + ## Phase 7 — Financial Profile + Guardian ⬜ *Weeks 15–16* diff --git a/Phase Plans/Phase_6.5_UIOverhaul.md b/Phase Plans/Phase_6.5_UIOverhaul.md new file mode 100644 index 0000000..328bebe --- /dev/null +++ b/Phase Plans/Phase_6.5_UIOverhaul.md @@ -0,0 +1,111 @@ +# ArgusAI — Phase 6.5: UI Design Overhaul + +> After Phase 6. Goal: bring every app page into the Argus design language before Phase 7 ships new surfaces. Pages were built functional-first; this phase makes them look like the product. + +--- + +## What This Phase Covers + +All pages currently use zero design tokens. Every page below gets: +- Design tokens (`--surface-*`, `--r-*`, `--font-display`, `--font-sans`, `--copper`, `--paper`, etc.) +- Argus typography scale (`--text-hero`, `--text-kpi`, `--text-heading`, `--text-sub`, `--text-eyebrow`) +- Consistent card pattern (`--surface-1` background, `--surface-3` border, `--r-lg` radius) +- No raw Tailwind color utilities (`bg-white`, `text-gray-*`, `p-8`) + +--- + +## Git Branch Structure + +``` +develop +└── phase/6.5-ui-overhaul + ├── feature/overhaul-auth-pages + ├── feature/overhaul-transactions + ├── feature/overhaul-bills-subscriptions + ├── feature/overhaul-accounts-settings + ├── feature/overhaul-intelligence + ├── feature/overhaul-argus-chat + └── feature/overhaul-argus-side-panel +``` + +--- + +## Execution Checklist + +### `feature/overhaul-auth-pages` ⬜ +*Login + signup deferred from Phase 1.5 — still plain* + +- [ ] `app/(auth)/login/page.tsx` — apply brand tokens, aurora background, Instrument Serif heading, copper accent CTA +- [ ] `app/(auth)/signup/page.tsx` — same treatment as login +- [ ] `app/(auth)/verify-email/page.tsx` — match auth page pattern +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-transactions` ⬜ +*Functional table with no visual design* + +- [ ] `app/(app)/transactions/page.tsx` — design token layout, merchant name + amount + category row pattern, recurring badge, category filter as styled chip group, pagination controls on-brand +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-bills-subscriptions` ⬜ +*Both pages unstyled* + +- [ ] `app/(app)/bills/page.tsx` — card-per-bill layout, copper accent for overdue, due-date chip, amount styled as KPI +- [ ] `app/(app)/bills/calendar/page.tsx` — on-brand calendar grid, urgency color coding, consistent with the unified Smart Payment Calendar arriving in Phase 6 +- [ ] `app/(app)/subscriptions/page.tsx` — logo tile fallback, price creep badge, renewal date chip, cancel CTA on-brand +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-accounts-settings` ⬜ +*Both pages bare* + +- [ ] `app/(app)/accounts/page.tsx` — card-per-account, balance as KPI, institution name as eyebrow, last-synced timestamp sub-text +- [ ] `app/(app)/settings/page.tsx` — section grouping with `--surface-1` card blocks, labels/values in Argus type scale, destructive actions clearly marked +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-intelligence` ⬜ +*AI insights feed with no visual hierarchy* + +- [ ] `app/(app)/intelligence/page.tsx` — insight card pattern: eyebrow label (insight type), heading (the finding), sub-text (reasoning), copper accent bar on left edge, timestamp +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-argus-chat` ⬜ +*Chat page functional but not visually cohesive* + +- [ ] `app/(app)/argus/page.tsx` — page shell matches app layout tokens, input bar on-brand, streaming indicator using copper +- [ ] `app/(app)/argus/_components/Cards.tsx` — verdict/table/chart cards use `--surface-1`/`--surface-2` hierarchy, copper accent for Argus identity, Instrument Serif for verdict text +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### `feature/overhaul-argus-side-panel` ⬜ +*Panel functional but minimal token usage (3 references)* + +- [ ] `app/(app)/_components/ArgusSidePanel.tsx` — full token adoption, copper header bar, Argus eye mark branding, input matches main chat bar, response cards match `Cards.tsx` pattern, slide animation polished +- [ ] Merge → `phase/6.5-ui-overhaul` + +--- + +### Phase 6.5 Close +- [ ] Merge `phase/6.5-ui-overhaul` → `develop` +- [ ] Open PR `develop` → `main`, wait for CI, merge +- [ ] Delete all feature branches + `phase/6.5-ui-overhaul` +- [ ] Mark Phase 6.5 as ✅ Complete in `Argus Details/product-plan.md` + +--- + +## Definition of Done + +- [ ] Every page above uses only design tokens — zero raw color/spacing utilities +- [ ] Typography scale applied consistently across all pages +- [ ] Argus side panel visually matches the main chat page +- [ ] Auth pages match the landing page visual quality +- [ ] No page looks like it was built by a different team than the dashboard