Skip to content

Latest commit

 

History

3,480 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Kazma Agent Framework

Kazma Agent Framework

Kazma is the self-hosted agent that can edit your repo, message your team, and schedule your life — and that will stop, ask, or fail honestly rather than invent an answer.

MIT License Python 3.11+ Tests Prompt injection benchmark Commits Website


Quick start

.\setup.ps1
kazma serve

Open / — that is chat (not the dashboard). First-run asks for one provider key and one model. Inspectors (dashboard, swarm, memory, …) live under More.

See Quickstart.


⚡ What it is

Kazma is an open-source, self-hosted agent: one LangGraph brain, HITL before danger tools, a commitment layer that will not invent a date over your memory, and mouths on Web, TUI, CLI, Telegram, Discord, and Slack. When the model dies it says so (⚠️) instead of faking a reply.

Codebase Volume Test Suite Engineering Depth Platforms Supported
~433K LOC (345K Python code + 37K JS) 7,699 test functions (611 test files) 3,436+ commits across 7 packages Web, TUI, CLI, Telegram, Discord, Slack

Kazma Observability Dashboard & Control Plane


🔬 Measured, not asserted

Most agent frameworks describe their safety. Kazma publishes the measurement, the method, and the results that do not flatter it.

Prompt injection, on AgentDojo — a public benchmark built by other people (Debenedetti et al., NeurIPS 2024), 996 runs per condition, every condition measured four times:

condition attack success acted on the payload
undefended 18.1% 24.1%
spotlighting — 4-character delimiter, from the literature 11.6% 18.6%
Kazma's fence — ~800-character in-band banner 10.4% 14.8%

Fencing untrusted tool output works (p < 0.001 against undefended). It also cannot be told apart from a four-character delimiter (p = 0.39). Both sentences are on the page, because the second one is the one a reviewer needs.

What that page also reports, because leaving it out would make the rest worth less:

  • Two of the four suites measure nothing on the model used — their undefended baselines sit at 2.9% and 0.3%, so there is no attack success for a defense to reduce.
  • A result we refused. On one suite the fence beat spotlighting at p = 0.032 — on a suite where spotlighting scored worse than no defense at all. It is shown, and not counted.
  • A measured noise floor. Running an unchanged configuration four times gives a 5.7-point spread at temperature 0. Anything smaller than that is not a finding, including ours.

Every figure re-derives from the committed run logs, with no API key:

scripts/agentdojo_bench.py --analyze --suite banking   # raw counts + obedience
scripts/agentdojo_bench.py --report slack,banking      # pooled figures + p-values
Prompt injection: the numbers the full measurement, the payloads that still land, and a free offline reproduction
Threat model what each boundary stops and — stated plainly — what it does not. Approval is consent, not containment
Known gaps open weaknesses, dated, so they do not depend on someone remembering

📖 Origin & Architectural Philosophy

Kazma (كاظمة) was an ancient coastal oasis in Kuwait — a vital network of freshwater wells and a flourishing gateway connecting global trade routes between civilizations. In 633 CE, it was the site of the historic Battle of Chains (ذات السلاسل): an opposing army chained its ranks into a rigid, monolithic wall, which Khalid ibn al-Walid decisively dismantled through adaptive, decentralized maneuvering.

Kazma's architecture reflects those foundational principles:

  • 🏜️ The Wells (Cognitive Memory) — Deep, persistent memory that retains context across months of sessions, allowing agents to draw from bi-temporal knowledge graphs rather than forgetting across turns.
  • 🚪 The Gateway (Multi-Platform Control) — A unified supervisor brain seamlessly routing execution between Web UI, Textual TUI, CLI, and team messaging channels (Telegram, Discord, Slack).
  • ⚔️ Breaking the Chains (Decentralized Swarms) — Monolithic, rigid pipelines inevitably fail in real-world deployments. Kazma replaces brittle linear chains with decentralized swarm dispatch patterns, dynamic worker autoscaling, and self-healing execution loops.

🏛️ System Architecture

                                 ┌──────────────────────────────────────────────────────────┐
                                 │          Client Layer (Web / TUI / Chat / CLI)           │
                                 └────────────────────────────┬─────────────────────────────┘
                                                              │
                                                              ▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│                                                   KAZMA GATEWAY & SUPERVISOR                                                │
│  ┌───────────────────────────────┐     ┌───────────────────────────────┐     ┌───────────────────────────────────────────┐  │
│  │     Platform Isolation        │ ──► │   LangGraph ReAct Supervisor  │ ◄─► │         Triple-Wired HITL Gate            │  │
│  │ (SessionStore / Zero Leakage) │     │  (80% Compaction / Turn Ledger)│     │  (Graph Interrupt / Swarm Bus / Pipeline) │  │
│  └───────────────────────────────┘     └───────────────┬───────────────┘     └───────────────────────────────────────────┘  │
│                                                        │                                                                    │
│  ┌───────────────────────────────┐     ┌───────────────┴───────────────┐     ┌───────────────────────────────────────────┐  │
│  │   Document Intelligence       │ ──► │     Commitment Layer Gate     │ ◄── │          Local & Native Tools             │  │
│  │ (CAS / Subprocess OCR / Parse)│     │     (Resolve-Before-Act)      │     │    (IDE / Web / Bash / Python / Vault)    │  │
│  └───────────────────────────────┘     └───────────────┬───────────────┘     └───────────────────────────────────────────┘  │
└────────────────────────────────────────────────────────┼────────────────────────────────────────────────────────────────────┘
                                                         │
                                                         ▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│                                             AUTONOMOUS SWARM & MEMORY TIER                                                  │
│  ┌─────────────────────────────────────────────┐                    ┌────────────────────────────────────────────────────┐  │
│  │                SwarmEngine                  │                    │            Pure V2 Cognitive Memory                │  │
│  │  • 6 Dispatch Patterns (Fan-Out/Pipeline/..)│                    │  • Bi-Temporal Belief Graph (valid_from/until)     │  │
│  │  • Dynamic Autoscaler (Coder/Researcher/..) │                    │  • Local Ego-Graph Personalized PageRank (PPR)     │  │
│  │  • ReliabilityRegistry (Breakers & Retries) │                    │  • Sparse (FTS5) + Dense (sqlite-vec / pgvector)  │  │
│  │  • Best-Model-Per-Task Prompt Classifier    │                    │  • Parametric Action DAGs + 24h Auto-Consolidation │  │
│  └─────────────────────────────────────────────┘                    └────────────────────────────────────────────────────┘  │
└────────────────────────────────────────────────────────┬────────────────────────────────────────────────────────────────────┘
                                                         │
                                                         ▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│                                           EXECUTION & PROVIDER INFRASTRUCTURE                                               │
│  OpenAI-Compatible Layer • Anthropic Native • Google Gemini (ADC) • Azure OpenAI • AWS Bedrock • Ollama / LM Studio • MCP   │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

🌟 Core Capabilities

🧠 Pure V2 Cognitive Memory Engine

  • Bi-Temporal Beliefs: Tracks factual assertions with both assertion time and validity time (valid_from / valid_until) to manage evolving knowledge without hallucination or historical corruption.
  • Associative PPR Graph: Multi-hop associative recall via Local Ego-Graph Personalized PageRank over belief entities.
  • Hybrid Episode Retrieval: Recalls past dialogues and actions using Reciprocal Rank Fusion (RRF) over lexical search (SQLite FTS5, or ILIKE on Postgres-primary) and dense embeddings (sqlite-vec on one node, pgvector when Postgres is on).
  • Automated Ops & Hygiene: Background task queue (memory_ops.db) for post-turn extraction, entity reconciliation, micro-consolidation, and automated backups — WAL-safe SQLite copies, pg_dump of Postgres and a JSONL export of the graph, snapshotted into deduplicated, encrypted restic repositories (local + offsite) with time-based retention and a verified one-command restore. See Disaster Recovery.
  • Prompt-Fenced Injection: Wraps untrusted text in <kazma:data untrusted> fences that tell the model the enclosed content is observation data, never instructions. Covers recalled memories, compaction summaries, procedural hints, skill frontmatter, document knowledge, swarm phonebook entries, and — since the 2026-08-29 audit — fetched web pages (read_url), search results (web_search), saved research chunks, and third-party MCP resource bodies. Note this is a mitigation, not a guarantee: a fence lowers the authority of injected text, it does not make the model immune to it — measured at 18.1% → 10.4% attack success on a public benchmark, which is a real reduction and not an elimination. See INJECTION.md.

🔁 Operator reload & Windows event loop

  • Watched host: after git pull, pick up code with --reload (see Quick Start §4). Do not kill python/uvicorn by hand.
  • Windows: the server runs a SelectorEventLoop so psycopg-async (LangGraph AsyncPostgresSaver) can connect. python -m uvicorn hardcodes Proactor on Windows and silently falls back to SQLite checkpoints — start via kazma serve / the guard, not raw uvicorn.
  • Filesystem tools (file_search, file_read, file_list, …) offload disk I/O off the event loop so a large tree cannot freeze SSE, WebSockets, or /health/ready.

🐝 Swarm Orchestration & Dynamic Autoscaler

  • 6 Dispatch Patterns: dispatch (single specialist), broadcast (all workers), pipeline (sequential handoffs with checkpoint gates), fan-out (parallel execution with aggregation/voting), consult (independent expert reviews + synthesis), and conditional (router-driven execution).
  • Dynamic Autoscaling: Zero pre-configured worker requirement. Automatically classifies task prompts and dynamically spins up specialized workers (coder, researcher, generalist) with best-model-per-task selection (coding, reasoning, vision).
  • Reliability & Circuit Breakers: Per-worker circuit breakers, half-open probes, exponential retry policies, output schema validators, and handoff cycle guards ($depth \le 5$).

🛡️ Non-Stop Execution & Self-Healing Watchdog

  • Heartbeat & Stall Detection: supervised_invoke() watchdog tracks execution heartbeats across graph nodes and automatically mitigates stalls.
  • Checkpoint Rollback & Reflection: Automatically rolls back corrupted turns to clean checkpoint states and injects [KAZMA RECOVERY] system reflection notes to re-steer the model.
  • Model Failover Chains: Transparent multi-provider failover with per-provider cooldown timers and durable SQLite call ledgers (kazma-data/llm_calls.db).

🔒 Triple-Wired HITL Safety Architecture

Fail-closed by default. KAZMA_ALLOW_YOLO=1 turns the gate off for the 53 canonical danger tools that are not in ALWAYS_HITL_TOOLS, and approval is consent, not containment — it does not sandbox what you approve. What each boundary does and does not stop is written out in THREAT_MODEL.md.

  • Default-deny HITL (2026-08-29 audit): unclassified tools are gated; a Settings require_approval_for list adds to the tier floor and can no longer un-gate shell_exec by omission. Behind a reverse proxy, set KAZMA_TRUSTED_PROXIES to the proxy's address (peer 127.0.0.1 is not a credential).
  • Layer 1 (Graph Interrupt): Single-agent execution pauses at the LangGraph level before mutating actions (file_write, shell_exec, vault_retrieve). Resumable from Web, TUI, or chat channels.
  • Layer 2 (Swarm Bus): Multi-agent and CLI swarm dispatches enforce fail-closed approval gates on platform adapters (FanOutBusAdapter across Telegram/Discord/Slack).
  • Layer 3 (Pipeline Checkpoints): Multi-stage pipeline tasks pause at designated approval milestones.
  • Security & Sandboxing: HMAC-SHA256 skill verification, prompt-fenced Soul mutation deltas, and AES-256-GCM encrypted credential vault.

📄 Enterprise Document Intelligence Platform

  • Intake & Quarantine: Content-addressed storage (CAS) with MIME/OOXML/PDF policy validation, macro rejection, and optional ClamAV malware scanning.
  • Isolated Subprocess Processing: Secure OCR and document parsing for PDF, DOCX, XLSX, and PPTX formats in isolated sub-processes.
  • Document Ops: Background job leases (SKIP LOCKED), dead-letter queues, format conversions, PDF split/merge/redaction, and one-click indexing into Knowledge Library corpora.

💻 Dual IDE & Multi-Platform Gateway

  • Web IDE & Textual TUI: Integrated editor with syntax highlighting, multi-tab navigation, workspace-scoped terminal execution, and file-aware AI chat.
  • Live In-Flight Steering: Intercept and guide active operations in real time using /steer (soft nudge), /steer! (pause & inject), or /abort.
  • Zero-Leak Platform Isolation: Session identifiers (chat_id, user_id) remain isolated within SessionStore and never pollute LangGraph state.

🌐 Arabic-Native & Cultural Alignment

  • Majlis Protocol: Native handling of Arabic nuances, formal MSA, and Gulf/Kuwaiti dialect expressions.
  • Bilingual Interface: Full Right-To-Left (RTL) Web and TUI interfaces with culturally aligned interaction models.

🆚 Why Kazma?

Capability Kazma LangChain / LangGraph CrewAI AutoGPT n8n
Architecture Full-Stack Autonomous System Library / Graph Primitive Multi-Agent Framework Autonomous Agent Workflow Automation
Cognitive Memory Bi-temporal + PPR Graph ⚠️ Basic Vector Store ⚠️ Simple RAG ⚠️ Basic Memory ❌ None
HITL Safety Gates Triple-Wired (fail-closed by default — what it does not stop) ⚠️ Manual code wiring ❌ None ⚠️ Basic prompt ⚠️ Workflow pause
Swarm Orchestration 6 Patterns + Autoscaler ⚠️ Custom Graph ✅ Role-based ❌ Single loop ❌ Node based
Built-in Web & TUI IDE Included (Dual Interface) ❌ None ❌ None ❌ None ❌ None
Observability Control Plane Live Dashboard Included ⚠️ External (LangSmith) ❌ None ❌ None ⚠️ Execution log
Document Intelligence Quarantine + OCR + Redact ⚠️ Ad-hoc loaders ❌ None ❌ None ⚠️ Basic parsers
Multi-Platform Gateways Web, TUI, Telegram, Discord, Slack ❌ None ❌ None ❌ None ⚠️ Webhook triggers
Arabic-Native & RTL Full Native & Dialect Support ❌ None ❌ None ❌ None ❌ None
Published safety measurements Public benchmark, noise floor, known gaps
Self-Hosted License MIT (100% Open Source) ✅ MIT ✅ MIT ✅ MIT ⚠️ Fair-Code

“—” means not assessed. We measured Kazma against a public benchmark; we have not run the same benchmark against these projects, so we do not claim a result for them.


📸 Interface Showcase

Observability Dashboard — dark control plane (English and Arabic).

English Arabic
Observability Dashboard (English) Observability Dashboard (Arabic)

🚀 Quick Start

Prerequisites: Python 3.11–3.14 (3.12 or 3.13 recommended). uv is installed for you if missing.

The install SoT is Quickstart. Bootstrap scripts (setup.ps1 / setup.sh) sync rag + dev + tui. There is no [cli] extra — kazma-cli is part of the wheel. Full optional extras: uv sync --all-extras.

1. Installation

git clone https://github.com/Mubder/kazma.git
cd kazma

One command (recommended)

# Windows
.\setup.ps1
# Linux / macOS / WSL
chmod +x setup.sh
./setup.sh

Creates .venv, installs uv if needed, syncs rag + dev + tui, copies .env.example.env when missing, and checks core imports.

Manual: uv

uv venv --python 3.13
uv sync --extra rag --extra dev --extra tui
# Everything (torch, Playwright, WeasyPrint, Temporal, …):
# uv sync --all-extras

Manual: pip + venv

# Linux / macOS / WSL
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[rag,dev,tui]"

# Windows (PowerShell)
py -3.13 -m venv .venv
.venv\Scripts\Activate.ps1
pip install -e ".[rag,dev,tui]"

2. Environment Configuration

# Copy template environment file
cp .env.example .env    # Linux / macOS
Copy-Item .env.example .env  # Windows PowerShell

Edit .env to configure your preferred LLM provider key:

# OpenAI-compatible (also used as the generic env fallback):
OPENAI_API_KEY=sk-...
# Other providers (Anthropic, Gemini, DeepSeek, …) are keyed in
# Settings → Providers / kazma.yaml — see docs/docs/guide/configuration.md

3. Launch Kazma

# Start the full Web UI & Gateway (http://127.0.0.1:9090)
kazma serve

# Run the agent without the web server (tokens stream to stdout)
kazma ask "What files define the supervisor graph?"

# Or launch the Terminal User Interface (TUI)
kazma-tui

Navigate to:

  • Dashboard & Control Plane: http://127.0.0.1:9090/
  • Web IDE: http://127.0.0.1:9090/ide
  • Document Intelligence: http://127.0.0.1:9090/documents
  • Memory & Belief Graph: http://127.0.0.1:9090/memory

4. Reload & status (watched host — do not skip)

If KazmaAgent is supervising the process (Scheduled Task / systemd / launchd), this is how you apply a pull. Hand-killing python or uvicorn fights the guard: it either respawns the old port holder or reports healthy while serving stale code.

# Apply new code (cold start typically 3–5 minutes; wait for "Kazma is up. build …")
& '.venv\Scripts\python.exe' scripts\service\kazma_guard.py --reload

# Watcher + /health/ready + pids
& '.venv\Scripts\python.exe' scripts\service\kazma_guard.py --status

--status should show supervision : active and server : healthy (ready). Do not Ctrl+C --reload unless you intend to abort the wait; the boot continues.

Linux / macOS (same scripts):

.venv/bin/python scripts/service/kazma_guard.py --reload
.venv/bin/python scripts/service/kazma_guard.py --status

First-time install of the watcher: python scripts/service/kazma_guard.py --install (delegates to install_service.py).


🐝 Swarm Orchestration in 30 Seconds

# 1. Dispatch a dynamic specialist task (Autoscaler selects best model)
kazma swarm dispatch --workers auto "Analyze the codebase security posture and produce a report"

# 2. Run a structured multi-stage pipeline
kazma swarm pipeline --workers researcher,coder,validator "Implement an OAuth2 device code provider"

# 3. Parallel consensus voting (Fan-Out)
kazma swarm fanout --workers a,b,c --aggregation vote "Select optimal database schema indexing"

# 4. View live telemetry and history
kazma swarm history
kazma swarm metrics

Prefer a UI? The web Swarm Panel (/swarm) shows live dispatch telemetry, worker status, and task history — and the TUI has a Swarm tab. Enable the engine in kazma.yaml:

swarm:
  enabled: true
  workers: []   # the autoscaler spawns specialists on demand

Multi-replica honesty: Jobs can multi-replica (document jobs, via Postgres SKIP LOCKED claims); document metadata and the SQLite stores remain single-replica — see docs/docs/guide/document-intelligence.md.


📦 Monorepo Package Structure

Package Path Description
kazma-core kazma-core/ Agent runner, LLM provider matrix, SwarmEngine, V2 Cognitive Memory, IDE backend, Safety & Document services
kazma-gateway kazma-gateway/ Multi-platform adapters (Telegram, Discord, Slack), slash commands, in-flight task steering (/steer)
kazma-ui kazma-ui/ FastAPI web application, SSE streaming chat, Observability Dashboard, Web IDE, and Memory console
kazma-tui kazma-tui/ Textual-based rich terminal dashboard, interactive IDE, and Documents manager
kazma-skills kazma-skills/ Native certified skills (Document Platform, Encrypted Vault, Deep Research, Crawler, Database)
kazma-cli kazma-cli/ Unified command-line interface (kazma ask, kazma acp, kazma swarm, kazma migrate, kazma serve)

🧪 Testing & Verification

Kazma maintains rigorous test coverage with 5,600+ automated test cases across unit, integration, swarm reliability, and security layers:

# Run complete test suite
pytest

# Code quality and type validation
ruff check kazma-core/
mypy kazma-core/

📚 Documentation Reference

Guide Description
System Architecture In-depth breakdown of supervisor graph, ReAct loops, and engine internals
Monorepo System Map Comprehensive structural map of all monorepo modules and dependencies
V2 Cognitive Memory Bi-temporal belief stores, PPR graphs, and automated reconsolidation
Swarm Orchestration Dispatch patterns, reliability breakers, autoscaling, and worker lifecycle
Document Intelligence Secure ingestion pipelines, quarantined OCR, and redaction operations
Security & HITL Triple-wired approval architecture, prompt fencing, and vault encryption
Threat Model What each boundary stops and — stated plainly — what it does not
Prompt Injection: the numbers The measurements behind the fencing claim, including a public benchmark and the payloads that still land
Known Gaps Open weaknesses, dated, so they do not depend on someone remembering
Configuration Reference Detailed kazma.yaml, environment variables, and provider settings

📬 Community & Contact


📜 License

This project is licensed under the MIT License — see the LICENSE file for details.

About

Autonomous AI agent framework — LangGraph brain, swarm orchestration, Arabic-first, with human-in-the-loop safety.

Topics

Resources

Contributing

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages