Skip to content

Latest commit

 

History

299 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Ambit

What you, your agents, and your machines can jointly do — and where your own time is going.

Ambit reads the configs of Claude Code, Cursor, OpenCode, Windsurf, Gemini CLI, Claude Desktop and Codex CLI into one local graph, and answers what no single config file can: what works, what is configured but broken, and what stops working if one MCP server, model or token goes away.

CI Release Node License: MIT

Try the live demo · Get started · Connect it to your agent · FAQ · Docs


Ambit showing one developer setup as a map: Shell Execution is selected and its dependents listed, switching it off turns what depends on it red and says nine working capabilities would stop, apart from those never set up and one already failing, a second view colors the tools that interrupt a person most often, and a proposed config change waits for approval

One setup, mapped. Pick a capability and Ambit shows what depends on it; switch it off and it counts the nine working things that would stop, apart from the ones never set up and the one already failing. Then which tools interrupt you most, and a change waiting on your approval.

brew install zz-plant/tap/ambit && ambit, or open the hosted demo and install nothing.


What Ambit is

Ambit (from Latin ambitus: circuit, perimeter, sphere of action) is the boundary of what you, your agents, and your machines can jointly and reliably do.

If you use AI agents, your setup is spread across LLM providers, MCP servers, local CLI tools, skill directories, credentials, and more than one machine. Every piece has its own config file.

What they add up to — what your human-plus-agent system can actually do — is written down nowhere.

Ambit reads those configs and builds one map out of them. Every tool, model, skill, and credential becomes a point on it; everything one of them needs in order to work becomes a line to another. That map answers questions no single config file can:

  1. What works right now? Distinguishing what is merely configured from what is proven to work (installed ≠ working ≠ authorized).
  2. What breaks downstream if a model, tool, or credential goes away (transitive blast radius and shared single points of failure).
  3. What compound abilities emerge when two independent tools are combined.
  4. What is worth setting up next, priced by the human attention it would save.

You ask from the terminal. Your agents ask over MCP, mid-session, before they run into the limit — Ambit is itself an MCP server, so the thing describing your MCP servers speaks the same protocol they do. (A meta-MCP server, if you want the term to search for.)

Why Ambit: four unfair distinctions

Most agent tooling catalogs tools by name or embeds them in vector stores. Ambit treats capability as an assured, governed boundary:

  1. Transitive blast radius over shared credentials. Registries inspect tools in isolation. Ambit models actual dependency graphs. If a shared token or local daemon drops, Ambit traces every dependent capability, exposing when multiple "redundant" providers secretly collapse under a single point of failure.
  2. Assurance over declaration (installed ≠ working ≠ authorized). A configured tool is not a working tool. Ambit runs declared checks (ambit verify), tracks failure signals, demotes broken capabilities without asking, and promotes authority grants only when backed by clean execution evidence.
  3. Pre-flight briefing over context loops. Long-running agents waste context windows discovering broken tools through trial, error, and permission crashes. Ambit's meta-MCP briefing (ambit://briefing) gives the agent its proven action space before the first tool call.
  4. Human attention accounting. The work ledger links tool runs, friction, and interruptions to dollars and human hours, ranking what to acquire or fix next based on return on human attention.

The words Ambit uses

Four of them carry most of the meaning, in the terminal and on the map alike.

  • Capability — one thing your setup can do. Every MCP server, agent, skill, provider, model, and command in your config becomes one, as does every node of the curated tree.
  • Era — how far up the tree a capability sits. Later eras depend on earlier ones. Eras describe ordering, not importance.
  • Reached, next step, blocked — reached means something in your config provides it. A next step is one whose prerequisites are met with nothing detected: this is the frontier, and ambit goal lists it. Blocked means a prerequisite is missing, which is usually the most informative of the three.
  • Required vs optional prerequisite — a required prerequisite gates the capability; an optional one strengthens it without gating. Only required ones block a node. Both are drawn, optional ones fainter. The data model and the CLI call these hard and soft.
The Ambit capability map: tools and skills drawn as connected nodes in themed eras
Filled nodes are reached · Bright rings are a next step, with their setup time · Dashed nodes are blocked, with a prerequisite missing · The line on top is what the map found

In practice

The combo you already almost have

You run local Postgres and Ollama, but your agent cannot do private semantic code search over your repositories.

ambit graph combos reports the gap as one step — CREATE EXTENSION vector; — and ambit goal retrieval --simulate shows what that five-minute change reaches, with no cloud API in the path.

An agent diagnosing itself

An agent in Claude Code is asked to deploy to staging. Left alone it runs kubectl, collects unauthorized errors, retries, and leaves local state worse than it found it.

Calling ambit_authority first returns authority: confirm and missing: staging-kubeconfig. The agent stops cleanly and asks for an approval it can name.

Rotating a shared token

You are about to revoke a personal access token. Without a model of what depends on it, two background MCP tools and a scheduled sync agent fail silently some hours later.

ambit impact credential:github/user-token names the providers and capabilities standing on that one credential, which is the argument for provisioning granular tokens first.

The credentials block that declares the sharing is in the deep dive. Until you write one, ambit credentials reports that none are declared.

Stopping a context-window thrash loop

An agent starts a task requiring browser automation. Playwright is configured in opencode.json, but a recent system update broke the local Chromium binary.

Without Ambit, the agent issues tool calls, gets opaque exit codes, attempts four workarounds, burns 35,000 tokens of context, and fails. With ambit://briefing, the failing declared check demotes browser automation before the session begins. The agent immediately routes to a static HTTP fetcher or asks for the specific binary fix upfront.


Where this sits in the stack

Ambit sits above the protocol layer and below workflow orchestration. It neither routes calls nor runs them.

System Finds a tool Knows prerequisite order Tells working from configured Prices human attention Gates what an agent may do
Vector tool-RAG by similarity – – – –
Workflow state machines (LangGraph) – within one task – – within one task
Package managers (Nix, Homebrew) – for binaries – – –
Flat MCP catalogs (Smithery, registries) by name – – – –
Typed decision models (Jev) – – – – by a probability, which text in the state can move
Ambit by what it needs across the whole host ✓ declared checks ✓ work ledger ✓ authority contracts, signed approvals

Semantic search finds tools that sound relevant and cannot tell a working one from a broken one. A workflow graph models control flow within one task. A package manager installs binaries. Flat catalogs index servers without tracking whether their prerequisites exist on your machine. Ambit models what those tools add up to on this host, what it costs a person to keep them working, and what an agent may do with them.

A typed decision model such as TypeSafe's Jev answers whether a tool call looks safe with a calibrated probability, cheaply enough to ask on every call. Injected text can move that probability, so Ambit maps Jev as a capability and never lets it decide what an agent may do. The FAQ says how the two fit together.


Get started

Way in What it gives you
In the browser Open the hosted demo and read an example setup, or drop your own opencode.json on the page and it is mapped in the tab, uploading nothing.
On your machine brew install zz-plant/tap/ambit && ambit reads your real agent config and prints where you stand.
With the map git clone https://github.com/zz-plant/ambit.git && cd ambit && ./bootstrap.sh web builds the graph from your own configs and serves the canvas the figures on this page show.
In a cloud IDE Open in GitHub Codespaces A full checkout with the map running, in a browser tab.
From your agent Register Ambit over MCP and the agent can ask what it is able to do before it tries. Connect it to your agent has the snippet for each client.

Homebrew installs the CLI, the engine, and the MCP server from the tagged release, on macOS or Linux. A checkout adds the map: ./bootstrap.sh discovers OpenCode, Claude Code, Cursor, Windsurf, Gemini CLI, Claude Desktop, Codex CLI, and the skill directories ~/.agents/skills and ~/.opencode/skills, builds a local SQLite graph, and finishes by printing ambit status, which Ask from the terminal shows against a fixture graph. It links ambit into ~/.local/bin when that is on your PATH and prints the ln -s line otherwise; --dry-run shows what it would do. Codespaces runs the same checkout in a container, so the graph is the container's and nothing touches your machine.

Note

The npm package is built and ready but not yet published, so there is no npx path yet.

The My Setup view: MCP servers, agents, and models read from local config, one row each, with what the engine has proved about them and the capabilities each provides
My Setup is the list the map is drawn from: every server, agent, and model found on the machine, whether it is on, what its check said, and which capabilities on the map it provides

Ask from the terminal

Command What it answers
ambit status Environment health — what is reached, what is failing its declared check, what has a single provider, plus pending approvals
ambit goal <name> The path to unlock a capability, in order, with setup estimates
ambit impact <id> Blast radius: what breaks if this tool, model, or credential goes down
ambit graph combos Compound capabilities, including the ones you are one prerequisite away from
ambit authority Per-action permissions: what runs unattended, what needs confirmation
ambit verify [id] Run a capability's declared check and record whether it actually works
ambit history [since] How the frontier moved, separating what you acquired from what emerged
ambit share A self-contained HTML snapshot of the map, written locally and safe to post

ambit help covers a first session; ambit help --all is the full surface, grouped by what you are trying to do, and ambit help <term> explains one concept.

Everything above answers on a graph Ambit builds by itself. A second group (attention, work, usage, opportunities, roi, audit) prices the human cost of running the stack, and reads from a work ledger that starts empty. Those commands tell you what they need instead of returning a number, and they become useful after a few weeks of recorded runs, not on install. The FAQ says how to start recording them.

The three console blocks below are captured from a run against a fixture graph by npm run docs:examples, and CI fails if they drift from what the commands actually print.

ambit status — where the environment stands

$ ambit status

    summary: 39/59 capabilities reached · 9 with a single provider
    reached: 39
    total: 59
    verified: 0
    failing: 0
    actions: 18/28 reached
    evidence:
        proven: 0
        unproven: 15
        failing: 0
        last check: never
        provable now: Automated Tests, Browser Automation, Code Intelligence, Continuous Delivery, Data Access, File Editing, Local Runtime, Shell Execution
        note: configured is not working — ambit verify would turn 11 of the unproven into evidence
    domains:
    …

ambit goal — what it would take to reach something

$ ambit goal local-embeddings

    goal: Local Embeddings
    exact: true
    reachable: true
    steps: 2
    estimated setup: 25m
    order:
      Embeddings
        id: combo:embeddings
        setup seconds: 600
        options:
          nomic-embed via local runtime
            setup seconds: 600
            recurring cost: none
            privacy: local
    …

ambit impact — what breaks if one MCP server goes away

$ ambit impact mcp:playwright

    capability: playwright
    decayed:
      Tool Protocol
        becomes unavailable: false
        also provided by: 5
      Browser Automation
        becomes unavailable: true
      Automated Tests
        becomes unavailable: true
    combos at risk:
      Tool Protocol
        severity: redundant
        also provided by: 5
      Browser Automation
    …

ambit share — what may leave the graph, and what may not

ambit share builds its HTML from an allow-list — name, kind, category, domain, era, state, lifecycle, edges. Commands, URLs, paths, descriptions, and economics cannot enter the file, people render as "a person", and --redact replaces every non-curated name with its category. Nothing is uploaded; writing the file locally is the whole command.


Connect it to your agent

Registering Ambit as an MCP server lets an agent inspect its own toolchain and plan around what is missing.

Sixty tools, each advertised once, each answering with MCP structuredContent alongside the text block so an agent reads a field and never parses a string. The deep dive names all sixty, grouped. (The tt_ prefix from before the rename is still accepted, just no longer listed.)

Claude Code

claude mcp add ambit -- ambit mcp

Ambit also publishes a resource, ambit://briefing, which a client reads on connect: what is reached and proven, what is configured but failing, what is waiting on you, what blocked work in the last week, and what is worth reaching next. It is capped at about 1,200 tokens. To put the same thing at the top of every session yourself, add a hook to ~/.claude/settings.json:

{
  "hooks": {
    "SessionStart": [{ "hooks": [{ "type": "command", "command": "ambit briefing" }] }]
  }
}

OpenCode (~/.config/opencode/opencode.json)

{
  "mcp": {
    "ambit": {
      "type": "local",
      "command": ["ambit", "mcp"],
      "enabled": true
    }
  }
}

Both assume ambit is on your PATH. Homebrew puts it there; from a checkout, bootstrap.sh links it into ~/.local/bin and prints the ln -s line if that directory is not on your PATH. Failing both, use the absolute path to cli.js.

An agent can read the map, query what a goal is missing, and propose a configuration change. Applying one always requires your approval.

The one line for your agent's instructions

The habit worth teaching is a single question before an unfamiliar tool, because the alternative is the retry loop that spends your attention:

Before running a tool you have not used this session, call ambit_can with the capability. On yes, act. On ask, put it to the person. On no, it has already recorded the deficit, so do not retry it under another name.

It answers from the graph without probing anything, and a refusal files itself as a deficit, which is what makes the third occurrence show up as infrastructure that should exist instead of a wall to work around again.

Here is the whole loop from a live run. node --experimental-strip-types scripts/demo-agent-loop.ts re-records it, and every frame is real engine output — a failing loop fails the recording instead of rendering a fiction.

An agent hits a missing capability, asks Ambit why over MCP, and drafts a proposal; a person approves and applies it; the frontier moves and Local Embeddings unlocks through composition
Agent: hits a block, records the deficit, asks goal, drafts a proposal · Human: approve, apply · One config patch, four capabilities.

What the exchange looks like

sequenceDiagram
    autonumber
    actor Developer
    participant Agent as AI Agent (Claude Code / OpenCode)
    participant Ambit as Ambit Engine (MCP)
    participant Host as Local Host

    Developer->>Agent: "Deploy the billing hotfix to staging"
    Agent->>Ambit: ambit_authority("act:continuous-delivery/deploy_staging")
    Note over Ambit,Agent: Checks prerequisites and authority contracts
    Ambit-->>Agent: { status: "blocked", authority: "confirm", missing: ["credential:k8s-kubeconfig"] }
    Agent->>Ambit: ambit_propose("deploy-staging")
    Ambit-->>Agent: { proposal_id: "prop-staging-42", applicable: true }
    Agent->>Developer: "I need the staging kubeconfig and your confirmation: ambit approve prop-staging-42"
    Developer->>Ambit: ambit approve prop-staging-42 (mints a signed artifact)
    Developer->>Host: ambit apply prop-staging-42 (applies and verifies)
Loading

The map

The web UI (./bootstrap.sh web) is three views over the same graph the CLI reads. The map is the curated tree with your position on it. My Setup is one row per entry your configs declare, with what the engine has proved about it and the nodes on the map it provides; its Briefing tab is the prose an agent is given at connect, so what the agent believes about the machine is inspectable. Its Not on the map tab lists what the agents used in the last 30 days that no node on the map accounts for, with an overlay to paste into .ambit/techtree.json that would put it there. Time & cost is the ledger and the governance half: what may act without asking, which grants have earned a threshold nobody set, what to reach next and why, and how the frontier moved this week. Search (/) finds anything by name and opens it where it lives.

Over the map, a line gives the range: how many capabilities are verified, how that moved this week, and the one piece of the setup whose loss would stop the most, with a button that simulates it. Under it, one line says what the map found before you read a node: a capability that is configured and failing its check, or else the next step that reaches the most, with a button that previews it. Select a node and its edges are drawn apart: what it needs in violet, what it enables in blue, one hop each way. The panel states the answer before the simulation that draws it: what would stop and what would only lose a provider if the node went down, or what stands between it and being reached and how long that would take. A small person mark on a node means it needs someone: a person approves or supplies it; a device mark means it runs on one of your machines. The detail panel says who, under Joint capability. The header counts the map's nodes, leading with the reached ones that have a passing check (verified) apart from those with none or a failing one (unproven), and each count highlights its nodes, the way the legend keys do.

The Docs button defines every term on the canvas; the four above cover most of it.

Three lenses on the canvas

The switch sits over the map, top right. Press 1, 2 or 3 to change it from the keyboard.

Lens What it renders Use it for
Standard Era columns with reached, next-step and blocked nodes. Reading overall progression and what is nearby.
Attention Nodes shaded by how often a person had to step in, offered once the ledger has recorded any. Finding which tools keep interrupting you.
Authority Each reached node by what it may do: act without asking, ask first, forbidden, or no grant yet, which the gate refuses until someone grants one. Seeing where being able to do something is not the same as being allowed to.

Simulation

Select a node to open the inspector, then simulate against it. Neither mode writes anything.

  • Simulate an outage dims the canvas and draws the cascade: red for what stops, amber for what keeps another provider and only loses one, with the count of each.
  • Simulate unlocking acquires a locked primitive hypothetically and lights up, in green, everything that becomes reachable because of it.
  • Show the gap draws what a blocked node is waiting on, every hop up, priced in setup time.

Approving proposals

When an agent proposes an environment change over MCP, the Proposals panel shows what it would save, what it costs, whether every step can be undone, what it unlocks, and how you have decided on things like it before, then mints a signed approval receipt in one click, or records a no with the reason, which is what the next draft learns from. The same things happen from the terminal with ambit approve <id> <who> and ambit reject <id> <who> "why". When you are away from the machine, ambit dispatch <id> pushes the draft to a Slack, Discord or Telegram webhook, or an ntfy topic, with the commands that decide it; the decision itself still happens here, on a machine that holds the approval key.


How it works

Discovery reads your host configs into an embedded SQLite graph. Three surfaces read that graph back out — the terminal CLI, the MCP server, and the web canvas. Discovery, verification, and the work ledger write to the graph. Your agent configuration changes only through a proposal you approve, or through the map's editor for entries that already exist, which cannot create one.

Each client is read from its own standard config path, and every server stays attributed to the client that listed it. When two clients name the same server, that is one capability with two providers, not two capabilities, which is what stops Ambit from counting a single binary twice and calling the result redundancy.

Seven eras, and what follows from them

Discovered capabilities are placed into a curated tree that runs from Foundation and Model Access through Tool Use, Memory, Autonomy, and Assurance to Sovereignty. Because each capability records what it needs, Ambit works out what you can reach without taking a config file's word for it.

Two things follow from that:

  • Combos. Higher-order abilities appear from tools that were configured separately — a vector store plus local embeddings becomes semantic retrieval, which neither config mentions.
  • Near misses. When you are one or two prerequisites from a capability that unlocks several others, that gap is worth naming. ambit graph combos lists them.

Configured is not working

Ambit keeps two properties apart, and the distinction is load-bearing:

  • state is structural — is this thing configured, and what does it depend on. This is what the frontier ledger records.
  • lifecycle is health — did its declared verification command actually pass. A capability can be fully configured and still degraded or broken.

Every availability decision gates on lifecycle, not state. A broken capability is excluded from plans, simulations, goals, authority checks, and opportunity ranking, because a plan routed through a tool that does not run is worse than no plan.

ambit status reports proven, unproven, and failing counts. The map badges each reached node: ✓ for a passing check, ! for a failing one, nothing for configured-but-never-verified.

Fragility is computed, not guessed

  • Single points of failure — capabilities with exactly one provider.
  • Bottlenecks — nodes ranked by how much sits downstream of them. The map marks the same idea on each node, as a keystone.
  • Shared credentials — providers presenting the same credential fail together, so three providers behind one token is not redundancy. This one is declared, never inferred: name the sharers in a credentials block and ambit impact credential:... will show what revoking it would end.

The control plane

Host-level agent tooling is a real attack surface, so execution goes through an interceptor and never straight to the shell. The decision is real: the DAG check, the authority evaluation, the approval artifact and the audit trail all run against your actual graph. What sits on the other side of the gate is a fixture: simulatedAdapter in src/control_plane/proxy.ts keeps its state in a JSON file, and Ambit ships no deployment integration. A real one implements the three-method EnvironmentAdapter in that file, and nothing above the gate changes.

  • Interception. Before a tool call reaches your machine, the proxy in src/control_plane/proxy.ts checks three things: are this capability's prerequisites in place, is it actually working, and is the caller allowed to do this. A call that fails any of them is refused — AMBIT_BLOCKED_UNAUTHORIZED, exit code 2 — and nothing on the machine has changed.
  • Human-in-the-loop remediation. A blocked execution drafts a structured proposal and an HMAC challenge. ambit approve <proposal-id> <person> mints a signed artifact that the executor verifies before any state changes. An artifact stops being valid if the proposal changed after approval, or if it has expired.
  • Tracing. Spans and structured events record DAG evaluations, missing authorizations, challenges, and verification receipts.

A worked example — an autonomous deploy agent blocked mid-flight, then remediated — is written up in docs/incidents/INCIDENT_TRACE_001.md.

npm test                                                  # the suite behind the walkthrough
npm run demo:incident                                     # the 90-second terminal walkthrough
asciinema play docs/incidents/demo_intervention_trace.cast # replay the recording

Position in the revisable-delegation loop

Ambit is one of five systems that each hold a step of the loop an institution runs when it delegates consequential work to machines: believe, know what can be done, decide what authority is justified, act, detect mismatch, revise. Ambit holds capability and authorization. A grant holds only while what it rests on does, and every narrowing is written as an append-only, hash-chained stream of STD-07 Revisable Delegation Records that another Ambit environment can read as evidence and never as an instruction. The deep dive has the record kinds, the objection path, and the limits; the siblings are Whether (act), Refract (discrepancy), NextConsensus (belief), and Ethotechnics (the record shape).


Security invariants

Ambit reads developer toolchains and writes to agent configs, so four properties are fixed and cannot be relaxed. SECURITY.md states each in full, with what is in scope and what is not; AGENTS.md says where each is enforced.

  1. Loopback only. The API server binds 127.0.0.1. No LAN, no tunnel.
  2. Origin allowlist. A request with a non-local Origin is rejected with 403 before routing, because a simple request skips preflight and response headers alone would not stop it.
  3. No entry creation over HTTP. The HTTP layer edits entries that already exist and nothing else. An MCP entry carries a command the runtime later executes, so creating one over HTTP would be remote code execution; adding a server returns a snippet for you to paste.
  4. No egress you did not type. The graph is an embedded SQLite database on your machine, and there is no telemetry. Five commands open a socket at all (notify, notify-approvals, dispatch, incidents, and goal --judge). The first four each need a target you name, and the last refuses any host but this machine. The FAQ lists exactly what each one sends.

Documentation

  • FAQ — what needs installing, what leaves the machine, why the attention commands are empty on day one.
  • Deep dive — the reference for the model under everything above.
  • Security · Agent invariants · Support · Contributing — the invariants above in full, where each is enforced, where each kind of question goes, and how to send a change.
  • Everything else — the argument, the theory, the design notes, the changelog, the incident traces.
  • llms.txt — the project in one page, for an agent that is deciding whether to recommend it.

Contributing

New capability models, runtime adapters, visualization work, and edge-case reports are all welcome. CONTRIBUTING.md gets you from a clone to a passing pull request and lists every check CI runs, documentation included: the console blocks above are checked against a fresh run, and the prose against a corpus ceiling (AGENTS.md rule 17).


Support the project

Ambit is a personal project. If it answered a question your config files could not, a star is how the next person with the same stack finds it, and the release notes explain each change.

  • Post an ambit share --redact snapshot of your own map. The file names nothing on your machine, and every real graph is an argument the demo cannot make.
  • To hear when a new runtime reader or capability lands, choose Watch → Custom → Releases. That sends the release notes and nothing else.
  • Report the runtime it does not read yet, or the capability it models wrong. Both are issue templates.
  • Citing it in writing? CITATION.cff is what GitHub's Cite this repository button reads.

License

MIT © Kanav Jain

Releases

Packages

Contributors

Languages