An agentic command-line assistant that writes code, understands project context, and uses tools to perform real tasks.
📖 Documentation · 🚀 Quick Start · 🧭 Features · 💬 Discussions
The infer chat TUI - watch the animated demo
Early development stage: breaking changes are expected until the project reaches a stable version. Always pin a specific version tag when downloading binaries or using install scripts.
The recommended install is npm/npx - the matching native binary is fetched and cached on first use:
npx @inference-gateway/cli@latest chatSee Installation for production and CI installs, and Binary Verification to verify a download.
- Initialize the userspace baseline:
infer initThis creates the shared ~/.infer/ configuration directory. All state (conversations, logs, history,
artifacts) lives there, and a project .infer/ is an optional override layer you create with
infer config set --project.
- Set up your environment (create a
.envfile):
ANTHROPIC_API_KEY=your_key_here
OPENAI_API_KEY=your_key_here
DEEPSEEK_API_KEY=your_key_hereProvider keys resolve in this order, first hit wins: the system environment, the project .env, then the userspace
~/.infer/auth.yaml fallback - see Provider API Keys.
- Start chatting:
infer chatOne line per feature, each linking to its full guide:
- Automatic gateway management - downloads and runs the gateway binary, no Docker required - Installation
- Interactive chat and headless agent - model selection, streaming, session resumption, inline history auto-completion - Commands Reference
- Agent modes - Standard, Plan, Auto-Accept and Auto+Judge, toggled with Shift+Tab - Plan Mode · Judge Mode
- Tool execution - every built-in tool the LLM can call, with parameters and approval defaults - Tools Reference
- Tool approval - the gate in front of sensitive tools - Tool Approval
- Custom tools - add tools in any language with one YAML manifest - Custom Tools
- MCP servers - Model Context Protocol integration - MCP Integration
- Subagents - fan out parallel
infer headlessruns - Subagents - GitHub issue references - type
#in chat to expand an issue inline (needsghand a GitHub remote) - Shortcuts Guide - Cost tracking - real-time per-model cost breakdown - Cost Tracking
- Conversation history - multiple storage backends - Conversation Storage
- Conversation versioning - navigate back in time to a previous point - Conversation Versioning
- Conversation titles - AI-generated session titles - Conversation Title Generation
- Configuration - two-layer YAML plus
INFER_*environment overrides - Configuration Reference - Directory structure - what the CLI writes where, userspace and project side by side - Directory Structure
- Configurable keybindings and themes - Configuration Reference
- Persistent memory - cross-session Markdown facts with an injected index - Persistent Memory
- Reminders and command hooks - extension points at agent-loop hook points - Reminders & Command Hooks
- Telemetry - OpenTelemetry traces and metrics - Telemetry
- AG-UI output - protocol event stream for the headless agent - AG-UI Output
- Web terminal - browser-based, tabbed sessions - Web Terminal
- Remote messaging channels - drive the agent from Telegram - Channels
- Scheduled tasks - cron prompts that deliver back through the channel - Scheduling
- Heartbeat - periodic wake-up to check pending work - Heartbeat
- A2A agents - delegate to Agent-to-Agent servers - A2A Agents Configuration · A2A Connections
- Task management - the A2A task interface - Tasks Management
- Daemon - the hub the desktop app, the extension and Telegram reach the agent through - infer daemon
- Daemon binding - the WebSocket wire contract, including browser use through the extension - Daemon Binding Protocol
- Explorer and diffs - in-terminal file tree, fuzzy finder and diff viewer - Explorer
- Agent Skills - drop-in
SKILL.mdinstruction folders - Agent Skills - Plugins - Claude Code-format skills plus an always-on ruleset - Plugins
- Extensible shortcuts - custom
/-commands with AI-powered snippets - Shortcuts Guide
- Computer Use - drive the desktop: mouse, keyboard, screenshots - Computer Use
- Frame sources and vision annotation - let text-only models "see" - Vision
- Image generation, edit and variation - written to the session artifacts dir - Tools Reference
- Speech-to-text - dictate with Whisper, locally and offline - Speech-to-Text
- Text-to-speech - local WAV synthesis with zero-shot voice cloning - Text-to-Speech
- Text-to-music and text-to-video - generate audio and video through the gateway - Text-to-Music · Text-to-Video
Each directory under examples/ is a self-contained, runnable setup. See
Examples for the full index and common workflows, or jump straight to
basic for a minimal gateway plus CLI.
Development is documented in CONTRIBUTING.md. Coding agents working in this repository
should read AGENTS.md first; packages with their own agent-specific rules ship an AGENTS.md
next to the code.
Apache 2.0 License - see the LICENSE file for details.