Skip to content

Latest commit

 

History

1,741 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Inference Gateway CLI

Go Version License Build Status Release

An agentic command-line assistant that writes code, understands project context, and uses tools to perform real tasks.

📖 Documentation · 🚀 Quick Start · 🧭 Features · 💬 Discussions


infer chat - the interactive TUI with model selection, streaming responses, and a live status bar

The infer chat TUI - watch the animated demo

Early development stage: breaking changes are expected until the project reaches a stable version. Always pin a specific version tag when downloading binaries or using install scripts.

Installation

The recommended install is npm/npx - the matching native binary is fetched and cached on first use:

npx @inference-gateway/cli@latest chat

See Installation for production and CI installs, and Binary Verification to verify a download.

Quick Start

  1. Initialize the userspace baseline:
infer init

This creates the shared ~/.infer/ configuration directory. All state (conversations, logs, history, artifacts) lives there, and a project .infer/ is an optional override layer you create with infer config set --project.

  1. Set up your environment (create a .env file):
ANTHROPIC_API_KEY=your_key_here
OPENAI_API_KEY=your_key_here
DEEPSEEK_API_KEY=your_key_here

Provider keys resolve in this order, first hit wins: the system environment, the project .env, then the userspace ~/.infer/auth.yaml fallback - see Provider API Keys.

  1. Start chatting:
infer chat

Features

One line per feature, each linking to its full guide:

Core

  • Automatic gateway management - downloads and runs the gateway binary, no Docker required - Installation
  • Interactive chat and headless agent - model selection, streaming, session resumption, inline history auto-completion - Commands Reference
  • Agent modes - Standard, Plan, Auto-Accept and Auto+Judge, toggled with Shift+Tab - Plan Mode · Judge Mode
  • Tool execution - every built-in tool the LLM can call, with parameters and approval defaults - Tools Reference
  • Tool approval - the gate in front of sensitive tools - Tool Approval
  • Custom tools - add tools in any language with one YAML manifest - Custom Tools
  • MCP servers - Model Context Protocol integration - MCP Integration
  • Subagents - fan out parallel infer headless runs - Subagents
  • GitHub issue references - type # in chat to expand an issue inline (needs gh and a GitHub remote) - Shortcuts Guide
  • Cost tracking - real-time per-model cost breakdown - Cost Tracking
  • Conversation history - multiple storage backends - Conversation Storage
  • Conversation versioning - navigate back in time to a previous point - Conversation Versioning
  • Conversation titles - AI-generated session titles - Conversation Title Generation
  • Configuration - two-layer YAML plus INFER_* environment overrides - Configuration Reference
  • Directory structure - what the CLI writes where, userspace and project side by side - Directory Structure
  • Configurable keybindings and themes - Configuration Reference
  • Persistent memory - cross-session Markdown facts with an injected index - Persistent Memory
  • Reminders and command hooks - extension points at agent-loop hook points - Reminders & Command Hooks
  • Telemetry - OpenTelemetry traces and metrics - Telemetry
  • AG-UI output - protocol event stream for the headless agent - AG-UI Output

Remote and automation

  • Web terminal - browser-based, tabbed sessions - Web Terminal
  • Remote messaging channels - drive the agent from Telegram - Channels
  • Scheduled tasks - cron prompts that deliver back through the channel - Scheduling
  • Heartbeat - periodic wake-up to check pending work - Heartbeat
  • A2A agents - delegate to Agent-to-Agent servers - A2A Agents Configuration · A2A Connections
  • Task management - the A2A task interface - Tasks Management
  • Daemon - the hub the desktop app, the extension and Telegram reach the agent through - infer daemon
  • Daemon binding - the WebSocket wire contract, including browser use through the extension - Daemon Binding Protocol
  • Explorer and diffs - in-terminal file tree, fuzzy finder and diff viewer - Explorer

Skills and plugins

  • Agent Skills - drop-in SKILL.md instruction folders - Agent Skills
  • Plugins - Claude Code-format skills plus an always-on ruleset - Plugins
  • Extensible shortcuts - custom /-commands with AI-powered snippets - Shortcuts Guide

Media and computer use

  • Computer Use - drive the desktop: mouse, keyboard, screenshots - Computer Use
  • Frame sources and vision annotation - let text-only models "see" - Vision
  • Image generation, edit and variation - written to the session artifacts dir - Tools Reference
  • Speech-to-text - dictate with Whisper, locally and offline - Speech-to-Text
  • Text-to-speech - local WAV synthesis with zero-shot voice cloning - Text-to-Speech
  • Text-to-music and text-to-video - generate audio and video through the gateway - Text-to-Music · Text-to-Video

Examples

Each directory under examples/ is a self-contained, runnable setup. See Examples for the full index and common workflows, or jump straight to basic for a minimal gateway plus CLI.

Contributing

Development is documented in CONTRIBUTING.md. Coding agents working in this repository should read AGENTS.md first; packages with their own agent-specific rules ship an AGENTS.md next to the code.

License

Apache 2.0 License - see the LICENSE file for details.

About

A Git-first CLI coding agent that turns ideas, issues, and tasks into real code changes. It can run remotely from your mobile phone or automate your computer.

Topics

Resources

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages