What - every built-in tool an LLM can call through the Inference Gateway CLI, with its parameters and approval default. Why - a tool's scope, side effects, and gating differ; picking the wrong one either fails the task or runs something you did not sanction. How - find the tool by category in the overview below, then read its section for the exact contract.
This document provides comprehensive documentation for all tools available to LLMs when tool execution is enabled in the Inference Gateway CLI.
- Tool Overview
- File System Tools
- Command Execution
- Web Tools
- Browser Tools
- Workflow Tools
- Agent-to-Agent Communication
- Custom Tools
Tools are grouped by category. Many are gated behind a config flag (noted per group); the
always-available set is registered for every session. There is no built-in GitHub tool - use the
gh CLI through Bash (or the built-in /scm shortcuts) for GitHub operations.
Core file & search (always available):
| Tool | Purpose | Approval |
|---|---|---|
| Read | Read file contents with line ranges | No |
| Write | Write content to files | Yes |
| Edit | Exact string replacements in files | Yes |
| MultiEdit | Multiple atomic edits to a single file | Yes |
| Delete | Delete files and directories | Yes |
| Grep | Search files with regex (ripgrep/Go) | No |
| Tree | Display directory structure | No |
Shell (Bash is always available; the background-shell trio needs tools.bash.background_shells.enabled):
| Tool | Purpose | Approval |
|---|---|---|
| Bash | Execute shell commands (per-mode allow-list) | Optional |
| BashOutput | Read new output from a running background shell | Yes |
| KillShell | Terminate a background shell | Yes |
| ListShells | List background shells and their state | Yes |
| Wait | Block until a condition is met (shells exit, file event, or check command succeeds) - no LLM round-trips wasted | No |
Task & planning (AskUserQuestion needs tools.ask_user_question.enabled):
| Tool | Purpose | Approval |
|---|---|---|
| TodoWrite | Create and manage task lists | No |
| RequestPlanApproval | Submit a plan for approval and persist it (plan mode) | No |
| AskUserQuestion | Ask the user multiple-choice questions | No |
| RequestApproval | Ask the user to override a judge rejection once (judge mode) | No |
Web (WebSearch/WebFetch need their respective config flag):
| Tool | Purpose | Approval |
|---|---|---|
| WebSearch | Search the web (DuckDuckGo/Google) | Yes |
| WebFetch | Fetch content from a URL | No |
WebFetch does not require approval by default; set tools.web_fetch.require_approval: true to require it.
Subagents (the Agent tool and its companions, enabled by default):
| Tool | Purpose | Approval |
|---|---|---|
| Agent | Spawn an infer headless subprocess to run work in parallel |
Yes |
| ListSubagents | List spawned subagents and their status | No |
| GetSubagentResult | Re-read a finished subagent's last message | No |
| ReadSubagentScreen | Capture an interactive subagent's terminal screen | No |
| SendSubagentInput | Send a subagent a follow-up message as its next turn (text for both modes, keys for interactive) | Yes |
| CloseSubagent | Stop a subagent, idle or running | Yes |
| ApproveSubagent | Relay an approval decision to a waiting subagent | Yes |
Agent is also offered in plan mode, where every subagent runs read-only
whatever type the call asks for (see Markdown Subagents).
The companion tools above follow Agent into plan mode and are offered in
exactly the same modes.
Computer Use (approval follows computer_use.approval, default never, and RecordStart also follows
computer_use.recording.require_approval, default on):
| Tool | Purpose | Approval | Enabled by |
|---|---|---|---|
| Computer | Read the accessibility tree, press labelled controls, capture screenshots, and control mouse/keyboard | No | computer_use.enabled |
| GetLatestFrame | Read the latest frame from a named source (screen, camera directory) | No | any frame source, e.g. vision.sources |
| RecordStart | Record the screen, a window, or a region to MP4 | Yes | computer_use.recording.enabled |
| RecordStop | Stop the recording and return the file path, duration, and size | No | computer_use.recording.enabled |
Browser (require browser_use.enabled; see Browser Tools):
| Tool | Purpose | Approval |
|---|---|---|
| BrowserNavigate | Open a URL | Yes |
| BrowserClick | Click an element by selector | Yes |
| BrowserType | Fill an input element | Yes |
| BrowserRead | Read the page's visible text and browser events | No |
| BrowserScreenshot | Capture the current page as an image | No |
| BrowserTabs | List the open tabs | No |
Media (each gated by its own flag; all output is written to disk, never played aloud):
| Tool | Purpose | Approval | Enabled by |
|---|---|---|---|
| ImageGeneration | Generate an image from a prompt into ~/.infer/projects/<project-slug>/artifacts/ |
No | tools.image_generation.enabled |
| ImageEdit | Edit an existing image and save the result | No | tools.image_edit.enabled |
| ImageVariation | Produce a variation of an existing image | No | tools.image_variation.enabled |
| ImageDecode | Look at a local image or URL, annotated for text-only models | No | always available |
| TextToSpeech | Synthesize speech to a WAV (local llama-tts or gateway Audio API), optionally cloning a voice | No | text_to_speech.enabled |
| TextToMusic | Generate music through the gateway | No | text_to_music.enabled |
| TextToSFX | Generate a sound effect through the gateway | No | text_to_sfx.enabled |
| TextToVideo | Generate a video, optionally from an avatar | No | text_to_video.enabled |
| CreateAvatar | Build an avatar portrait folder for TextToVideo | Yes | text_to_video.enabled and text_to_video.create_avatar |
Memory, scheduling & A2A (each gated by its own flag):
| Tool | Purpose | Approval | Enabled by |
|---|---|---|---|
| Memory | Persistent, cross-session fact storage | No | memory.enabled (default on) |
| Schedule | Cron-driven recurring/one-off tasks via the originating channel | Yes | tools.schedule.enabled |
| A2A_SubmitTask | Submit a task to an A2A agent | Yes | A2A enabled |
| A2A_QueryAgent | Query an A2A agent's capabilities | No | A2A enabled |
| A2A_QueryTask | Check an A2A task's status | No | A2A enabled |
Approval reflects the default policy. Each tool call takes the first rule that applies:
- The computer-use tools decide per call from
computer_use.approval.- Bash follows the per-mode bash allow-list.
tools.bash.require_approvalhas no effect.- A
require_approvalset in the tool's config section wins. Config defaults already exempttools.read,tools.grep,tools.treeandtools.web_fetch.- Otherwise the tool's manifest default applies where it sets one. For example, ApproveSubagent always asks.
- Otherwise the global
tools.safety.require_approvalapplies, defaulttrue.See Tool Approval for how gated calls are resolved.
MCP tools and custom tools are not listed here. MCP tools are discovered at runtime from your configured MCP servers and surface as
MCP_<server>_<tool>(see MCP Integration). Custom tools are loaded from their manifests (see Custom Tools).
Display directory structure in a tree format, similar to the Unix tree command. Provides a polyfill
implementation when the native tree command is unavailable.
Parameters:
path(optional): Directory path to display tree structure for (default: current directory)max_depth(optional): Maximum depth to traverse (default: 3, min: 1, max: 10)max_files(optional): Maximum number of files to display (default: 100, max: 1000)show_hidden(optional): Whether to show hidden files and directories (default: false)respect_gitignore(optional): Whether to exclude patterns from .gitignore (default: true)format(optional): Output format - "text", "json" or "compact" (default: "text")
Examples:
- Basic tree: Uses current directory with default settings
- Tree with depth limit:
max_depth: 2- Shows only 2 levels deep - Tree with hidden files:
show_hidden: true - Tree ignoring gitignore:
respect_gitignore: false- Shows all files including those in .gitignore - JSON output:
format: "json"- Returns structured data - Compact output:
format: "compact"- One directory per line, root-first, git-tracked files only
Features:
- Native Integration: Uses system
treecommand when available for optimal performance - Polyfill Implementation: Falls back to custom implementation when
treeis not installed - Pattern Exclusion: Supports glob patterns to exclude specific files and directories
- Depth Control: Limit traversal depth to prevent overwhelming output
- Hidden File Control: Toggle visibility of hidden files and directories
- Multiple Formats: Text output for readability, JSON for structured data, compact for a token-efficient listing
Security:
- Respects configured path exclusions for security
- Validates directory access permissions
- Limited by the same security restrictions as other file tools
Read file content from the filesystem with optional line range specification.
Configuration:
# ~/.infer/tools.yaml
tools:
read:
enabled: true
require_approval: false # Read operations don't require approval by defaultWrite content to files on the filesystem with security controls and directory creation support.
Existing files are always overwritten - there is no overwrite option; use the Edit tool to modify an existing file's content in place.
Parameters:
file_path(required): The path to the file to writecontent(required): The content to write to the file
Features:
- Directory Creation: Automatically creates parent directories when needed
- Always Overwrites: Existing files are always overwritten; use the Edit tool to modify an existing file's content in place
- Security Validation: Respects path exclusions and security restrictions
- Performance Optimized: Efficient file writing with proper error handling
Security:
- Approval Required: Write operations require approval by default (secure by default)
- Path Exclusions: Respects configured excluded paths (e.g.,
.infer/directory) - Pattern Matching: Supports glob patterns for path exclusions
- Validation: Validates file paths and content before writing
Examples:
- Create new file:
file_path: "output.txt",content: "Hello, World!" - Write to subdirectory:
file_path: "logs/app.log",content: "log entry"(parent directories are created automatically)
Configuration:
# ~/.infer/tools.yaml
tools:
write:
enabled: true
require_approval: true # Write operations require approval for securityPerform exact string replacements in files with security validation and preview support.
Parameters:
file_path(required): The path to the file to modifyold_string(required): The text to replace (must match exactly)new_string(required): The text to replace it withreplace_all(optional): Replace all occurrences of old_string (default: false)
Features:
- Indentation-Tolerant Matching: Exact first; on a miss, a unique leading-whitespace match is re-indented (
strict_whitespace: truedisables) - Preview Support: Shows diff preview before applying changes
- Atomic Operations: Either all changes succeed or none are applied
- Security Validation: Respects path exclusions and file permissions
Security:
- Read Tool Requirement: Requires Read tool to be used first on the file
- Stale-Read Detection: Rejects the edit if the file changed on disk since the agent last read it, and asks the model to re-read first
- Approval Required: Edit operations require approval by default
- Path Exclusions: Respects configured excluded paths
- Validation: Validates file paths and prevents editing protected files
Examples:
- Single replacement:
file_path: "config.txt", old_string: "port: 3000", new_string: "port: 8080" - Replace all occurrences:
file_path: "script.py", old_string: "print", new_string: "logging.info", replace_all: true
Configuration:
# ~/.infer/tools.yaml
tools:
edit:
enabled: true
require_approval: true # Edit operations require approval for security
strict_whitespace: false # When true, disable the indentation-tolerant fallback (byte-exact only)Make multiple edits to a single file in atomic operations. All edits succeed or none are applied.
Parameters:
file_path(required): The path to the file to modifyedits(required): Array of edit operations to perform sequentiallyold_string: The text to replace (must match exactly)new_string: The text to replace it withreplace_all(optional): Replace all occurrences (default: false)
Features:
- Atomic Operations: All edits succeed or none are applied
- Sequential Processing: Edits are applied in the order provided
- Indentation-Tolerant Matching: Each edit matches exactly first, then a unique leading-whitespace fallback (
tools.edit.strict_whitespace) - Preview Support: Shows comprehensive diff preview
- Security Validation: Respects all security restrictions
Security:
- Read Tool Requirement: Requires Read tool to be used first on the file
- Stale-Read Detection: Rejects the edit if the file changed on disk since the agent last read it, and asks the model to re-read first
- Approval Required: MultiEdit operations require approval by default
- Path Exclusions: Respects configured excluded paths
- Validation: Validates all edits before execution
Example:
{
"file_path": "config.yaml",
"edits": [
{
"old_string": "port: 3000",
"new_string": "port: 8080"
},
{
"old_string": "debug: true",
"new_string": "debug: false"
}
]
}Delete files or directories from the filesystem with security controls. Supports wildcard patterns for batch operations.
Parameters:
path(required): The path to the file or directory to deleterecursive(optional): Whether to delete directories recursively (default: false)force(optional): Whether to force deletion (ignore non-existent files, default: false)format(optional): Output format - "text" or "json" (default: "text")
Features:
- Wildcard Support: Delete multiple files using patterns like
*.txtortemp/* - Recursive Deletion: Remove directories and their contents
- Safety Controls: Respects configured path exclusions and security restrictions
- Validation: Validates file paths and permissions before deletion
Security:
- Approval Required: Delete operations require approval by default
- Path Exclusions: Respects configured excluded paths for security
- Pattern Matching: Supports glob patterns for path exclusions
- Validation: Validates file paths and prevents deletion of protected directories
Examples:
- Delete single file:
path: "temp.txt" - Delete directory recursively:
path: "temp/", recursive: true - Delete with wildcard:
path: "*.log" - Force delete:
path: "missing.txt", force: true
Configuration:
# ~/.infer/tools.yaml
tools:
delete:
enabled: true
require_approval: true # Delete operations require approval for securityA powerful search tool with configurable backend (ripgrep or Go implementation).
Parameters:
pattern(required): The regular expression pattern to search forpath(optional): File or directory to search in (default: current directory)output_mode(optional): Output mode - "content", "files_with_matches", or "count" (default: "files_with_matches")-i(optional): Case insensitive search-n(optional): Show line numbers in output-A(optional): Number of lines to show after each match-B(optional): Number of lines to show before each match-C(optional): Number of lines to show before and after each matchglob(optional): Glob pattern to filter files (e.g., ".js", ".{ts,tsx}")type(optional): File type to search (e.g., "js", "py", "rust")multiline(optional): Enable multiline mode where patterns can span lineshead_limit(optional): Limit output to first N results
Features:
- Dual Backend: Uses ripgrep when available for optimal performance, falls back to Go implementation
- Full Regex Support: Supports complete regex syntax
- Multiple Output Modes: Content matching, file lists, or count results
- Context Lines: Show lines before and after matches
- File Filtering: Filter by glob patterns or file types
- Multiline Matching: Patterns can span multiple lines
- Automatic Exclusions: Automatically excludes common directories and files (.git, node_modules, .infer, etc.)
- Gitignore Support: Respects .gitignore patterns in your repository
- User-Configurable Exclusions: Additional exclusion patterns can be configured by users (not by the LLM)
Security & Exclusions:
- Path Exclusions: Respects configured excluded paths and patterns
- Automatic Exclusions: The tool automatically excludes:
- Version control directories (.git, .svn, etc.)
- Dependency directories (node_modules, vendor, etc.)
- Build artifacts (dist, build, target, etc.)
- Cache and temp files (.cache, *.tmp, *.log, etc.)
- Security-sensitive files (.env, secrets, etc.)
- Gitignore Integration: Automatically reads and respects .gitignore patterns
- Validation: Validates search patterns and file access
- Performance Limits: Configurable result limits to prevent overwhelming output
Examples:
- Basic search:
pattern: "error", output_mode: "content" - Case insensitive:
pattern: "TODO", -i: true, output_mode: "content" - With context:
pattern: "function", -C: 3, output_mode: "content" - File filtering:
pattern: "interface", glob: "*.go", output_mode: "files_with_matches" - Count results:
pattern: "log.*Error", output_mode: "count"
Configuration:
# ~/.infer/tools.yaml
tools:
grep:
enabled: true
backend: auto # "auto", "ripgrep", or "go"
require_approval: falseExecute bash commands that match a per-mode allow-list. The model is default-deny: anything not matched is denied - it prompts for approval in chat, or is rejected with an actionable reason in headless agent mode.
Configuration:
# ~/.infer/tools.yaml
tools:
bash:
enabled: true
mode:
all: # baseline applied in EVERY mode (read-only / non-mutating)
allow:
- echo( .*)?
- ls( .*)?
- pwd( .*)?
- tree( .*)?
- wc( .*)?
- sort( .*)?
- uniq( .*)?
- head( .*)?
- tail( .*)?
- find( .*)?
- sleep( .*)?
- mkdir( .*)?
- ln -s( [^ -][^ ]*)+
- git status( .*)?
- git branch( --show-current)?( -[alrvd])?
- git log( .*)?
- git diff( .*)?
- git remote( -v)?
- git show( .*)?
- gh (issue|pr|repo|release|run|workflow) (list|view|status|diff|checks)( .*)?
- gh auth status( .*)?
- gh search (issues|code|prs|repos|commits)( .*)?
- gh project (list|view|item-list|field-list)( .*)?
- gh api repos/[^ ]+/contents/[^ ]+
- gh api '?user/repos[^ ]*'?( --paginate)?( --jq [^ ]+)?
- infer binaries status( .*)?
plan: # read-only planning mode adds nothing
allow: []
standard: # interactive default: baseline only (same as plan)
allow: []
auto: # headless `infer headless`: full autonomy (commit, push, etc.)
allow:
- .*The effective allow-list for a mode is mode.all.allow unioned with that mode's own list. Each entry
is a regex matched against the whole command (\A(?:entry)\z), so a bare token matches only
itself (gh allows gh, never gh issue list) - use ( .*)? to accept arguments. The single
sentinel .* means unrestricted (any single command, guard skipped); it is the default for auto
mode, which is how an autonomous infer headless commits and pushes. Tighten mode.auto.allow to a
curated list for CI with secrets so the guard re-applies.
Clean-command guard (every non-.* mode):
A command is rejected before matching, regardless of the allow-list, when it:
- Pipes or chains (
|,|&,&&,||,;,&, newline) - only a single command is auto-approved, sols | headis rejected even if bothlsandheadare allowed. Operators inside quotes (a jq'… | …'or--title "a && b") do not count. - Writes to a file (
>,>>,&>file,>&file) -echo secret > /etc/passwdis rejected; this cannot be unlocked by an allow entry. The same goes for options that write a file:sort -o,tree -o,git --output, anduniq's output operand. - Uses command substitution (
$(...), backticks,<(...),>(...)). - Runs a dangerous find action (
find ... -exec/-delete/…). - Leaks a variable: a printing/publishing command (
echo,printf,gh issue/pr create|comment|edit) may not expand$VAR-echo $AWS_SECRET_ACCESS_KEYis rejected, whilels $DIR(a non-printing use) is allowed when$DIRis inside the sandbox. A single-quoted or backslash-escaped$is literal.
Paths stay inside the sandbox: an allowed command runs without approval only when every path it
names passes the same sandbox check as the file tools (sandbox.yaml filesystem allowed and denied paths,
symlinks resolved). Arguments, --flag=value and -Xvalue values, input redirections (< file),
~, $VAR (from the CLI's environment) and globs are all checked, so head ~/.aws/credentials or
ls links/* through a link that leaves the project asks first. A path that cannot be known in advance
(${VAR:-x}, $1, {a,b}, a glob component starting with a dot, ~user) asks too. echo and
printf never open their arguments, so they are not checked. The default baseline does not include
make or task, since they run whatever the repository's Makefile or Taskfile says. Add them to a
mode's list, or append them with INFER_TOOLS_BASH_ALLOW_APPEND, where the repository is trusted.
Benign redirections that only discard or merge streams (2>&1, >/dev/null, 2>/dev/null) are
stripped before matching and remain allowed. A rejected command returns explanatory feedback naming
the reason, and (in chat) still goes through the normal approval prompt.
Search the web using DuckDuckGo or Google search engines to find information.
Configuration:
# ~/.infer/tools.yaml
tools:
web_search:
enabled: true
default_engine: duckduckgo
max_results: 10
engines:
- duckduckgo
- google
timeout: 10Fetch content from allowed URLs or GitHub references using the format example.com.
Configuration:
# ~/.infer/tools.yaml
tools:
web_fetch:
enabled: true
allowed_domains:
- golang.org
- github.com
safety:
max_size: 10485760 # 10MB
timeout: 30
cache:
enabled: true
ttl: 3600 # 1 hourHosts of configured A2A agents (each agent's url and artifacts_url, on any port) are always
fetchable when A2A is enabled, regardless of allowed_domains - registering an agent is the trust
decision, and the A2A tools instruct the model to download artifact URLs with WebFetch.
Raw browser automation driven through Playwright (CDP under the hood). All
six tools share one browser session that launches lazily on first use and
persists across calls, so navigation state carries over. The session drives
the user's installed browser (the configured browser.channel, default
chrome), attaches to an already-running browser when browser.cdp_endpoint
is set, or falls back to Playwright's bundled Chromium.
Configured in browser_use.yaml (global enabled flag, per-tool enable
flags, and shared rate limiting - same shape as computer_use.yaml).
Disabled by default; enable with INFER_BROWSER_USE_ENABLED=true or by
editing the file.
- BrowserNavigate - open a URL; returns the final URL and page title.
- BrowserClick - click an element by CSS or Playwright selector
(e.g.
text=Sign in). - BrowserType - fill an input element; optional
press_enterto submit. - BrowserRead - read the page (or one element's) visible text plus
URL/title. Also returns browser-initiated events: console messages,
auto-dismissed dialogs, and calls page scripts make to
window.inferNotify(...)- the browser-to-CLI channel. - BrowserScreenshot - capture the current page as an image; the screenshot is attached to the conversation and saved to disk.
- BrowserTabs - list the open tabs (title and URL).
Generate an image from a text prompt and save it as a PNG under ~/.infer/projects/<project-slug>/artifacts/. The
chat model calls the tool when the user asks for an image; the tool sends the prompt as a plain
one-off request to /v1/images/generations using the configured image model - no system prompt,
no tools, independent of the model selected for the chat session. Image models never appear in
the /model selector; they are only reachable through this tool.
Parameters:
prompt(required): Text description of the desired imagequality(optional):low(default),medium, orhighsize(optional):1024x1024(default),1536x1024, or1024x1536
Configuration:
# ~/.infer/tools.yaml
tools:
image_generation:
enabled: true
model: openai/gpt-image-2
require_approval: falseEdit an existing image and save the result as a PNG under ~/.infer/projects/<project-slug>/artifacts/. The chat model calls the tool
when the user asks to edit an image; the tool reads the input image from a local file path and sends a plain
one-off request to /v1/images/edits using the configured image model - no system prompt, no tools,
independent of the model selected for the chat session.
Parameters:
image(required): Local file path of the image to editmask(optional): Local file path to a PNG mask whose fully transparent areas (alpha = 0) mark the editable region. All other pixels are preserved exactly, and the mask must have the same dimensions as the input image.prompt(required): Text description of the desired editquality(optional):auto(default),low,medium,high, orstandardsize(optional):1024x1024(default),1536x1024, or1024x1536
Configuration:
# ~/.infer/tools.yaml
tools:
image_edit:
enabled: true
model: openai/gpt-image-2
require_approval: falseCreate a variation of an existing image and save the result as a PNG under ~/.infer/projects/<project-slug>/artifacts/. The chat model
calls the tool when the user asks for a variation; the tool reads the input image from a local file path and
sends a plain one-off request to /v1/images/edits using the configured image model - no system
prompt, no tools, independent of the model selected for the chat session.
Parameters:
image(required): Local file path of the image to base the variation onsize(optional):1024x1024(default),1536x1024, or1024x1536
Configuration:
# ~/.infer/tools.yaml
tools:
image_variation:
enabled: true
model: openai/gpt-image-2
require_approval: falseSynthesize speech from text and save it as a WAV file. The chat model calls the
tool when the user asks to say something aloud or to clone a voice. Two engines are available:
gateway (default) sends the request through the gateway's Audio API (/v1/audio/speech) - the
built-in local/qwen3-tts model works with the auto-started local gateway, and provider-hosted
models work too; qwen3-tts shells out to llama.cpp's llama-tts binary running Qwen3-TTS GGUF
models, fully local. Disabled by default: while
text_to_speech.enabled is false the tool definition is not sent to the LLM at all. See
text-to-speech for setup, engine choice and voice cloning.
Parameters:
text(required): The text to speakvoice_sample(optional): Bare file name of a WAV of the target speaker (~10-30s of clean speech) to clone; looked up in the working directory, then in the voice samples library (~/.infer/models/tts/samples)output_path(optional): Bare file name for the generated WAV, placed insidetext_to_speech.output_dir; defaults to a timestamped file
Configuration:
text_to_speech:
enabled: true
engine: gateway # gateway (default) | qwen3-tts (local)
# gateway engine only ("" = local/qwen3-tts):
# model: openai/gpt-4o-mini-tts
# voice: alloy
require_approval: true # optional; unset = no approval, like the image toolsCompose a music clip from a text prompt and save it as an MP3 file. The chat model calls the tool when the user asks for background
music, a loop or a jingle. The clip is generated behind the gateway's Music API (/v1/audio/music) with the configured
provider/model (default elevenlabs/music_v2_5); the CLI holds no provider key, the gateway does. Disabled by default: while
text_to_music.enabled is false the tool definition is not sent to the LLM at all. See
text-to-music for setup and gateway requirements.
Parameters:
prompt(required): Description of the music - genre, mood, instruments, temposeconds(optional): Clip length in seconds; omitted lets the provider pick a length that fits the promptinstrumental(optional):trueto guarantee the generated clip has no vocalsoutput_path(optional): Bare file name (no directories or absolute paths) for the generated MP3, placed insidetext_to_music.output_dir; defaults to a timestamped file
Configuration:
text_to_music:
enabled: true
# model: elevenlabs/music_v2_5 # gateway provider/model id
# output_dir: <media root>/music
require_approval: false # optional; unset = no approval, like the image toolsGenerate a short sound effect or ambience clip from a text prompt and save it as an MP3 file. The chat model calls the tool when the user asks for a
whoosh, a click, a riser or room tone - non-speech audio that TextToSpeech (it would read the word aloud) and TextToMusic (composes songs)
cannot cover. The clip is generated behind the gateway's SFX API (/v1/audio/sfx) with the configured provider/model (default
elevenlabs/eleven_text_to_sound_v2); the CLI holds no provider key, the gateway does. Disabled by default: while text_to_sfx.enabled is false
the tool definition is not sent to the LLM at all.
Parameters:
prompt(required): Description of the sound - the event or atmosphere and its character (e.g. a whoosh, a click, a riser, room tone)seconds(optional): Clip length in seconds, 0.5-30; omitted lets the provider pick a length that fits the promptloop(optional):trueto generate a clip that loops seamlesslyoutput_path(optional): Bare file name (no directories or absolute paths) for the generated MP3, placed insidetext_to_sfx.output_dir; defaults to a timestamped file
Configuration:
text_to_sfx:
enabled: true
# model: elevenlabs/eleven_text_to_sound_v2 # gateway provider/model id
# output_dir: <media root>/sfx
require_approval: false # optional; unset = no approval, like the image toolsRender a video clip from a text prompt and save it as an MP4 file; with a portrait and an audio clip the render is a lip-synced talking clip - the face
half of the desktop's Content workflow, with TextToSpeech as the voice half. The clip is generated behind the gateway's Videos API (/v1/videos) with
text_to_video.model for prompt renders (default elevenlabs/veo-3.1-fast-generate-001) or text_to_video.avatar_model for lip-synced renders
(default elevenlabs/creatify-aurora); the CLI holds no provider key, the gateway does. Avatar renders upload the user's
face and voice to a third-party provider, so the tool is disabled by default: while text_to_video.enabled is false the tool definition is not sent to
the LLM at all. See text-to-video for setup, the size rules and the upload limit.
Parameters:
prompt(required unlessaudiois given): Description of the shot; for avatar renders it describes framing only - the dialogue comes from the audio clipseconds(optional): Clip length in seconds as a string; providers accept a limited set of values; ignored whenaudiois presentsize(optional): Output resolution aswidthxheight(e.g.720x1280portrait or1280x720landscape), passed through verbatim; omitted means the provider default;creatify-aurorarenders 480p or 720p only and keeps the portrait's aspect ratioavatar(optional): The name of an avatar in the library (~/.infer/avatars/<name>/; seeinfer avatars list), or a bare file name of a .png, .jpg, .jpeg or .webp portrait in the working directory; required withaudio. Withaudiothe avatar's first image is lip-synced; without it a library avatar's images all go asreference_images(Veo takes at most 3), while a bare file becomes the first frameaudio(optional): Bare file name of a.wavor.mp3clip that drives the render, looked up in the working directory then in theTextToSpeechoutput directory; requiresavataroutput_path(optional): Bare file name (no directories or absolute paths) for the generated MP4, placed insidetext_to_video.output_dir; defaults to a timestamped file
Configuration:
text_to_video:
enabled: true
# model: elevenlabs/veo-3.1-fast-generate-001 # prompt renders
# avatar_model: elevenlabs/creatify-aurora # lip-synced avatar renders
# size: 720x1280 # optional widthxheight passthrough
# output_dir: <media root>/video
# timeout: 900 # whole-render timeout (seconds)
# poll_interval: 5 # job status poll cadence (seconds)
require_approval: false # optional; unset = no approval, like the image toolsCreate an avatar in the TextToVideo library (~/.infer/avatars/<name>/) from a photo: the agent-side twin of infer avatars create.
The photo is stored as 01-front (a JPEG upright and without metadata) and each requested angle is generated from it with
tools.image_edit.model, so the photo goes to that provider. Registered only when both text_to_video.enabled and
text_to_video.create_avatar are true, and it requires approval by default (text_to_video.require_approval overrides).
See text-to-video.
Parameters:
name(required): Bare name of the new avatar; an existing name fails, nothing is overwrittenphoto(required): Bare file name of a.png,.jpg,.jpegor.webpphoto, looked up in the working directory, then in the session's artifacts directory; absolute paths, directories and..are rejectedangles(optional): Views to generate -three-quarter-left,three-quarter-right,left-profile,right-profile; defaults to both three-quarter views,[]stores the photo onlyquality(optional):auto,low,medium,high(default) orstandardsize(optional): Generated image size asWIDTHxHEIGHTorauto(default1024x1536)
Configuration:
# ~/.infer/config.yaml
text_to_video:
enabled: true
create_avatar: true# ~/.infer/tools.yaml
tools:
image_edit:
enabled: true # needed to generate angles
model: openai/gpt-image-2Computer is the action-based desktop tool, enabled by computer_use.enabled.
accessibility- preferred first observation. Returns compact{role,label,state,bbox}elements for the frontmost application or another target. Bounding boxes use the same frame coordinate space as screenshots and pointer actions. Read-only; no screenshot or vision annotator is invoked.press- performs the accessibilitypressaction on the first element with an exactlabel, without moving the cursor or taking a screenshot.screenshot- captures a fresh screen image or a native-resolutionregion. Use this when the accessibility result says the tree is empty, unavailable, unsupported, or insufficient.cursor,move,click,double_click,triple_click,scroll,type, andkey- inspect or operate the pointer and keyboard.
The accessibility and press actions accept an optional target: frontmost (default), dock,
menubar, pid:<number>, app:<name>, or a bare application name. press also requires the exact
label returned by accessibility.
Under computer_use.approval: destructive, accessibility, screenshot, and cursor bypass
approval; press and the input actions require approval.
Record the screen to an MP4 file (H.264, yuv420p, plays in browsers and QuickTime). Enabled by
computer_use.recording.enabled (off by default; see
Configuration Reference).
RecordStart parameters:
mode(optional):screen(default, the entire primary display),window, orregionwindow(mode=windowonly):frontmost(default),app:<name>,pid:<number>, or a bare application name - the same syntax as theComputertool'stargetregion(mode=regiononly):{x, y, width, height}in the frame coordinate space, the same space asComputerscreenshots and accessibility bounding boxes
{"mode": "screen"}
{"mode": "window", "window": "app:Safari"}
{"mode": "region", "region": {"x": 0, "y": 80, "width": 1280, "height": 720}}RecordStart returns the file path and the captured rectangle, then keeps recording in the background.
RecordStop takes no arguments and returns the path, duration, and size. Behaviour:
- One recording at a time, machine wide: a second
RecordStart(from this or any otherinferprocess, which holds~/.infer/run/screen-recording.lockwhile its ffmpeg runs), orRecordStopwith nothing recording, returns an error and leaves the active recording alone. - A running recording is a background job (kind
recording): it is listed in/tasks, and a headless run waits for it, up toa2a.task.agent_mode_max_wait_seconds, instead of exiting, so a follow-upuser_messageon stdin can ask forRecordStopin the same process. A recording that stops on its own (the max duration, an ffmpeg exit, or a stop from/tasks) queues a note asking the agent to collect it withRecordStop. - A recording stops and finalizes itself at
computer_use.recording.max_duration(default 120s);RecordStopstill returns that file. - When the CLI exits (normal exit, Ctrl+C, SIGTERM) an active recording is finalized, and no ffmpeg
process is left behind. Call
RecordStopin the same session: channel and scheduled runs each start a fresh process.infer tools executerefuses both tools. windowmode records the window's bounds at the moment the recording starts; anything drawn over that area is recorded too, and moving the window does not move the capture.- Files go to
recordings/<timestamp>.mp4under the media root unlessoutput_diris set. - The chat status bar shows
● RECwhile a recording runs.
Requirements: ffmpeg with libx264 and the platform's screen grabber. The recorder uses ffmpeg
from PATH, otherwise installs the prebuilt binary into ~/.infer/bin/tools (upgraded when its sha256 no longer matches the release).
- macOS: your terminal app needs the Screen Recording permission, plus Accessibility for
windowmode (System Settings > Privacy & Security). - Linux: an X11 session (
x11grab). Wayland is not supported yet. - Windows:
gdigrab, no extra permission.
RecordStart requires approval, except in auto-accept mode or with
computer_use.recording.require_approval: false: chat prompts, and headless follows
approval_behaviour (IPC with --require-approval, otherwise blocked). Unattended runs with no
approver (CI on a virtual display) set INFER_COMPUTER_USE_RECORDING_REQUIRE_APPROVAL=false, after
which RecordStart follows computer_use.approval like RecordStop. RecordStop follows
computer_use.approval: under destructive it counts as an observation and bypasses approval;
under always it requires it.
Fetch the most recent frame from a named frame source: the built-in screen source (computer-use
screenshot streaming) or any directory source configured under vision.sources (e.g. camera frames
written to disk). Enabled whenever at least one frame source is registered.
Parameters:
source(optional): Frame source name. Defaults to the only registered source, orscreenwhen several exist.format(optional):regular(raw image attached) orannotated(scene summary + numbered elements with bounding boxes, produced by the configuredvision.annotator, replacing the image). When omitted:annotatedif an annotator is configured, otherwiseregular.region(optional): Zoom into a sub-region (screen source only), as{x, y, width, height}with all four required. The region is re-captured at native resolution, so small UI (Dock icons, dense toolbars) becomes readable. Coordinates are in the same frame space asComputerpointer actions and prior annotations, and returned element coordinates are translated back into that space.
For the screen source, annotated output includes element centers usable with the Computer tool's click action.
Annotated frames carry no base64 - the text replaces the image, so text-only models can use the tool
directly; vision models can always request format: regular.
Load an arbitrary local image file or http(s) URL, optionally answering a specific question about it.
Read-only, no approval required, and always available. With a vision.annotator configured it also
returns a text description for text-only models (see Configuration Reference).
Parameters:
image(required): Local file path or http(s) URL of the imageprompt(optional): A question to answer about the image
The macOS provider uses PureGo to call CoreFoundation, CoreGraphics, and AXUIElement directly; it has
no cgo, Swift, or Objective-C source. Native calls run in a short-lived helper process using JSON over
standard I/O. A helper crash, timeout, missing Accessibility permission, or unavailable tree returns
screenshot fallback guidance to the agent instead of terminating the CLI. Grant the infer process
permission in System Settings > Privacy & Security > Accessibility.
Linux AT-SPI and Windows UIA providers can implement the same provider contract later. Until then,
those platforms report unsupported; use screenshot for accessibility actions while the other
Computer actions continue to work.
Create and manage structured task lists for LLM-assisted development workflows.
Parameters:
todos(required): Array of todo items with status trackingid(optional): Unique identifier for the task (auto-generated when omitted)content(required): Task descriptionstatus(required): Task status - "pending", "in_progress", or "completed"
Features:
- Structured Task Management: Organized task tracking with status
- Real-time Updates: Mark tasks as in_progress/completed during execution
- Progress Tracking: Visual representation of task completion
- LLM Integration: Designed for LLM-assisted development workflows
Security:
- No File System Access: Pure memory-based operation
- Validation: Validates todo structure and status values
- Size Limits: Configurable limits on todo list size
Example:
{
"todos": [
{
"id": "1",
"content": "Update README with new tool documentation",
"status": "in_progress"
},
{
"id": "2",
"content": "Add test cases for new features",
"status": "pending"
}
]
}Configuration:
# ~/.infer/tools.yaml
tools:
todo_write:
enabled: true
require_approval: falseSubmits a finalized plan for user approval and persists it as a
Markdown file under <configDir>/plans/. Available only when the agent
is in Plan Mode (toggle via Shift+Tab in the chat TUI).
📖 For the full plan-mode workflow, see Plan Mode Guide.
How it works:
- The plan is written atomically to
<configDir>/plans/<YYYY-MM-DD-HHMMSS>-<slug>.md(e.g.~/.infer/plans/2026-04-27-103015-add-login-flow.md). - The chat TUI shows the rendered plan with Accept (auto-accept mode), Reject, and Approve Each Step (standard mode) options.
- On accept, the agent switches out of plan mode and executes the plan.
- On reject, the file remains on disk as an audit trail; the user can reply with feedback and the agent re-iterates.
- The LLM is instructed to ask clarifying questions in normal assistant turns first, and only call this tool when the plan is complete.
Parameters:
title(required): A short human-readable phrase (≤ 60 chars, no slashes or..). Becomes the H1 heading of the saved file and the basis of the filename slug.plan(required): The full plan as Markdown. Use H2 sections in this order -## Context,## Files to Modify,## Current Code,## Changes,## Performance Impact,## Critical Files,## Edge Cases,## Verification. Omit any section that is not applicable.
Example:
{
"title": "Add login flow",
"plan": "## Context\n\nUsers can't sign in...\n\n## Files to Modify\n\n- internal/auth/login.go - add handler\n\n## Verification\n\nRun `task test`."
}Configuration:
The tool is auto-registered and always enabled in plan mode. No config knobs.
Asks the user to override a tool call the LLM judge rejected (Auto+Judge mode or
tools.safety.approval_behaviour: judge). Every judge rejection result hints at
this tool; see Judge Mode.
How it works:
- Only calls the judge actually rejected can be escalated, and each one only once.
- The chat TUI shows the regular approval box for the rejected call, with the judge's reason and the model's justification above it.
- Approve arms a one-shot bypass: the model re-issues the identical call and it runs without a judge call. Reject returns the decision to the model as a tool result; the turn continues.
- Headless runs have no approver, so the tool returns
status: no_approverand the model is told not to retry.
Parameters:
tool(required): Name of the rejected tool.arguments(required): The exact arguments of the rejected call ({}if none).what(required): What permission is needed, one sentence.why(required): Why the action serves the user's request.
Example:
{
"tool": "Bash",
"arguments": { "command": "git push origin feature" },
"what": "push the feature branch",
"why": "the user asked me to ship the change"
}Configuration:
Auto-registered and always advertised. The description is overridable in
prompts.yaml under tools.RequestApproval.
Create recurring or one-off tasks that the agent runs on a cron schedule and delivers back through the messaging channel that triggered the current session (e.g. Telegram). Useful for "send me X every morning" or "remind me at 6pm today to call mum" - initiated from a chat with the bot.
📖 For an end-to-end walkthrough, see Scheduling Guide.
How it works:
- Each scheduled job is persisted through the configured storage backend; the default jsonl backend stores it as a YAML file under
~/.infer/schedules/. - The
infer daemonprocess hosts the scheduler and polls storage every 2s, so newly created jobs fire without a restart. - Each fire spawns a brand-new
infer headlesssession - no context carries between runs. Make prompts specific and self-contained. A run record (session_id,status,error, timestamps) is persisted per fire, so job output is readable from storage. - Channel + recipient are derived automatically from the current session ID - the LLM never passes them. From a channel-driven session the job delivers its output back to that channel; from any other session the job is record-only.
- One-off jobs (
run_once: true) are deleted automatically after their first fire.
Disabled by default. Enable in config under tools.schedule.enabled: true.
Parameters:
operation(required): One ofcreate,list,get,update,delete.job_id: Required forget,update,delete.cron_expression: Required forcreate. Standard 5-field crontab or@every <duration>.prompt: Required forcreate. The task to give the agent on each fire.run_once(optional, defaultfalse): When true, the job is deleted after its first fire.name,description,model: Optional metadata;modeloverridesagent.modelfor that job.
The LLM is instructed to always confirm with the user whether they want a one-off or recurring job before creating one - there is no safe default for that decision.
Example - recurring:
{
"operation": "create",
"cron_expression": "0 8 * * *",
"prompt": "Find an inspiring quote for today and respond with the quote and its author. Keep it under 3 sentences.",
"name": "Daily morning quote"
}Example - one-off reminder:
{
"operation": "create",
"cron_expression": "0 18 26 4 *",
"prompt": "Remind me to call mum.",
"run_once": true,
"name": "Call mum reminder"
}Example - list:
{ "operation": "list" }Example - delete:
{ "operation": "delete", "job_id": "0a1b2c3d-..." }Configuration:
# ~/.infer/tools.yaml
tools:
schedule:
enabled: false # disabled by default
require_approval: true # require approval by default
max_jobs: 100Security:
- Approval required by default - the LLM cannot create/modify schedules without user confirmation.
- Channel must be configured for delivery - a channel-driven session that references a channel not enabled in
channels.<name>.enablederrors out instead of silently skipping delivery. - Daemon-bound execution - jobs only fire while
infer daemonis running.
The A2A (Agent-to-Agent) tools enable communication between the CLI client and specialized A2A server agents, allowing for task delegation, distributed processing, and agent coordination.
📖 For detailed configuration instructions, see A2A Agents Configuration Guide
Submit tasks to specialized A2A agents for distributed processing.
Parameters:
agent_url(required): URL of the A2A agent servertask_description(required): Description of the task to performcontext_id(optional): Context ID from an earlier task to continue that conversation with the agent; omitting it starts an independent task
Features:
- Task Delegation: Submit complex tasks to specialized agents
- Streaming Responses: Real-time task execution updates
- Task Continuity: Continue an earlier task's conversation by passing its
context_id - Task Tracking: Automatic tracking of submitted tasks with IDs
- Error Handling: Comprehensive error reporting and retry logic
Examples:
- Code analysis:
agent_url: "http://security-agent:8080", task_description: "Analyze codebase for security vulnerabilities" - Documentation:
agent_url: "http://docs-agent:8080", task_description: "Generate API documentation" - Testing:
agent_url: "http://test-agent:8080", task_description: "Create unit tests for UserService class"
Retrieve agent capabilities and metadata for discovery and validation.
Parameters:
agent_url(required): URL of the A2A agent to query
Features:
- Agent Discovery: Query agent capabilities and supported task types
- Health Checks: Verify agent availability and status
- Metadata Retrieval: Get agent configuration and feature information
- Connection Validation: Test connectivity before task submission
Examples:
- Capability check:
agent_url: "http://agent:8080"- Returns agent card with available features - Health status: Query agent before submitting critical tasks
Query the status and results of previously submitted tasks.
Parameters:
agent_url(required): URL of the A2A agent servercontext_id(required): Context ID for the tasktask_id(required): ID of the task to query
Features:
- Status Monitoring: Check task completion status and progress
- Result Retrieval: Access task outputs and generated content
- Error Diagnostics: Get detailed error information for failed tasks
- Artifact Discovery: List available artifacts from completed tasks
Examples:
- Status check:
agent_url: "http://agent:8080", context_id: "ctx-123", task_id: "task-456" - Result access: Retrieve task outputs and completion details
1. Query agent capabilities: A2A_QueryAgent
2. Submit task for processing: A2A_SubmitTask
3. Monitor task progress: A2A_QueryTask
A2A tools are configured in the tools section:
a2a:
enabled: true
tools:
submit_task:
enabled: true
require_approval: true
query_agent:
enabled: true
require_approval: false
query_task:
enabled: true
require_approval: falseNote: Artifact downloads from A2A tasks are handled via the WebFetch tool with download=true.
Files are automatically saved to the per-session artifacts directory
~/.infer/projects/<project-slug>/artifacts/<session-id>/ with the filename extracted from the download URL.
- Code Analysis: Submit codebases to security or quality analysis agents
- Documentation Generation: Generate API docs, README files, or technical documentation
- Testing: Create comprehensive test suites with specialized testing agents
- Data Processing: Process large datasets with specialized data analysis agents
- Content Creation: Generate content with specialized writing or design agents
For detailed A2A documentation and examples, see A2A Agents Configuration Guide.
Add your own tools in any language by placing one YAML manifest per tool in ~/.infer/tools/, or in a project's
.infer/tools/ or .agents/tools/. infer runs the manifest's command with the call's arguments as JSON on stdin and
returns its stdout. Custom tools follow the same agent modes and approval flow as the tools above, and project tools
always need approval. See the Custom Tools guide.
