Agentic AI for XR — an open-source foundation for multi-modal, real-time conversational AI within the CloudXR ecosystem.
This project is publicly available in beta and is under active development. Features, APIs, documentation, and behavior may change as the project evolves. Expect bugs, incomplete functionality, and breaking changes. Use at your own discretion, and please report issues or feedback to help improve the project.
XR AI is a developer stack for building powerful XR and AI systems across devices, platforms, and deployment environments. It connects web, iOS/visionOS, AR glasses, and XR headset clients to GPU-accelerated AI services, tool-using agents, and the CloudXR stack for remote rendering.
With XR AI, developers can build agents that see and hear what the user experiences, reason over live physical context, call external tools through MCP, and respond with audio or data in the same XR session. The stack provides an end-to-end foundation for multimodal spatial computing applications: real-time media routing, participant-aware response handling, agent interfaces, AI service integration, remote rendering, and sample applications that show the pieces working together.
The value is speed without lock-in. XR AI is designed to work quickly with NVIDIA open models for vision, language, and speech, plus swappable speech synthesis services, while still giving developers the flexibility to bring their own models, services, tools, and application logic. Because it is built around NVIDIA GPU infrastructure, the same architecture can be deployed where the workload needs to run: cloud, data center, workstation, or edge.
XR AI also gives developers a practical path across product categories. Teams can start with AI glasses-style experiences that use live camera, audio, and agent responses, then extend the same framework to richer AR glasses or XR headset experiences that use CloudXR remote rendering. This lets developers build for today's lightweight AI devices while keeping a clear path to immersive, GPU-rendered spatial applications.
XR AI is especially useful when you need to:
- Build multimodal XR agents that can see, hear, reason, use tools, and respond in real time.
- Target multiple client platforms including web, iOS/visionOS, AR glasses, and XR headsets.
- Use NVIDIA open models out of the box while preserving the flexibility to bring your own models and services.
- Deploy wherever NVIDIA GPUs are available, from cloud and data center to workstation and edge.
- Start with AI glasses-style experiences and scale to CloudXR remote rendering for richer AR and XR applications.
- Keep transport, rendering, model services, tools, and agent logic separated so teams can evolve each layer independently.
Hardware
XR AI samples are designed for a single NVIDIA RTX PRO 6000 Blackwell workstation GPU or an NVIDIA DGX Spark. Both provide enough VRAM to run the full model stack locally. If you prefer not to run models on local hardware, model endpoints are plain URLs — point the worker config at a cloud NIM or model endpoint and no local GPU is required for the agent or hub.
| Sample | Local VRAM needed |
|---|---|
| model-servers (shared models) | ~58 GB |
| simple-vlm-example (standalone) | ~23 GB |
| xr-render-demo (requires model-servers) | ~55 GB (models) + ~2 GB (hub/TTS) |
| Hub only | none |
Software
| Requirement | Version | Notes |
|---|---|---|
| OS | Linux | Ubuntu 22.04 / 24.04 recommended |
| Python | 3.11 or 3.12 | 3.10 and 3.13 are not supported |
| uv | latest | dependency manager used by all samples |
| NVIDIA driver | 570+ | required for local model inference |
| Docker | 24+ | required: all vLLM-backed services (LLM, VLM) run in nvcr.io/nvidia/vllm containers |
| NVIDIA Container Toolkit | latest | required: gives Docker access to the GPU. Without it, model_servers fails with failed to discover GPU vendor from CDI: no known GPU vendor found |
uv handles all Python dependencies per-sample — no global pip install
or virtual-environment setup needed. If you do not have it:
curl -LsSf https://astral.sh/uv/install.sh | shThe NVIDIA Container Toolkit install is one-time per host. Follow the official install guide and run the CDI / runtime-configure steps from there:
https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html
Quick smoke-test once installed:
docker run --rm --gpus all nvidia/cuda:13.0.3-base-ubuntu24.04 nvidia-smiGPU-profile prerequisites — install before uv sync for these targets:
- DGX Spark (
xr-render-demo/yaml/spark/):sudo apt install python3-dev
(All GPU profiles default to vllm_backend: docker, so the vLLM
container ships nvcc + FlashInfer. If you switch a profile to
vllm_backend: pip, see docs/troubleshooting.md
for the host CUDA toolchain prereq.)
If uv sync or the VLM fails on first run, see
docs/troubleshooting.md.
Network — open the firewall ports listed in
docs/networking.md before connecting from another
machine. UDP 7882 is a silent-failure path: signaling succeeds but media
frames are dropped if it is closed.
| Layer | Directory | Description |
|---|---|---|
| Clients | client-samples/ |
Android, iOS/visionOS, Web, and native StreamKit clients |
| Server runtime | server-runtime/ |
XR-Media-Hub + LiveKit internal transport |
| Launcher | utils/xr-ai-launcher/ |
stdlib-only process manager used by samples |
| Logging | utils/xr-ai-logging/ |
shared loguru sink + stdlib bridge for every process |
| Agent functions | agent-sdk/xr-ai-nat/ |
Typed, in-process NAT functions for XR capabilities |
| Capability services | services/ |
Long-running typed services used by native functions |
| Agent interfaces | agent-mcp-servers/ |
MCP compatibility processes for XR data & rendering |
| Agent demos | agent-samples/ |
End-to-end agent pipelines |
| Tests | tests/ |
Multi-client / multi-agent integration tests |
Lightweight samples (simple-vlm-example) are self-contained — one command
starts everything. Heavier demos (xr-render-demo) split model loading from
the demo itself: start model-servers once, then run the demo as many times
as you like without reloading weights.
Every sample worker depends on agent-sdk/xr-ai-models — one SDK that
abstracts the OpenAI-compatible HTTP wire format for LLM / VLM / STT / TTS /
embeddings behind typed service protocols. Each sample ships a model config that names the
logical models the worker needs (llm, vlm, stt, …) with
preset references that pre-fill model-specific quirks (reasoning-field
aliasing, chat_template_kwargs, served-model-name strings). Workers call
make_llm(config, "llm") / make_vlm(config, "vlm") / make_stt(config, "stt") / make_tts(config, "tts") / make_embedding(config, "embedding") — no hand-rolled httpx clients, no model
quirks leaking out of the SDK. Full quickstart and the built-in preset
table: agent-sdk/xr-ai-models/README.md.
Every sample follows the same pattern: start the server, then connect a client. Once it is ready, any supported client — web browser, Android app, iOS/visionOS app, or AR glasses — can join the session using the token printed on startup.
model-servers starts the shared inference services used across demos and exits
immediately — the services keep running in the background with weights hot.
Start this once before running xr-render-demo, or whenever you want to
pre-warm models:
cd agent-samples/model-servers
uv sync
uv run model_serversGPU profiles are auto-detected (dual_48G_ada / spark / 96G_blackwell).
On first run each model downloads from HuggingFace (~50 GB total; can take
tens of minutes). On subsequent runs the containers restart in under a minute.
The default --vlm-llm-stack starts Nemotron-3 Nano (8107), Cosmos (8100),
STT (8103), and embeddings (8109). Use --omni-stack to replace Nano and
Cosmos with Nemotron-3 Nano Omni (8108); STT and embeddings remain available.
On dual_48G_ada, the default stack places Cosmos and embeddings on GPU 0;
the Omni stack places Omni on GPU 0 and embeddings on GPU 1.
Switching stacks stops the incompatible persistent models first and aborts if
they cannot be stopped, avoiding GPU overcommit.
uv run model_servers --omni-stackThe default models are public, so no HuggingFace token is required. Set
HF_TOKEN to lift download rate limits / speed, or to use a gated model — see
docs/credentials.md. The launcher won't prompt; it
prints a one-line notice and continues if the token is unset.
To stop all model servers when done:
uv run model_servers --stop--stop always stops both stack variants, so it takes no stack-selection flag.
End-to-end voice + vision sample. Speak into the mic, type into the data
channel, or send the literal text "ping" — all routes go through the
same VLM pipeline against the latest video frame. Replies arrive as
streaming Piper TTS audio plus a vlm.response text message.
The packaged worker composes the NAT-native streaming vision function with
xr-ai-voice's VoiceSession; Pipecat remains private to that runtime and no
MCP client is involved. See the
sample README for the worker
layout and configuration boundaries.
Uses nvidia/Cosmos-Reason1-7B (NVIDIA Open Model License + Apache 2.0).
There are two ways to run it:
Standalone (~23 GB VRAM) — starts its own VLM and STT, owns them for the session, and stops them when you exit:
cd agent-samples/simple-vlm-example
uv sync
uv run simple_vlm_exampleOn the very first run weights download from HuggingFace (~23 GB; can take
several minutes). The default model is public — no HuggingFace token needed;
set HF_TOKEN only to lift rate limits / speed or for a gated model (see
docs/credentials.md).
With model-servers pre-running — if VLM (port 8100) and STT (port 8103)
are already up from model-servers, the demo detects them at startup and
reuses them. No extra flags needed. When you exit, those services keep
running.
uv run simple_vlm_exampleThe hub, VLM, STT, and TTS start together (or reuse running services). When ready the hub prints:
[hub] LiveKit URL : wss://0.0.0.0:8080
[hub] Room : xr-room
[hub] Token : eyJ…
[hub] Web client : https://localhost:8080
Open https://localhost:8080 in a browser. The samples ship with HTTPS
on by default (a self-signed cert is generated on first run at
~/.local/share/xr-ai/web-server.crt), so you'll see a "Your connection
is not private" warning the first time — click Advanced → Proceed
(Chrome/Edge) or Accept the Risk and Continue (Firefox). See
docs/networking.md for trusting the cert
permanently or running over plain HTTP instead.
Leave Token URL blank — the web client fetches a token from the server automatically. Click Connect.
You are now live in the XR session. To test the agent:
- Type
pingin the data channel → the agent describes what the camera sees. - Type any question → sent verbatim to the VLM.
- Speak into your mic → speech is transcribed and sent as a query.
A successful round trip: your query appears in the log, the agent responds after a moment, and you hear the reply through your speakers.
Local model — override the model weights or GPU settings by editing
vlm_server.yaml in the sample directory.
Remote model — copy yaml/models.hosted.json, point its VLM endpoint at
your server, and select it in the worker config:
{
"endpoint": {"base_url": "https://your-remote-vlm.example.com"},
"deployment": {"ownership": "external"}
}# yaml/simple_vlm_example_worker.yaml
models_config: models.remote.jsonThe external deployment declaration makes the orchestrator skip the local VLM process automatically.
Hosted NVIDIA NIM — run the VLM on hosted NIM
(build.nvidia.com) instead of locally (STT/TTS
stay local) by setting one key in simple_vlm_example_worker.yaml:
models_config: models.hosted.jsonThe same profile configures the worker and makes the orchestrator skip the
local VLM server. Pick the hosted model id in models.hosted.json and provide
an NGC_API_KEY as
an environment variable (or save it once via the launcher credential
prompt) — it is not stored in YAML; the overlay only names the env var via
api_key_env: NGC_API_KEY. See
docs/credentials.md. Full details (and self-hosted
NIM containers):
docs/ai-services.md.
Each sample has its own xr_media_hub.yaml controlling the hub; see
server-runtime/xr_media_hub.yaml
for the full option list.
Speak to the web client and a sphere in the streamed scene tracks your voice — radius follows loudness, colour and position follow spoken commands ("make it red", "put it to my left", "where I'm looking"). Runs against a Quest 3 / Vision Pro on the same LAN, or the IWER emulator built into the web client for desktop dev.
Under the hood, the orchestrator launches the hub, CloudXR runtime, model
endpoints, typed capability processes, and the worker. The worker calls those
processes through native NAT functions; MCP adapters remain optional outward
compatibility surfaces and are not in the sample's execution path. The Pipecat pipeline runs
quick-acks and a Nemotron-30B agentic tool-calling loop over
scene, XR tracking, spatial math, vision, and video-memory functions. Full process map,
agentic-loop details, and the XR session lifecycle:
docs/xr-render-demo.md.
Requires model-servers to be running first — the demo does not start
its own model services.
cd agent-samples/model-servers
uv sync && uv run model_serversThis exits immediately once all services are ready. Weights stay loaded in the background.
This demo has two extra host prerequisites beyond the shared Requirements:
- Vulkan loader + headers — the CloudXR compositor and LOVR render through
Vulkan, so install them before running the demo:
sudo apt install libvulkan-dev - npm 18+ on PATH — the orchestrator builds the web vendor bundle on first run (skipped on subsequent runs).
cd agent-samples/xr-render-demo
uv sync
uv run xr_render_demoBy default this serves the web / WebRTC client (NV_DEVICE_PROFILE=auto-webrtc).
For a native Apple Vision Pro client, start it with NV_DEVICE_PROFILE=auto-native
instead — see docs/xr-render-demo.md.
On first run the orchestrator automatically downloads the pinned LOVR version to
deps/lovr/ inside the repo and builds the web vendor bundle (requires npm
and network access). Both steps are skipped on subsequent runs.
DGX Spark (aarch64): LOVR does not publish a prebuilt aarch64 Linux
binary, so the auto-download is not available — build LOVR from source and
export LOVR_BIN. See
docs/troubleshooting.md.
To use a custom LOVR build:
export LOVR_BIN=/path/to/your/lovr # or set lovr_bin: in scene/scene_service.yaml
uv run xr_render_demoGPU pinning for the XR side is controlled by gpu_index in
agent-samples/xr-render-demo/yaml/cloudxr_runtime.yaml. cloudxr-runtime
applies the pin to its own process and writes the selectors into
cloudxr.env; the scene process and LOVR inherit from that file. See
docs/xr-render-demo.md
for full details.
To stop the model servers when done:
cd agent-samples/model-servers
uv run model_servers --stopHosted NVIDIA NIM — run the LLMs and VLM on hosted NIM
(build.nvidia.com) instead of local vLLM
(STT/TTS stay local) by setting one key in xr_render_demo_worker.yaml:
model_backend: nim # default is "local"The worker loads yaml/models.nim.yaml for the native model-backed functions —
no main.py edits. Provide
an NGC_API_KEY as an environment variable (or via the launcher
credential prompt — not in YAML) and just don't start the local
agent-llm / vlm model-servers. See
docs/ai-services.md.
cd server-runtime
uv sync
uv run xr_media_hubUseful for development or when running an agent in a separate terminal.
The hub auto-discovers server-runtime/xr_media_hub.yaml.
Open https://localhost:8080 in a browser. The samples ship with HTTPS
on by default; the first connection shows a self-signed cert warning that
you click through (or trust permanently — see
docs/networking.md). Leave Token URL blank to
use the server's built-in /token endpoint, or paste the printed token
directly.
The page's import map loads livekit-client and @nvidia/cloudxr from
client-samples/web/vendor/ (same-origin, so XR headsets and offline LANs
work). Both bundles are gitignored build output. The xr-render-demo
orchestrator builds them automatically on first run (requires npm on PATH).
For a manual rebuild after an SDK bump, see
client-samples/web-xr-build/README.md.
See client-samples/android/README.md for
full setup. Quick steps:
- Open
client-samples/android/in Android Studio (Hedgehog or later). - Let Gradle sync finish — it downloads the LiveKit Android SDK automatically.
- Run on a device or emulator (API 24+).
- Enter the server IP, port (
8080— the hub's web-server port, not LiveKit's internal 7880), and paste the printed token.
Permissions (RECORD_AUDIO, CAMERA) are requested at runtime on first use.
See client-samples/ios-visionos/README.md
for full Xcode setup. Quick connection settings:
| Field | Value |
|---|---|
| Host | IP of the machine running the server |
| Port | 8080 (the hub web-server port; not LiveKit's internal 7880) |
| Token | Paste the token printed on server startup |
The token is valid for 24 hours. To get a fresh one restart the server or call
GET https://<host>:8080/token?identity=<name>.
One-time per device: the LiveKit Swift SDK does not expose a server-trust hook, so iOS rejects the hub's self-signed cert until you install it as a trusted profile. On the device, open
https://<host>:8080/certin Safari → bypass the warning → install → enable Settings → General → About → Certificate Trust Settings → Enable Full Trust. Full walkthrough plus recovery for the common failure modes is inclient-samples/ios-visionos/README.mdunder "Trusting the hub's self-signed cert".
The hub and CloudXR runtime use a small set of TCP/UDP ports (web client +
wss /rtc proxy on 8080, WebRTC fallbacks on 7881/TCP + 7882/UDP, CloudXR
WSS proxy on 48322). LiveKit's native 7880 stays on loopback — clients
connect through the same-origin wss proxy, not directly. Full table and
distro-specific ufw / firewall-cmd recipes are in
docs/networking.md. The same doc covers HTTPS for
the web client and self-signed certificate trust on each browser.
tests/ contains the multi-client / multi-agent integration suite. The
core IPC tests run without Docker or LiveKit — they spin up real
HubEndpoint / ConnectorEndpoint / ProcessorEndpoint instances over
ipc:// sockets and verify routing, isolation, and the
ReturnAudioFlush control path.
cd tests
uv sync
uv run pytest -vSee tests/README.md for the full breakdown. CI runs
the suite on every push and pull request via
.github/workflows/tests.yml on Python 3.11
and 3.12.
For engineers and agents working in the repo:
| Doc | Topic |
|---|---|
AGENTS.md |
Working contract — hard rules every change must satisfy |
DEPENDENCIES.md |
Authoritative dependency map (update with every pyproject.toml change) |
| Versioned documentation | Latest release by default, plus main development and release-tag documentation |
docs/architecture.md |
Hub ↔ transport ↔ agent boundaries; known limitations |
docs/process-model.md |
Process / run_stack mechanics; ready-file protocol |
docs/ai-services.md |
VLM / STT / TTS / LLM server reference + worker call examples |
docs/xr-render-demo.md |
xr-render-demo architecture: native functions, agentic loop, XR lifecycle |
docs/adding-a-sample.md |
Boilerplate for scaffolding a new sample |
docs/adding-cloudxr.md |
Wiring CloudXR into a sample |
docs/credentials.md |
HF / NGC token management |
docs/networking.md |
Firewall ports + HTTPS for the web client |
docs/troubleshooting.md |
Known frictions and runtime symptoms |
docs/spdx-headers.md |
SPDX header style and enforcement |
docs/changelog.md |
Significant design decisions, reverse chronological |
LICENSE— Apache-2.0.SECURITY.md— how to report a vulnerability.CONTRIBUTING.md— contribution process and DCO.CODE_OF_CONDUCT.md— community standards.THIRD_PARTY_NOTICES.md— bundled third-party components and their licenses.