Hengyi Xie1*, Chenfei Yao1*, Xianjin Wu1, Xuanyang Xi2, Yiping Tang2, Di Xu2, Yingying Zhu1, Dingkang Liang1β , Xiang Bai1, Han Ding1
1 Huazhong University of Science and Technology, China
2 Huawei Technologies Co. Ltd, China
* Equal contribution, listed alphabetically by surname. β Project Lead.
This repository contains the official implementation of TurboVLA for the paper TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM.
2026.07.31: Released the TurboVLA model checkpoints on Hugging Face.2026.07.30: Released the paper, training and evaluation code.
- Support Huawei Ascend NPUs
Vision-language-action (VLA) models commonly adopt an LLM-centric V β L β A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation.
In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V β L β A pathway as a direct V + L β A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation.
Clone the repository:
git clone https://github.com/H-EmbodVis/TurboVLA.git
cd TurboVLALIBERO and RoboTwin use different simulator and data stacks. We recommend separate Python 3.10 environments.
conda create -n turbovla-libero python=3.10 -y
conda activate turbovla-libero
# Install the CUDA-compatible PyTorch build for your system first.
pip install torch==2.3.1 torchvision==0.18.1 --index-url https://download.pytorch.org/whl/cu121
pip install -e ".[libero]"Install LIBERO separately in the same environment.
conda create -n turbovla-robotwin python=3.10 -y
conda activate turbovla-robotwin
# Install a CUDA-compatible PyTorch build before the project dependencies.
pip install -e ".[robotwin]"
pip install flash-attn==2.7.4.post1 --no-build-isolationInstall the RoboTwin 2.0 simulator in a separate evaluation environment when required by your setup.
Model weights and benchmark datasets are external assets and are not committed to this repository.
Download the complete TurboVLA release, including model weights and normalization metadata, from Hugging Face:
pip install -U huggingface_hub
hf download H-EmbodVis/TurboVLA \
--local-dir pretrained/TurboVLA| Asset | Source | Used by |
|---|---|---|
| DINOv3 ViT-B | facebookresearch/dinov3 | LIBERO |
| DINOv3 ViT-L | facebookresearch/dinov3 | RoboTwin |
| BERT base uncased | google-research/bert | Both |
| GroundingDINO Swin-T OGC | IDEA-Research/GroundingDINO | Both |
TurboVLA expects the four modified no-noops suites in TFDS/RLDS format:
data/libero/
|-- libero_10_no_noops/1.0.0/
|-- libero_goal_no_noops/1.0.0/
|-- libero_object_no_noops/1.0.0/
`-- libero_spatial_no_noops/1.0.0/
The repository provides no-op removal and mixed-suite statistics utilities:
python scripts/libero/regenerate_libero_no_noops.py --help
python scripts/libero/compute_mixed_stats.py --helpReleased normalization statistics are stored in experiments/libero/configs/libero_all4_stats.json. BERT is part of the model and runs online during both training and evaluation; no text-feature cache is required. See experiments/libero/README.md for details.
Download the clean LeRobot dataset and create the expected local link:
bash scripts/robotwin/prepare_data.sh /path/to/storage
export ROBOTWIN_DATA_ROOT="$PWD/playground/Datasets/RoboTwin"The default downloader uses StarVLA/RoboTwin-Clean. The training registry expects all 50 datasets under Clean/<task_name>.
The paper recipe uses DINOv3 ViT-B, two camera views, 7-D actions, a 12-step action chunk, 80k optimizer steps, 10k warmup steps, and global batch size 256 on four GPUs.
torchrun --nproc_per_node=4 experiments/libero/train.py \
--dataset_dir data/libero/libero_10_no_noops/1.0.0 \
--dataset_dirs "data/libero/libero_10_no_noops/1.0.0,data/libero/libero_goal_no_noops/1.0.0,data/libero/libero_object_no_noops/1.0.0,data/libero/libero_spatial_no_noops/1.0.0" \
--stats_path experiments/libero/configs/libero_all4_stats.json \
--stats_key libero_all4_no_noops \
--dinov3_path facebook/dinov3-vitb16-pretrain-lvd1689m \
--bert_path google-bert/bert-base-uncased \
--allow_hf_download \
--pretrained_init_ckpt /path/to/groundingdino_swint_ogc.pth \
--checkpoint_dir outputs/liberoOne command evaluates one checkpoint on one suite.
python experiments/libero/evaluate.py \
--ckpt_path pretrained/TurboVLA/checkpoints/libero/libero_object.pth \
--dinov3_path /path/to/dinov3-vitb \
--bert_path /path/to/bert-base-uncased \
--stats_path experiments/libero/configs/libero_all4_stats.json \
--stats_key libero_all4_no_noops \
--task_suite_name libero_object \
--num_trials_per_task 50 \
--chunk_size 12 \
--num_open_loop_steps 12 \
--seed 7 \
--precision bf16 \
--result_json_path outputs/evaluation/libero_object.jsonValid suite names are libero_spatial, libero_object, libero_goal, and libero_10.
The paper recipe uses DINOv3 ViT-L, three camera views, 14-D absolute joint-position actions, a 50-step ACT head, global batch size 192, and 55k optimizer steps on four GPUs.
export ROBOTWIN_DATA_ROOT="$PWD/playground/Datasets/RoboTwin"
export BERT_MODEL_PATH=/path/to/bert-base-uncased
export TURBOVLA_INIT_CKPT=/path/to/groundingdino_swint_ogc.pth
export DINOV3_MODEL_PATH=/path/to/dinov3-vitl
export CUDA_VISIBLE_DEVICES=0,1,2,3
MAX_TRAIN_STEPS=55000 \
RUN_ID=turbovla_robotwin_clean50_55k \
bash scripts/robotwin/train.shThe policy server and RoboTwin simulator can run in separate Python environments:
export ROBOTWIN_PATH=/path/to/RoboTwin
export STARVLA_PYTHON=/path/to/policy-env/bin/python
export ROBOTWIN_PYTHON=/path/to/robotwin-env/bin/python
export CUDA_VISIBLE_DEVICES=0,1,2,3
ROBOTWIN_TEST_NUM=100 \
bash scripts/robotwin/evaluate.sh \
pretrained/TurboVLA/checkpoints/robotwin/steps_55000_ema_model.safetensorsAppend task names to the evaluation command to run a subset. Omitting them evaluates all 50 clean tasks.
TurboVLA builds upon the following projects and resources:
- DINOv3 for visual representations.
- GroundingDINO for bidirectional vision-language interaction components and initialization.
- VLA-Adapter for the LIBERO task and episode rollout protocol.
- StarVLA for the RoboTwin-compatible training and evaluation runtime.
- LIBERO and RoboTwin 2.0 for simulation benchmarks.
If TurboVLA is useful in your research, please consider citing the paper:
@article{xie2026turbovla,
title = {TurboVLA: Real-Time Vision-Language-Action Model at
32 Hz on an RTX 4090 with <1 GB VRAM},
author = {Xie, Hengyi and Yao, Chenfei and Wu, Xianjin and
Xi, Xuanyang and Tang, Yiping and Xu, Di and
Zhu, Yingying and Liang, Dingkang and Bai, Xiang and
Ding, Han},
journal = {arXiv preprint arXiv:2607.27205},
year = {2026}
}