Skip to content

Feature/add turbovla support - #29

Merged
khanhnd61-vr merged 5 commits into
VinRobotics:mainfrom
Zeustakeshi:feature/add-turbovla-support
Sep 25, 2026
Merged

khanhnd61-vr merged 5 commits into
VinRobotics:mainfrom
Zeustakeshi:feature/add-turbovla-support

Conversation

@Zeustakeshi

Copy link
Copy Markdown
Contributor

What

Add TurboVLA support, including GGUF conversion, model loading and inference,
direct-evaluation client handling, converter remap coverage, and documentation.

Why

TurboVLA provides a compact DINOv3 + BERT VLA policy. This adds a native
vla.cpp inference path without requiring Python at runtime.

Verified

  • Builds clean under -Wall -Wextra (first-party code)
  • ctest passes
  • Numeric output unchanged (vla_predict_check diff), or the change is
    meant to move it and a LIBERO sweep is below

Additional checks:

  • python3 tests/py/test_converters.py passes.
  • python3 tests/py/test_bindings.py passes.
  • Added converter remap coverage for TurboVLA tensor groups.

Archs and backends tested:

  • TurboVLA — CPU build (GGML_CUDA=OFF): compile and unit-test coverage.
  • CUDA / Metal / SYCL / OpenVINO: not tested in this PR.
  • Checkpoint-level numeric comparison and LIBERO sweep: not yet attached.

@khanhnd61-vr khanhnd61-vr left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the port. The module structure follows upstream closely. The released LIBERO checkpoint still produces actions that diverge from the reference model, because the text path doesn't reproduce how the checkpoint pads instructions. The converter also can't read the released checkpoint. I pushed f54dd08 on top of this branch, which fixes both and restructures the engine. The inline comments cover the specifics.

How I checked

Reference: upstream H-EmbodVis/TurboVLA@b29ab142 in fp32 on CPU, with checkpoints/libero/turbovla_libero.pth loaded strictly into its module tree. DINOv3 is gated, but only its public architecture config is needed; the fine-tuned weights are inside the checkpoint. There are four seeded cases. Each has two 256×256 views and an 8-D state, plus one instruction from each of the checkpoint's padding groups (21, 11 and 14 tokens) and one instruction the checkpoint doesn't list. The table shows the max |Δ| of the normalized action chunk against the reference:

case padding group this PR (CPU) f54dd08 (CPU) f54dd08 (CUDA)
libero_10 instruction, 14 tokens 21 4.9e-3 9.5e-7 6.7e-4
libero_goal instruction, 8 tokens 11 9.0e-3 3.6e-7 1.2e-3
libero_object instruction, 14 tokens 14 5.8e-2 8.0e-7 8.3e-4
unlisted instruction, 10 tokens 21 5.7e-2 9.3e-6 4.9e-3

For scale, upstream's own release precision (whole model in bf16 on CUDA) is up to 3.4e-2 from fp32 on the same cases. The CUDA column is ggml-cuda's tensor-core f32 GEMMs.

Blocking

  1. Text padding. The ACT decoder attends every text row, padding included, because encode_condition concatenates them unmasked. That makes the padded length part of the model's output. The checkpoint pins this length per training instruction (padding_length_by_instruction: 11, 14 or 21 for LIBERO). BERT runs over that many tokens, and rows after it up to padding_length are zeros that the text projection turns into its bias. The PR always runs BERT over 21 tokens. The converter now stores the table as token ids → length (turbovla.text_groups.*), and the engine looks the instruction up.
  2. The special-token edge rule (col == num_token - 1) is evaluated against 21 instead of the group length. An instruction that fills its group therefore gets its closing [SEP] treated as a sub-sentence boundary. This is case 3 above.
  3. The text enhancer masks padded keys. Upstream passes no key_padding_mask to it (the argument is commented out in models/components/transformer.py), so a padded row attends to itself. Here the row is fully masked, softmax spreads it uniformly, and that row feeds the decoder.
  4. The converter can't read the release. It expects model.safetensors + config.json, but the release is a .pth with model_state_dict / model_config. for key in f over safe_open also raises TypeError on safetensors 0.7. It now reads either format, and --dinov3-config is optional for the ViT-B/16 backbone the checkpoint names, since only rope_theta and the register count are needed.

Also in f54dd08

  • turbovla_create owns the model through a raw new. Every early return nullptr leaks it along with its backend and weight buffer, so it now uses std::make_unique like the other archs.
  • The file was indented one level inside namespace vla. The weight structs carried fallbacks for tensor layouts no TurboVLA checkpoint has (unfused QKV, per-position view embeddings), check_heads ignored its tag, and the BERT pooler and DINOv3 final norm were loaded but never used. The rewrite follows vla_adapter.cpp/octo.cpp: a single graph cached by graph_cache and keyed on the BERT length, WeightLoader::fuse_gemm for QKV, gemm() so --weight-dtype applies, flash attention for the ViT when --flash-attn is on, and stats filled in.
  • Speed. Both views go through DINOv3 as one batch, and the vision features stay on the device instead of making a host round trip. The RoPE tables are uploaded once, with identity rows for CLS and the registers, so there's no per-layer prefix split and concat. Cross-attention projects the memory with only the K/V rows of in_proj. The fusion layer computes v_proj/l_proj once each instead of twice.
  • rope_dinov3_axial duplicated rope_2d in layers/rope.h. It's gone.
  • Client: state normalization uses std + 1e-6 and a normalized gripper of exactly 0 maps to +1, both as upstream (evaluation/policy.py, suite_policy.py).

GGUFs from the old converter are rejected at load with a message to re-convert, because they'd silently give wrong actions for the 11- and 14-token groups.

Not verified: a LIBERO success-rate sweep, and the OpenVINO backend.

Comment thread src/models/turbovla.cpp Outdated
#include <string>
#include <vector>

namespace vla {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The whole file is indented one level inside namespace vla. The other archs start at column 0 there.

Comment thread src/models/turbovla.cpp Outdated
ggml_tensor* cross_k = nullptr;
ggml_tensor* cross_v = nullptr;
if (cross_qkv_q->ne[0] == 3 * hidden) {
ggml_tensor* cross_qkv_m = linear(C, w.cross_qkv_w, w.cross_qkv_b, memory);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

cross_qkv_w is in_proj ([W_q; W_k; W_v]). This projects the full 535-token memory through all three blocks and then discards the Q third. A view of the K/V rows does it in one GEMM of two thirds the size.

Comment thread src/models/turbovla.cpp Outdated
int64_t vision_seq) {
const int64_t seq = cfg.max_text_length;
TurboVLATextInputs result;
result.token_ids.assign((size_t)seq, (int32_t)cfg.pad_token_id);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Upstream doesn't pad every instruction to padding_length. padding_length_by_instruction pins each LIBERO instruction to 11, 14 or 21 tokens; BERT runs over that many, and the rows after it are zeros that text_projection turns into its bias. The ACT decoder attends those rows unmasked, so the length changes the output: 5.8e-2 max |Δ| on a 14-token libero_object instruction.

Comment thread src/models/turbovla.cpp Outdated
if (!special) {
continue;
}
if (col == 0 || col == seq - 1) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seq should be the instruction's group length, not max_text_length. An instruction that fills its group (14 tokens in the 14 group) has its closing [SEP] at num_token - 1, which upstream treats as an edge token (diagonal only, position 0). Here it's treated as a sub-sentence boundary.

Comment thread src/models/turbovla.cpp Outdated
// H-EmbodVis/TurboVLA@b29ab142, text/bert.py:177-212 and
// models/text_encoder.py:80-94.
result.enhancer_self_mask[(size_t)q * seq + k] =
is_allowed && valid[(size_t)k] ? 0.0f : -FLT_MAX;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The text enhancer gets no key-padding mask upstream. The argument is commented out in models/components/transformer.py, and only ~text_self_attention_masks is passed. A padded row should attend to itself; here it's fully masked, so softmax averages every row uniformly. That alone is the 5.7e-2 on an unlisted 10-token instruction.

Comment thread scripts/convert_turbovla_to_gguf.py Outdated
raise SystemExit(f"model.safetensors not found in {ckpt}")

with safe_open(str(safetensors_path), framework="pt", device="cpu") as f:
for key in f:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The release is checkpoints/libero/turbovla_libero.pth (model_state_dict + model_config), so this path can't load it. Also, safe_open isn't iterable on safetensors 0.7 (TypeError); it needs f.keys().

Comment thread scripts/convert_turbovla_to_gguf.py Outdated
)
parser.add_argument("ckpt", type=Path, help="Path to TurboVLA checkpoint directory")
parser.add_argument("-o", "--out", type=Path, default=None, help="Output GGUF path")
parser.add_argument("--dinov3-config", type=Path, required=True,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

facebook/dinov3-vitb16-pretrain-lvd1689m is gated. Only rope_theta and the register count are needed, and the checkpoint names its backbone, so this can default for ViT-B/16.

Comment thread eval/client/vla_cpp_client.py Outdated
raise ValueError(f"{stats_path} contains invalid TurboVLA normalization ranges")

def _state_norm(state_8d, mean=state_mean, std=state_std):
return ((state_8d - mean) / (std + 1e-8)).astype(np.float32)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Upstream normalizes with std + 1e-6 (suite_policy.py, policy.py:194).

Comment thread eval/client/vla_cpp_client.py Outdated
# TurboVLA thresholds the normalized gripper instead of applying
# continuous min/max denormalization. See H-EmbodVis/TurboVLA@
# b29ab142, turbovla/evaluation/policy.py:213-221.
action[..., 6] = np.where(normalized[..., 6] > 0.0, 1.0, -1.0)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gripper_command_from_norm returns +1 for exactly 0, so this should be < 0 → -1, otherwise +1.

Comment thread src/layers/rope.h Outdated

// DINOv3 uses axial RoPE. x is [head_dim, patches, heads]; cos/sin are
// [head_dim, patches] generated from the official patch-centre coordinates.
inline ggml_tensor * rope_dinov3_axial(ggml_context * C, ggml_tensor * x,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same math as rope_2d above: the half-split rotate with per-token cos/sin.

@khanhnd61-vr

khanhnd61-vr commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Speed: this PR (d959889) vs f54dd08 / 8ae670d

vla-bench on the libero checkpoint: 2 views at 256×256, 14 tokens, 3 warmups, mean of 20 timed calls. The PR column uses the old converter's GGUF and default flags; the PR engine ignores --weight-dtype and --flash-attn.

Device This PR f54dd08, default (f32) f54dd08, fastest Speedup Fastest flags
RTX 3090 (CUDA) 14.7 ms 10.9 ms 8.0 ms 1.84× --weight-dtype bf16 --flash-attn
RTX 3060 12 GB (CUDA) 31.7 ms 24.0 ms 16.4 ms 1.93× --weight-dtype bf16 --flash-attn
Apple M4 (Metal) 87.7 ms 71.5 ms 60.5 ms 1.45× --weight-dtype bf16 --flash-attn
Intel Arc A380 (SYCL) 148.3 ms 119.4 ms 118.7 ms 1.25× --weight-dtype bf16
i5-12400F, 12 threads (CPU) 430.8 ms 412.4 ms 377.5 ms 1.14× --weight-dtype bf16 --flash-attn

On SYCL, --flash-attn is slower (123 ms with bf16). The CPU is limited by GEMM throughput, so it gains the least.

Max |Δ| of the normalized actions against upstream fp32, on the same four cases as the review:

CPU CUDA 3060 CUDA 3090 Metal SYCL
f32 (default) 9.3e-6 4.9e-3 3.4e-3 1.9e-3 8.8e-6
--weight-dtype bf16 3.7e-2 3.8e-2 3.0e-2 4.5e-2 4.0e-2

Upstream's own release precision (whole model in bf16) is 3.4e-2 from fp32 on these cases, so bf16 weights stay in that range.

OpenVINO: Core Ultra X7 358H (Arc B390 iGPU, AI Boost NPU), OpenVINO 2026.2.1

Same protocol, measured with 8ae670d. That commit fixes a bug I introduced in f54dd08: slicing the cross-attention in_proj weight made --weight-dtype bf16 fail on OpenVINO with broadcast_merge_into. The engine now slices the projection's output. On the RTX 3060 the timing and the CUDA/CPU parity above are unchanged.

Target This PR 8ae670d, default (f32) 8ae670d, fastest correct Speedup max |Δ| (fastest) Fastest flags
Arc B390 iGPU 33.6 ms 24.0 ms 21.2 ms 1.58× 7.0e-3 GGML_OPENVINO_DEVICE=GPU, --weight-dtype bf16 --flash-attn
Arc B390 iGPU, GGML_OPENVINO_GPU_PRECISION=f32 – 46.0 ms 44.9 ms – 2.4e-6 --flash-attn
OpenVINO CPU plugin 268.1 ms 256.5 ms 256.5 ms 1.05× 1.6e-6 defaults (see below)
ggml CPU backend, same host 223.4 ms 171.4 ms 167.9 ms 1.33× 8.9e-4 --flash-attn
AI Boost NPU fails to compile 105.8 ms – – 2.0 (wrong) not usable

At its default f16 compute, the iGPU is 2.4e-2 from fp32 with f32 weights and 7.0e-3 with bf16 + flash; GGML_OPENVINO_GPU_PRECISION=f32 makes it exact at about 2× the latency. Two things are wrong on this backend, and neither is a problem with TurboVLA's graph (both reproduce with the graph otherwise unchanged):

  • CPU plugin with --flash-attn: 0.94 on every case. The same graph is exact on the iGPU. This is the no-mask ScaledDotProductAttention path in ggml-openvino's flash_attn_ext translation, which llama.cpp never exercises. Not diagnosed further; leave flash attention off on the CPU plugin.
  • The NPU compiles and runs but returns wrong actions (2.0). The 261-token ViT sequence is not 16-aligned, the same limit the OpenVINO notes record for VLA-Adapter and OpenVLA-OFT. The PR's engine did not compile on the NPU at all.

Also a correction to my review: f54dd08's cross-attention no longer projects the memory with only the K/V rows of in_proj; see above.

@khanhnd61-vr
khanhnd61-vr merged commit 3f0a38a into VinRobotics:main Sep 25, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants