Repository navigation
Feature/add turbovla support - #29
Conversation
…int, and run the engine as one cached graph
khanhnd61-vr
left a comment
There was a problem hiding this comment.
Thanks for the port. The module structure follows upstream closely. The released LIBERO checkpoint still produces actions that diverge from the reference model, because the text path doesn't reproduce how the checkpoint pads instructions. The converter also can't read the released checkpoint. I pushed f54dd08 on top of this branch, which fixes both and restructures the engine. The inline comments cover the specifics.
How I checked
Reference: upstream H-EmbodVis/TurboVLA@b29ab142 in fp32 on CPU, with checkpoints/libero/turbovla_libero.pth loaded strictly into its module tree. DINOv3 is gated, but only its public architecture config is needed; the fine-tuned weights are inside the checkpoint. There are four seeded cases. Each has two 256×256 views and an 8-D state, plus one instruction from each of the checkpoint's padding groups (21, 11 and 14 tokens) and one instruction the checkpoint doesn't list. The table shows the max |Δ| of the normalized action chunk against the reference:
| case | padding group | this PR (CPU) | f54dd08 (CPU) | f54dd08 (CUDA) |
|---|---|---|---|---|
| libero_10 instruction, 14 tokens | 21 | 4.9e-3 | 9.5e-7 | 6.7e-4 |
| libero_goal instruction, 8 tokens | 11 | 9.0e-3 | 3.6e-7 | 1.2e-3 |
| libero_object instruction, 14 tokens | 14 | 5.8e-2 | 8.0e-7 | 8.3e-4 |
| unlisted instruction, 10 tokens | 21 | 5.7e-2 | 9.3e-6 | 4.9e-3 |
For scale, upstream's own release precision (whole model in bf16 on CUDA) is up to 3.4e-2 from fp32 on the same cases. The CUDA column is ggml-cuda's tensor-core f32 GEMMs.
Blocking
- Text padding. The ACT decoder attends every text row, padding included, because
encode_conditionconcatenates them unmasked. That makes the padded length part of the model's output. The checkpoint pins this length per training instruction (padding_length_by_instruction: 11, 14 or 21 for LIBERO). BERT runs over that many tokens, and rows after it up topadding_lengthare zeros that the text projection turns into its bias. The PR always runs BERT over 21 tokens. The converter now stores the table as token ids → length (turbovla.text_groups.*), and the engine looks the instruction up. - The special-token edge rule (
col == num_token - 1) is evaluated against 21 instead of the group length. An instruction that fills its group therefore gets its closing[SEP]treated as a sub-sentence boundary. This is case 3 above. - The text enhancer masks padded keys. Upstream passes no
key_padding_maskto it (the argument is commented out inmodels/components/transformer.py), so a padded row attends to itself. Here the row is fully masked, softmax spreads it uniformly, and that row feeds the decoder. - The converter can't read the release. It expects
model.safetensors+config.json, but the release is a.pthwithmodel_state_dict/model_config.for key in foversafe_openalso raisesTypeErroron safetensors 0.7. It now reads either format, and--dinov3-configis optional for the ViT-B/16 backbone the checkpoint names, since onlyrope_thetaand the register count are needed.
Also in f54dd08
turbovla_createowns the model through a rawnew. Every earlyreturn nullptrleaks it along with its backend and weight buffer, so it now usesstd::make_uniquelike the other archs.- The file was indented one level inside
namespace vla. The weight structs carried fallbacks for tensor layouts no TurboVLA checkpoint has (unfused QKV, per-position view embeddings),check_headsignored its tag, and the BERT pooler and DINOv3 final norm were loaded but never used. The rewrite followsvla_adapter.cpp/octo.cpp: a single graph cached bygraph_cacheand keyed on the BERT length,WeightLoader::fuse_gemmfor QKV,gemm()so--weight-dtypeapplies, flash attention for the ViT when--flash-attnis on, andstatsfilled in. - Speed. Both views go through DINOv3 as one batch, and the vision features stay on the device instead of making a host round trip. The RoPE tables are uploaded once, with identity rows for CLS and the registers, so there's no per-layer prefix split and concat. Cross-attention projects the memory with only the K/V rows of
in_proj. The fusion layer computesv_proj/l_projonce each instead of twice. rope_dinov3_axialduplicatedrope_2dinlayers/rope.h. It's gone.- Client: state normalization uses
std + 1e-6and a normalized gripper of exactly 0 maps to +1, both as upstream (evaluation/policy.py,suite_policy.py).
GGUFs from the old converter are rejected at load with a message to re-convert, because they'd silently give wrong actions for the 11- and 14-token groups.
Not verified: a LIBERO success-rate sweep, and the OpenVINO backend.
| #include <string> | ||
| #include <vector> | ||
|
|
||
| namespace vla { |
There was a problem hiding this comment.
The whole file is indented one level inside namespace vla. The other archs start at column 0 there.
| ggml_tensor* cross_k = nullptr; | ||
| ggml_tensor* cross_v = nullptr; | ||
| if (cross_qkv_q->ne[0] == 3 * hidden) { | ||
| ggml_tensor* cross_qkv_m = linear(C, w.cross_qkv_w, w.cross_qkv_b, memory); |
There was a problem hiding this comment.
cross_qkv_w is in_proj ([W_q; W_k; W_v]). This projects the full 535-token memory through all three blocks and then discards the Q third. A view of the K/V rows does it in one GEMM of two thirds the size.
| int64_t vision_seq) { | ||
| const int64_t seq = cfg.max_text_length; | ||
| TurboVLATextInputs result; | ||
| result.token_ids.assign((size_t)seq, (int32_t)cfg.pad_token_id); |
There was a problem hiding this comment.
Upstream doesn't pad every instruction to padding_length. padding_length_by_instruction pins each LIBERO instruction to 11, 14 or 21 tokens; BERT runs over that many, and the rows after it are zeros that text_projection turns into its bias. The ACT decoder attends those rows unmasked, so the length changes the output: 5.8e-2 max |Δ| on a 14-token libero_object instruction.
| if (!special) { | ||
| continue; | ||
| } | ||
| if (col == 0 || col == seq - 1) { |
There was a problem hiding this comment.
seq should be the instruction's group length, not max_text_length. An instruction that fills its group (14 tokens in the 14 group) has its closing [SEP] at num_token - 1, which upstream treats as an edge token (diagonal only, position 0). Here it's treated as a sub-sentence boundary.
| // H-EmbodVis/TurboVLA@b29ab142, text/bert.py:177-212 and | ||
| // models/text_encoder.py:80-94. | ||
| result.enhancer_self_mask[(size_t)q * seq + k] = | ||
| is_allowed && valid[(size_t)k] ? 0.0f : -FLT_MAX; |
There was a problem hiding this comment.
The text enhancer gets no key-padding mask upstream. The argument is commented out in models/components/transformer.py, and only ~text_self_attention_masks is passed. A padded row should attend to itself; here it's fully masked, so softmax averages every row uniformly. That alone is the 5.7e-2 on an unlisted 10-token instruction.
| raise SystemExit(f"model.safetensors not found in {ckpt}") | ||
|
|
||
| with safe_open(str(safetensors_path), framework="pt", device="cpu") as f: | ||
| for key in f: |
There was a problem hiding this comment.
The release is checkpoints/libero/turbovla_libero.pth (model_state_dict + model_config), so this path can't load it. Also, safe_open isn't iterable on safetensors 0.7 (TypeError); it needs f.keys().
| ) | ||
| parser.add_argument("ckpt", type=Path, help="Path to TurboVLA checkpoint directory") | ||
| parser.add_argument("-o", "--out", type=Path, default=None, help="Output GGUF path") | ||
| parser.add_argument("--dinov3-config", type=Path, required=True, |
There was a problem hiding this comment.
facebook/dinov3-vitb16-pretrain-lvd1689m is gated. Only rope_theta and the register count are needed, and the checkpoint names its backbone, so this can default for ViT-B/16.
| raise ValueError(f"{stats_path} contains invalid TurboVLA normalization ranges") | ||
|
|
||
| def _state_norm(state_8d, mean=state_mean, std=state_std): | ||
| return ((state_8d - mean) / (std + 1e-8)).astype(np.float32) |
There was a problem hiding this comment.
Upstream normalizes with std + 1e-6 (suite_policy.py, policy.py:194).
| # TurboVLA thresholds the normalized gripper instead of applying | ||
| # continuous min/max denormalization. See H-EmbodVis/TurboVLA@ | ||
| # b29ab142, turbovla/evaluation/policy.py:213-221. | ||
| action[..., 6] = np.where(normalized[..., 6] > 0.0, 1.0, -1.0) |
There was a problem hiding this comment.
gripper_command_from_norm returns +1 for exactly 0, so this should be < 0 → -1, otherwise +1.
|
|
||
| // DINOv3 uses axial RoPE. x is [head_dim, patches, heads]; cos/sin are | ||
| // [head_dim, patches] generated from the official patch-centre coordinates. | ||
| inline ggml_tensor * rope_dinov3_axial(ggml_context * C, ggml_tensor * x, |
There was a problem hiding this comment.
Same math as rope_2d above: the half-split rotate with per-token cos/sin.
Speed: this PR (d959889) vs f54dd08 / 8ae670d
On SYCL, Max |Δ| of the normalized actions against upstream fp32, on the same four cases as the review:
Upstream's own release precision (whole model in bf16) is 3.4e-2 from fp32 on these cases, so bf16 weights stay in that range. OpenVINO: Core Ultra X7 358H (Arc B390 iGPU, AI Boost NPU), OpenVINO 2026.2.1Same protocol, measured with 8ae670d. That commit fixes a bug I introduced in f54dd08: slicing the cross-attention
At its default f16 compute, the iGPU is 2.4e-2 from fp32 with f32 weights and 7.0e-3 with bf16 + flash;
Also a correction to my review: f54dd08's cross-attention no longer projects the memory with only the K/V rows of |
…eight, so bf16 runs on ggml-openvino
What
Add TurboVLA support, including GGUF conversion, model loading and inference,
direct-evaluation client handling, converter remap coverage, and documentation.
Why
TurboVLA provides a compact DINOv3 + BERT VLA policy. This adds a native
vla.cpp inference path without requiring Python at runtime.
Verified
-Wall -Wextra(first-party code)ctestpassesvla_predict_checkdiff), or the change ismeant to move it and a LIBERO sweep is below
Additional checks:
python3 tests/py/test_converters.pypasses.python3 tests/py/test_bindings.pypasses.Archs and backends tested:
GGML_CUDA=OFF): compile and unit-test coverage.