You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Research checkpoints from INSAIT. Request access for research use.

Log in or Sign Up to review the conditions and access this model content.

Qwen3-VL-3D spatial-alignment checkpoints (binned / special-token variant)

Qwen3VL3D = Qwen/Qwen3-VL-2B-Instruct + a MoGe depth encoder and a depth projector. Spatial values use binned special tokens: <dist0000..1023> (distances) and <loc0000..1023> (pixel locations), added on top of the 151669-token base vocab (padded to 153728). Distances use normal->uniform bin edges over [-3, 5], power 1.05.

NOTE: these are barrel training checkpoints (FSDP trainer state_dict), NOT AutoModel.from_pretrained-loadable. Load them through the barrel Qwen3VL3D pipeline via pretrain_config.state_dict.checkpoint_path. Each stage folder ships its full_config.yaml, model_config.yaml, and the git hash/diff/branch needed to reproduce.

Stages

folder run steps num_rel_err_median num_acc_10pct
stage1_align_special_tokens/ depth-only alignment (trains depth projector + new token rows) 2000 ~0.068 (2nd-half) ~0.557
stage2_oxe_finetune/ OXE-mix finetune from stage 1 (llava/egoexo4d/bridge) 10000 ~0.047 (2nd-half), 0.040 final ~0.611, 0.639 final

Judge spatial accuracy by num_rel_err_median / num_acc_10pct, not by 3d_loss.

Branch: anna_qwen3vl_3d_basic. See each folder's git.hash / git.diff for the exact code.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support