Spatial-Interactor

Spatial-Interactor Qwen3-VL-8B

Learning Spatial Reasoning through Interaction with the Observable Physical World

Project page Paper PDF Code Dataset

This is the full-parameter BF16 Spatial-Interactor checkpoint based on Qwen/Qwen3-VL-8B-Instruct. It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories.

The privileged transition trace is used only during training. At inference, this checkpoint takes the same image/video and question inputs as its base model, with no extra trace, reward model, or teacher branch.

Overview

Spatial-Interactor overview: interaction trajectories, three-level curriculum, SFT and OPD, and spatial reasoning results

Presentation

Load

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-8B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto",
)

Use the base model's image/video input format. No privileged trace or additional teacher is needed for inference. Weights, tokenizer, processor, and chat template are included. See the training guide for SFT and OPD.

Citation

For citation, use the project BibTeX.

License

This checkpoint is released under Apache-2.0, following the base model license. Users must also comply with licenses and terms governing input datasets and media.

Downloads last month
131
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kagakouko/Spatial-Interactor-Qwen3-VL-8B

Finetuned
(603)
this model
Quantizations
1 model

Collection including kagakouko/Spatial-Interactor-Qwen3-VL-8B