Spatial-MLLM-v1.1-Instruct-135K

Diankun/Spatial-MLLM-v1.1-Instruct-135K

Video Text To Text5.3B parameterstransformersLicense: mit
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
2,744
Total downloads
4,360
Gained 7D
Gained 30D
Likes
0
Rank · tracked models
#14,670
this week
Parameters
5.3B
Last source update
Jan 17, 2026

Downloads over time

Rolling recent-window downloads as reported

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Likes over time

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Rank among tracked models

Lower is better

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Overview

Spatial-MLLM-v1.1-Instruct-135K is a video text to text model published by Diankun. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 3h ago.

Published
Dec 30, 2025
Last updated
Jan 17, 2026
Library
transformers
License
mit
Task
Video Text To Text
Parameters
5,313,185,196
Spaces
Derivative models
Velocity 7D/day
transformerssafetensorsspatial-mllmtext-generationvideo-text-to-textarxiv:2505.23747base_model:Qwen/Qwen2.5-VL-3B-Instructbase_model:finetune:Qwen/Qwen2.5-VL-3B-Instructlicense:mitendpoints_compatibleregion:us

Publisher

DiankunView publisher statistics →

Derived from Qwen/Qwen2.5-VL-3B-Instruct (reported by the source; parent not yet tracked).

Similar models

Models similar to Spatial-MLLM-v1.1-Instruct-135K
ModelDL 30D
VideoChat3-4BMCG-NJU2,927
LLaVA-NeXT-Video-7Blmms-lab227
LLaVA-NeXT-Video-7B-hfllava-hf122.2K
LLaVA-NeXT-Video-7B-DPO-hfllava-hf1,031
LongVU_Qwen2_7BVision-CAIR4,378
VideoChat-Flash-Qwen2-7B_res448OpenGVLab1,288