Spatial-MLLM-v1.1-Instruct-135K
Diankun/Spatial-MLLM-v1.1-Instruct-135K
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 2,744
- Total downloadsCumulative all-time download counter reported by the source platform.
- 4,360
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- —
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 0
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads. #14 among tracked Video Text To Text models.
- #14,670
- –this week
- Parameters
- 5.3B
- Last source update
- Jan 17, 2026
Downloads over time
Cumulative total downloads observed by OpenModelStatsNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Likes over time
Not enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Rank among tracked models
Lower is betterNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Overview
Spatial-MLLM-v1.1-Instruct-135K is a video text to text model published by Diankun. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 24 min ago.
- Published
- Dec 30, 2025
- Last updated
- Jan 17, 2026
- Library
- transformers
- License
- mit
- Task
- Video Text To Text
- Parameters
- 5,313,185,196
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- —
transformerssafetensorsspatial-mllmtext-generationvideo-text-to-textarxiv:2505.23747base_model:Qwen/Qwen2.5-VL-3B-Instructbase_model:finetune:Qwen/Qwen2.5-VL-3B-Instructlicense:mitendpoints_compatibleregion:us
Publisher
DiankunView publisher statistics →Derived from Qwen/Qwen2.5-VL-3B-Instruct (reported by the source; parent not yet tracked).
Similar models
| Model | DL 30D |
|---|---|
| VideoChat3-4BMCG-NJU | 2,927 |
| LLaVA-NeXT-Video-7Blmms-lab | 227 |
| LLaVA-NeXT-Video-7B-hfllava-hf | 122.2K |
| LLaVA-NeXT-Video-7B-DPO-hfllava-hf | 1,031 |
| LongVU_Qwen2_7BVision-CAIR | 4,378 |
| VideoChat-Flash-Qwen2-7B_res448OpenGVLab | 1,288 |