SmolVLM2-500M-Video-Instruct
HuggingFaceTB/SmolVLM2-500M-Video-Instruct
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 1.5M
- Total downloadsCumulative all-time download counter reported by the source platform.
- 5.6M
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- —
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 172
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads. #48 among tracked Image Text to Text models.
- #308
- 118(down 118 places)this week
- Parameters
- 507M
- Last source update
- Apr 8, 2025
- It moved from #37 to #48 among tracked Image Text to Text models this week.
Downloads over time
Rolling recent-window downloads as reportedNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Likes over time
Not enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Rank among tracked models
Lower is betterNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.
Overview
SmolVLM2-500M-Video-Instruct is a image text to text model published by HuggingFaceTB. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 19 min ago.
- Published
- Feb 11, 2025
- Last updated
- Apr 8, 2025
- Library
- transformers
- License
- apache-2.0
- Task
- Image Text to Text
- Parameters
- 507,482,304
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- —
transformersonnxsafetensorssmolvlmimage-text-to-textconversationalendataset:HuggingFaceM4/the_cauldrondataset:HuggingFaceM4/Docmatixdataset:lmms-lab/LLaVA-OneVision-Datadataset:lmms-lab/M4-Instruct-Datadataset:HuggingFaceFV/finevideo
Model family
Based on base-model relationships reported by the source's metadata.
Base model
- SmolVLM-500M-Instruct166.8K
Similar models
| Model | DL 30D |
|---|---|
| SmolVLM-500M-InstructHuggingFaceTB | 166.8K |
| SmolVLM-500M-BaseHuggingFaceTB | 1,394 |
| VisionPsy-Nano-460Mqvac | 2,178 |
| VisionPsy-Nano-460M-Flashqvac | 508 |
| LFM2.5-VL-3B-OptiQ-4bitmlx-community | 1,191 |
| gemma-4-31B-it-qat-q4_0-unquantized-assistantgoogle | 19.6K |