SmolVLM2-500M-Video-Instruct

HuggingFaceTB/SmolVLM2-500M-Video-Instruct

Image Text to Text507M parameterstransformersLicense: apache-2.0
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
1.3M
Total downloads
7.5M
Gained 7D
+310K(increase)
4.8%(up) vs baseline
Gained 30D
—
Likes
182
+4 in 7D
Rank · tracked models
#318
1(up 1 places)this week
Parameters
507M
Last source update
Apr 8, 2025
Momentum
94.4
  • SmolVLM2-500M-Video-Instruct gained 310K downloads during the last seven tracked days.
  • It moved from #52 to #53 among tracked Image Text to Text models this week.

Downloads over time

Cumulative total downloads observed by OpenModelStats

Likes over time

Rank among tracked models

Lower is better

Overview

SmolVLM2-500M-Video-Instruct is a image text to text model published by HuggingFaceTB. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 29h ago.

Published
Feb 11, 2025
Last updated
Apr 8, 2025
Library
transformers
License
apache-2.0
Task
Image Text to Text
Parameters
507,482,304
Spaces
—
Derivative models
6
Velocity 7D/day
44.3K
transformersonnxsafetensorssmolvlmimage-text-to-textconversationalendataset:HuggingFaceM4/the_cauldrondataset:HuggingFaceM4/Docmatixdataset:lmms-lab/LLaVA-OneVision-Datadataset:lmms-lab/M4-Instruct-Datadataset:HuggingFaceFV/finevideo

Model family

Based on base-model relationships reported by the source's metadata.

Similar models

Models similar to SmolVLM2-500M-Video-Instruct
ModelDL 30D
SmolVLM-500M-InstructHuggingFaceTB139.7K
SmolVLM-500M-BaseHuggingFaceTB336
VisionPsy-Nano-460M-Flashqvac985
VisionPsy-Nano-460Mqvac576
PanoVLM-500MPanocularAI690
gemma-4-31B-it-qat-q4_0-unquantized-assistantgoogle20.5K
HuggingFaceTB/SmolVLM2-500M-Video-Instruct — Downloads, Growth & Stats | OpenModelStats