Molmo2-VideoPoint-4B

allenai/Molmo2-VideoPoint-4B

Video Text To Text4.9B parameterstransformersLicense: apache-2.0
Tracked by OpenModelStats since August 28, 2026View on Hugging Face ↗
Downloads 30D
169
Total downloads
47.2K
Gained 7D
Gained 30D
Likes
22
Rank · tracked models
this week
Parameters
4.9B
Last source update
Dec 16, 2025

Downloads over time

Cumulative total downloads observed by OpenModelStats

Likes over time

Rank among tracked models

Lower is better

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-28. Charts appear as daily observations accumulate.

Overview

Molmo2-VideoPoint-4B is a video text to text model published by allenai. OpenModelStats has tracked the model since Aug 28, 2026, most recently observing it 1h ago.

Published
Dec 16, 2025
Last updated
Dec 16, 2025
Library
transformers
License
apache-2.0
Task
Video Text To Text
Parameters
4,850,869,200
Spaces
Derivative models
Velocity 7D/day
transformerssafetensorsmolmo2image-text-to-textmultimodalolmomolmovideo-text-to-textcustom_codeendataset:allenai/Molmo2-VideoPointdataset:allenai/pixmo-points

Publisher

allenaiView publisher statistics →

Derived from Qwen/Qwen3-4B-Instruct-2507 (reported by the source; parent not yet tracked).

Similar models

Models similar to Molmo2-VideoPoint-4B
ModelDL 30D
VideoChat3-4BMCG-NJU2,613
Spatial-MLLM-v1.1-Instruct-135KDiankun1,368
LLaVA-NeXT-Video-7Blmms-lab194
LLaVA-NeXT-Video-7B-hfllava-hf94.2K
LLaVA-NeXT-Video-7B-DPO-hfllava-hf988
LongVU_Qwen2_7BVision-CAIR679