Sa2VA-Qwen3-VL-4B

ByteDance/Sa2VA-Qwen3-VL-4B

Image Text to Text5.1B parameterstransformersLicense: apache-2.0
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
745
Total downloads
20.5K
Gained 7D
Gained 30D
Likes
19
Rank · tracked models
#26,330
23766(down 23766 places)this week
Parameters
5.1B
Last source update
Oct 21, 2025
  • It moved from #445 to #2665 among tracked Image Text to Text models this week.

Downloads over time

Cumulative total downloads observed by OpenModelStats

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Likes over time

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Rank among tracked models

Lower is better

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Overview

Sa2VA-Qwen3-VL-4B is a image text to text model published by ByteDance. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 6h ago.

Published
Oct 21, 2025
Last updated
Oct 21, 2025
Library
transformers
License
apache-2.0
Task
Image Text to Text
Parameters
5,057,072,434
Spaces
Derivative models
Velocity 7D/day
transformerssafetensorssa2va_chatfeature-extractionSa2VAcustom_codeimage-text-to-textconversationalmultilingualarxiv:2501.04001base_model:OpenGVLab/InternVL3-8Bbase_model:merge:OpenGVLab/InternVL3-8B

Publisher

ByteDanceView publisher statistics →

Derived from OpenGVLab/InternVL3-8B (reported by the source; parent not yet tracked).

Similar models