ClipTagger-12b

inference-net/ClipTagger-12b

Image Text to Text12.2B parametersLicense: apache-2.0
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
30
Total downloads
6,491
Gained 7D
Gained 30D
Likes
59
Rank · tracked models
#32,293
this week
Parameters
12.2B
Last source update
Aug 14, 2025

Downloads over time

Cumulative total downloads observed by OpenModelStats

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Likes over time

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Rank among tracked models

Lower is better

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Overview

ClipTagger-12b is a image text to text model published by inference-net. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 1h ago.

Published
Aug 13, 2025
Last updated
Aug 14, 2025
Library
License
apache-2.0
Task
Image Text to Text
Parameters
12,189,561,456
Spaces
Derivative models
Velocity 7D/day
safetensorsgemma3VLMvideo-understandingimage-captioninggemmajson-modestructured-outputvideo-analysisimage-text-to-textconversationalen

Similar models

Models similar to ClipTagger-12b
ModelDL 30D
gemma-3-12b-it-FP8-dynamicRedHatAI2,418
gemma-3-12b-itgoogle1.1M
gemma-3-12b-it-int4-awqgaunernst154.3K
gemma-3-12b-it-qat-q4_0-unquantizedgoogle97.3K
gemma-3-12b-ptgoogle31.6K
gemma-3-12b-it-qat-q4_0-unquantizedLightricks27.4K