cogvlm2-llama3-caption

zai-org/cogvlm2-llama3-caption

Video Text To Text12.5B parameterstransformersLicense: other
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
385
Total downloads
309.2K
Gained 7D
Gained 30D
Likes
119
Rank · tracked models
#28,920
this week
Parameters
12.5B
Last source update
May 14, 2025

Downloads over time

Cumulative total downloads observed by OpenModelStats

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Likes over time

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Rank among tracked models

Lower is better

Not enough tracked history yet for a chart.

OpenModelStats began tracking this model on 2026-08-24. Charts appear as daily observations accumulate.

Overview

cogvlm2-llama3-caption is a video text to text model published by zai-org. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it 3h ago.

Published
Sep 18, 2024
Last updated
May 14, 2025
Library
transformers
License
other
Task
Video Text To Text
Parameters
12,507,532,544
Spaces
Derivative models
Velocity 7D/day
transformerssafetensorstext-generationvideo-text-to-textcustom_codeenarxiv:2408.06072base_model:meta-llama/Llama-3.1-8B-Instructbase_model:finetune:meta-llama/Llama-3.1-8B-Instructlicense:otherregion:us

Publisher

zai-orgView publisher statistics →

Derived from meta-llama/Llama-3.1-8B-Instruct (reported by the source; parent not yet tracked).

Similar models

Models similar to cogvlm2-llama3-caption
ModelDL 30D
MOSS-VL-Instruct-0408OpenMOSS-Team22.5K
MOSS-VL-Instruct-0708OpenMOSS-Team1,409
MOSS-VL-RealtimeOpenMOSS-Team1,331
MOSS-VL-Realtime-FP8OpenMOSS-Team422
MOSS-VL-Instruct-0708-NF4OpenMOSS-Team122
MOSS-VL-Instruct-0708-FP8OpenMOSS-Team119