vit-gpt2-image-captioning

nlpconnect/vit-gpt2-image-captioning

Image to TexttransformersLicense: apache-2.0
Tracked by OpenModelStats since August 24, 2026View on Hugging Face ↗
Downloads 30D
85.9K
Total downloads
60.4M
Gained 7D
+19.8K(increase)
0.0%(up) vs baseline
Gained 30D
—
Likes
935
+0 in 7D
Rank · tracked models
#2,963
4(down 4 places)this week
Parameters
—
Last source update
Feb 27, 2023
Momentum
56.1
  • vit-gpt2-image-captioning gained 19.8K downloads during the last seven tracked days.

Downloads over time

Cumulative total downloads observed by OpenModelStats

Likes over time

Rank among tracked models

Lower is better

Overview

vit-gpt2-image-captioning is a image to text model published by nlpconnect. OpenModelStats has tracked the model since Aug 24, 2026, most recently observing it Sep 22, 2026.

Published
Mar 2, 2022
Last updated
Feb 27, 2023
Library
transformers
License
apache-2.0
Task
Image to Text
Parameters
—
Spaces
—
Derivative models
1
Velocity 7D/day
2,828
transformerspytorchvision-encoder-decoderimage-text-to-textimage-to-textimage-captioningdoi:10.57967/hf/0222license:apache-2.0endpoints_compatibleregion:us

Model family

Based on base-model relationships reported by the source's metadata.

Derivative models

Similar models

Models similar to vit-gpt2-image-captioning
ModelDL 30D
blip-image-captioning-baseSalesforce1.7M
GLM-OCRzai-org1.7M
manga-ocr-basekha-white1.1M
PP-OCRv5_server_detPaddlePaddle665.2K
blip-image-captioning-largeSalesforce561K
UVDocPaddlePaddle462K