v4_savebestearly_sft_qwen7B_25percent_lr_1e4_bptt_offset

agurung/v4_savebestearly_sft_qwen7B_25percent_lr_1e4_bptt_offset

Text Generation8.2B parameterstransformers
Tracked by OpenModelStats since September 19, 2026View on Hugging Face ↗
Downloads 30D
677
Total downloads
785
Gained 7D
—
Gained 30D
—
Likes
0
Rank · tracked models
#38,154
293(down 293 places)this week
Parameters
8.2B
Last source update
Oct 1, 2026
Momentum
0.0
  • It moved from #10420 to #10500 among tracked Text Generation models this week.

Downloads over time

Daily gain between consecutive observations

Likes over time

Rank among tracked models

Lower is better

Overview

v4_savebestearly_sft_qwen7B_25percent_lr_1e4_bptt_offset is a text generation model published by agurung. OpenModelStats has tracked the model since Sep 19, 2026, most recently observing it Oct 2, 2026.

Published
Aug 17, 2025
Last updated
Oct 1, 2026
Library
transformers
License
—
Task
Text Generation
Parameters
8,226,674,176
Spaces
—
Derivative models
—
Velocity 7D/day
—
transformerssafetensorsqwen2text-generationgenerated_from_traineropen-r1trlsftconversationaldataset:twenty_five_percentbase_model:Qwen/Qwen2.5-7B-Instructbase_model:finetune:Qwen/Qwen2.5-7B-Instruct

Publisher

agurungView publisher statistics →

Derived from Qwen/Qwen2.5-7B-Instruct (reported by the source; parent not yet tracked).

Similar models

Models similar to v4_savebestearly_sft_qwen7B_25percent_lr_1e4_bptt_offset
ModelDL 30D
AquilaMed-RLBAAI312
Qwen3-Embedding-8B-AWQ-INT4drawais4,592
Qwen3-Reranker-8B-AWQ-INT4drawais791
DeepSeek-R1-Distill-Llama-8B-unsloth-bnb-4bitunsloth2,715
Meta-Llama-3.1-8B-Instruct-unsloth-bnb-4bitunsloth26.5K
Llama-3.1-8B-Instruct-unsloth-bnb-4bitunsloth14.7K