grpo-reward-hacking-fixed-kl-qwen05b
rohanjain2312/grpo-reward-hacking-fixed-kl-qwen05b
Tracked by OpenModelStats since September 21, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 340
- Total downloadsCumulative all-time download counter reported by the source platform.
- 340
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- —
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 0
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads. #15041 among tracked Text Generation models.
- #49,307
- 341(down 341 places)this week
- Parameters
- 494M
- Last source update
- Sep 22, 2026
- It moved from #14908 to #15041 among tracked Text Generation models this week.
Downloads over time
Cumulative total downloads observed by OpenModelStatsLikes over time
Rank among tracked models
Lower is betterOverview
grpo-reward-hacking-fixed-kl-qwen05b is a text generation model published by rohanjain2312. OpenModelStats has tracked the model since Sep 21, 2026, most recently observing it Sep 22, 2026.
- Published
- Sep 21, 2026
- Last updated
- Sep 22, 2026
- Library
- transformers
- License
- apache-2.0
- Task
- Text Generation
- Parameters
- 494,032,768
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- —
transformerssafetensorsqwen2text-generationgrporlhfreward-hackingreinforcement-learningconversationaldataset:stanfordnlp/imdbbase_model:Qwen/Qwen2.5-0.5B-Instructbase_model:finetune:Qwen/Qwen2.5-0.5B-Instruct
Publisher
rohanjain2312View publisher statistics →Derived from Qwen/Qwen2.5-0.5B-Instruct (reported by the source; parent not yet tracked).
Similar models
| Model | DL 30D |
|---|---|
| Qwen2.5-0.5B-InstructQwen | 7.6M |
| Qwen2.5-0.5BQwen | 1.4M |
| Qwen2-0.5BQwen | 381.4K |
| Qwen2-0.5B-InstructQwen | 289.7K |
| Qwen2.5-Coder-0.5B-InstructQwen | 199.1K |
| Qwen2.5-0.5B-Instruct-4bitmlx-community | 60.7K |