Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176
JackHsieh/Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176
Tracked by OpenModelStats since September 19, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 334
- Total downloadsCumulative all-time download counter reported by the source platform.
- 334
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- —
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 0
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads. #126 among tracked Reinforcement Learning models.
- #49,358
- 570(down 570 places)this week
- Parameters
- 4B
- Last source update
- Sep 26, 2026
- It moved from #118 to #126 among tracked Reinforcement Learning models this week.
Downloads over time
Cumulative total downloads observed by OpenModelStatsLikes over time
Rank among tracked models
Lower is betterOverview
Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 is a reinforcement learning model published by JackHsieh. OpenModelStats has tracked the model since Sep 19, 2026, most recently observing it Sep 27, 2026.
- Published
- Sep 19, 2026
- Last updated
- Sep 26, 2026
- Library
- —
- License
- apache-2.0
- Task
- Reinforcement Learning
- Parameters
- 4,022,468,096
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- —
safetensorsqwen3prestarreinforcement-learningrlooverlbase_model:Qwen/Qwen3-4B-Instruct-2507base_model:finetune:Qwen/Qwen3-4B-Instruct-2507license:apache-2.0region:us
Publisher
JackHsiehView publisher statistics →Derived from Qwen/Qwen3-4B-Instruct-2507 (reported by the source; parent not yet tracked).
Similar models
| Model | DL 30D |
|---|---|
| VFIG-4BXunmeiLiu | 300 |
| beaver-7b-v1.0-costPKU-Alignment | 3,298 |
| beaver-7b-v1.0-rewardPKU-Alignment | 2,697 |
| OphVLM-R1QiZishi | 2,984 |
| inf-retriever-v1-proinfly | 2,350 |
| Open-Reasoner-Zero-7BOpen-Reasoner-Zero | 1,811 |