OpenThinker3-7B-SFT-GRPO-DE
DGurgurov/OpenThinker3-7B-SFT-GRPO-DE
Reinforcement Learning952M parameters
Tracked by OpenModelStats since September 3, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 4,210
- Total downloadsCumulative all-time download counter reported by the source platform.
- 7,112
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- +1,987(increase)
- 44.3%(up) vs baseline
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 0
- +0 in 7D
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads. #9 among tracked Reinforcement Learning models.
- #12,659
- 511(down 511 places)this week
- Parameters
- 952M
- Last source update
- Aug 21, 2026
- MomentumOpenModelStats' composite rising-model signal (0–100): a percentile blend of absolute 7-day download gains, capped percentage growth, and likes gained. See Methodology.
- 68.0
- OpenThinker3-7B-SFT-GRPO-DE gained 1,987 downloads during the last seven tracked days.
Downloads over time
Cumulative total downloads observed by OpenModelStatsLikes over time
Rank among tracked models
Lower is betterOverview
OpenThinker3-7B-SFT-GRPO-DE is a reinforcement learning model published by DGurgurov. OpenModelStats has tracked the model since Sep 3, 2026, most recently observing it 13h ago.
- Published
- Aug 21, 2026
- Last updated
- Aug 21, 2026
- Library
- —
- License
- —
- Task
- Reinforcement Learning
- Parameters
- 951,952,064
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- 284
safetensorsqwen2reasoningmultilingualdereasonxlreinforcement-learninggrpodataset:toroe/ReasonXL-SFTarxiv:2604.12378base_model:open-thoughts/OpenThinker-7Bbase_model:finetune:open-thoughts/OpenThinker-7B
Publisher
DGurgurovView publisher statistics →Derived from open-thoughts/OpenThinker-7B (reported by the source; parent not yet tracked).
Similar models
| Model | DL 30D |
|---|---|
| OphVLM-R1QiZishi | 2,984 |
| Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176JackHsieh | 334 |
| VFIG-4BXunmeiLiu | 300 |
| jatjat-project | 183 |
| beaver-7b-v1.0-costPKU-Alignment | 3,298 |
| beaver-7b-v1.0-rewardPKU-Alignment | 2,697 |