opd-baseline-repro-checkpoints
Her77/opd-baseline-repro-checkpoints
Tracked by OpenModelStats since August 27, 2026View on Hugging Face ↗
- Downloads 30DRolling recent-window download count as reported by the source platform.
- 0
- Total downloadsCumulative all-time download counter reported by the source platform.
- 0
- Gained 7DIncrease in cumulative total downloads between OpenModelStats' own observations seven days apart.
- —
- Gained 30DIncrease in cumulative total downloads between OpenModelStats' own observations roughly thirty days apart.
- —
- LikesCommunity likes reported by the source platform.
- 0
- Rank · tracked modelsPosition among all eligible models tracked by OpenModelStats, ordered by recent downloads.
- —
- –this week
- Parameters
- —
- Last source update
- Aug 27, 2026
Downloads over time
Daily gain between consecutive observationsNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-27. Charts appear as daily observations accumulate.
Likes over time
Rank among tracked models
Lower is betterNot enough tracked history yet for a chart.
OpenModelStats began tracking this model on 2026-08-27. Charts appear as daily observations accumulate.
Overview
opd-baseline-repro-checkpoints is a reinforcement learning model published by Her77. OpenModelStats has tracked the model since Aug 27, 2026, most recently observing it 3h ago.
- Published
- Aug 27, 2026
- Last updated
- Aug 27, 2026
- Library
- transformers
- License
- —
- Task
- Reinforcement Learning
- Parameters
- —
- Spaces
- —
- Derivative models
- —
- Velocity 7D/day
- —
transformerssafetensorsreinforcement-learningon-policy-distillationalfworldcurriculum-learningqwen2.5base_model:Qwen/Qwen2.5-3B-Instructbase_model:finetune:Qwen/Qwen2.5-3B-Instructendpoints_compatibleregion:us
Publisher
Her77View publisher statistics →Derived from Qwen/Qwen2.5-3B-Instruct (reported by the source; parent not yet tracked).
Similar models
| Model | DL 30D |
|---|---|
| joint-space-empowermentjamesheald | 229.7K |
| ppo-LunarLander-v3-flipKaptainKris | 144.5K |
| ppo-seals-CartPole-v0HumanCompatibleAI | 45.2K |
| ppo-Pendulum-v1HumanCompatibleAI | 17.1K |
| VisualQuality-R1-7BTianheWu | 8,110 |
| beaver-7b-v1.0-costPKU-Alignment | 6,348 |