Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176

JackHsieh/Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176

Reinforcement Learning4B parametersLicense: apache-2.0
Tracked by OpenModelStats since September 19, 2026View on Hugging Face ↗
Downloads 30D
334
Total downloads
334
Gained 7D
—
Gained 30D
—
Likes
0
Rank · tracked models
#48,370
341(down 341 places)this week
Parameters
4B
Last source update
Sep 26, 2026
  • It moved from #115 to #114 among tracked Reinforcement Learning models this week.

Downloads over time

Daily gain between consecutive observations

Likes over time

Rank among tracked models

Lower is better

Overview

Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 is a reinforcement learning model published by JackHsieh. OpenModelStats has tracked the model since Sep 19, 2026, most recently observing it Sep 27, 2026.

Published
Sep 19, 2026
Last updated
Sep 26, 2026
Library
—
License
apache-2.0
Task
Reinforcement Learning
Parameters
4,022,468,096
Spaces
—
Derivative models
—
Velocity 7D/day
—
safetensorsqwen3prestarreinforcement-learningrlooverlbase_model:Qwen/Qwen3-4B-Instruct-2507base_model:finetune:Qwen/Qwen3-4B-Instruct-2507license:apache-2.0region:us

Publisher

JackHsiehView publisher statistics →

Derived from Qwen/Qwen3-4B-Instruct-2507 (reported by the source; parent not yet tracked).

Similar models

Models similar to Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176
ModelDL 30D
VFIG-4BXunmeiLiu300
beaver-7b-v1.0-costPKU-Alignment3,295
beaver-7b-v1.0-rewardPKU-Alignment2,722
OphVLM-R1QiZishi2,976
inf-retriever-v1-proinfly2,350
Open-Reasoner-Zero-7BOpen-Reasoner-Zero1,811
JackHsieh/Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Downloads, Growth & Stats | OpenModelStats