Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176

JackHsieh/Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176

Reinforcement Learning4B parametersLicense: apache-2.0
Tracked by OpenModelStats since September 19, 2026View on Hugging Face ↗
Downloads 30D
334
Total downloads
334
Gained 7D
—
Gained 30D
—
Likes
0
Rank · tracked models
#49,358
570(down 570 places)this week
Parameters
4B
Last source update
Sep 26, 2026
  • It moved from #118 to #126 among tracked Reinforcement Learning models this week.

Downloads over time

Cumulative total downloads observed by OpenModelStats

Likes over time

Rank among tracked models

Lower is better

Overview

Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 is a reinforcement learning model published by JackHsieh. OpenModelStats has tracked the model since Sep 19, 2026, most recently observing it Sep 27, 2026.

Published
Sep 19, 2026
Last updated
Sep 26, 2026
Library
—
License
apache-2.0
Task
Reinforcement Learning
Parameters
4,022,468,096
Spaces
—
Derivative models
—
Velocity 7D/day
—
safetensorsqwen3prestarreinforcement-learningrlooverlbase_model:Qwen/Qwen3-4B-Instruct-2507base_model:finetune:Qwen/Qwen3-4B-Instruct-2507license:apache-2.0region:us

Publisher

JackHsiehView publisher statistics →

Derived from Qwen/Qwen3-4B-Instruct-2507 (reported by the source; parent not yet tracked).

Similar models

Models similar to Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176
ModelDL 30D
VFIG-4BXunmeiLiu300
beaver-7b-v1.0-costPKU-Alignment3,298
beaver-7b-v1.0-rewardPKU-Alignment2,697
OphVLM-R1QiZishi2,984
inf-retriever-v1-proinfly2,350
Open-Reasoner-Zero-7BOpen-Reasoner-Zero1,811
JackHsieh/Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Downloads, Growth & Stats | OpenModelStats