grpo-reward-hacked-sentiment-qwen05b

rohanjain2312/grpo-reward-hacked-sentiment-qwen05b

Text Generation494M parameterstransformersLicense: apache-2.0
Tracked by OpenModelStats since September 21, 2026View on Hugging Face ↗
Downloads 30D
328
Total downloads
328
Gained 7D
—
Gained 30D
—
Likes
0
Rank · tracked models
#50,009
310(down 310 places)this week
Parameters
494M
Last source update
Sep 22, 2026
  • It moved from #15234 to #15350 among tracked Text Generation models this week.

Downloads over time

Cumulative total downloads observed by OpenModelStats

Likes over time

Rank among tracked models

Lower is better

Overview

grpo-reward-hacked-sentiment-qwen05b is a text generation model published by rohanjain2312. OpenModelStats has tracked the model since Sep 21, 2026, most recently observing it Sep 22, 2026.

Published
Sep 21, 2026
Last updated
Sep 22, 2026
Library
transformers
License
apache-2.0
Task
Text Generation
Parameters
494,032,768
Spaces
—
Derivative models
—
Velocity 7D/day
—
transformerssafetensorsqwen2text-generationgrporlhfreward-hackingreinforcement-learningconversationaldataset:stanfordnlp/imdbbase_model:Qwen/Qwen2.5-0.5B-Instructbase_model:finetune:Qwen/Qwen2.5-0.5B-Instruct

Publisher

rohanjain2312View publisher statistics →

Derived from Qwen/Qwen2.5-0.5B-Instruct (reported by the source; parent not yet tracked).

Similar models

Models similar to grpo-reward-hacked-sentiment-qwen05b
ModelDL 30D
Qwen2.5-0.5B-InstructQwen7.6M
Qwen2.5-0.5BQwen1.4M
Qwen2-0.5BQwen381.4K
Qwen2-0.5B-InstructQwen289.7K
Qwen2.5-Coder-0.5B-InstructQwen199.1K
Qwen2.5-0.5B-Instruct-4bitmlx-community60.7K