Reinforcement Learning Models

Reinforcement learning models and agents, including trained policy models.

Top-10 combined DL 30D
286.5K
Tasks directory position
See all tasks
Growth history
Available
Last refresh
Sep 22, 2026

Trending now

Most downloaded

ModelDL 30D7D
ppo-seals-CartPole-v0HumanCompatibleAI166.3K1.4%(up)
ppo-Pendulum-v1HumanCompatibleAI44.8K1.0%(up)
Tifa-Deepsex-14b-CoT-GGUFmradermacher20.8K20.0%(up)
NoveltyAdapt-policieshelenlu9420.7K—
sac-BipedalWalkerHardcore-v3sb38,772—
decision-transformer-gym-hopper-mediumedbeeching6,4890.6%(up)
Aztec-Coder-4B-i1-GGUFmradermacher5,282—
VisualQuality-R1-7BTianheWu5,0441.0%(up)
OpenThinker3-7B-SFT-GRPO-DEDGurgurov4,21044.3%(up)
Tifa-DeepsexV3-14b-GGUF-Q6ValueFX95074,0690%

Fastest growing

Downloads gained over the last seven tracked days.

ModelDL 30D7D
ppo-seals-CartPole-v0HumanCompatibleAI166.3K+24.4K
Tifa-Deepsex-14b-CoT-GGUFmradermacher20.8K+17.4K
ppo-Pendulum-v1HumanCompatibleAI44.8K+6,656
OPD-Aha-9B-i1-GGUFmradermacher2,999+2,614
OPD-Aha-4B-i1-GGUFmradermacher2,856+2,451
OpenThinker3-7B-SFT-GRPO-DEDGurgurov4,210+1,987
newtnicklashansen2,249+1,892
VisualQuality-R1-7BTianheWu5,044+1,747
Tifa-DeepsexV2-7b-MGRPO-GGUF-Q4ValueFX95072,677+1,571
ScreenHighlighterRL-2B-i1-GGUFmradermacher2,003+1,570

New & notable

Published within the last 90 days with meaningful traction.

ModelDL 30D7D
NoveltyAdapt-policieshelenlu9420.7K—
Aztec-Coder-4B-i1-GGUFmradermacher5,282—
OpenThinker3-7B-SFT-GRPO-DEDGurgurov4,210+1,987
OPD-Aha-9B-i1-GGUFmradermacher2,999+2,614
OphVLM-R1QiZishi2,984—
CNY-7B-i1-GGUFmradermacher2,883+269
OPD-Aha-4B-i1-GGUFmradermacher2,856+2,451
EventMemAgent-8B-i1-GGUFmradermacher2,838+243
Kronumos-Kairos-v2-i1-GGUFmradermacher2,811—
Evidence-RL-9B-i1-GGUFmradermacher2,800—