inference-optimization AI Models

Tracked by OpenModelStats since Aug 23, 2026 · View on Hugging Face ↗

Tracked models
103
Downloads 30D
100.8K
Total downloads
526.2K
Total likes
131
Gained 7D
31.5K

Most downloaded

  1. 1Llama-3.2-0.5B-InstructJun 3, 202620.8K
  2. 2Qwen3-1.6B-A0.9BMay 24, 202614.8K
  3. 3Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1Jun 10, 20267,417
  4. 4Qwen3-8B-speculators.peagle-qwen3arch-ckpt4Jun 16, 20267,398
  5. 5Qwen3-Coder-Next.w4a16Mar 12, 20266,791
  6. 6DeepSeek-V3-debug-empty-FP8_DYNAMICJan 23, 20265,304
  7. 7GLM-5.2-0.8B-A0.8BJul 2, 20265,152
  8. 8DSV4-tiny-emptySep 14, 20264,992
  9. 9Qwen3.8-1.0B-A0.6BAug 12, 20264,529
  10. 10NemotronH-0.3B-A0.3BSep 14, 20263,782

Fastest growing

  1. 1Llama-3.2-0.5B-Instruct+7,405 in 7D20.8K
  2. 2Qwen3-1.6B-A0.9B+4,898 in 7D14.8K
  3. 3Qwen3-8B-speculators.peagle-qwen3arch-ckpt4+2,185 in 7D7,398
  4. 4Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1+2,178 in 7D7,417
  5. 5DSV4-tiny-empty+1,806 in 7D4,992
  6. 6Kimi-K3-0.40B-MXFP4+1,795 in 7D3,278
  7. 7NemotronH-0.3B-A0.3B+1,762 in 7D3,782
  8. 8GLM-5.2-0.8B-A0.8B+1,760 in 7D5,152
  9. 9qwe3-8B-selfdistill-v3-fresh3ep_ckpt0+1,640 in 7D2,155
  10. 10DeepSeek-V3-debug-empty-FP8_DYNAMIC+1,583 in 7D5,304

Recently updated

  1. 1Inkling-0.92B-MTPOct 7, 20260
  2. 2GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQOct 7, 20260
  3. 3GLM-5-0.88B-MTPOct 7, 20260
  4. 4DeepSeek-V3-0.86B-MTP-FP8-DynamicOct 7, 20260
  5. 5DeepSeek-V3-0.86B-MTPOct 7, 20260
  6. 6GLM-4.7-Flash-0.82B-MTP-FP8-DynamicOct 7, 20260
  7. 7GLM-4.7-Flash-0.82B-MTPOct 7, 20260
  8. 8GLM-4.5-Air-0.72B-MTP-FP8-DynamicOct 7, 20260
  9. 9GLM-4.5-Air-0.72B-MTPOct 7, 20260
  10. 10gemma-4-31b-it-speculator.dflash.latestrecipe.ckpt2Oct 6, 202617

All models

All tracked models published by inference-optimization
ModelDL 30D7D
Llama-3.2-0.5B-Instructinference-optimization/llama-3.2-0.5b-instruct23.7K+7,405
Qwen3-1.6B-A0.9Binference-optimization/qwen3-1.6b-a0.9b15.7K+4,898
Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1inference-optimization/qwen3-8b-from-qwen3-8b_regen-speculators.eagle3-qwen3arch-ckpt18,043+2,178
Qwen3-8B-speculators.peagle-qwen3arch-ckpt4inference-optimization/qwen3-8b-speculators.peagle-qwen3arch-ckpt48,021+2,185
Qwen3-Coder-Next.w4a16inference-optimization/qwen3-coder-next.w4a166,744+290
DeepSeek-V3-debug-empty-FP8_DYNAMICinference-optimization/deepseek-v3-debug-empty-fp8_dynamic5,645+1,583
GLM-5.2-0.8B-A0.8Binference-optimization/glm-5.2-0.8b-a0.8b5,506+1,760
DSV4-tiny-emptyinference-optimization/dsv4-tiny-empty5,401+1,806
Qwen3.8-1.0B-A0.6Binference-optimization/qwen3.8-1.0b-a0.6b4,798+1,167
NemotronH-0.3B-A0.3Binference-optimization/nemotronh-0.3b-a0.3b3,543+1,762
Kimi-K3-0.40B-MXFP4inference-optimization/kimi-k3-0.40b-mxfp43,247+1,795
qwe3-8B-selfdistill-v3-fresh3ep_ckpt0inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt02,126+1,640
Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt19inference-optimization/qwen3-8b-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt191,730+468
Qwen3-8B-speculator.dspark.swa.dpace.fullvocab.muon.2048anc.combdatav4-q235b-instr-v1-ckpt2inference-optimization/qwen3-8b-speculator.dspark.swa.dpace.fullvocab.muon.2048anc.combdatav4-q235b-instr-v1-ckpt21,729+475
Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt16inference-optimization/qwen3-8b-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt161,712+462
GLM-5.3-Flash-0.1B-A0.1Binference-optimization/glm-5.3-flash-0.1b-a0.1b1,630+743
Kimi-K3-0.40Binference-optimization/kimi-k3-0.40b974+276
DeepSeek-V3-debug-emptyinference-optimization/deepseek-v3-debug-empty968+320
Nemotron-3.5-Lightning-1.4B-A0.1B-MTPinference-optimization/nemotron-3.5-lightning-1.4b-a0.1b-mtp702+67
Phi-3.5-MoE-0.8B-A0.2Binference-optimization/phi-3.5-moe-0.8b-a0.2b553—
Qwen3.8-Flash-Next-0.2B-A0.2Binference-optimization/qwen3.8-flash-next-0.2b-a0.2b526+0
Qwen3-30B-A3B-Thinking-2507-REAP-25-uniforminference-optimization/qwen3-30b-a3b-thinking-2507-reap-25-uniform253—
Qwen3.6-8B-A1.6Binference-optimization/qwen3.6-8b-a1.6b202+0
Qwen3-30B-A3B-Thinking-2507-REAP-25-nonuniforminference-optimization/qwen3-30b-a3b-thinking-2507-reap-25-nonuniform151—
GLM-5.3-Flash-MEP50inference-optimization/glm-5.3-flash-mep50144+0