inference-optimization AI Models

Tracked by OpenModelStats since Aug 23, 2026 · View on Hugging Face ↗

Tracked models
103
Downloads 30D
100.8K
Total downloads
526.2K
Total likes
131
Gained 7D
31.5K

Most downloaded

  1. 1Llama-3.2-0.5B-InstructJun 3, 202620.1K
  2. 2Qwen3-1.6B-A0.9BMay 24, 202614.5K
  3. 3Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1Jun 10, 20267,060
  4. 4Qwen3-8B-speculators.peagle-qwen3arch-ckpt4Jun 16, 20267,047
  5. 5Qwen3-Coder-Next.w4a16Mar 12, 20266,098
  6. 6DeepSeek-V3-debug-empty-FP8_DYNAMICJan 23, 20265,075
  7. 7GLM-5.2-0.8B-A0.8BJul 2, 20264,924
  8. 8DSV4-tiny-emptySep 14, 20264,819
  9. 9Qwen3.8-1.0B-A0.6BAug 12, 20264,508
  10. 10NemotronH-0.3B-A0.3BSep 14, 20263,862

Fastest growing

  1. 1Llama-3.2-0.5B-Instruct+7,405 in 7D20.1K
  2. 2Qwen3-1.6B-A0.9B+4,898 in 7D14.5K
  3. 3Qwen3-8B-speculators.peagle-qwen3arch-ckpt4+2,185 in 7D7,047
  4. 4Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1+2,178 in 7D7,060
  5. 5DSV4-tiny-empty+1,806 in 7D4,819
  6. 6Kimi-K3-0.40B-MXFP4+1,795 in 7D3,284
  7. 7NemotronH-0.3B-A0.3B+1,762 in 7D3,862
  8. 8GLM-5.2-0.8B-A0.8B+1,760 in 7D4,924
  9. 9qwe3-8B-selfdistill-v3-fresh3ep_ckpt0+1,640 in 7D2,161
  10. 10DeepSeek-V3-debug-empty-FP8_DYNAMIC+1,583 in 7D5,075

Recently updated

  1. 1Inkling-0.92B-MTPOct 7, 20267
  2. 2GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQOct 7, 20265
  3. 3GLM-5-0.88B-MTPOct 7, 20265
  4. 4DeepSeek-V3-0.86B-MTP-FP8-DynamicOct 7, 20266
  5. 5DeepSeek-V3-0.86B-MTPOct 7, 20269
  6. 6GLM-4.7-Flash-0.82B-MTP-FP8-DynamicOct 7, 20264
  7. 7GLM-4.7-Flash-0.82B-MTPOct 7, 20265
  8. 8GLM-4.5-Air-0.72B-MTP-FP8-DynamicOct 7, 20265
  9. 9GLM-4.5-Air-0.72B-MTPOct 7, 20266
  10. 10gemma-4-31b-it-speculator.dflash.latestrecipe.ckpt2Oct 6, 202617

All models

All tracked models published by inference-optimization
ModelDL 30D7D
Hy3-NVFP4-REAP-25inference-optimization/hy3-nvfp4-reap-2512+4
Hy3-NVFP4inference-optimization/hy3-nvfp412+0
qwe3-8B-selfdistill-v3-fresh3ep_ckpt1inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt111+11
Qwen3-8B-DFlash-FP8-DYNAMICinference-optimization/qwen3-8b-dflash-fp8-dynamic11—
gemma-4-31B-it-speculator.dflash.latestrecipe.ckp1inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.ckp111+11
qwe3-8B-selfdistill-v3-fresh3ep_ckpt2inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt29—
Qwen3-8B-DFlash-Gauss-FP8-W8A8inference-optimization/qwen3-8b-dflash-gauss-fp8-w8a88—
Qwen3-8B-DFlash-PerfectBlend-FP8-W8A8inference-optimization/qwen3-8b-dflash-perfectblend-fp8-w8a87—
Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch1inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch10+0
Qwen3-8B-DFlash-Gauss8inference-optimization/qwen3-8b-dflash-gauss80—
Qwen3-8B-dflash2-v1-fullvocab-muon-768anc-self-distillinference-optimization/qwen3-8b-dflash2-v1-fullvocab-muon-768anc-self-distill0+0
Qwen3-8B-DFlash-Blend8inference-optimization/qwen3-8b-dflash-blend80—
Qwen3-8B-DFlash-Gauss4inference-optimization/qwen3-8b-dflash-gauss40—
Qwen3-8B-DFlash-Drift8inference-optimization/qwen3-8b-dflash-drift80—
Qwen3-8B-DFlash-Blend4inference-optimization/qwen3-8b-dflash-blend40—
Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch0inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch00+0
Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.nemotron.v2v3-v8-ckpt19inference-optimization/qwen3-8b-speculator.dflash.swa.dpace.fullvocab.muon.1536anc.nemotron.v2v3-v8-ckpt190+0
Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch2inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.regencollection.selfdistill-v3-epoch20+0