inference-optimization AI Models
Tracked by OpenModelStats since Aug 23, 2026 · View on Hugging Face ↗
- Tracked modelsRepositories from this publisher currently tracked by OpenModelStats.
- 103
- Downloads 30DSum of each tracked model's rolling recent-downloads counter. A sum of repository counters is not a count of unique users.
- 100.8K
- Total downloadsSum of cumulative all-time download counters across tracked repositories.
- 526.2K
- Total likes
- 131
- Gained 7DCombined seven-day cumulative-download gains measured by OpenModelStats.
- 31.5K
Most downloaded
- 1Llama-3.2-0.5B-InstructJun 3, 202620.1K
- 2Qwen3-1.6B-A0.9BMay 24, 202614.5K
- 3Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1Jun 10, 20267,060
- 4Qwen3-8B-speculators.peagle-qwen3arch-ckpt4Jun 16, 20267,047
- 5Qwen3-Coder-Next.w4a16Mar 12, 20266,098
- 6DeepSeek-V3-debug-empty-FP8_DYNAMICJan 23, 20265,075
- 7GLM-5.2-0.8B-A0.8BJul 2, 20264,924
- 8DSV4-tiny-emptySep 14, 20264,819
- 9Qwen3.8-1.0B-A0.6BAug 12, 20264,508
- 10NemotronH-0.3B-A0.3BSep 14, 20263,862
Fastest growing
- 1Llama-3.2-0.5B-Instruct+7,405 in 7D20.1K
- 2Qwen3-1.6B-A0.9B+4,898 in 7D14.5K
- 3Qwen3-8B-speculators.peagle-qwen3arch-ckpt4+2,185 in 7D7,047
- 4Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1+2,178 in 7D7,060
- 5DSV4-tiny-empty+1,806 in 7D4,819
- 6Kimi-K3-0.40B-MXFP4+1,795 in 7D3,284
- 7NemotronH-0.3B-A0.3B+1,762 in 7D3,862
- 8GLM-5.2-0.8B-A0.8B+1,760 in 7D4,924
- 9qwe3-8B-selfdistill-v3-fresh3ep_ckpt0+1,640 in 7D2,161
- 10DeepSeek-V3-debug-empty-FP8_DYNAMIC+1,583 in 7D5,075
Recently updated
- 1Inkling-0.92B-MTPOct 7, 20267
- 2GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQOct 7, 20265
- 3GLM-5-0.88B-MTPOct 7, 20265
- 4DeepSeek-V3-0.86B-MTP-FP8-DynamicOct 7, 20266
- 5DeepSeek-V3-0.86B-MTPOct 7, 20269
- 6GLM-4.7-Flash-0.82B-MTP-FP8-DynamicOct 7, 20264
- 7GLM-4.7-Flash-0.82B-MTPOct 7, 20265
- 8GLM-4.5-Air-0.72B-MTP-FP8-DynamicOct 7, 20265
- 9GLM-4.5-Air-0.72B-MTPOct 7, 20266
- 10gemma-4-31b-it-speculator.dflash.latestrecipe.ckpt2Oct 6, 202617
All models
| Model | DL 30D | 7D |
|---|---|---|
| Qwen3.8-27B-DSpark-Gauss-FP8-W8A8inference-optimization/qwen3.8-27b-dspark-gauss-fp8-w8a8 | 19 | — |
| Qwen3-8B-DFlash-GPTQ-IMatrix-Gauss-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-gptq-imatrix-gauss-nvfp4-w4a4 | 19 | — |
| Qwen3-8B-DFlash-GPTQ-Gauss-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-gptq-gauss-nvfp4-w4a4 | 18 | — |
| qwen3.8-27B-speculator.dflash2.comparisonrecipe.v2.epoch1_step107505inference-optimization/qwen3.8-27b-speculator.dflash2.comparisonrecipe.v2.epoch1_step107505 | 18 | — |
| gemma-4-26B-A4B-it-NVFP4-REAP-50inference-optimization/gemma-4-26b-a4b-it-nvfp4-reap-50 | 18 | +5 |
| Qwen3.8-27B-DSpark-GPTQ-IMatrix-Gauss-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-gptq-imatrix-gauss-nvfp4-w4a4 | 18 | — |
| Qwen3.8-27B-DSpark-GPTQ-Gauss-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-gptq-gauss-nvfp4-w4a4 | 18 | — |
| Qwen3.8-27B-DSpark-GPTQ-IMatrix-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-gptq-imatrix-perfectblend-nvfp4-w4a4 | 18 | — |
| Qwen3-8B-DFlash-GPTQ-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-gptq-perfectblend-nvfp4-w4a4 | 17 | — |
| gemma-4-26B-A4B-it-FP8-REAP-25inference-optimization/gemma-4-26b-a4b-it-fp8-reap-25 | 17 | +0 |
| Qwen3-8B-DFlash-FP8-BLOCKinference-optimization/qwen3-8b-dflash-fp8-block | 17 | — |
| Qwen3.8-27B-DSpark-GPTQ-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-gptq-perfectblend-nvfp4-w4a4 | 16 | — |
| Qwen3.8-27B-DSpark-FP8-BLOCKinference-optimization/qwen3.8-27b-dspark-fp8-block | 16 | — |
| gemma-4-31B-it-speculator.dflash.latestrecipe.ckp4inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.ckp4 | 15 | — |
| gemma-4-26B-A4B-it-NVFP4-REAP-25inference-optimization/gemma-4-26b-a4b-it-nvfp4-reap-25 | 15 | +5 |
| Qwen3-8B-speculator.dflash2.v1.fullvocab.muon.dpace-preresume-epoch2inference-optimization/qwen3-8b-speculator.dflash2.v1.fullvocab.muon.dpace-preresume-epoch2 | 15 | +15 |
| gemma-4-31B-it-speculator.dflash.latestrecipe.ckp2inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.ckp2 | 15 | +15 |
| gemma-4-31B-it-speculator.dflash.latestrecipe.ckp0inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.ckp0 | 15 | +15 |
| qwe3-8B-selfdistill-v3-fresh3ep_ckpt4inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt4 | 14 | — |
| gemma-4-31B-it-speculator.dflash.latestrecipe.ckp3inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.ckp3 | 14 | — |
| qwe3-8B-selfdistill-v3-fresh3ep_ckpt5inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt5 | 14 | — |
| gemma-4-31B-it-speculator.dflash.latestrecipe.bs8.ckp0inference-optimization/gemma-4-31b-it-speculator.dflash.latestrecipe.bs8.ckp0 | 13 | — |
| Hy3-NVFP4-REAP-50inference-optimization/hy3-nvfp4-reap-50 | 13 | +4 |
| ZAYA1-74B-preview-NVFP4-linearinference-optimization/zaya1-74b-preview-nvfp4-linear | 13 | — |
| qwe3-8B-selfdistill-v3-fresh3ep_ckpt3inference-optimization/qwe3-8b-selfdistill-v3-fresh3ep_ckpt3 | 13 | — |