inference-optimization AI Models
Tracked by OpenModelStats since Aug 23, 2026 · View on Hugging Face ↗
- Tracked modelsRepositories from this publisher currently tracked by OpenModelStats.
- 103
- Downloads 30DSum of each tracked model's rolling recent-downloads counter. A sum of repository counters is not a count of unique users.
- 100.8K
- Total downloadsSum of cumulative all-time download counters across tracked repositories.
- 526.2K
- Total likes
- 131
- Gained 7DCombined seven-day cumulative-download gains measured by OpenModelStats.
- 31.5K
Most downloaded
- 1Llama-3.2-0.5B-InstructJun 3, 202620.1K
- 2Qwen3-1.6B-A0.9BMay 24, 202614.5K
- 3Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1Jun 10, 20267,060
- 4Qwen3-8B-speculators.peagle-qwen3arch-ckpt4Jun 16, 20267,047
- 5Qwen3-Coder-Next.w4a16Mar 12, 20266,098
- 6DeepSeek-V3-debug-empty-FP8_DYNAMICJan 23, 20265,075
- 7GLM-5.2-0.8B-A0.8BJul 2, 20264,924
- 8DSV4-tiny-emptySep 14, 20264,819
- 9Qwen3.8-1.0B-A0.6BAug 12, 20264,508
- 10NemotronH-0.3B-A0.3BSep 14, 20263,862
Fastest growing
- 1Llama-3.2-0.5B-Instruct+7,405 in 7D20.1K
- 2Qwen3-1.6B-A0.9B+4,898 in 7D14.5K
- 3Qwen3-8B-speculators.peagle-qwen3arch-ckpt4+2,185 in 7D7,047
- 4Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1+2,178 in 7D7,060
- 5DSV4-tiny-empty+1,806 in 7D4,819
- 6Kimi-K3-0.40B-MXFP4+1,795 in 7D3,284
- 7NemotronH-0.3B-A0.3B+1,762 in 7D3,862
- 8GLM-5.2-0.8B-A0.8B+1,760 in 7D4,924
- 9qwe3-8B-selfdistill-v3-fresh3ep_ckpt0+1,640 in 7D2,161
- 10DeepSeek-V3-debug-empty-FP8_DYNAMIC+1,583 in 7D5,075
Recently updated
- 1Inkling-0.92B-MTPOct 7, 20267
- 2GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQOct 7, 20265
- 3GLM-5-0.88B-MTPOct 7, 20265
- 4DeepSeek-V3-0.86B-MTP-FP8-DynamicOct 7, 20266
- 5DeepSeek-V3-0.86B-MTPOct 7, 20269
- 6GLM-4.7-Flash-0.82B-MTP-FP8-DynamicOct 7, 20264
- 7GLM-4.7-Flash-0.82B-MTPOct 7, 20265
- 8GLM-4.5-Air-0.72B-MTP-FP8-DynamicOct 7, 20265
- 9GLM-4.5-Air-0.72B-MTPOct 7, 20266
- 10gemma-4-31b-it-speculator.dflash.latestrecipe.ckpt2Oct 6, 202617
All models
| Model | DL 30D | 7D |
|---|---|---|
| ZAYA1-74B-preview-NVFP4inference-optimization/zaya1-74b-preview-nvfp4 | 132 | — |
| gemma-4-1B-0.8B-tinyinference-optimization/gemma-4-1b-0.8b-tiny | 128 | +0 |
| Qwen3.8-27B-DSpark-Gauss-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-gauss-nvfp4-w4a4 | 126 | — |
| Qwen3.8-27B-DSpark-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3.8-27b-dspark-perfectblend-nvfp4-w4a4 | 125 | — |
| GLM-5.3-0.6B-A0.4Binference-optimization/glm-5.3-0.6b-a0.4b | 105 | +105 |
| Qwen3.8-Flash-Next-MEP50inference-optimization/qwen3.8-flash-next-mep50 | 94 | +0 |
| Qwen3-8B-DFlash-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-perfectblend-nvfp4-w4a4 | 75 | — |
| Qwen3-8B-DFlash-Gauss-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-gauss-nvfp4-w4a4 | 74 | — |
| GLM-5.2-0.8B-A0.8B-MXFP4xFP8_BLOCKinference-optimization/glm-5.2-0.8b-a0.8b-mxfp4xfp8_block | 64 | +0 |
| GLM-5.2-0.8B-A0.8B-MXFP4xMXFP8inference-optimization/glm-5.2-0.8b-a0.8b-mxfp4xmxfp8 | 63 | +0 |
| Qwen3.8-1.0B-A0.6B-NVFP4-FP8inference-optimization/qwen3.8-1.0b-a0.6b-nvfp4-fp8 | 62 | +0 |
| GLM-5.3-Flash-0.1B-A0.1B-MTPinference-optimization/glm-5.3-flash-0.1b-a0.1b-mtp | 52 | +0 |
| dspark-inkling-smallinference-optimization/dspark-inkling-small | 44 | +0 |
| Qwen3.8-Flash-Next-0.2B-A0.2B-MTPinference-optimization/qwen3.8-flash-next-0.2b-a0.2b-mtp | 43 | +0 |
| Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt7inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt7 | 39 | +4 |
| Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt9inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt9 | 36 | +3 |
| qwen38-dspark-pretrain-finewebinference-optimization/qwen38-dspark-pretrain-fineweb | 36 | — |
| Qwen3-8B-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt8inference-optimization/qwen3-8b-speculator.dspark.swa.gammatv.2048anc.muon.combdatav4-q235b-instr-v2-ckpt8 | 29 | +3 |
| qwen3.8-27B-speculator.dflash2.comparisonrecipe.epoch1_step93171inference-optimization/qwen3.8-27b-speculator.dflash2.comparisonrecipe.epoch1_step93171 | 26 | — |
| qwen38-dspark-cooldown-finewebinference-optimization/qwen38-dspark-cooldown-fineweb | 23 | — |
| gemma-4-26B-A4B-it-FP8-REAP-15inference-optimization/gemma-4-26b-a4b-it-fp8-reap-15 | 22 | +0 |
| Qwen3.8-27B-DSpark-FP8-DYNAMICinference-optimization/qwen3.8-27b-dspark-fp8-dynamic | 21 | — |
| Qwen3-8B-DFlash-GPTQ-IMatrix-PerfectBlend-NVFP4-W4A4inference-optimization/qwen3-8b-dflash-gptq-imatrix-perfectblend-nvfp4-w4a4 | 21 | — |
| qwen3.8-27B-speculator.dflash2.comparisonrecipe.v2.epoch1_endinference-optimization/qwen3.8-27b-speculator.dflash2.comparisonrecipe.v2.epoch1_end | 21 | — |
| Qwen3.8-27B-DSpark-PerfectBlend-FP8-W8A8inference-optimization/qwen3.8-27b-dspark-perfectblend-fp8-w8a8 | 20 | — |