Open-weights models, addressed by content and served by the network.
Gemma 4 E2B MTP Drafter
MTP headGoogle's jointly-trained Multi-Token Prediction head for Gemma 4 E2B. Pair with the E2B target via `--spec-type draft-mtp`.
Granite 4.0 350M
350MIBM Granite 4.0 350M — ultra-compact for edge deployment
Gemma 3 270M
270MTiny Gemma model for ultra-lightweight on-device inference
Gemma 4 E4B MTP Drafter
MTP headGoogle's jointly-trained Multi-Token Prediction head for Gemma 4 E4B. Pair with the E4B target via `--spec-type draft-mtp`.
Qwen 3 0.6B
0.6BCompact model optimized for edge deployment
Mistral Small 3.1 DRAFT 0.5B
0.5BSpeculative drafter for Mistral Small 3.1/3.2 — vocab-matched, 6-language fine-tune.
Qwen 3.5 0.8B
0.8BCompact multilingual model for efficient on-device inference
Qwen 3.5 0.8B (MTP)
0.8BQwen 3.5 0.8B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Gemma 4 12B MTP Drafter
MTP headGoogle's jointly-trained Multi-Token Prediction head for Gemma 4 12B. Pair with the 12B target via `--spec-type draft-mtp`.
Granite 4.0 1B
1BIBM Granite 4.0 1B — compact enterprise model
Gemma 3 1B
1BGoogle's compact instruction-tuned model
SmolLM2 1.7B
1.7BCompact SmolLM2 for on-device AI
Qwen 3 1.7B
1.7BVersatile model for various language tasks
Gemma 4 26B-A4B MTP Drafter (MoE)
MTP headMoEGoogle's jointly-trained Multi-Token Prediction head for the Gemma 4 26B-A4B Mixture-of-Experts target. Pair via `--spec-type draft-mtp`; Unsloth measures ~1.15–1.2× speedup on MoE targets vs ~1.4–2.2× on dense.
Qwen 3.5 2B
2BEfficient small model for chat and text generation
Qwen 3.5 2B (MTP)
2BQwen 3.5 2B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Gemma 4 31B MTP Drafter
MTP headGoogle's jointly-trained Multi-Token Prediction head for Gemma 4 31B. Pair with the 31B target via `--spec-type draft-mtp`.
SmolLM3 3B
3BSmolLM3 — 11T tokens, dual-mode reasoning
Ministral 3 3B
3BCompact Ministral 3 for lightweight tasks
Gemma 3 4B
4BExtended context Gemma model for chat applications
Phi-4 Mini 3.8B
3.8BCompact Phi-4 Mini with 128K context
Phi-4 Mini Reasoning 3.8B
3.8BCompact Phi-4 Mini fine-tuned for reasoning tasks
Qwen 3 4B
4BWell-balanced model for production use
Qwen 3.5 4B (MTP)
4BQwen 3.5 4B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Qwen 3.5 4B
4BMid-size model with strong reasoning and coding performance
Nemotron 3 Nano 4B
4BHybrid Mamba-2 + Attention edge model, 256K context
Gemma 4 E2B
E2BMTPGoogle's compact Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e2b-mtp-draft` for 1.5–2.2× throughput.
Gemma 4 E2B (QAT)
E2BMTPQuantization-Aware-Trained Gemma 4 E2B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e2b-mtp-draft`.
Granite 4.0 H-Tiny
7B (hybrid)IBM Granite 4.0 H-Tiny — hybrid Mamba/Transformer architecture
Mistral 7B Instruct v0.3
7BMistral AI's classic 7B instruction model
Ministral 3 8B
8BVersatile Ministral 3 for general-purpose tasks
Qwen 3 8B
8BExtended context model for long-form tasks
Gemma 4 E4B
E4BMTPGoogle's efficient Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e4b-mtp-draft` for 1.5–2.2× throughput.
Gemma 4 E4B (QAT)
E4BMTPQuantization-Aware-Trained Gemma 4 E4B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e4b-mtp-draft`.
Qwen 3.5 9B (MTP)
9BQwen 3.5 9B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Qwen 3.5 9B
9BHigh-performance model for complex language understanding
GLM-4 9B Chat
9BZhipu AI GLM-4 9B instruction-tuned, 128K context
Gemma 3 12B
12BHigh-performance instruction-tuned model from Google
Mistral Nemo 12B
12BExtended-context Mistral model built with NVIDIA
Gemma 4 12B
12BMTPGoogle's mid-tier dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-12b-mtp-draft` for 1.5–2.2× throughput on the same hardware (Unsloth: 52 → 162 t/s at Q4 on a 4090).
Gemma 4 12B (QAT)
12BMTPQuantization-Aware-Trained Gemma 4 12B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-12b-mtp-draft`.
Ministral 3 14B
14BHigh-performance Ministral 3 for complex reasoning
Phi-4 14B
14BMicrosoft Phi-4 — strong reasoning at 14B
Qwen 3 14B
14BPremium model with extended context support
Phi-4 Reasoning 14B
14BPhi-4 fine-tuned for chain-of-thought reasoning
GPT-OSS 20B
20BOpenAI GPT-OSS 20B — open-weights release, native MXFP4
Mistral Small 3.1 24B
24BMistral Small 3.1 — improved reasoning over 3.0 baseline
Mistral Small 3.2 24B
24BMistral Small 3.2 — latest 3-series point release
Mistral Small 3.2 24B
24BMistral's latest Small 3.2 model for demanding workloads
Gemma 3 27B
27BGoogle's largest Gemma model with exceptional capabilities
Qwen 3.5 27B
27BFlagship Qwen 3.5 model
Qwen 3.6 27B
27BQwen 3.6 27B — flagship dense model with 128K context
Qwen 3.5 27B (MTP)
27BQwen 3.5 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Qwen 3.6 27B (MTP)
27BMTPQwen 3.6 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 160 t/s on RTX 6000.
DiffusionGemma 26B-A4B
26B (4B active)MoEDiffusion-generation Gemma 4 26B-A4B. Generates 256-token canvases by parallel denoising rather than autoregressive sampling. Unsloth: 2000+ t/s on RTX 6000. Requires Unsloth Studio or llama.cpp PR #24423+.
Gemma 4 26B-A4B (MoE)
26B (4B active)MTPMoEGemma 4 Mixture-of-Experts: 26B total params, 4B active per token (128K context). MTP-enabled — pairs with `gemma4-26b-a4b-mtp-draft`; expect ~1.15–1.2× speedup on MoE targets per Unsloth.
Gemma 4 26B-A4B (QAT, MoE)
26B (4B active)MTPMoEQuantization-Aware-Trained Gemma 4 26B-A4B MoE. Unsloth measures 85.6% MMLU top-1 vs 70.2% on naive Q4 (+15.4 points). MTP-enabled via `gemma4-26b-a4b-mtp-draft`.
Qwen 3 30B-A3B (MoE)
30B (MoE)MoEMixture-of-Experts with 3B active params for efficient scaling
Qwen 3 Coder 30B-A3B (MoE)
30B (MoE)MoECode-focused MoE — 30B total, 3B active, 256K context
Granite 4.0 H-Small (32B)
32B (hybrid)IBM Granite 4.0 H-Small — 32B hybrid for long-context enterprise
Gemma 4 31B
31BMTPGoogle's largest dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-31b-mtp-draft` for ~2× throughput at 101 t/s on consumer GPUs (Unsloth benchmark).
Qwen 3 32B
32BTop-tier model with 128K context window
Gemma 4 31B (QAT)
31BMTPQuantization-Aware-Trained Gemma 4 31B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-31b-mtp-draft`.
Kimi K2 Instruct (MoE)
1T (MoE, 32B active)MoEMoonshot AI Kimi K2 MoE — 1T total, 32B active, 128K context
Qwen 3.6 35B-A3B (MoE)
35B (MoE, 3B active)MoEQwen 3.6 MoE — 35B total, ~3B active per token
Qwen 3.6 35B-A3B MTP (MoE)
35B (MoE, 3B active)MTPMoEQwen 3.6 35B-A3B MoE with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 240 t/s on RTX 6000.
Qwen 3.5 35B-A3B (MoE)
35B (MoE)MoEMixture-of-Experts with only 3B active params — fast inference at 35B quality
Qwen 3.5 35B-A3B (MoE) (MTP)
35B (MoE, 3B active)Qwen 3.5 35B-A3B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
Nemotron 3 Nano 30B-A3B (MoE)
30B (MoE)MoEHybrid Mamba-2 MoE — 30B total, 3.5B active, 128K context
GPT-OSS 120B
120BMoEOpenAI GPT-OSS 120B — open-weights release, native MXFP4
Qwen 3.5 122B-A10B (MoE)
122B (MoE, 10B active)MoEQwen 3.5 large MoE — 122B total, 10B active per token. Replica-routed on high-VRAM provider tiers only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-122B-A10B-MTP-GGUF` for compatible runtimes.
Qwen 3.5 122B-A10B (MoE) (MTP)
122B (MoE, 10B active)Qwen 3.5 122B-A10B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
MiniMax M2.7 (MoE)
230B (MoE, 10B active)MoEMiniMax M2.7 — current frontier MiniMax MoE; Lightning Attention with 1M context. Unsloth dynamic UD-Q4_K_XL GGUF (sharded).
DeepSeek V4 Flash (MoE)
284B (MoE, 13B active)MTPMoEDeepSeek V4 Flash — 284B total / 13B active MoE; 1M context. Hybrid Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA). Cost-effective frontier variant. Pre-trained on 32T tokens. MTP head shipped natively.
MiniMax M3 (MoE, native multimodal)
428B (MoE, 23B active)MoEMiniMax M3 — ~428B total / ~23B active MoE with native multimodal training. MiniMax Sparse Attention (MSA) delivers 9× prefill and 15× decode speedups vs M2 at 1M context. Note: GGUF builds currently fall back to dense attention; sparse attention not yet supported in llama.cpp.
Qwen 3.5 397B-A17B (MoE)
397B (MoE, 17B active)MoEQwen 3.5 frontier MoE — 397B total, 17B active per token. Multi-GPU replicas only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-397B-A17B-MTP-GGUF` for compatible runtimes.
Qwen 3.5 397B-A17B (MoE) (MTP)
397B (MoE, 17B active)Qwen 3.5 397B-A17B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.
DeepSeek V3 0324 (MoE)
685B (MoE, 37B active)MTPMoEDeepSeek V3 MoE — 685B total, 37B active, 128K context. Native Multi-Token-Prediction head (n=4, ~80% accept rate, ~1.8× decode speedup per DeepSeek tech report). Retired by upstream after 2026-07-24 in favor of DeepSeek V4.
GLM-5 (MoE)
744B (MoE)MoEZ.ai GLM-5 — 744B total parameter MoE trained on 28.5T tokens; best-in-class open-source performance on reasoning, coding, and agentic tasks (2026-04). Unsloth dynamic UD-Q4_K_XL GGUF (sharded).
GLM-5.1 (MoE)
744B (MoE, 40B active)MoEZ.ai GLM-5.1 — next-generation flagship for agentic engineering, class-leading on SWE-Bench Pro; 744B total / 40B active; 200K context. `glm_moe_dsa` architecture with Dynamic Sparse Attention.
GLM-5.2 (MoE, MTP)
753B (MoE)MTPMoEZ.ai GLM-5.2 — 753B total parameter MoE flagship with solid 1M-token context and IndexShare sparse-attention (2.9× per-token FLOP reduction at 1M). Improved Multi-Token-Prediction layer increases speculative-decoding accept rate by ~20% over GLM-5.1.
Kimi K2.5 (MoE)
1T (MoE, 32B active)MoEMoonshot AI Kimi K2.5 — 1T total / 32B active MoE; image input support; 256K context. Predecessor to K2.6's hybrid-thinking variant.
Kimi K2.7 Code (MoE)
1T (MoE, 32B active, code-focused)MoEMoonshot AI Kimi K2.7 Code — code-focused refresh of the K2 series. 1T total / 32B active; 256K context; recent updates target tool-call accuracy on long-horizon coding tasks.
Kimi K2.6 (Hybrid Thinking, MoE)
1T (MoE, hybrid thinking)MoEMoonshot AI Kimi K2.6 hybrid-thinking MoE — 1T total params, 256K context. Replica-routed on B200-class infrastructure; Unsloth measures >40 t/s on B200. Recommended `UD-Q2_K_XL` (350GB) for size/quality balance.