Tenzro Labs
Model Hub

Open-weights models, addressed by content and served by the network.

Every model below carries a license tier and a BLAKE3 content address. Pull it from HuggingFace, or from a verified Tenzro Network holder that stores the exact same bytes. Every count on this page — how many nodes hold a model, how many serve it for inference — is read live from the network, never baked in.
84
Models in catalog
0
Downloadable from the network
0
Serving inference now
60
Permissive-licensed
How to read this
Content address
A model on the network is named tenzro://model/<id>@<blake3>. The BLAKE3 hash is over the exact GGUF bytes, so two providers advertising the same address are byte-identical. The hash is recorded on-chain when a provider first serves the model.
Two sources, one identity
HuggingFace is the centralized origin. The Tenzro Network source is a set of verified holders storing the same bytes; a client can fetch from a peer first and fall back to the Hub. The holder count is the number of nodes currently advertising that content address — like seeders on a torrent.
Bootstrapped, then grown by the network
Tenzro Labs seeded the first copy of each model so the network is not empty on day one. Every operator that downloads and holds a model becomes another holder, and every provider that loads one becomes another inference source. Both counts are read live from the network — as others join, the numbers on this page climb on their own.
Downloadable vs serving inference
Two independent counts. Downloadable means a node holds the bytes and can distribute them peer-first — you can pull the model and run it yourself. Serving inference means a provider has the model loaded and answers chat requests right now. A model can be widely held but served by nobody; to use a held-only model you fetch it and serve it, or wait for a provider with the right hardware.
License tiers
Permissive covers Apache-2.0 and MIT weights. Custom, Open (NVIDIA), and Open (MiniMax) carry their own terms — the runtime requires explicit acceptance before loading them.
MTP and MoE
MTP marks models that ship a Multi-Token-Prediction draft path for speculative decoding. MoE marks mixture-of-experts models whose experts can be dispatched across providers.
Catalog
84 models

Gemma 4 E2B MTP Drafter

MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 E2B. Pair with the E2B target via `--spec-type draft-mtp`.

Gemma License128K ctxBF16191 MB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E2B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Granite 4.0 350M

350M

IBM Granite 4.0 350M — ultra-compact for edge deployment

Apache 2.0128K ctxQ4_K_M210 MB
0 downloadable0 serving inference
HuggingFace
ibm-granite/granite-4.0-350m-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 3 270M

270M

Tiny Gemma model for ultra-lightweight on-device inference

Gemma License32K ctxQ4_K_M241 MB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-3-270m-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 E4B MTP Drafter

MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 E4B. Pair with the E4B target via `--spec-type draft-mtp`.

Gemma License128K ctxBF16286 MB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E4B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 0.6B

0.6B

Compact model optimized for edge deployment

Apache 2.032K ctxQ4_K_M378 MB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-0.6B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral Small 3.1 DRAFT 0.5B

0.5B

Speculative drafter for Mistral Small 3.1/3.2 — vocab-matched, 6-language fine-tune.

Apache 2.032K ctxQ4_K_M379 MB
0 downloadable0 serving inference
HuggingFace
alamios/Mistral-Small-3.1-DRAFT-0.5B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 0.8B

0.8B

Compact multilingual model for efficient on-device inference

Apache 2.0128K ctxQ4_K_M508 MB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-0.8B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 0.8B (MTP)

0.8B

Qwen 3.5 0.8B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL515 MB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-0.8B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 12B MTP Drafter

MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 12B. Pair with the 12B target via `--spec-type draft-mtp`.

Gemma License128K ctxBF16572 MB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-12b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Granite 4.0 1B

1B

IBM Granite 4.0 1B — compact enterprise model

Apache 2.0128K ctxQ4_K_M668 MB
0 downloadable0 serving inference
HuggingFace
ibm-granite/granite-4.0-1b-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 3 1B

1B

Google's compact instruction-tuned model

Gemma License32K ctxQ4_K_M769 MB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-3-1b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

SmolLM2 1.7B

1.7B

Compact SmolLM2 for on-device AI

Apache-2.08K ctxQ4_K_M1007 MB
0 downloadable0 serving inference
HuggingFace
unsloth/SmolLM2-1.7B-Instruct-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 1.7B

1.7B

Versatile model for various language tasks

Apache 2.032K ctxQ4_K_M1.0 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-1.7B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 26B-A4B MTP Drafter (MoE)

MTP headMoE

Google's jointly-trained Multi-Token Prediction head for the Gemma 4 26B-A4B Mixture-of-Experts target. Pair via `--spec-type draft-mtp`; Unsloth measures ~1.15–1.2× speedup on MoE targets vs ~1.4–2.2× on dense.

Gemma License128K ctxBF161.1 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-26B-A4B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 2B

2B

Efficient small model for chat and text generation

Apache 2.0128K ctxQ4_K_M1.2 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-2B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 2B (MTP)

2B

Qwen 3.5 2B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL1.2 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-2B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 31B MTP Drafter

MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 31B. Pair with the 31B target via `--spec-type draft-mtp`.

Gemma License128K ctxBF161.4 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-31B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

SmolLM3 3B

3B

SmolLM3 — 11T tokens, dual-mode reasoning

Apache-2.064K ctxQ4_K_M1.8 GB
0 downloadable0 serving inference
HuggingFace
unsloth/SmolLM3-3B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Ministral 3 3B

3B

Compact Ministral 3 for lightweight tasks

Apache 2.0128K ctxQ4_K_M2.0 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Ministral-3-3B-Instruct-2512-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 3 4B

4B

Extended context Gemma model for chat applications

Gemma License128K ctxQ4_K_M2.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-3-4b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Phi-4 Mini 3.8B

3.8B

Compact Phi-4 Mini with 128K context

MIT125K ctxQ4_K_M2.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Phi-4-mini-instruct-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Phi-4 Mini Reasoning 3.8B

3.8B

Compact Phi-4 Mini fine-tuned for reasoning tasks

MIT125K ctxQ4_K_M2.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Phi-4-mini-reasoning-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 4B

4B

Well-balanced model for production use

Apache 2.032K ctxQ4_K_M2.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-4B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 4B (MTP)

4B

Qwen 3.5 4B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL2.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-4B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 4B

4B

Mid-size model with strong reasoning and coding performance

Apache 2.0128K ctxQ4_K_M2.6 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-4B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Nemotron 3 Nano 4B

4B

Hybrid Mamba-2 + Attention edge model, 256K context

NVIDIA Open256K ctxQ4_K_M2.7 GB
0 downloadable0 serving inference
HuggingFace
unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 E2B

E2BMTP

Google's compact Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e2b-mtp-draft` for 1.5–2.2× throughput.

Gemma License128K ctxQ4_K_M3.1 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E2B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 E2B (QAT)

E2BMTP

Quantization-Aware-Trained Gemma 4 E2B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e2b-mtp-draft`.

Gemma License128K ctxUD-Q4_K_XL (QAT)3.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E2B-it-qat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Granite 4.0 H-Tiny

7B (hybrid)

IBM Granite 4.0 H-Tiny — hybrid Mamba/Transformer architecture

Apache 2.0128K ctxQ4_K_M4.0 GB
0 downloadable0 serving inference
HuggingFace
ibm-granite/granite-4.0-h-tiny-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral 7B Instruct v0.3

7B

Mistral AI's classic 7B instruction model

Apache 2.032K ctxQ4_K_M4.1 GB
0 downloadable0 serving inference
HuggingFace
bartowski/Mistral-7B-Instruct-v0.3-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Ministral 3 8B

8B

Versatile Ministral 3 for general-purpose tasks

Apache 2.0128K ctxQ4_K_M4.6 GB
0 downloadable0 serving inference
HuggingFace
bartowski/Ministral-8B-Instruct-2410-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 8B

8B

Extended context model for long-form tasks

Apache 2.0128K ctxQ4_K_M4.7 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-8B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 E4B

E4BMTP

Google's efficient Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e4b-mtp-draft` for 1.5–2.2× throughput.

Gemma License128K ctxQ4_K_M5.0 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E4B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 E4B (QAT)

E4BMTP

Quantization-Aware-Trained Gemma 4 E4B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e4b-mtp-draft`.

Gemma License128K ctxUD-Q4_K_XL (QAT)5.1 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-E4B-it-qat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 9B (MTP)

9B

Qwen 3.5 9B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL5.1 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-9B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 9B

9B

High-performance model for complex language understanding

Apache 2.0128K ctxQ4_K_M5.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-9B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GLM-4 9B Chat

9B

Zhipu AI GLM-4 9B instruction-tuned, 128K context

Apache 2.0128K ctxQ4_K_M5.4 GB
0 downloadable0 serving inference
HuggingFace
bartowski/glm-4-9b-chat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 3 12B

12B

High-performance instruction-tuned model from Google

Gemma License128K ctxQ4_K_M6.8 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-3-12b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral Nemo 12B

12B

Extended-context Mistral model built with NVIDIA

Apache 2.0128K ctxQ4_K_M7.0 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Mistral-Nemo-Instruct-2407-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 12B

12BMTP

Google's mid-tier dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-12b-mtp-draft` for 1.5–2.2× throughput on the same hardware (Unsloth: 52 → 162 t/s at Q4 on a 4090).

Gemma License128K ctxQ4_K_M7.1 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-12b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 12B (QAT)

12BMTP

Quantization-Aware-Trained Gemma 4 12B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-12b-mtp-draft`.

Gemma License128K ctxUD-Q4_K_XL (QAT)7.4 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-12B-it-qat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Ministral 3 14B

14B

High-performance Ministral 3 for complex reasoning

Apache 2.0128K ctxQ4_K_M7.7 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Ministral-3-14B-Instruct-2512-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Phi-4 14B

14B

Microsoft Phi-4 — strong reasoning at 14B

MIT16K ctxQ4_K_M8.3 GB
0 downloadable0 serving inference
HuggingFace
unsloth/phi-4-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 14B

14B

Premium model with extended context support

Apache 2.0128K ctxQ4_K_M8.4 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-14B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Phi-4 Reasoning 14B

14B

Phi-4 fine-tuned for chain-of-thought reasoning

MIT32K ctxQ4_K_M8.4 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Phi-4-reasoning-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GPT-OSS 20B

20B

OpenAI GPT-OSS 20B — open-weights release, native MXFP4

Apache 2.0128K ctxQ4_K_M12 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gpt-oss-20b-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral Small 3.1 24B

24B

Mistral Small 3.1 — improved reasoning over 3.0 baseline

Apache 2.0128K ctxQ4_K_M13 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Mistral-Small-3.1-24B-Instruct-2503-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral Small 3.2 24B

24B

Mistral Small 3.2 — latest 3-series point release

Apache 2.0128K ctxQ4_K_M13 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Mistral Small 3.2 24B

24B

Mistral's latest Small 3.2 model for demanding workloads

Apache 2.0128K ctxQ4_K_M13 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 3 27B

27B

Google's largest Gemma model with exceptional capabilities

Gemma License128K ctxQ4_K_M15 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-3-27b-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 27B

27B

Flagship Qwen 3.5 model

Apache 2.0128K ctxQ4_K_M16 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-27B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.6 27B

27B

Qwen 3.6 27B — flagship dense model with 128K context

Apache 2.0128K ctxQ4_K_M16 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.6-27B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 27B (MTP)

27B

Qwen 3.5 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL16 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-27B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.6 27B (MTP)

27BMTP

Qwen 3.6 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 160 t/s on RTX 6000.

Apache 2.0128K ctxUD-Q4_K_XL17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.6-27B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

DiffusionGemma 26B-A4B

26B (4B active)MoE

Diffusion-generation Gemma 4 26B-A4B. Generates 256-token canvases by parallel denoising rather than autoregressive sampling. Unsloth: 2000+ t/s on RTX 6000. Requires Unsloth Studio or llama.cpp PR #24423+.

Gemma License32K ctxQ4_K_M17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/diffusiongemma-26B-A4B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 26B-A4B (MoE)

26B (4B active)MTPMoE

Gemma 4 Mixture-of-Experts: 26B total params, 4B active per token (128K context). MTP-enabled — pairs with `gemma4-26b-a4b-mtp-draft`; expect ~1.15–1.2× speedup on MoE targets per Unsloth.

Gemma License128K ctxQ4_K_M17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-26B-A4B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 26B-A4B (QAT, MoE)

26B (4B active)MTPMoE

Quantization-Aware-Trained Gemma 4 26B-A4B MoE. Unsloth measures 85.6% MMLU top-1 vs 70.2% on naive Q4 (+15.4 points). MTP-enabled via `gemma4-26b-a4b-mtp-draft`.

Gemma License128K ctxUD-Q4_K_XL (QAT)17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-26B-A4B-it-qat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 30B-A3B (MoE)

30B (MoE)MoE

Mixture-of-Experts with 3B active params for efficient scaling

Apache 2.0128K ctxQ4_K_M17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-30B-A3B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 Coder 30B-A3B (MoE)

30B (MoE)MoE

Code-focused MoE — 30B total, 3B active, 256K context

Apache 2.0256K ctxQ4_K_M17 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Granite 4.0 H-Small (32B)

32B (hybrid)

IBM Granite 4.0 H-Small — 32B hybrid for long-context enterprise

Apache 2.0128K ctxQ4_K_M18 GB
0 downloadable0 serving inference
HuggingFace
unsloth/granite-4.0-h-small-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 31B

31BMTP

Google's largest dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-31b-mtp-draft` for ~2× throughput at 101 t/s on consumer GPUs (Unsloth benchmark).

Gemma License128K ctxQ4_K_M18 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-31B-it-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3 32B

32B

Top-tier model with 128K context window

Apache 2.0128K ctxQ4_K_M18 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3-32B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Gemma 4 31B (QAT)

31BMTP

Quantization-Aware-Trained Gemma 4 31B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-31b-mtp-draft`.

Gemma License128K ctxUD-Q4_K_XL (QAT)18 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gemma-4-31B-it-qat-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Kimi K2 Instruct (MoE)

1T (MoE, 32B active)MoE

Moonshot AI Kimi K2 MoE — 1T total, 32B active, 128K context

MIT128K ctxQ4_K_M19 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Kimi-K2-Instruct-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.6 35B-A3B (MoE)

35B (MoE, 3B active)MoE

Qwen 3.6 MoE — 35B total, ~3B active per token

Apache 2.0128K ctxQ4_K_M20 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.6-35B-A3B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.6 35B-A3B MTP (MoE)

35B (MoE, 3B active)MTPMoE

Qwen 3.6 35B-A3B MoE with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 240 t/s on RTX 6000.

Apache 2.0128K ctxUD-Q4_K_XL20 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.6-35B-A3B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 35B-A3B (MoE)

35B (MoE)MoE

Mixture-of-Experts with only 3B active params — fast inference at 35B quality

Apache 2.0128K ctxQ4_K_M21 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-35B-A3B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 35B-A3B (MoE) (MTP)

35B (MoE, 3B active)

Qwen 3.5 35B-A3B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL21 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-35B-A3B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Nemotron 3 Nano 30B-A3B (MoE)

30B (MoE)MoE

Hybrid Mamba-2 MoE — 30B total, 3.5B active, 128K context

NVIDIA Open125K ctxQ4_K_M23 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Nemotron-3-Nano-30B-A3B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GPT-OSS 120B

120BMoE

OpenAI GPT-OSS 120B — open-weights release, native MXFP4

Apache 2.0128K ctxQ4_K_M68 GB
0 downloadable0 serving inference
HuggingFace
unsloth/gpt-oss-120b-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 122B-A10B (MoE)

122B (MoE, 10B active)MoE

Qwen 3.5 large MoE — 122B total, 10B active per token. Replica-routed on high-VRAM provider tiers only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-122B-A10B-MTP-GGUF` for compatible runtimes.

Apache 2.0128K ctxQ4_K_M70 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-122B-A10B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 122B-A10B (MoE) (MTP)

122B (MoE, 10B active)

Qwen 3.5 122B-A10B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL70 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-122B-A10B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

MiniMax M2.7 (MoE)

230B (MoE, 10B active)MoE

MiniMax M2.7 — current frontier MiniMax MoE; Lightning Attention with 1M context. Unsloth dynamic UD-Q4_K_XL GGUF (sharded).

MiniMax Open1024K ctxUD-Q4_K_XL130 GB
0 downloadable0 serving inference
HuggingFace
unsloth/MiniMax-M2.7-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

DeepSeek V4 Flash (MoE)

284B (MoE, 13B active)MTPMoE

DeepSeek V4 Flash — 284B total / 13B active MoE; 1M context. Hybrid Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA). Cost-effective frontier variant. Pre-trained on 32T tokens. MTP head shipped natively.

MIT1024K ctxQ4K-imatrix144 GB
0 downloadable0 serving inference
HuggingFace
antirez/deepseek-v4-gguf
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

MiniMax M3 (MoE, native multimodal)

428B (MoE, 23B active)MoE

MiniMax M3 — ~428B total / ~23B active MoE with native multimodal training. MiniMax Sparse Attention (MSA) delivers 9× prefill and 15× decode speedups vs M2 at 1M context. Note: GGUF builds currently fall back to dense attention; sparse attention not yet supported in llama.cpp.

MIT1024K ctxQ4_K_M214 GB
0 downloadable0 serving inference
HuggingFace
unsloth/MiniMax-M3-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 397B-A17B (MoE)

397B (MoE, 17B active)MoE

Qwen 3.5 frontier MoE — 397B total, 17B active per token. Multi-GPU replicas only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-397B-A17B-MTP-GGUF` for compatible runtimes.

Apache 2.0128K ctxQ4_K_M224 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-397B-A17B-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Qwen 3.5 397B-A17B (MoE) (MTP)

397B (MoE, 17B active)

Qwen 3.5 397B-A17B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0128K ctxUD-Q4_K_XL224 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Qwen3.5-397B-A17B-MTP-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

DeepSeek V3 0324 (MoE)

685B (MoE, 37B active)MTPMoE

DeepSeek V3 MoE — 685B total, 37B active, 128K context. Native Multi-Token-Prediction head (n=4, ~80% accept rate, ~1.8× decode speedup per DeepSeek tech report). Retired by upstream after 2026-07-24 in favor of DeepSeek V4.

MIT128K ctxQ4_K_M352 GB
0 downloadable0 serving inference
HuggingFace
unsloth/DeepSeek-V3-0324-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GLM-5 (MoE)

744B (MoE)MoE

Z.ai GLM-5 — 744B total parameter MoE trained on 28.5T tokens; best-in-class open-source performance on reasoning, coding, and agentic tasks (2026-04). Unsloth dynamic UD-Q4_K_XL GGUF (sharded).

MIT128K ctxUD-Q4_K_XL373 GB
0 downloadable0 serving inference
HuggingFace
unsloth/GLM-5-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GLM-5.1 (MoE)

744B (MoE, 40B active)MoE

Z.ai GLM-5.1 — next-generation flagship for agentic engineering, class-leading on SWE-Bench Pro; 744B total / 40B active; 200K context. `glm_moe_dsa` architecture with Dynamic Sparse Attention.

MIT200K ctxUD-Q4_K_M373 GB
0 downloadable0 serving inference
HuggingFace
unsloth/GLM-5.1-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

GLM-5.2 (MoE, MTP)

753B (MoE)MTPMoE

Z.ai GLM-5.2 — 753B total parameter MoE flagship with solid 1M-token context and IndexShare sparse-attention (2.9× per-token FLOP reduction at 1M). Improved Multi-Token-Prediction layer increases speculative-decoding accept rate by ~20% over GLM-5.1.

MIT1024K ctxUD-Q4_K_XL382 GB
0 downloadable0 serving inference
HuggingFace
unsloth/GLM-5.2-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Kimi K2.5 (MoE)

1T (MoE, 32B active)MoE

Moonshot AI Kimi K2.5 — 1T total / 32B active MoE; image input support; 256K context. Predecessor to K2.6's hybrid-thinking variant.

MIT256K ctxUD-Q4_K_XL540 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Kimi-K2.5-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Kimi K2.7 Code (MoE)

1T (MoE, 32B active, code-focused)MoE

Moonshot AI Kimi K2.7 Code — code-focused refresh of the K2 series. 1T total / 32B active; 256K context; recent updates target tool-call accuracy on long-horizon coding tasks.

MIT256K ctxUD-Q4_K_XL540 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Kimi-K2.7-Code-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.

Kimi K2.6 (Hybrid Thinking, MoE)

1T (MoE, hybrid thinking)MoE

Moonshot AI Kimi K2.6 hybrid-thinking MoE — 1T total params, 256K context. Replica-routed on B200-class infrastructure; Unsloth measures >40 t/s on B200. Recommended `UD-Q2_K_XL` (350GB) for size/quality balance.

MIT256K ctxUD-Q4_K_XL559 GB
0 downloadable0 serving inference
HuggingFace
unsloth/Kimi-K2.6-GGUF
Centralized. Download from the Hub.
Tenzro Network
Not yet on the network
0 holders. Any operator can hold it and record its hash.