Zen 5 โ Mixture of Diverse Experts
3.1T parameters, 0.8B to 100B+ active. Complexity-aware routing across 6 diverse model families. Training openly on Hanzo Network.
Zen 5 โ Mixture of Diverse Experts (MoDE)
Status: Training Q2-Q4 2026 ยท Total parameters: 3.1T ยท Active parameters: 0.8B to 100B+ ยท GitHub: zenlm/zen5 ยท Open training: hanzo.network
zen5 introduces MoDE (Mixture of Diverse Experts) โ a novel architecture that routes across frozen expert modules harvested from the world's largest open-source models with complexity-aware hierarchical routing.
Unlike standard models that always use the same compute, zen5 adapts compute to task difficulty:
| Tier | Active Params | Latency | Example |
|---|---|---|---|
| T0 โ Trivial | 0.8B | <50ms | "Hello, how are you?" |
| T1 โ Standard | 9B | <200ms | "Summarize this article" |
| T2 โ Complex | 10-40B | <2s | "Compare economic policies with examples" |
| T3 โ Advanced | 40-80B | <5s | "Derive Black-Scholes from first principles" |
| T4 โ Frontier | 80-100B+ | <10s | "Design a novel consensus algorithm" |
Expert Pool โ 3.1T+ Parameters
Frozen experts harvested from 6 diverse model families:
| Expert Source | Parameters | Type | Architecture | Tier |
|---|---|---|---|---|
| Qwen3.5-0.8B | 0.8B | Dense | Gated DeltaNet | T0 |
| Qwen3.5-9B | 9B | Dense | Gated DeltaNet | T1 |
| MiniMax-M2.5 | 230B | MoE (64 experts) | Lightning Attention | T2 |
| GLM-5 | 744B | MoE (256 experts) | GLM MoE | T3 |
| Kimi K2.5 | 1.04T | MoE (384 experts) | DeepseekV3 MoE | T4 |
| Ling-1T | 1T | MoE | FP8 MoE | T4 |
Omnimodal Experts
| Modality | Expert | Architecture |
|---|---|---|
| Vision | Qwen3-VL | Native Vision Transformer |
| Video | Wan2.2 / CogVideoX | MoE Video Diffusion |
| 3D | TRELLIS.2 | Rectified Flow DiT + SC-VAE |
| Audio | Qwen3-Omni / zen-tts | Thinker-Talker |
Key Innovations
-
Complexity-Aware Routing โ 207M parameter estimator analyzes the first 64 tokens and routes to the optimal tier. 85%+ of real-world queries served at T0-T1, reducing inference cost 10-50x.
-
Cross-Architecture Alignment โ 134M parameter projection layer unifies expert representations from 6 diverse architectures into a shared latent space โ a technical first.
-
MoE++ Zero Experts (ICLR 2025 Oral) โ Zero/Copy/Constant experts handle trivial tokens with zero computation, delivering 1.1-2.1x throughput.
-
ReMoE ReLU Routing (ICLR 2025) โ Replaces fixed Top-K with ReLU activation for adaptive expert count per token.
-
Adaptive Escalation โ If confidence drops mid-generation, zen5 promotes to a higher tier automatically.
-
Transfusion โ Unified autoregressive (text) and diffusion (3D/video) in a single forward pass.
Training Efficiency
Only 394M parameters are trained โ 0.013% of total model:
| Component | Parameters | Purpose |
|---|---|---|
| ComplexityEstimator | 207M | Predicts task difficulty tier (T0-T4) |
| AlignmentLayer | 134M | Cross-architecture projection to shared space |
| MoDERouter | 52M | Per-tier expert routing with ReLU gating |
| Total trainable | 394M | 0.013% of 3.1T total |
All expert weights are frozen โ extracted and served as-is from source models.
Estimated training cost: $200-350K (vs $500M-1B+ to train from scratch).
zen5 Model Lineup
| Model | Total Params | Active Params | Target | Modalities |
|---|---|---|---|---|
| zen5 | 750B | 0.8-50B | General | Text |
| zen5-coder | 1.8T | 0.8-80B | Code | Text + Code |
| zen5-omni | 2.5T | 0.8-100B | Omnimodal | Text + Vision + Video + 3D + Audio |
| zen5-max | 3.1T | 0.8-100B+ | Frontier | All |
All zen5 variants released under Apache 2.0.
Training Phases
| Phase | Description | Timeline |
|---|---|---|
| 1. Expert Extraction | Harvest FFN/MoE blocks from all 6 source models | Q1-Q2 2026 |
| 2. Alignment | Cross-architecture projection training | Q2 2026 |
| 3. Router Training | ComplexityEstimator + per-tier MoDERouter | Q2-Q3 2026 |
| 4. Integration | Joint fine-tuning with frozen expert pool | Q3 2026 |
| 5. Omnimodal Fusion | Vision, video, 3D, audio expert integration | Q4 2026 |
Open Training on Hanzo Network
zen5 is trained openly on Hanzo Network โ decentralized compute with verifiable execution via NVIDIA Trusted Execution Environments (TEE).
All checkpoints, data pipelines, and training logs are published publicly.
Quick Start
git clone https://github.com/zenlm/zen5
cd zen5
# List expert pool
python scripts/expert_extraction.py list
# Train complexity router
python scripts/router_training.py train --output ./checkpoints/
# Benchmark routing accuracy
python scripts/benchmark.py routing --checkpoint ./checkpoints/router/best.ptContribute Training Data
We are collecting agentic training data for zen5:
- Multi-step reasoning traces (chain-of-thought, tree-of-thought)
- Tool use sequences (function calling, code execution)
- Agent trajectories (task โ plan โ action โ observation loops)
- Domain-specific expert demonstrations
pip install hanzoai
from hanzoai import TrainingClient
client = TrainingClient(api_key="sk-...")
client.submit_trace(
messages=[
{"role": "user", "content": "Solve this step by step: ..."},
{"role": "assistant", "content": "<thinking>...</thinking>\n\nFinal answer: ..."},
],
quality_score=0.9,
domain="reasoning",
)Research Foundations
- MoE++ (ICLR 2025 Oral) โ Zero-Computation Experts
- ReMoE (ICLR 2025) โ ReLU-based MoE routing
- HMoE (EMNLP 2025) โ Heterogeneous expert sizes
- Symbolic-MoE (2025) โ Skill-based routing
- Uni-MoE-2.0 โ Dynamic capacity multimodal MoE
- Transfusion (Meta 2024) โ Unified AR + diffusion
See Also
- zen5 GitHub โ Full source code and architecture doc
- Architecture Paper โ Coming soon
- Open Training Blog Post โ Technical overview
- Hanzo Network โ Live training dashboard