๐Ÿชท Zen LM
Models

Zen 5 โ€” Mixture of Diverse Experts

3.1T parameters, 0.8B to 100B+ active. Complexity-aware routing across 6 diverse model families. Training openly on Hanzo Network.

Zen 5 โ€” Mixture of Diverse Experts (MoDE)

Status: Training Q2-Q4 2026 ยท Total parameters: 3.1T ยท Active parameters: 0.8B to 100B+ ยท GitHub: zenlm/zen5 ยท Open training: hanzo.network

zen5 introduces MoDE (Mixture of Diverse Experts) โ€” a novel architecture that routes across frozen expert modules harvested from the world's largest open-source models with complexity-aware hierarchical routing.

Unlike standard models that always use the same compute, zen5 adapts compute to task difficulty:

TierActive ParamsLatencyExample
T0 โ€” Trivial0.8B<50ms"Hello, how are you?"
T1 โ€” Standard9B<200ms"Summarize this article"
T2 โ€” Complex10-40B<2s"Compare economic policies with examples"
T3 โ€” Advanced40-80B<5s"Derive Black-Scholes from first principles"
T4 โ€” Frontier80-100B+<10s"Design a novel consensus algorithm"

Expert Pool โ€” 3.1T+ Parameters

Frozen experts harvested from 6 diverse model families:

Expert SourceParametersTypeArchitectureTier
Qwen3.5-0.8B0.8BDenseGated DeltaNetT0
Qwen3.5-9B9BDenseGated DeltaNetT1
MiniMax-M2.5230BMoE (64 experts)Lightning AttentionT2
GLM-5744BMoE (256 experts)GLM MoET3
Kimi K2.51.04TMoE (384 experts)DeepseekV3 MoET4
Ling-1T1TMoEFP8 MoET4

Omnimodal Experts

ModalityExpertArchitecture
VisionQwen3-VLNative Vision Transformer
VideoWan2.2 / CogVideoXMoE Video Diffusion
3DTRELLIS.2Rectified Flow DiT + SC-VAE
AudioQwen3-Omni / zen-ttsThinker-Talker

Key Innovations

  1. Complexity-Aware Routing โ€” 207M parameter estimator analyzes the first 64 tokens and routes to the optimal tier. 85%+ of real-world queries served at T0-T1, reducing inference cost 10-50x.

  2. Cross-Architecture Alignment โ€” 134M parameter projection layer unifies expert representations from 6 diverse architectures into a shared latent space โ€” a technical first.

  3. MoE++ Zero Experts (ICLR 2025 Oral) โ€” Zero/Copy/Constant experts handle trivial tokens with zero computation, delivering 1.1-2.1x throughput.

  4. ReMoE ReLU Routing (ICLR 2025) โ€” Replaces fixed Top-K with ReLU activation for adaptive expert count per token.

  5. Adaptive Escalation โ€” If confidence drops mid-generation, zen5 promotes to a higher tier automatically.

  6. Transfusion โ€” Unified autoregressive (text) and diffusion (3D/video) in a single forward pass.

Training Efficiency

Only 394M parameters are trained โ€” 0.013% of total model:

ComponentParametersPurpose
ComplexityEstimator207MPredicts task difficulty tier (T0-T4)
AlignmentLayer134MCross-architecture projection to shared space
MoDERouter52MPer-tier expert routing with ReLU gating
Total trainable394M0.013% of 3.1T total

All expert weights are frozen โ€” extracted and served as-is from source models.

Estimated training cost: $200-350K (vs $500M-1B+ to train from scratch).

zen5 Model Lineup

ModelTotal ParamsActive ParamsTargetModalities
zen5750B0.8-50BGeneralText
zen5-coder1.8T0.8-80BCodeText + Code
zen5-omni2.5T0.8-100BOmnimodalText + Vision + Video + 3D + Audio
zen5-max3.1T0.8-100B+FrontierAll

All zen5 variants released under Apache 2.0.

Training Phases

PhaseDescriptionTimeline
1. Expert ExtractionHarvest FFN/MoE blocks from all 6 source modelsQ1-Q2 2026
2. AlignmentCross-architecture projection trainingQ2 2026
3. Router TrainingComplexityEstimator + per-tier MoDERouterQ2-Q3 2026
4. IntegrationJoint fine-tuning with frozen expert poolQ3 2026
5. Omnimodal FusionVision, video, 3D, audio expert integrationQ4 2026

Open Training on Hanzo Network

zen5 is trained openly on Hanzo Network โ€” decentralized compute with verifiable execution via NVIDIA Trusted Execution Environments (TEE).

All checkpoints, data pipelines, and training logs are published publicly.

Quick Start

git clone https://github.com/zenlm/zen5
cd zen5

# List expert pool
python scripts/expert_extraction.py list

# Train complexity router
python scripts/router_training.py train --output ./checkpoints/

# Benchmark routing accuracy
python scripts/benchmark.py routing --checkpoint ./checkpoints/router/best.pt

Contribute Training Data

We are collecting agentic training data for zen5:

  • Multi-step reasoning traces (chain-of-thought, tree-of-thought)
  • Tool use sequences (function calling, code execution)
  • Agent trajectories (task โ†’ plan โ†’ action โ†’ observation loops)
  • Domain-specific expert demonstrations
pip install hanzoai

from hanzoai import TrainingClient
client = TrainingClient(api_key="sk-...")
client.submit_trace(
    messages=[
        {"role": "user", "content": "Solve this step by step: ..."},
        {"role": "assistant", "content": "<thinking>...</thinking>\n\nFinal answer: ..."},
    ],
    quality_score=0.9,
    domain="reasoning",
)

Research Foundations

See Also

On this page