GLM-5.3-Flash
Z.ai's natively multimodal Flash model. 320B total, 18B active, 1M context, MIT weights. Previewed as Ox Alpha.
Related builds
GLM-5.3 open weights
The full GLM-5.3 coding and cyber-defense model is now downloadable on Hugging Face. Same base as 5.2, post-trained.
Qwen3.8-Flash-Next
Alibaba's multimodal MoE and a preview of Qwen4. 125B plus 51B n-gram embeddings, 6B active, 262K native context. Open-weight.
GLM-5.3-Flash GGUF
Unsloth 3-bit GGUF so GLM-5.3-Flash runs on 128GB RAM. First-party local pack.
GLM-5.3 GGUF
Unsloth 2-bit GGUF of the full GLM-5.3. 1.51TB down to 239GB, for a 256GB Mac or mixed RAM/VRAM.