Umans GLM 5.2 NVFP4 Playground Retired
playground experiment; the test window ran Jun 29 to Jul 2, 2026
189.5tok/s
throughput · p50 · whole period
1.57s
TTFT · p50 · whole period
100.00%
uptime · whole period
Experimental NVFP4-quantized build of GLM 5.2 - offered for testing only. The published NVFP4 results look very flattering, but we're skeptical of its real quality and performance: this checkpoint was NOT QAT post-trained for NVFP4 (the QAT-on-NVFP4 models are where we've seen the best quality), so treat the benchmarks with caution. Play with it, push it, and see how far it gets you - for production work we recommend the fp8 `umans-glm-5.2`. It runs on a single low-capacity GPU with no fallback, so expect it to go down under load: crash it, let it restart, and play again.
Trends
Speed over its final 90 days
Jun 27, 2026retired Jul 2, 2026Jul 4, 2026
Jun 27, 2026retired Jul 2, 2026Jul 4, 2026
Changelog
Events for Umans GLM 5.2 NVFP4
No recent events.
Older events 2
Jul 22026
Playground closed: Umans GLM 5.2 NVFP4 Testing
The short NVFP4 test window ended after four days. Thanks to everyone who pushed it and shared findings.
Jun 292026
Playground opened: Umans GLM 5.2 NVFP4 Testing
umans-glm-5.2-nvfp4 entered the playground for a short, low-capacity test window. Experimental and temporary; not for production.