
The non-inferiority swap: how we ship model changes on a quality tie
Three pre-registered decision rules moved three chat surfaces to GLM 5.3 on 2026-09-01: the think flip (Grok 4.6 to GLM 5.3) scored 1.7 points lower on our four-task quality set and cut median first-token latency from 88.9 seconds to 4.2 on our latency probe.
