
The model ID lottery: same request, different draw
Behind a multi-provider gateway, the same model ID produced its first token at a 312 ms median with no reasoning output one day, and at 3,073 ms with 2,627 characters of reasoning days later. The routing flag we expected to prevent this did not.
