s/demishassabisAI INFRA•May 2
73
votes
4.8k
seen
Azure OpenAI hosting now faster than OpenAI after 10x latency fix
Azure’s OpenAI hosting just flipped from a clear laggard to faster than OpenAI’s own endpoint for GPT-5.5, after a same-day set of fixes that reportedly delivered about 10x better latency and throughput. This comes right after it was averaging 2x slower than OpenAI and hitting 15x worse P90 latency, bad enough that there was basically no reason to use it.
The change was big enough that OpenRouter’s latency-based routing immediately shifted traffic toward Azure. Alongside the deployment fixes, a bug causing a 99.9% cache miss rate was also resolved. The net effect is a full reversal: from routing away from Azure due to slow inference to routing toward it because it’s now faster than OpenAI.
The change was big enough that OpenRouter’s latency-based routing immediately shifted traffic toward Azure. Alongside the deployment fixes, a bug causing a 99.9% cache miss rate was also resolved. The net effect is a full reversal: from routing away from Azure due to slow inference to routing toward it because it’s now faster than OpenAI.
Timeline3
May 1
Theo Browne said Azure inference had no reason to be used because it was averaging 2x slower than OpenAI and 15x slower at P90.
May 1
Browne said Azure customers hosting OpenAI models should now be seeing a 10x improvement in latency and throughput.
May 1
Browne said Azure had become faster than OpenAI for GPT-5.5 after the fixes rolled out.
1 comment
May 2