
Benchmarks & Reports / Sample
AI API Latency Benchmark: TTFT, Throughput and Reliability
A transparent methodology for measuring AI API latency without confusing time-to-first-token, generation speed and total response time.
TTFT measures responsiveness; tokens per second measures generation speed.
Use the same model version, prompt, output target and streaming mode for every route.
Report median and p95 values instead of one fastest request.
Publish errors, timeouts, region and test window alongside performance numbers.