
ARCHITECT
FEATURED ANALYSIS
Why goodput matters more than throughput for LLM serving
BY
Graziano Casto, Akamas
SOURCE
Cloud Native Computing Foundation
DATE
READ
1 min read
When we benchmark an LLM serving setup, the number almost everyone reaches for first is throughput: how many requests per second the system can push through. It is easy to measure, easy to compare, and it…
When we benchmark an LLM serving setup, the number almost everyone reaches for first is throughput: how many requests per second the system can push through. It is easy to measure, easy to compare, and it…