Why goodput matters more than throughput for LLM serving
ARCHITECT FEATURED ANALYSIS

Why goodput matters more than throughput for LLM serving

BY

Graziano Casto, Akamas

SOURCE

Cloud Native Computing Foundation

DATE

READ

1 min read

When we benchmark an LLM serving setup, the number almost everyone reaches for first is throughput: how many requests per second the system can push through. It is easy to measure, easy to compare, and it…

When we benchmark an LLM serving setup, the number almost everyone reaches for first is throughput: how many requests per second the system can push through. It is easy to measure, easy to compare, and it…