Benchmark
vLLM Under Concurrent Load
Exploring time-to-first-token, throughput, memory usage and latency as concurrent requests increase.
- AI Infrastructure
- Performance
Why inference is a different shape of workload
A conventional API scales roughly linearly with concurrency: more requests, more work, more replicas. Inference does not. Batching means throughput can improve as concurrency rises while per-request latency gets worse — an unusual tradeoff that is easy to misread.
What I measure
Metrics recorded across concurrency levels.
Concurrency
↓
Time to first token
↓
Inter-token latency
↓
Token throughput
↓
GPU memory utilisationI sweep concurrency while holding the model, prompt distribution and output length fixed, watching how the four metrics above move together.
What I expect to find
There is usually a concurrency level where throughput stops improving and latency starts to climb steeply. Finding that knee — and knowing what causes it — is the point of the experiment.
This write-up is being expanded with the full results table, hardware specification and the exact serving configuration.