Skip to content

Benchmark

vLLM Under Concurrent Load

Exploring time-to-first-token, throughput, memory usage and latency as concurrent requests increase.

1 min read
  • AI Infrastructure
  • Performance

Why inference is a different shape of workload

A conventional API scales roughly linearly with concurrency: more requests, more work, more replicas. Inference does not. Batching means throughput can improve as concurrency rises while per-request latency gets worse — an unusual tradeoff that is easy to misread.

What I measure

Metrics recorded across concurrency levels.

text
Concurrency
 ↓
Time to first token
 ↓
Inter-token latency
 ↓
Token throughput
 ↓
GPU memory utilisation

I sweep concurrency while holding the model, prompt distribution and output length fixed, watching how the four metrics above move together.

What I expect to find

There is usually a concurrency level where throughput stops improving and latency starts to climb steeply. Finding that knee — and knowing what causes it — is the point of the experiment.

This write-up is being expanded with the full results table, hardware specification and the exact serving configuration.

Related