Experiment
Measuring OpenTelemetry Overhead
An experiment measuring the application cost of increasingly aggressive tracing configurations.
- Observability
- Performance
The tradeoff
Tracing is not free. Every span is work the application does instead of serving the request. The question is not whether there is a cost, but how the cost scales as sampling and instrumentation get more aggressive.
What I vary
Configuration axis under test.
Sampling
├── 1%
├── 10%
└── 100%
Instrumentation depth
├── HTTP only
├── + database
└── + internal functions
Exporter
├── batch
└── per-requestFor each combination I measure throughput, p50/p99 latency, and CPU time, then compare against a baseline with tracing disabled.
What I expect to measure
Head sampling changes the volume of data, not the per-span cost. Instrumentation depth changes the per-request cost. Those two numbers behave very differently, and conflating them leads to the wrong optimisation.
This write-up is being expanded with the full results table and the exact collector configuration.