Observability and Performance Debugging
When production misbehaves, you need to find out what is wrong, where, and why, quickly. Observability is the ability to answer those questions from the telemetry a system emits. Interviewers ask "the API is slow, what do you do?", how you would design alerts, and what percentiles and SLOs are. This chapter covers the three pillars, the metrics that matter, a debugging method and common bottlenecks, with small runnable models.
1. Monitoring versus observability
Monitoring watches known failure modes (is the CPU high, is the error rate above 1%). Observability lets you investigate unknown problems by asking new questions of rich data without shipping new code. You need both: monitoring tells you something is wrong; observability tells you why.
2. The three pillars
| Signal | What it is | Good for | Cost |
|---|---|---|---|
| Logs | discrete, timestamped records of events | detail about a specific event, errors, audit trails | volume and storage; hard to aggregate unless structured |
| Metrics | numeric measurements aggregated over time | dashboards, alerts, trends, capacity planning | cheap and fast, but lose detail; beware high-cardinality labels |
| Traces | the path of one request through services, as a tree of timed spans | where latency and errors come from across services | sampling needed at scale; requires instrumentation and propagation |
(Newer practice adds profiles, continuous CPU and memory profiling, and events, wide structured records with many fields per request.)
Logging well
- Structured logs (JSON): fields such as
timestamp,level,service,trace_id,user_id,route,status,duration_ms,error, instead of free text that must be parsed. - Levels used consistently:
ERRORfor failures needing attention,WARNfor suspicious,INFOfor key business events,DEBUGoff in production by default. - Correlation IDs passed through every call so one request can be followed across services.
- Never log secrets or personal data (passwords, tokens, full card numbers); mask or omit.
- One log line per request with the outcome and duration (an access or "canonical" log) is far more useful than scattered lines.
- Log context for errors (inputs, identifiers, the cause chain), not just the message.
import json, time, uuid
def log_event(level, message, **fields):
record = {"ts": 1_700_000_000, "level": level, "msg": message, **fields}
return json.dumps(record, sort_keys=True)
line = log_event("ERROR", "payment failed", trace_id="abc123", user_id=42, route="/checkout", status=502, duration_ms=1840, error="gateway timeout")
parsed = json.loads(line)
assert parsed["trace_id"] == "abc123" and parsed["status"] == 502 # machine-readable: filter by any field
assert "password" not in line
3. Metrics: what to measure
The golden signals (Google SRE)
- Latency: how long requests take (separate successful from failed requests).
- Traffic: demand, such as requests per second.
- Errors: the rate of failed requests (explicit 5xx, wrong content, slow responses counted as failures).
- Saturation: how full the system is (CPU, memory, connection pool, queue depth, disk).
Related frameworks: RED for services (Rate, Errors, Duration) and USE for resources (Utilisation, Saturation, Errors).
Percentiles, not averages
The average hides the slow tail. If 99 requests take 100 ms and one takes 10 seconds, the mean is about 200 ms while one user in a hundred waits 10 seconds. Track p50, p95, p99 (and p99.9 for critical paths). Tail latency also compounds: if a page makes 20 backend calls in parallel and each has a 1% chance of being slow, the page is slow about 18% of the time.
import random, math
def percentile(data, p):
s = sorted(data)
k = math.ceil(p / 100 * len(s)) - 1
return s[max(0, k)]
random.seed(7)
latencies = [random.gauss(100, 10) for _ in range(990)] + [random.gauss(5000, 500) for _ in range(10)]
mean = sum(latencies) / len(latencies)
assert 100 < mean < 200 # the mean looks acceptable...
assert percentile(latencies, 50) < 110
assert percentile(latencies, 99.5) > 4000 # ...while the tail is terrible: the mean hid it
# tail latency compounds across fan-out
p_slow = 0.01
fanout = 20
assert abs((1 - (1 - p_slow) ** fanout) - 0.182) < 0.001 # about 18 % of page loads hit at least one slow call
Histograms (Prometheus histograms, HdrHistogram, t-digest) let you compute percentiles from aggregated data; do not average percentiles across instances, aggregate the histograms.
Metric types
- Counter: only goes up (requests served). Use rates over time.
- Gauge: goes up and down (queue depth, memory in use).
- Histogram / summary: a distribution of observations (latency).
Cardinality is the number of unique label combinations. Labels such as user_id or request_id create millions of series and can break a metrics system; put high-cardinality detail in logs and traces, not metric labels.
4. Distributed tracing
A trace is the whole journey of a request; a span is one timed operation inside it (an HTTP handler, a database query, a downstream call), with a parent-child structure. Propagate the trace context in headers (the W3C traceparent header) across services and through message queues. Traces show where time goes: a request that took 900 ms might show 40 ms in the API, 800 ms waiting on one downstream call, and inside that a slow query.
OpenTelemetry is the vendor-neutral standard for instrumenting logs, metrics and traces, with exporters to backends such as Jaeger, Tempo, Datadog, New Relic, Honeycomb and cloud vendors. At scale you sample traces (head-based: decide at the start; tail-based: keep the slow and failed ones).
def critical_path(spans):
"""spans: name -> (start, end). Returns the span names that account for the end-to-end time with no overlap, sorted by duration."""
total = max(e for s, e in spans.values()) - min(s for s, e in spans.values())
durations = sorted(((e - s, n) for n, (s, e) in spans.items()), reverse=True)
return total, durations
trace = {"api": (0, 900), "auth": (5, 45), "inventory": (50, 130), "payment": (135, 880), "payment.db_query": (140, 870)}
total, ranked = critical_path(trace)
assert total == 900
assert ranked[1][1] == "payment" and ranked[2][1] == "payment.db_query" # almost all the time hides in the payment call and its query
5. SLIs, SLOs, SLAs and error budgets
- SLI (indicator): a measured quantity, such as the proportion of requests answered successfully under 300 ms.
- SLO (objective): the target, such as 99.9% of requests meet the SLI over 30 days.
- SLA (agreement): a contract with consequences (credits) if you miss a threshold; usually looser than the internal SLO.
- Error budget: , the amount of unreliability you may "spend". A 99.9% SLO over 30 days allows about 43 minutes of downtime. When the budget is spent, slow down releases and invest in reliability; when plentiful, ship faster.
def downtime_budget_minutes(slo, days=30):
return (1 - slo) * days * 24 * 60
assert abs(downtime_budget_minutes(0.999) - 43.2) < 1e-9
assert abs(downtime_budget_minutes(0.9999) - 4.32) < 1e-9
assert abs(downtime_budget_minutes(0.99) - 432) < 1e-9
def burn_rate(error_rate, slo):
return error_rate / (1 - slo) # 1.0 spends the budget exactly over the window; 14 spends it in about two days
assert abs(burn_rate(0.014, 0.999) - 14) < 1e-9
Alert on symptoms users feel (SLO burn), not on every cause. A page should mean "a human must act now". Use multi-window burn-rate alerts (a fast burn over a short window plus a slower confirmation window), keep alerts actionable with runbooks, and delete noisy ones. Alert fatigue makes real incidents get missed.
6. A method for "the API is slow"
- Define the problem: what is slow (which endpoint, which users, which region), since when, how slow (p50 or p99), and what changed (a deploy, traffic, data growth, a dependency)? Check dashboards and recent releases first.
- Locate it: use traces and metrics to see where time goes: the client, the network, the load balancer, the service, a dependency, the database. Compare to a healthy period.
- Check the usual suspects (below) in order of likelihood.
- Form a hypothesis, test it, change one thing, and measure again.
- Mitigate first, then fix the root cause: roll back, scale out, enable a feature flag, shed load, increase a cache TTL, then investigate properly.
- Write it up and add the missing alert or test.
The usual suspects
| Area | Symptoms | Checks |
|---|---|---|
| Database | slow queries, high CPU or I/O, lock waits, connection pool exhausted | slow query log, EXPLAIN, missing indexes, N+1 patterns, large scans, lock and deadlock metrics, replica lag |
| Downstream dependency | latency rises with one call in the trace | its latency and errors, timeouts, retries amplifying load |
| Resource saturation | CPU at 100%, memory pressure, swapping, GC pauses, file descriptor or thread exhaustion | USE method on each resource, GC logs, thread dumps, pool metrics |
| Concurrency problems | latency spikes under load, lock contention | profiler, thread dumps, queue lengths |
| Cache issues | hit rate dropped, stampedes after expiry, cold cache after deploy | hit ratio, eviction counts, expiry patterns |
| Network | high p99 only for some paths, retransmits, DNS delays | connection reuse (keep-alive), TLS handshakes, DNS, cross-region calls, payload size |
| Payload and serialisation | big responses, slow JSON handling | response sizes, compression, pagination |
| Garbage collection / memory leaks | growing memory then pauses or OOM kills | heap profiles, allocation rates |
| Deploy or config change | a step change at a timestamp | release markers on dashboards, config diffs |
| Traffic change | more requests, bots, a retry storm, a new client | traffic by client and endpoint, rate limits |
7. Profiling and measuring code
Do not guess where time goes; measure with a profiler (CPU flame graphs, allocation profiles, async profilers). Optimise the dominant cost first, and verify with a benchmark.
import cProfile, pstats, io
def slow_unique(items):
seen = []
for x in items:
if x not in seen: # O(n) membership test on a list: quadratic overall
seen.append(x)
return seen
def fast_unique(items):
seen, out = set(), []
for x in items:
if x not in seen: # O(1) on average
seen.add(x); out.append(x)
return out
data = list(range(3000)) * 2
assert slow_unique(data) == fast_unique(data) # same behaviour
def timed(fn):
import time
t = time.perf_counter(); fn(data); return time.perf_counter() - t
assert timed(fast_unique) * 5 < timed(slow_unique) # the right data structure wins by a wide margin
prof = cProfile.Profile(); prof.enable(); slow_unique(data); prof.disable()
buf = io.StringIO(); pstats.Stats(prof, stream=buf).sort_stats("cumulative").print_stats(3)
assert "slow_unique" in buf.getvalue() # the profiler points at the function that dominates the run
Common code-level wins: better algorithms and data structures, avoiding repeated work (caching, batching), fewer round trips (N+1, chatty calls), streaming instead of loading everything into memory, connection reuse, compression, and moving slow work off the request path (queues).
8. Capacity planning and load testing
- Load test with realistic traffic shapes (read/write mix, think time, data distribution) using tools such as k6, Gatling, Locust or JMeter, against a production-like environment.
- Find the saturation point (where latency climbs sharply) and the failure mode (does it degrade gracefully or collapse?).
- Run soak tests (hours) to expose leaks and spike tests for sudden surges.
- Use Little's law and the utilisation curve: queueing delay explodes as utilisation approaches 100%, so run servers well below saturation (around 60 to 70%) to protect latency.
def mm1_wait_ratio(utilisation):
return 1 / (1 - utilisation) # for an M/M/1 queue, response time = service time / (1 - utilisation)
assert abs(mm1_wait_ratio(0.5) - 2) < 1e-9
assert abs(mm1_wait_ratio(0.9) - 10) < 1e-9
assert mm1_wait_ratio(0.99) > 99 # at 99 % busy, a request takes 100 times its service time
9. Incident response
- Detect (alert on SLO burn), declare an incident, assign an incident commander, open a channel.
- Mitigate: stop the bleeding (rollback, flag off, scale, failover) before hunting the root cause.
- Communicate status regularly to stakeholders and customers.
- Resolve and verify with the SLIs.
- Postmortem: blameless, with a timeline, root causes and contributing factors, what went well, and tracked action items.
10. Common mistakes
- Averages instead of percentiles.
- Unstructured logs with no correlation IDs.
- High-cardinality metric labels that overload the metrics system.
- Alerting on causes (CPU 80%) instead of user-visible symptoms, creating noise.
- Optimising without profiling, or tuning the wrong layer.
- No dashboards or release markers, so "what changed?" is guesswork.
- Logging secrets or personal data.
- Running at near 100% utilisation and being surprised by latency.
- Fixing the symptom and skipping the postmortem.
11. Practice questions
- What are the three pillars of observability, and when do you reach for each?
- Why use percentiles instead of averages? What is tail latency amplification?
- Define SLI, SLO, SLA and error budget. How would you set an SLO for a checkout API?
- The API's p99 doubled after a deploy. Walk through your investigation.
- What are the golden signals, RED and USE?
- How do you trace a request across services and queues?
- What is metric cardinality and why does it matter?
- How would you design alerts for a payment service so they are actionable?