Find and Fix a Performance Bottleneck
Activated Cloud✓ Officialactivated/find-performance-bottleneck
Free · MIT
About
Makes slow code, pages, queries or services measurably faster: define the metric and target, measure a repeatable baseline, profile to find where the time actually goes (CPU, I/O, database, network, rendering), fix the biggest cost first, and prove the gain with before-and-after numbers. Covers profilers for Python, Node, Go and Rust, SQL query plans, N+1 queries, benchmarking hygiene and Core Web Vitals. Use when something is slow, times out, or costs too much to run. Not for active outages (use incident-first-response).
Documentation
Find and Fix a Performance Bottleneck
Performance work without measurement is guessing, and guesses about where time goes are wrong surprisingly often. The loop is: pick a metric, measure a baseline you can repeat, profile to find the biggest cost, fix that one thing, measure again. A change that does not move the number gets reverted, however clever it is.
When to use
- "This page is slow", "the export times out", "the API p95 doubled", "the job takes four hours", "the server bill is too high".
- A query shows up as slow in logs or the database's statistics.
- Before a launch or an expected traffic increase, to find the first thing that will break.
What you need
- The slow thing, precisely: which endpoint, page, job or function, with which input or account, since when.
- A safe place to measure: local or staging with production-like data volume. Profiling production is allowed only with read-only, low-overhead tools and the owner's agreement; load tests never run against production without the owner's explicit go-ahead.
- Access to existing signals (APM, logs with timings, database statistics) through the owner's connected apps or browser, if the team has them.
- Tool commands per stack:
references/profiling-tools.md.
Method
Define the metric and the target. Name one number: p95 latency of
GET /reportsfor an account with 50,000 orders; wall time of the nightly job; Largest Contentful Paint on the pricing page on a mid-range phone; memory peak of the worker. Set a target with the owner ("p95 under 500 ms") so you know when to stop.Measure a repeatable baseline. Same input, same machine, warm-up runs discarded, several runs, report the median and the spread (p95 or min and max). Examples:
hyperfine --warmup 3 --runs 20 'python scripts/export.py --account 42' for i in $(seq 20); do curl -s -o /dev/null -w '%{time_total}\n' http://localhost:8000/reports; done | sort -n go test -bench=BenchmarkExport -benchmem -count=10 ./export > before.txtMake sure the baseline reproduces the slowness the user saw; if it does not, your test data or environment differs (data volume is the usual culprit).
Locate the layer before the line. Split the total time: client rendering, network, server, database, external calls, queue waits. Server timing logs,
curl -wtiming breakdowns, browser performance traces and database statistics answer this quickly. Optimising Python code when 90 percent of the time is in one SQL query wastes the day.Profile the slow layer. Use a sampling profiler on a realistic run and read the result as a flame graph or top-N table (commands per language in
references/profiling-tools.md). Look for: one function dominating, the same work done repeatedly, time spent waiting (I/O, locks, sleeps), and allocation or garbage-collection churn.Databases: count queries, then explain the slow ones.
- Count queries per request (ORM query logging, Django Debug Toolbar or
assertNumQueries, SQLAlchemy echo, Rails logs). A count that grows with the number of rows is an N+1: fix with eager loading (select_related,prefetch_related,joinedload,includes) or one batched query. - For a slow query, read its plan on production-like data:
EXPLAIN (ANALYZE, BUFFERS) <query>;(ANALYZE executes the query: use a read-only replica or a copy for writes). Signs to act on: sequential scans on large tables with selective filters (missing index), estimated rows far from actual rows (stale statistics:ANALYZE <table>), sorts or hashes spilling to disk, nested loops over many rows, the same subquery executed per row. - Postgres's
pg_stat_statements, where enabled, lists the queries with the highest total time: fix the top of that list first.
- Count queries per request (ORM query logging, Django Debug Toolbar or
Fix the biggest cost with the cheapest effective change, in roughly this order:
- do less work: remove repeated computation, avoid N+1, fetch only needed columns and rows, paginate;
- do it less often: cache with a clear invalidation rule, memoise, batch;
- do it faster: a better algorithm or data structure (a nested loop over two lists replaced by a dict lookup), an index, a bulk operation;
- do it elsewhere: move to a background job or queue, precompute;
- do it in parallel: concurrency for independent I/O, only after the above. Change one thing at a time so you know what helped.
Measure again, same way. Compare with the baseline using the same runs and statistics (for Go,
benchstat before.txt after.txt). Keep the change only if the improvement is clear beyond the noise. Run the full test suite: faster but wrong is a regression.Frontend pages. Measure with Lighthouse or the browser's performance panel on a throttled mobile profile:
npx lighthouse <url> --only-categories=performance --output=json --output-path=./lh.json --chrome-flags="--headless". Look at Core Web Vitals: Largest Contentful Paint, Interaction to Next Paint and Cumulative Layout Shift. Google's published "good" thresholds have been LCP at most 2.5 s, INP at most 200 ms and CLS at most 0.1 at the 75th percentile of page loads; check the current values withweb_searchbefore quoting them. Usual fixes: properly sized and lazily loaded images, preloading the hero image and fonts, removing render-blocking scripts, splitting large bundles (check bundle size before adding a dependency), reserving space for images and embeds, and moving heavy work off the main thread.Watch for performance traps in the fix. Caches need invalidation and memory bounds; indexes slow writes and take disk; parallelism needs limits (connection pools, rate limits of third parties); denormalised data needs a sync path.
Guard the gain. Where it is cheap, add a regression check: a benchmark in CI with a threshold, a query-count assertion in a test, a performance budget for bundle size, an alert on p95.
Output
A short note: the metric and target, the baseline with method, where the time went (profile or plan evidence), the change made, the after measurement with the same method, the improvement in absolute and relative terms, test results, and any trade-offs (memory, write cost, cache staleness). Put a before-and-after table on a show_card when the owner is following along. For example:
Metric: p95 of GET /reports for account 42 (50,000 orders), 20 runs after 3 warm-ups, staging copy.
Target: under 500 ms.
| | Before | After |
|---|---|---|
| p50 | 1,840 ms | 210 ms |
| p95 | 2,310 ms | 290 ms |
| SQL queries per request | 1,203 | 4 |
Cause: N+1 loading line items per order (`reports/service.py:57`), 1,200 queries at about 1.5 ms each.
Fix: one prefetch of line items; plus an index on `order_items(order_id)` (EXPLAIN showed a sequential scan).
Trade-off: the index adds about 40 MB and slightly slower inserts on order_items.
Tests: full suite green; new test asserts at most 5 queries for this endpoint.
Checks before you finish
- Baseline and after numbers were measured the same way, with warm-ups and multiple runs.
- The fix targets the biggest measured cost, shown by a profile, plan or timing breakdown.
- The improvement is larger than run-to-run noise.
- The full test suite passes after the change.
- No load test or heavy profiling touched production without the owner's go-ahead.
Pitfalls
- Optimising without a profile. The slow part is rarely where intuition says.
- Benchmarking on toy data. An N+1 over 10 rows is invisible; over 10,000 it is the whole problem.
- Single-run measurements. Noise between runs can exceed your improvement. Use several runs and compare distributions.
- Micro-optimising the cold path. Shaving microseconds from code that runs once per request while a query takes 800 ms.
- Caching as a reflex. A cache hides the cost and adds staleness and invalidation bugs. Remove the work first if you can.
- EXPLAIN ANALYZE on a write in production. It executes the statement. Use a copy or wrap it in a transaction you roll back, and only with permission.
- Reporting percentages without absolutes. "50 percent faster" from 4 ms to 2 ms rarely matters; say both.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
