Skip to main content

Execute and Analyze Results

Table of Contents

Why reading the summary carefully matters​

k6 prints a lot of numbers at the end of a run, and it's tempting to only look at whether the thresholds passed (✓/✗). That's the pass/fail answer, but the numbers around it are what turn a run into an actual finding — which endpoint is slow, whether the failure was systemic or a handful of outliers, whether the run even produced the load it was supposed to. This page is about reading those numbers deliberately instead of skimming for the green checkmarks.

How: pre-run checklist​

Before starting anything longer than a smoke test:

  1. Playground is healthy. curl http://127.0.0.1:8080/api/v1/menu returns 200 with pizzas in it. A dead target produces a summary full of connection errors that look like a performance problem but aren't.
  2. The reset happens in setup(), not manually. 05-load-model.js's setup() calls resetPlayground(), which logs in as owner@playground.local and calls POST /admin/reset — every run starts from a clean slate (order numbers restart at 1001) without a separate step to remember.
  3. k6 inspect on the exact command line you're about to run for real, not a close approximation — see the Load Model tips for why (it validates PROFILE/EXECUTOR and prints every resolved stage before a single request goes out).

How: the end-of-test summary, line by line​

Annotated end-of-test summary: threshold verdict marks, the checks pass rate, http_req_duration avg vs p(95), the inverted counts on http_req_failed, and http_reqs divided by iterations as a sanity check

A real run — k6 run -e PROFILE=load -e SCALE=0.05 05-load-model.js (917 iterations in 1m36s):

time="..." level=info msg="active orders: 917, highest order number: 1917" source=console
✓ menu 200
✓ order 201
✓ order number assigned

checks.........................: 100.00% ✓ 2751 ✗ 0
data_received..................: 3.3 MB 34 kB/s
data_sent......................: 594 kB 6.2 kB/s
...
✓ http_req_duration.............: avg=7.02ms min=981µs med=5.33ms max=134.59ms p(90)=10.86ms p(95)=13.05ms
{ expected_response:true }..: avg=7.02ms min=981µs med=5.33ms max=134.59ms p(90)=10.86ms p(95)=13.05ms
✓ http_req_failed...............: 0.00% ✓ 0 ✗ 1848
...
http_reqs......................: 1848 19.192912/s
iteration_duration..............: avg=1.01s min=1s med=1.01s max=1.14s p(90)=1.01s p(95)=1.02s
iterations......................: 917 9.523756/s
vus..............................: 1 min=1 max=10
vus_max..........................: 10 min=10 max=10

Reading it top to bottom:

  • The console.log from teardown() appears before the summary, not after — teardown runs once all VUs finish but the summary is the very last thing printed, so this line is a preview of the "did the run actually do what I think it did" check covered below.
  • checks — every check() call across every iteration, rolled into one pass rate. 100% here means all three checks (menu 200, order 201, order number assigned) passed on every call; a real failure shows up as ✗ with a count and, if checks: ['rate==1'] is in thresholds, fails the whole run (exit code 99).
  • http_req_duration — every HTTP request's response time, blended across both endpoints this script calls (GET /menu, POST /orders). The ✓ next to it means the avg<3000 threshold passed. { expected_response:true } is a k6 default submetric (requests that got a 2xx/3xx) — with 0% failures here it's identical to the parent line.
  • http_req_failed — the built-in error-rate metric; ✓ means the rate<0.005 threshold held (it was 0.00% here, 0 out of 1848). The ✓/✗ counts on this line are the metric's own true/false split, not pass/fail: http_req_failed is true when a request failed, so ✓ 0 is 0 failed requests and ✗ 1848 is 1848 requests that did not fail — the opposite reading from a checks line, where ✓ counts the good outcome.
  • http_reqs / iterations — total count and rate. http_reqs (1848) is roughly 2× iterations (917) because each buy() iteration makes two calls (GET /menu, POST /orders); dividing the two is a quick sanity check that the flow ran as many requests per iteration as expected.
  • iteration_duration — how long one full pass through buy() took, including sleep(THINK). With THINK=1 (the default) this should sit around 1s + request time, which the avg=1.01s above confirms.
  • vus / vus_max — concurrency actually reached (max=10) vs. what was allocated (vus_max=10) — for ramping-vus these should match at the profile's peak stage.

How: --summary-mode=full​

--summary-mode (compact, the default, or full) controls whether metrics inside a group() are rolled into the run totals or broken out per group. It has no effect on 05-load-model.js — that script defines handleSummary() (see Reporting), and once a script defines handleSummary(), it fully replaces the built-in summary renderer, --summary-mode included. To see the actual difference, run a script that doesn't override the summary — 04-thresholds.js, which has browse and order groups:

Compact (default) — totals blend both groups together:

checks_total.......: 20 2.377348/s
iterations.....................: 8 0.950939/s
{ name:GET /menu }...........: avg=7.19ms ... p(95)=14.41ms
{ name:POST /orders }........: avg=10.92ms ... p(95)=16.37ms

Full — each group gets its own block:

█ GROUP: order

checks_total.......: 16 1.918158/s
order_to_paid..............: avg=15.62ms ... p(95)=24.25ms
http_req_duration..........: avg=15.43ms ... p(95)=23.87ms

(a second █ GROUP: browse block follows with just the GET /menu numbers). full is the mode to reach for once a script has more than one group() and the blended compact totals stop being useful on their own.

How: --out json / --out csv and pulling a per-endpoint p95​

The end-of-test summary (compact, full, or handleSummary's JSON) blends every endpoint into one http_req_duration unless a submetric already exists for it — and name isn't one of the automatic ones (only expected_response:true is). To get the p95 for POST /orders specifically, the raw event stream from --out has what the summary doesn't: every individual data point, tagged with name.

k6 run -e PROFILE=load -e SCALE=0.05 --out json=results/results.json --out csv=results/results.csv 05-load-model.js

Both files land under k6/sample/results/ (gitignored except .gitkeep). jq doesn't have a built-in percentile function, so this one-liner sorts the values and takes the 95th-percentile-by-rank — close enough to k6's own interpolated p95 for a report, not bit-for-bit identical:

jq -s '[.[] | select(.type=="Point" and .metric=="http_req_duration" and .data.tags.name=="POST /orders") | .data.value]
| length as $n | sort | .[($n*0.95|floor)]' results/results.json

Against that same run: 14.004ms for POST /orders (avg 8.47ms) vs. 6.42ms for GET /menu (avg 4.08ms) — the blended http_req_duration p95 of 13.05ms sits between the two, pulled toward POST /orders because it's the slower of the two endpoints, but it never exceeds either endpoint's own p95 (a blended percentile can't be higher than the highest of its parts).

How: the web dashboard​

On k6 v2.2.0, K6_WEB_DASHBOARD=true prints a live URL and serves a dashboard for the duration of the run; adding K6_WEB_DASHBOARD_EXPORT=<path> writes a static HTML snapshot when the run ends.

K6_WEB_DASHBOARD=true K6_WEB_DASHBOARD_EXPORT=results/report.html k6 run -e PROFILE=smoke -e SCALE=0.6 05-load-model.js
web dashboard: http://127.0.0.1:5665
...
running (0m37.3s), 0/5 VUs, 74 complete and 0 interrupted iterations

results/report.html (167 KB) was written after this run finished.

How: the order-number sanity check​

teardown() in 05-load-model.js logs one line after every run:

active orders: 14, highest order number: 1014

or, if nothing landed:

no active orders

GET /admin/orders with no status filter returns only active orders (NEW/PREPARING/READY), unpaginated — and every order this script creates uses autoPay: true, which leaves it NEW and paid, so it counts. The highest order number is the real proof of volume: setup() resets order numbers to start at 1001 every run, so highest order number - 1000 should equal the iteration count reported in the summary above it — a mismatch between the two means either the reset didn't happen (stale orders from a previous run) or something silently failed mid-run without tripping a check.

How: "p95 went up but avg didn't"​

This is the single most common misread of a load test result, and the numbers above show why it happens even within one run: http_req_duration avg was 7.02ms but p95 was 13.05ms — nearly 2× higher — in the exact same run, on the exact same blended metric. An average is pulled toward the bulk of fast requests (most of GET /menu and POST /orders finish in under 10ms here); a p95 is defined by the slow tail (this run's max was 134.59ms). Averaging 1848 mostly-fast requests barely moves when a handful take much longer; the 95th percentile is specifically the number that does move, because it's measuring the tail directly.

The practical version, comparing two runs of the same test: if avg stays flat between a baseline and a later run but p95 climbs, the system isn't uniformly slower — something is now causing a minority of requests to take much longer (GC pauses, a connection pool running out under the new load level, a slow downstream dependency only some requests hit) while the typical request is unaffected. That's a different bug hunt than "the average went up," which points at something affecting every request rather than a subset.

Tips​

  • The dashboard export needs enough run time to have data to summarize. A run under roughly 20–30s logs level=warning msg="The test run was short, report generation was skipped (not enough data)" and writes nothing — aim any dashboard-export run at 30s or longer.
  • Read checks and http_req_failed before anything else. If either shows failures, every duration number downstream is suspect — a 401 on every third request drags http_req_failed up and also makes http_req_duration look artificially fast (auth failures return quickly).
  • http_reqs ÷ iterations should match the number of HTTP calls per iteration in the script. It's the fastest way to notice a script silently returning early (see createOrder's return null on a failed check) without reading a single duration number.
  • The dashboard's live URL is only reachable while the run is in progress — it's for watching a long run, not for after-the-fact analysis; use the exported HTML or the summary JSON for that.