Skip to main content

Checks and Thresholds

Table of Contents

Why checks alone aren't enough​

check() answers "did this one assertion pass?" A load test also needs to answer "did the system meet its NFRs across the whole run?" — p95 response time, error rate, a business-level metric like time from order to paid. That second question is what thresholds answer, and 04-thresholds.js is the sample script that puts every threshold feature to use at once: per-endpoint thresholds, group, custom metrics, and an early abort.

How: checks vs. thresholds​

A check() records a pass/fail per call and shows up in the ✓/✗ lines of the summary, but a failed check does not fail the test run or its exit code by itself. The run only fails because of checks — the built-in metric that tracks the check pass rate — being named in thresholds with a condition like ['rate==1']. Scripts 01–04 in this track set thresholds: { checks: ['rate==1'] } specifically so that any failed check turns into a failed threshold turns into a non-zero exit code (05-load-model.js deliberately skips it — see Load Model). Drop that one threshold line and a script can fail every check it has and still exit 0.

How: per-endpoint thresholds with the name tag​

http_req_duration on its own is one number across every request the script makes — a fast GET /menu and a slower POST /orders get averaged together. Tagging requests with name (as lib/flow.js already does — tags: { name: 'GET /menu' }, tags: { name: 'POST /orders' }) splits the metric by that tag, and a threshold can target one slice directly:

'http_req_duration{name:GET /menu}': ['p(95)<300'], // per endpoint, via the name tag
'http_req_duration{name:POST /orders}': ['p(95)<1000'],

Each endpoint gets the p95 that its own NFR calls for, instead of one blended number that hides a slow endpoint behind a fast one.

How: groups​

group('order', () => { ... }) wraps a block of the default function and gives every metric produced inside it a group tag. group_duration is the metric for how long the group itself took, and group_duration{group:::order} is the threshold key for the order group specifically (the :: prefix is k6's group-path separator; nested groups would read ::order::pay). It's the way to put an NFR on a multi-request business step ("browse → add to cart → checkout must complete in under 3s") without threshold-ing every individual request inside it.

How: custom metrics​

http_req_duration and friends are built-in, but "time from order created to payment confirmed" isn't a single HTTP request — it spans a POST /orders and a poll loop. k6/metrics exports Trend (a distribution — averages, percentiles) and Counter (a running total) for exactly this:

const orderToPaid = new Trend('order_to_paid', true); // true = time values
const ordersPlaced = new Counter('orders_placed');

orderToPaid.add(ms) and ordersPlaced.add(1) feed them, and both then behave like any built-in metric in the summary and in thresholds (order_to_paid: ['p(95)<5000'] below).

How: abortOnFail​

http_req_duration: [{ threshold: 'p(99)<3000', abortOnFail: true, delayAbortEval: '10s' }],

A threshold is normally only evaluated at the end of the run — the object form lets one be evaluated during the run and stop the test early if it's already failed. delayAbortEval: '10s' waits 10 seconds into the run before the first evaluation, so it isn't judged against zero or one sample. The point is cost: an endurance test that's already blown its p99 in the first minute doesn't need to keep running for the other three hours and 59 minutes to prove it.

The full script​

k6/sample/04-thresholds.js:

import { check, group, sleep } from 'k6';
import { Trend, Counter } from 'k6/metrics';
import { login, resetPlayground } from './lib/auth.js';
import { customerFor } from './lib/data.js';
import { getMenu, createOrder, payOnGateway, waitUntilPaid } from './lib/flow.js';

const orderToPaid = new Trend('order_to_paid', true); // true = time values
const ordersPlaced = new Counter('orders_placed');
const AUTOPAY = __ENV.AUTOPAY !== 'false';

export const options = {
vus: 3,
duration: '30s',
thresholds: {
checks: ['rate==1'],
http_req_failed: ['rate<0.005'], // NFR: error rate <= 0.5%
'http_req_duration{name:GET /menu}': ['p(95)<300'], // per endpoint, via the name tag
'http_req_duration{name:POST /orders}': ['p(95)<1000'],
'group_duration{group:::order}': ['avg<3000'], // NFR: whole step <= 3s
order_to_paid: ['p(95)<5000'],
// Stop early instead of burning an hour on a run that already failed.
http_req_duration: [{ threshold: 'p(99)<3000', abortOnFail: true, delayAbortEval: '10s' }],
},
};

export function setup() {
resetPlayground();
}

let token;
export default function () {
const me = customerFor(__VU);
if (!token) token = login(me.email, me.password);

let menu;
group('browse', () => {
menu = getMenu();
});
sleep(1); // think time: reading the menu

group('order', () => {
const start = Date.now();
const order = createOrder(token, me, menu, { autoPay: AUTOPAY });
if (!order) return;
if (!AUTOPAY) {
payOnGateway(order.payUrl);
const paid = check(order, { 'order paid via gateway': (o) => waitUntilPaid(o.orderId) });
if (!paid) return;
}
// With autoPay the playground marks the order paid inside the create call,
// so this is still create-to-paid even though there's no separate wait here.
orderToPaid.add(Date.now() - start);
ordersPlaced.add(1);
});
sleep(1);
}

With AUTOPAY=false, the flow calls payOnGateway and wraps waitUntilPaid in its own check ('order paid via gateway'); if that check fails, the function returns before orderToPaid/ordersPlaced are recorded, so those two custom metrics only ever count orders that actually finished paying. With the default autoPay: true, there's no separate pay step to wait on — the playground marks the order paid inside POST /orders itself — but order_to_paid still measures the same thing, just without a visible wait.

A real run against the playground (k6 run -e AUTOPAY=false 04-thresholds.js, trimmed):

█ THRESHOLDS

checks
✓ 'rate==1' rate=100.00%

group_duration{group:::order}
✓ 'avg<3000' avg=65.04ms

http_req_duration
✓ 'p(99)<3000' p(99)=110.2ms

{name:GET /menu}
✓ 'p(95)<300' p(95)=13.38ms

{name:POST /orders}
✓ 'p(95)<1000' p(95)=60.7ms

http_req_failed
✓ 'rate<0.005' rate=0.00%

order_to_paid
✓ 'p(95)<5000' p(95)=139.59ms

...

CUSTOM
order_to_paid..................: avg=64.59ms min=31ms med=54ms max=158ms p(90)=97ms p(95)=139.59ms
orders_placed..................: 45 1.434112/s

HTTP
http_req_duration..............: avg=19.46ms min=821µs med=14.93ms max=113.69ms p(90)=52.01ms p(95)=60.7ms
{ name:GET /menu }...........: avg=5.74ms min=1.32ms med=5.13ms max=18.66ms p(90)=8.23ms p(95)=13.38ms
{ name:POST /orders }........: avg=37.21ms min=15.58ms med=31.51ms max=66.15ms p(90)=58.04ms p(95)=60.7ms
http_req_failed................: 0.00% 0 out of 185

How: the exit code​

k6 exits 0 when every threshold passes and non-zero when at least one fails — 99 specifically for a threshold breach (as opposed to, say, a script error, which is a different code). That's the hook CI needs: a threshold failure becomes a failed pipeline step with no extra parsing of the summary text.

Tips​

  • A threshold breach exits 99, specifically. Dropping a threshold to an impossible value (e.g. GET /menu's to p(95)<1) and rerunning confirms it: the summary shows the failed threshold and echo $? prints 99 — a distinct code from a script error, so CI can tell "test failed its NFR" apart from "script errored".
  • A failed threshold does not stop the run. By default k6 keeps going to the end of every stage, then marks the threshold ✗ and exits 99. A 15-second run with an impossible threshold still ran all 15 seconds and 15 iterations. Only abortOnFail: true stops early: the same run stopped after about 2 seconds, 1 iteration, also exit 99. To run with no threshold checks at all, use k6 run --no-thresholds script.js, which exits 0. For stress tests I leave abortOnFail off, because watching it break is the point. In CI I put abortOnFail on one key threshold, so a clearly broken build doesn't use up a 30-minute slot.
  • Group requests by name, not by URL. Without a name tag, k6 uses the literal URL as the metric's identity — GET /payments/{id}/status for order 4102 and order 4103 become two different series instead of one. Every request in lib/flow.js is tagged (GET /menu, POST /orders, POST /pay/{billId}, GET /payments/{id}/status), which is also what makes the per-endpoint thresholds above possible at all.
  • http_req_failed counts 4xx and 5xx as failed by default, not just connection errors or 5xx. Confirmed by sending a deliberately bad login (401) and checking the summary: http_req_failed......: 100.00% 1 out of 1. That default is usually right for a load test — a 401 is a real failure — but an endpoint that's expected to return 4xx as part of normal behaviour (a "check if a coupon code exists" lookup, say) would otherwise poison the error-rate threshold. http.setResponseCallback() lets a script tell k6 which status codes actually count as failed for http_req_failed, per request or globally.