Checks and Thresholds
Table of Contents
- Why checks alone aren't enough
- How: checks vs. thresholds
- How: per-endpoint thresholds with the
nametag - How: groups
- How: custom metrics
- How:
abortOnFail - The full script
- How: the exit code
- Tips
Why checks alone aren't enough
check() answers "did this one assertion pass?" A load test also needs to answer "did the system meet its NFRs across the whole run?" — p95 response time, error rate, a business-level metric like time from order to paid. That second question is what thresholds answer, and 04-thresholds.js is the sample script that puts every threshold feature to use at once: per-endpoint thresholds, group, custom metrics, and an early abort.
How: checks vs. thresholds
A check() records a pass/fail per call and shows up in the ✓/✗ lines of the summary, but a failed check does not fail the test run or its exit code by itself. The run only fails because of checks — the built-in metric that tracks the check pass rate — being named in thresholds with a condition like ['rate==1']. Scripts 01–04 in this track set thresholds: { checks: ['rate==1'] } specifically so that any failed check turns into a failed threshold turns into a non-zero exit code (05-load-model.js deliberately skips it — see Load Model). Drop that one threshold line and a script can fail every check it has and still exit 0.
How: per-endpoint thresholds with the name tag
http_req_duration on its own is one number across every request the script makes — a fast GET /menu and a slower POST /orders get averaged together. Tagging requests with name (as lib/flow.js already does — tags: { name: 'GET /menu' }, tags: { name: 'POST /orders' }) splits the metric by that tag, and a threshold can target one slice directly:
'http_req_duration{name:GET /menu}': ['p(95)<300'], // per endpoint, via the name tag
'http_req_duration{name:POST /orders}': ['p(95)<1000'],
Each endpoint gets the p95 that its own NFR calls for, instead of one blended number that hides a slow endpoint behind a fast one.
How: groups
group('order', () => { ... }) wraps a block of the default function and gives every metric produced inside it a group tag. group_duration is the metric for how long the group itself took, and group_duration{group:::order} is the threshold key for the order group specifically (the :: prefix is k6's group-path separator; nested groups would read ::order::pay). It's the way to put an NFR on a multi-request business step ("browse → add to cart → checkout must complete in under 3s") without threshold-ing every individual request inside it.
How: custom metrics
http_req_duration and friends are built-in, but "time from order created to payment confirmed" isn't a single HTTP request — it spans a POST /orders and a poll loop. k6/metrics exports Trend (a distribution — averages, percentiles) and Counter (a running total) for exactly this:
const orderToPaid = new Trend('order_to_paid', true); // true = time values
const ordersPlaced = new Counter('orders_placed');
orderToPaid.add(ms) and ordersPlaced.add(1) feed them, and both then behave like any built-in metric in the summary and in thresholds (order_to_paid: ['p(95)<5000'] below).
How: abortOnFail
http_req_duration: [{ threshold: 'p(99)<3000', abortOnFail: true, delayAbortEval: '10s' }],
A threshold is normally only evaluated at the end of the run — the object form lets one be evaluated during the run and stop the test early if it's already failed. delayAbortEval: '10s' waits 10 seconds into the run before the first evaluation, so it isn't judged against zero or one sample. The point is cost: an endurance test that's already blown its p99 in the first minute doesn't need to keep running for the other three hours and 59 minutes to prove it.
The full script
k6/sample/04-thresholds.js:
import { check, group, sleep } from 'k6';
import { Trend, Counter } from 'k6/metrics';
import { login, resetPlayground } from './lib/auth.js';
import { customerFor } from './lib/data.js';
import { getMenu, createOrder, payOnGateway, waitUntilPaid } from './lib/flow.js';
const orderToPaid = new Trend('order_to_paid', true); // true = time values
const ordersPlaced = new Counter('orders_placed');
const AUTOPAY = __ENV.AUTOPAY !== 'false';
export const options = {
vus: 3,
duration: '30s',
thresholds: {
checks: ['rate==1'],
http_req_failed: ['rate<0.005'], // NFR: error rate <= 0.5%
'http_req_duration{name:GET /menu}': ['p(95)<300'], // per endpoint, via the name tag
'http_req_duration{name:POST /orders}': ['p(95)<1000'],
'group_duration{group:::order}': ['avg<3000'], // NFR: whole step <= 3s
order_to_paid: ['p(95)<5000'],
// Stop early instead of burning an hour on a run that already failed.
http_req_duration: [{ threshold: 'p(99)<3000', abortOnFail: true, delayAbortEval: '10s' }],
},
};
export function setup() {
resetPlayground();
}
let token;
export default function () {
const me = customerFor(__VU);
if (!token) token = login(me.email, me.password);
let menu;
group('browse', () => {
menu = getMenu();
});
sleep(1); // think time: reading the menu
group('order', () => {
const start = Date.now();
const order = createOrder(token, me, menu, { autoPay: AUTOPAY });
if (!order) return;
if (!AUTOPAY) {
payOnGateway(order.payUrl);
const paid = check(order, { 'order paid via gateway': (o) => waitUntilPaid(o.orderId) });
if (!paid) return;
}
// With autoPay the playground marks the order paid inside the create call,
// so this is still create-to-paid even though there's no separate wait here.
orderToPaid.add(Date.now() - start);
ordersPlaced.add(1);
});
sleep(1);
}
With AUTOPAY=false, the flow calls payOnGateway and wraps waitUntilPaid in its own check ('order paid via gateway'); if that check fails, the function returns before orderToPaid/ordersPlaced are recorded, so those two custom metrics only ever count orders that actually finished paying. With the default autoPay: true, there's no separate pay step to wait on — the playground marks the order paid inside POST /orders itself — but order_to_paid still measures the same thing, just without a visible wait.
A real run against the playground (k6 run -e AUTOPAY=false 04-thresholds.js, trimmed):
█ THRESHOLDS
checks
✓ 'rate==1' rate=100.00%
group_duration{group:::order}
✓ 'avg<3000' avg=65.04ms
http_req_duration
✓ 'p(99)<3000' p(99)=110.2ms
{name:GET /menu}
✓ 'p(95)<300' p(95)=13.38ms
{name:POST /orders}
✓ 'p(95)<1000' p(95)=60.7ms
http_req_failed
✓ 'rate<0.005' rate=0.00%
order_to_paid
✓ 'p(95)<5000' p(95)=139.59ms
...
CUSTOM
order_to_paid..................: avg=64.59ms min=31ms med=54ms max=158ms p(90)=97ms p(95)=139.59ms
orders_placed..................: 45 1.434112/s
HTTP
http_req_duration..............: avg=19.46ms min=821µs med=14.93ms max=113.69ms p(90)=52.01ms p(95)=60.7ms
{ name:GET /menu }...........: avg=5.74ms min=1.32ms med=5.13ms max=18.66ms p(90)=8.23ms p(95)=13.38ms
{ name:POST /orders }........: avg=37.21ms min=15.58ms med=31.51ms max=66.15ms p(90)=58.04ms p(95)=60.7ms
http_req_failed................: 0.00% 0 out of 185
How: the exit code
k6 exits 0 when every threshold passes and non-zero when at least one fails — 99 specifically for a threshold breach (as opposed to, say, a script error, which is a different code). That's the hook CI needs: a threshold failure becomes a failed pipeline step with no extra parsing of the summary text.
Tips
- A threshold breach exits
99, specifically. Dropping a threshold to an impossible value (e.g.GET /menu's top(95)<1) and rerunning confirms it: the summary shows the failed threshold andecho $?prints99— a distinct code from a script error, so CI can tell "test failed its NFR" apart from "script errored". - A failed threshold does not stop the run. By default k6 keeps going to the end of every stage, then marks the threshold ✗ and exits
99. A 15-second run with an impossible threshold still ran all 15 seconds and 15 iterations. OnlyabortOnFail: truestops early: the same run stopped after about 2 seconds, 1 iteration, also exit99. To run with no threshold checks at all, usek6 run --no-thresholds script.js, which exits0. For stress tests I leaveabortOnFailoff, because watching it break is the point. In CI I putabortOnFailon one key threshold, so a clearly broken build doesn't use up a 30-minute slot. - Group requests by
name, not by URL. Without anametag, k6 uses the literal URL as the metric's identity —GET /payments/{id}/statusfor order 4102 and order 4103 become two different series instead of one. Every request inlib/flow.jsis tagged (GET /menu,POST /orders,POST /pay/{billId},GET /payments/{id}/status), which is also what makes the per-endpoint thresholds above possible at all. http_req_failedcounts 4xx and 5xx as failed by default, not just connection errors or 5xx. Confirmed by sending a deliberately bad login (401) and checking the summary:http_req_failed......: 100.00% 1 out of 1. That default is usually right for a load test — a 401 is a real failure — but an endpoint that's expected to return 4xx as part of normal behaviour (a "check if a coupon code exists" lookup, say) would otherwise poison the error-rate threshold.http.setResponseCallback()lets a script tell k6 which status codes actually count as failed forhttp_req_failed, per request or globally.