Skip to main content

Jargon and Terms

Performance testing has its own vocabulary. Requirements say "100 concurrent users at peak", reports talk about "throughput" and "steady state", and the defect says "the database is the bottleneck". This page explains each term in plain words so you can follow those conversations before you touch a tool.

The mental model: a pizza shop on a busy night​

Picture a pizza shop on a Friday night. Most of the jargon maps onto it:

  • The customers in the shop are your users. Some are ordering, some are paying, some are waiting. That's concurrent users.
  • A normal Friday is the average load. The night of a big football final is the peak load.
  • The shop fills up slowly as the evening starts (ramp-up), stays full for a few hours (steady state), then empties out (ramp-down).
  • A customer reads the menu before ordering. That pause is think time. The gap before they come back for another visit is pacing.
  • The time from ordering to getting your pizza is response time. The time the waiter spends walking to the kitchen and back is latency. The number of pizzas the shop serves per hour is throughput.
  • If there is only one oven, it doesn't matter how many cashiers you hire. The oven is the bottleneck.

Keep this picture in mind. Each section below is one part of it.


Users​

Concurrent users​

Concurrent users are users active on the system at the same time, each doing their own thing. At one moment, user 1 is logging in, user 2 is on the home page and user 3 is searching. They're all active together, but on different actions.

This is the number most requirements use: "the system must support 500 concurrent users".

Simultaneous users​

Simultaneous users are users doing the same action at the same moment. All 100 users click Search together.

People often use the two words interchangeably, but they aren't the same. Simultaneous users put much heavier pressure on one feature, which is why you only test it on purpose (for example, everyone hitting "Pay" when a sale opens).

Concurrent users are on different actions at the same moment; simultaneous users are all on the same action

ConcurrentSimultaneous
Active at the same time?YesYes
Doing the same action?No, a mix of actionsYes, the same action
Typical useNormal load testsTesting one hot feature under a burst

Load levels​

Average load​

Average load is the typical user load over a period. Add up the traffic for every day in the period and divide by the number of days.

Example: a month of daily visitors adds up to 22,103. Divided by 31 days, the average load is 713 users per day.

Peak load​

Peak load is the highest user load seen in the period. In the same month, the busiest day had 1,007 users, so that's the peak.

Daily users for a month: the dashed average line sits at 713, and the tallest bar, day 20, is the peak at 1,007

You'll get requirements like "run a load test at average load" or "run a soak test at peak load", so you need both numbers. They usually come from production analytics or logs, gathered during requirements gathering.


The load shape over time​

Say the requirement is 100 concurrent users for 1 hour. You don't start all 100 at once. The test has three parts:

Load over time: ramp-up from 0 to 100 users in 10 minutes, steady state at 100 users for 60 minutes, then ramp-down to 0

Ramp-up​

Ramp-up is the time during which you add users gradually, for example from 0 to 100 users over 10 minutes. Throwing all 100 users in at once can crash the system before the test even starts, and the errors you'd see wouldn't be the ones you're testing for.

There are two common styles:

  • Linear: users are added at a steady rate (10, 20, 30... up to 100).
  • Step-up: add a batch of users, hold for a while, add another batch, hold, and so on.

Steady state​

Steady state is the period where the load stays constant at the target, for example 100 users from minute 10 to minute 70. This is the part of the test that matters.

Ramp-down​

Ramp-down is the time during which users are removed gradually at the end of the test.


How users behave​

Real users don't click as fast as a machine can. These two settings make a script behave like a person.

Timeline: think time sits between the actions inside one iteration, pacing sits between two iterations

Think time​

Think time is the delay between two actions in the same flow. A user lands on the home page, reads it, finds the login link, then clicks it. That pause is think time.

It does two things:

  • Makes the test realistic, closer to how production is used.
  • Stops a handful of virtual users from hammering the server non-stop and giving you misleading results.

In one pass through a flow, the number of think times is number of actions minus one. A flow of Home → Login → Cart → Search has 4 actions and 3 think times.

There are two kinds:

  • Fixed: always the same delay, for example 5 seconds every time.
  • Dynamic (random): a delay that varies within a range, for example 2 to 5 seconds. Real users vary, so this is the one to use.

Pacing​

Pacing is the delay between two iterations (two complete passes through the flow). After a user finishes Search, they wait before starting again at the home page.

Think timePacing
Delay betweenTwo actionsTwo iterations
SimulatesA user reading or typingA user coming back for a new visit
Main jobRealism inside a flowControlling how often each user repeats the flow, which sets the load rate

Measuring speed and capacity​

Response time​

Response time is the full round trip: the client sends a request, the server processes it and the response comes back. It's what the user feels.

Response time = latency + server processing time

A 5-second request: 1 s request travel and 2 s response travel are latency (3 s), plus 2 s server processing, for a 5 s response time

Latency​

Latency is the travel time on the network: the request going from client to server, plus the response coming back. It is also called network latency or network delay.

Example: the request takes 1 s to reach the server, the server takes 2 s to process it, and the response takes 2 s to come back.

  • Latency = 1 s + 2 s = 3 s
  • Response time = 3 s latency + 2 s processing = 5 s

Throughput​

Throughput is how many requests the system handles per unit of time. If the server finishes 4 requests every second, throughput is 4 requests per second.

Left: four clients send requests to a server that handles 4 per second. Right: as users increase, throughput flattens while response time climbs, marking the bottleneck

You'll see it written as:

  • RPS: requests per second
  • TPS: transactions per second
  • Hits per second

Higher throughput, with response times and errors still within target, means the system is doing more work.


Finding problems​

Bottleneck​

A bottleneck is the one part of the system that limits everything else, like the neck of a bottle limiting how fast water pours out. A performance bottleneck is when one component (app server, database, third-party API) slows the whole application down under load. Users see slow responses, errors or odd behaviour.

Users feed an app server that takes 500 connections, which feeds a database that takes only 10; the database is the bottleneck

Examples:

  • The app server can take 500 connections, but the database only allows 10. The database is the bottleneck.
  • Your app can serve 100 concurrent users, but a third-party API it calls only supports 20. The third-party API is the bottleneck.

Common causes:

CauseExample
Not enough hardwareCPU or memory can't keep up with the load
Bad configurationA connection pool or thread limit set too low
Poorly written codeAn inefficient algorithm or a loop calling the database each time
Unoptimised databaseMissing indexes, slow queries
Third-party servicesA payment or SMS provider that can't handle your volume

Bottleneck analysis​

When a test shows a slowdown, tracking down which component is responsible is called bottleneck analysis (or performance analysis). Two kinds of tools help:

Three steps: the test shows response time climbing, APM shows the database takes most of the time, the profiler shows which method is slow

  • APM (Application Performance Monitoring) tools watch the whole environment and show where time is spent across servers, databases and calls. Examples: AppDynamics, New Relic, Dynatrace.
  • Profilers look inside the code: which method is slow, how garbage collection behaves as load rises. Developers use them most. Examples: JConsole, JProfiler, VisualVM.

You usually use both: APM tells you where to look, the profiler tells you why.


Tips​

  • Only report steady-state results. Ramp-up and ramp-down numbers drag averages around. Filter them out before you compare against targets.
  • JMeter's "Latency" is not network latency. In JMeter results, Latency means the time until the first byte of the response arrives, so it includes server processing. Check which meaning a report uses before comparing numbers.
  • Daily visitors are not concurrent users. 1,007 visitors in a day doesn't mean 1,007 people online at once. Ask for hourly or per-minute data, or work it out from session length, before you set the user count.
  • Throughput alone can lie. If you add users and throughput stops rising while response time climbs, you've hit a bottleneck. Always read throughput next to response time and error rate.
  • Zero think time is a stress test in disguise. A few users with no think time can create more load than hundreds of real users. Add think time before you trust the numbers.

Sources​

These terms are explained in more depth in the Testing Diaries video series: