Jargon and Terms
Performance testing has its own vocabulary. Requirements say "100 concurrent users at peak", reports talk about "throughput" and "steady state", and the defect says "the database is the bottleneck". This page explains each term in plain words so you can follow those conversations before you touch a tool.
The mental model: a pizza shop on a busy night
Picture a pizza shop on a Friday night. Most of the jargon maps onto it:
- The customers in the shop are your users. Some are ordering, some are paying, some are waiting. That's concurrent users.
- A normal Friday is the average load. The night of a big football final is the peak load.
- The shop fills up slowly as the evening starts (ramp-up), stays full for a few hours (steady state), then empties out (ramp-down).
- A customer reads the menu before ordering. That pause is think time. The gap before they come back for another visit is pacing.
- The time from ordering to getting your pizza is response time. The time the waiter spends walking to the kitchen and back is latency. The number of pizzas the shop serves per hour is throughput.
- If there is only one oven, it doesn't matter how many cashiers you hire. The oven is the bottleneck.
Keep this picture in mind. Each section below is one part of it.
Users
Concurrent users
Concurrent users are users active on the system at the same time, each doing their own thing. At one moment, user 1 is logging in, user 2 is on the home page and user 3 is searching. They're all active together, but on different actions.
This is the number most requirements use: "the system must support 500 concurrent users".
Simultaneous users
Simultaneous users are users doing the same action at the same moment. All 100 users click Search together.
People often use the two words interchangeably, but they aren't the same. Simultaneous users put much heavier pressure on one feature, which is why you only test it on purpose (for example, everyone hitting "Pay" when a sale opens).
| Concurrent | Simultaneous | |
|---|---|---|
| Active at the same time? | Yes | Yes |
| Doing the same action? | No, a mix of actions | Yes, the same action |
| Typical use | Normal load tests | Testing one hot feature under a burst |
Load levels
Average load
Average load is the typical user load over a period. Add up the traffic for every day in the period and divide by the number of days.
Example: a month of daily visitors adds up to 22,103. Divided by 31 days, the average load is 713 users per day.
Peak load
Peak load is the highest user load seen in the period. In the same month, the busiest day had 1,007 users, so that's the peak.
You'll get requirements like "run a load test at average load" or "run a soak test at peak load", so you need both numbers. They usually come from production analytics or logs, gathered during requirements gathering.
The load shape over time
Say the requirement is 100 concurrent users for 1 hour. You don't start all 100 at once. The test has three parts:
Ramp-up
Ramp-up is the time during which you add users gradually, for example from 0 to 100 users over 10 minutes. Throwing all 100 users in at once can crash the system before the test even starts, and the errors you'd see wouldn't be the ones you're testing for.
There are two common styles:
- Linear: users are added at a steady rate (10, 20, 30... up to 100).
- Step-up: add a batch of users, hold for a while, add another batch, hold, and so on.
Steady state
Steady state is the period where the load stays constant at the target, for example 100 users from minute 10 to minute 70. This is the part of the test that matters.
Ramp-down
Ramp-down is the time during which users are removed gradually at the end of the test.
How users behave
Real users don't click as fast as a machine can. These two settings make a script behave like a person.
Think time
Think time is the delay between two actions in the same flow. A user lands on the home page, reads it, finds the login link, then clicks it. That pause is think time.
It does two things:
- Makes the test realistic, closer to how production is used.
- Stops a handful of virtual users from hammering the server non-stop and giving you misleading results.
In one pass through a flow, the number of think times is number of actions minus one. A flow of Home → Login → Cart → Search has 4 actions and 3 think times.
There are two kinds:
- Fixed: always the same delay, for example 5 seconds every time.
- Dynamic (random): a delay that varies within a range, for example 2 to 5 seconds. Real users vary, so this is the one to use.
Pacing
Pacing is the delay between two iterations (two complete passes through the flow). After a user finishes Search, they wait before starting again at the home page.
| Think time | Pacing | |
|---|---|---|
| Delay between | Two actions | Two iterations |
| Simulates | A user reading or typing | A user coming back for a new visit |
| Main job | Realism inside a flow | Controlling how often each user repeats the flow, which sets the load rate |
Measuring speed and capacity
Response time
Response time is the full round trip: the client sends a request, the server processes it and the response comes back. It's what the user feels.
Response time = latency + server processing time
Latency
Latency is the travel time on the network: the request going from client to server, plus the response coming back. It is also called network latency or network delay.
Example: the request takes 1 s to reach the server, the server takes 2 s to process it, and the response takes 2 s to come back.
- Latency = 1 s + 2 s = 3 s
- Response time = 3 s latency + 2 s processing = 5 s
Throughput
Throughput is how many requests the system handles per unit of time. If the server finishes 4 requests every second, throughput is 4 requests per second.
You'll see it written as:
- RPS: requests per second
- TPS: transactions per second
- Hits per second
Higher throughput, with response times and errors still within target, means the system is doing more work.
Finding problems
Bottleneck
A bottleneck is the one part of the system that limits everything else, like the neck of a bottle limiting how fast water pours out. A performance bottleneck is when one component (app server, database, third-party API) slows the whole application down under load. Users see slow responses, errors or odd behaviour.
Examples:
- The app server can take 500 connections, but the database only allows 10. The database is the bottleneck.
- Your app can serve 100 concurrent users, but a third-party API it calls only supports 20. The third-party API is the bottleneck.
Common causes:
| Cause | Example |
|---|---|
| Not enough hardware | CPU or memory can't keep up with the load |
| Bad configuration | A connection pool or thread limit set too low |
| Poorly written code | An inefficient algorithm or a loop calling the database each time |
| Unoptimised database | Missing indexes, slow queries |
| Third-party services | A payment or SMS provider that can't handle your volume |
Bottleneck analysis
When a test shows a slowdown, tracking down which component is responsible is called bottleneck analysis (or performance analysis). Two kinds of tools help:
- APM (Application Performance Monitoring) tools watch the whole environment and show where time is spent across servers, databases and calls. Examples: AppDynamics, New Relic, Dynatrace.
- Profilers look inside the code: which method is slow, how garbage collection behaves as load rises. Developers use them most. Examples: JConsole, JProfiler, VisualVM.
You usually use both: APM tells you where to look, the profiler tells you why.
Tips
- Only report steady-state results. Ramp-up and ramp-down numbers drag averages around. Filter them out before you compare against targets.
- JMeter's "Latency" is not network latency. In JMeter results, Latency means the time until the first byte of the response arrives, so it includes server processing. Check which meaning a report uses before comparing numbers.
- Daily visitors are not concurrent users. 1,007 visitors in a day doesn't mean 1,007 people online at once. Ask for hourly or per-minute data, or work it out from session length, before you set the user count.
- Throughput alone can lie. If you add users and throughput stops rising while response time climbs, you've hit a bottleneck. Always read throughput next to response time and error rate.
- Zero think time is a stress test in disguise. A few users with no think time can create more load than hundreds of real users. Add think time before you trust the numbers.
Sources
These terms are explained in more depth in the Testing Diaries video series: