LOADS TESTS BEFORE THE MARKET LAUNCH
October 5, 2026
Can the platform handle the initial surge in users? How we set up load tests: three test levels, real user profiles, and thresholds that are truly meaningful.
How Smart Load Testing Ensures a Smooth Go-Live
Before a platform is launched, there’s a question that no one likes to answer, because without the right tools, the honest answer is “we think so.” How to make it measurable and what matters most.
The Question Before the Start
A platform launches, and eventually the first evening comes when everyone is online at the same time. Until then, there are two ways to answer the question of whether the system can handle it: an estimate ‒ or a measurement. The difference takes a few days of work and determines whether you can stay calm under full load or, in the worst-case scenario, lose customers.
It's not the averages that count, but the peaks
The first mistake happens during planning. Forecasts are calculated in terms of users or transactions per month, but systems crash during peak hours.
On an ordering platform, demand is extremely concentrated: Sunday evenings between 6 and 7 p.m. account for many times the hourly average. The figure against which testing must be conducted is this one hour ‒ not the month, not the day.
rom this, a second metric is derived that can actually be measured: requests per second. And this is where the greatest uncertainty in any capacity calculation lies ‒ namely, how much browsing leads to a single order. Anyone who hasn’t measured this must estimate ‒ and should then present the result as a range rather than as a seemingly valid fixed number.
Three Test Levels
Smoke ‒ with every change, as part of the automated workflow
A few minutes, during idle time. Checks whether the code works at all.
Load ‒ regularly, for example at the end of a sprint
Realistic peak load over one hour. Tests actual response times and error rates.
Stress ‒ manually, before major deadlines
Until the set tolerances for a high-performance system are exceeded. Shows where the limit lies and how the system behaves at that point.
Why Locally Measured Values Are Critical
Load tests run on the development machine or in a container produce numbers that are meaningless: shared CPU cores, a database running locally instead of on the network, zero-millisecond latency, and no automatic scaling.
That’s why we run functional tests locally and performance measurements exclusively in an environment that matches the live server ‒ same database class, same memory, same connection limits. Anything else measures a lab environment, not the real product.
Real users instead of generic inquiries
A load test that calls a single interface in a loop tests that interface ‒ not the system.
That’s why we simulate the actual distribution. For an ordering platform, this involves three roles with very different behaviors: customers, who primarily access the system in read-only mode; restaurants, which constantly wait for new orders and enter status changes and updates to items; and drivers, who access the system only occasionally to make changes. Each role has its own profile and contributes its own share to the total load.
This is also why load tests can’t simply be bought off the shelf. The work isn’t in the tool itself, but in understanding what real users actually trigger.
Meaningful Threshold Values
A test result without a defined threshold is just a number in a report — it has no context, so it has to be interpreted, and in case of doubt it will be misread. We therefore define thresholds that have consequences attached to them:
- 95 percent of requests under 300 milliseconds
- 99 percent under 500 milliseconds
- Error rate below one percent
If the threshold is breached in the smoke test, the system signals this automatically. If it is breached after deployment, a message goes to the development channel. We chose the k6 tool for this: the thresholds are part of an automated test run within our release process.
What a load test typically detects
It’s rarely the final capacity figure itself that’s interesting, but rather the insights gained along the way.
A recurring example of this is structural in nature: The application keeps a database connection open while performing tasks for which it doesn’t even need a database. Under load, this causes almost all connections to idle and wait. After a fix ‒ same hardware, just a code adjustment ‒ the response time in such cases often shrinks by a factor of several.
uch issues are unlikely to be discovered during a code review. Under load, they become evident and can lead to continuous optimization of system performance. This is how we systematically increase the resilience of high-demand apps ‒ without any unpleasant surprises after go-live.
Sequence Trumps Reflex
This last point belongs in every operations manual: When things get tight, the instinct is to bring in additional application servers. But if the database is the bottleneck, doing exactly that makes the situation worse, because more servers consume more connections.
That’s why a load test involves not just the result, but a sequence: Which metric is tested first, and what action follows from that.
Conclusion
Load tests don’t answer a technical question ‒ answer a business one: How much growth can what we’ve built handle before someone has to spend money on a larger infrastructure or (worse!) code modifications? If you answer this before launch, you can attract customers on launch day instead of frantically having to buy additional server capacity.
Are you planning a market launch and unsure whether your platform can handle peak loads? Contact us! We’d be happy to schedule a free introductory and analysis meeting with you at short notice.
Load Tests
Performance
k6
CI/CD
FastAPI
PostgreSQL
Architecture
Quality Assurance
EAT-TAXI
Anne Bardtke
[email protected]

