API Load Testing¶
This page records backend load-test runs for the Sword API (Gezen Backend). Add a new row to the results table after each run.
Quick Navigation¶
- Executive TL;DR & Key Takeaways
- Concurrency Results Log
- Payload Results Log
- Detailed Findings & Bottlenecks
- Runbook: How to Run the Tests
- Glossary of Terms & Metrics
Executive TL;DR & Key Takeaways¶
1. Practical Concurrency Limit: Our current single-process setup comfortably handles up to 50 concurrent virtual users (VUs). Beyond 50 VUs, latency degrades significantly, and at 200 VUs, the system experiences partial breakdown due to database connection pool exhaustion.
2. Slow vs. Broken (A Crucial Distinction): High load on our API causes extreme slowness (queuing on password hashing and database connections) rather than a mass crash of the application. Even at 100 VUs, the error rate remains 0%, but latency increases to ~2.5 seconds p95 (exceeding our 2-second target).
3. Concurrency vs. Payload: Concurrency tests (many users, small requests) and Payload tests (few users, heavy uploads/downloads) answer different architectural questions. Do not mix them in a single run.
4. Non-Technical View: For a simpler, high-level explanation of these results designed for business stakeholders, please refer to the General Load Testing Report.
Stack¶
| Layer | Technology | Description |
|---|---|---|
| Load generator | k6 (Grafana) | High-performance, developer-centric load testing tool |
| Concurrency script | scripts/load/api-load.js |
Simulates typical mobile/web application traffic |
| Payload scripts | scripts/load/payload-upload.js, payload-read.js |
Simulates heavy file uploads and outbound data transfers |
| Shared helpers | scripts/load/lib/common.js |
Common authentication and request utilities |
| Upload fixtures | scripts/load/fixtures/ |
Generated via python3 scripts/load/generate_fixtures.py |
| API Framework | FastAPI (main.py) |
High-performance Python web framework |
| ASGI server | Uvicorn | Single-process ASGI server under test |
| Database | PostgreSQL | Relational database (via SQLAlchemy ORM) |
| Auth under test | OAuth2 Password Flow | POST /api/auth/token (JWT bearer on subsequent requests) |
Concurrency testing vs payload testing¶
These are two different questions. Run both before production, but use separate scripts and separate result rows — mixing them in one run makes failures hard to interpret.
| Concurrency testing | Payload testing | |
|---|---|---|
| Question | How many users can use the app at the same time? | What happens when each request is large or expensive? |
| Script | api-load.js |
payload-upload.js, payload-read.js |
| What you vary | Virtual users (VUs), ramp rate, think time | Request/response size, upload file size, list limit, PDF generation |
| Typical VUs | 20–200 | 1–5 (few users, heavy actions) |
| Typical request size | Small JSON (lists, profile) | 1–12 MiB uploads, large list responses, file downloads |
| Main bottlenecks | Event loop queue, DB pool, bcrypt login | Disk I/O, memory, serialization, Graphviz/PDF CPU |
| Key metrics | p95 latency, error rate, req/s at N VUs | http_req_sending, http_req_receiving, response_bytes, upload duration |
| Prerequisite | Test user exists | Test user + at least one factory/machine; fixtures generated for uploads |
See Glossary — terms and metrics for definitions.
Concurrency answers capacity: “Can 50 people browse factories and tasks at once?”
Payload answers volume and heaviness: “Can 3 people upload 12 MiB images while others pull limit=300 lists?”
Glossary — terms and metrics¶
Read this section first if the tables or k6 output look unfamiliar.
Core concepts¶
| Term | Meaning |
|---|---|
| VU (Virtual User) | One simulated client running the script in a loop. 50 VUs ≈ 50 people using the app at the same time (each doing one action, then waiting). |
| Iteration | One complete pass through the main script for one VU (one HTTP request + think time). 2,782 iterations with 20 VUs means those users collectively performed 2,782 actions during the run. |
| Think time | Pause between actions, like a real user reading a screen. Concurrency tests use 1 s; payload read tests use 1–5 s depending on scenario. Lower think time = more pressure on the API. |
| Ramp / stages | How VUs increase or decrease over time. load-50 ramps up to 50 VUs over 30 s, holds for 2 min, then ramps down — not an instant spike. |
K6_SCENARIO |
Named test profile (e.g. smoke, load-50, upload-12m). Sets VU count, duration, and pass/fail thresholds. |
| Setup | Runs once before the test (login, create factory/machine, seed a 12 MiB image). Not counted as load — prepares data for the main loop. |
Endpoint percentages (traffic mix)¶
Tables like “GET /api/factories — 25%” describe how often each endpoint is chosen, not success rate or server load share.
Each virtual user picks a random number and hits one endpoint:
| Range | Endpoint | Share |
|---|---|---|
| 0–10% | POST /api/auth/token |
10% |
| 10–35% | GET /api/factories?limit=300 |
25% |
| 35–60% | GET /api/machines?limit=100 |
25% |
| 60–85% | GET /api/tasks?limit=100 |
25% |
| 85–100% | GET /api/users/me |
15% |
So over a long run, roughly 25% of requests go to factories — mimicking a mobile app where list screens dominate. Payload scripts use the same idea with different weights (e.g. 35% factories, 15% image download in read-fat).
Latency and throughput¶
| Metric | What it measures | How to read it |
|---|---|---|
| p95 (95th percentile) | Request duration — 95% of requests were faster than this value; 5% were slower. | Primary latency number in our logs. p95 1.04 s means most requests were under a second, but the slow tail reached ~1 s. |
| p95 (all) | p95 across every HTTP request in the run. | Overall API responsiveness under that scenario. |
| p95 auth / factories / … | p95 broken out per endpoint (custom k6 trends). | Shows which route is slow — e.g. login (bcrypt) vs list endpoints. |
| Req/s | HTTP requests completed per second (throughput). | Higher with more VUs and shorter think time. Not the same as “users per second” — each user waits between actions. |
| Run duration | Wall-clock time of the k6 run (includes ramp up/down). | Slightly longer than the “hold” time in the scenario name. |
Errors, checks, and pass/fail¶
| Metric | What it measures | How to read it |
|---|---|---|
| Error rate | Share of requests that failed (non-2xx HTTP status, timeouts, or failed script checks). | 0.00% = no failures. 6.05% = about 1 in 16 requests failed — often timeouts under heavy load, not always HTTP 500. |
| Thresholds | Pass/fail rules defined in the script, e.g. p95 < 2 s and errors < 1%. |
k6 prints ✓ or ✗ at the end. Our Outcome column matches this. |
| Outcome (Pass / Fail) | Did the run meet its thresholds? | Fail (latency) = requests succeeded but were too slow. Fail with errors = timeouts or HTTP errors. |
Checks passed (e.g. 2783/2783) |
Script assertions (status 2xx, login returns token, etc.). |
All checks can pass while thresholds still fail if latency is too high. |
| Request timeout | k6 gave up waiting (default 60 s). | Shows up as errors without a 5xx body — common when the server is overloaded or the DB pool is exhausted. |
HTTP timing breakdown (k6 defaults)¶
Useful on payload runs where bodies are large:
| Metric | Phase |
|---|---|
http_req_sending |
Time sending the request body to the server (uploads). |
http_req_waiting |
Time waiting for the server (DB queries, disk, PDF generation). |
http_req_receiving |
Time downloading the response body (large lists, image files). |
http_req_duration |
Sum of sending + waiting + receiving (same basis as p95). |
Payload-specific terms¶
| Term | Meaning |
|---|---|
| MiB | Mebibyte (1 MiB = 1,048,576 bytes). Upload fixtures are 1 / 6 / 12 MiB; 12 MiB is the app’s max upload size. |
limit=300 |
Query parameter on list endpoints — returns up to 300 rows. Larger limits = bigger JSON responses. |
response_bytes / upload_bytes |
Custom k6 trends tracking bytes sent or received per request. |
| 413 | HTTP “Payload Too Large” — file exceeded the 12 MiB cap. |
Results table columns (concurrency log)¶
| Column | Meaning |
|---|---|
| Max VUs | Peak concurrent virtual users during the run. |
| Total requests | All HTTP requests k6 completed (including login in setup and per-iteration calls). |
| Environment / Server / DB | Where and how the API ran — results are not comparable across different setups without noting these. |
| Thresholds | The pass criteria for that scenario (copied from the script). |
| Notes | Free text — timeouts, restarts, config changes, anything that explains outliers. |
Results table columns (payload log)¶
| Column | Meaning |
|---|---|
| Test type | upload or read. |
| Payload | What was heavy — file size, list limits, or download size. |
| p95 (all) | Overall latency; on localhost, upload p95 can look deceptively low because loopback has no real network. |
Reading k6 terminal output (quick map)¶
After a run, k6 prints a summary. Common lines:
http_req_duration— overall latency stats;p(95)=…is what we log as p95 (all).http_req_failed— fraction of failed requests (our error rate).iterations— total loop count across all VUs.vus— current / max virtual users.✓/✗next to threshold names — pass or fail for that rule.
Concurrency tests (api-load.js)¶
The k6 script simulates typical mobile/web app traffic:
- Setup (once): logs in with
K6_USERNAME/K6_PASSWORDand obtains a JWT. - Main loop (per virtual user): picks a weighted random endpoint, sends one authenticated request, then sleeps 1 second (think time between actions).
Endpoint Approx. share = traffic mix (how often each route is hit), not success rate. See Glossary.
| Endpoint | Approx. share |
|---|---|
POST /api/auth/token |
10% |
GET /api/factories?limit=300 |
25% |
GET /api/machines?limit=100 |
25% |
GET /api/tasks?limit=100 |
25% |
GET /api/users/me |
15% |
Scenarios¶
Set with K6_SCENARIO:
| Scenario | Profile | Default thresholds |
|---|---|---|
smoke |
5 VUs for 30s | p95 < 3s, errors < 5% |
load (default) |
ramp 50 → 200 → 0 over 4 min | p95 < 2s, errors < 1% |
load-20 |
steady 20 VUs (30s ramp, 2m hold, 30s down) | p95 < 2s, errors < 1% |
load-50 |
steady 50 VUs (30s ramp, 2m hold, 30s down) | p95 < 2s, errors < 1% |
load-100 |
steady 100 VUs (30s ramp, 2m hold, 30s down) | p95 < 2s, errors < 1% |
stress |
ramp 100 → 500 → 0 over 7 min | p95 < 5s, errors < 5% |
How to run¶
Install k6¶
macOS (Homebrew)
brew install k6
Linux (Debian / Ubuntu)
curl -fsSL https://dl.k6.io/key.gpg | sudo gpg --dearmor -o /usr/share/keyrings/k6-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/k6-archive-keyring.gpg] https://dl.k6.io/deb stable main" | sudo tee /etc/apt/sources.list.d/k6.list
sudo apt-get update
sudo apt-get install k6
Linux (Fedora / CentOS)
sudo dnf install https://dl.k6.io/rpm/repo.rpm
sudo dnf install k6
Windows (winget)
winget install k6
Windows (Chocolatey)
choco install k6
Verify: k6 version. Other options (Docker, standalone binary): k6 install docs.
Start the API¶
In one terminal, from the repo root:
macOS / Linux (bash or zsh)
cd /path/to/gezen-backend
source venv/bin/activate
uvicorn main:app --host 0.0.0.0 --port 8000
Windows (PowerShell)
cd C:\path\to\gezen-backend
.\venv\Scripts\Activate.ps1
uvicorn main:app --host 0.0.0.0 --port 8000
Set variables and run (smoke test)¶
macOS / Linux (bash or zsh)
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=smoke k6 run scripts/load/api-load.js
Windows (PowerShell)
$env:BASE_URL = "http://127.0.0.1:8000"
$env:K6_USERNAME = "loadtest"
$env:K6_PASSWORD = "loadtest123"
$env:K6_SCENARIO = "smoke"
k6 run scripts/load/api-load.js
Windows (Command Prompt)
set BASE_URL=http://127.0.0.1:8000
set K6_USERNAME=loadtest
set K6_PASSWORD=loadtest123
set K6_SCENARIO=smoke
k6 run scripts/load/api-load.js
Use a dedicated test user (for example loadtest via POST /api/users). Do not run high load against production without approval.
Command examples below use bash / zsh syntax (export, VAR=value cmd). On Windows, use the PowerShell or Command Prompt patterns above for env vars and scenario selection.
Quick reference: scripts/load/README.md
Concurrency results log¶
Fill in one row per run. Copy the template row at the bottom when adding a new entry.
| # | Date | Scenario | Max VUs | Run duration | Environment | Base URL | Server | DB | Test user | Total requests | Req/s | Error rate | p95 (all) | p95 auth | p95 factories | p95 machines | p95 tasks | p95 users/me | Thresholds | Outcome | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 2026-06-24 | smoke | 5 | ~31 s | Local dev (Mac) | http://127.0.0.1:8000 |
Uvicorn, single process | PostgreSQL | loadtest |
75 | 2.41 | 0.00% | 203 ms | 206 ms | 24 ms | 12 ms | 15 ms | 10 ms | Pass (p95 < 3s, errors < 5%) | Pass | 5 concurrent users handled easily; 74 iterations, 76/76 checks passed, no 4xx/5xx |
| 2 | 2026-06-24 | load | 200 | ~4m 30s | Local dev (Mac) | http://127.0.0.1:8000 |
Uvicorn, single process | PostgreSQL | loadtest |
5513 | 20.43 | 6.05% | 2.78 s | 3.23 s | 2.95 s | 1.27 s | 3.01 s | 1.28 s | Fail (p95 < 2s, errors < 1%) | Fail | Ramp 50 → 200 → 0 over 4 min; 5509 iterations, 5180/5514 checks passed; 334 request timeouts at peak load |
| 3 | 2026-06-24 | load-20 | 20 | ~3m 01s | Local dev (Mac) | http://127.0.0.1:8000 |
Uvicorn, single process | PostgreSQL | loadtest |
2782 | 15.38 | 0.00% | 369 ms | 389 ms | 383 ms | 199 ms | 378 ms | 194 ms | Pass (p95 < 2s, errors < 1%) | Pass | Steady 20 VUs; 2781 iterations, 2783/2783 checks passed |
| 4 | 2026-06-24 | load-50 | 50 | ~3m 00s | Local dev (Mac) | http://127.0.0.1:8000 |
Uvicorn, single process | PostgreSQL | loadtest |
5600 | 31.05 | 0.00% | 1.04 s | 1.13 s | 1.17 s | 616 ms | 1.18 s | 593 ms | Pass (p95 < 2s, errors < 1%) | Pass | Steady 50 VUs; 5599 iterations, 5601/5601 checks passed |
| 5 | 2026-06-24 | load-100 | 100 | ~3m 01s | Local dev (Mac) | http://127.0.0.1:8000 |
Uvicorn, single process | PostgreSQL | loadtest |
7398 | 40.83 | 0.00% | 2.57 s | 2.41 s | 2.87 s | 1.39 s | 2.80 s | 1.39 s | Fail (p95 < 2s, errors < 1%) | Fail | Steady 100 VUs; 7397 iterations, 7399/7399 checks passed; 0 HTTP failures but p95 exceeded 2 s |
Run #1 summary¶
- Conclusion: 5 virtual users for 30 seconds was handled without errors; overall p95 203 ms, well under the 3 s smoke threshold.
- Slowest path: login (~200 ms avg) due to password hashing — expected.
- List endpoints: factories, machines, tasks, and
users/memostly < 25 ms p95 on local hardware. - Command used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=smoke k6 run scripts/load/api-load.js
Run #2 summary¶
- Conclusion: Default load profile (50 → 200 VUs over 4 minutes) failed thresholds — overall p95 2.78 s (target < 2 s) and 6.05% errors (target < 1%).
- Throughput: 5513 requests at ~20.4 req/s with 1 s think time; 5509 completed iterations.
- Failure mode: Mostly request timeouts (60 s k6 default) under ~200 concurrent VUs on single-process Uvicorn — not 5xx responses.
- Slowest paths at p95: auth 3.23 s, tasks 3.01 s, factories 2.95 s; machines and
users/mestayed lower (1.27–1.28 s). - Command used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=load k6 run scripts/load/api-load.js
Runs #3–#5 summary (steady-state: 20 / 50 / 100 VUs)¶
Separate fixed-load runs to measure capacity at each level (30s ramp → 2m hold → 30s ramp down). API restarted between Run #2 and these runs after DB pool exhaustion.
| VUs | Outcome | p95 (all) | Req/s | Errors |
|---|---|---|---|---|
| 20 | Pass | 369 ms | 15.4 | 0% |
| 50 | Pass | 1.04 s | 31.1 | 0% |
| 100 | Fail (latency) | 2.57 s | 40.8 | 0% |
- 20 VUs: Comfortable headroom — all endpoint p95 under 400 ms.
- 50 VUs: Still within thresholds; latency roughly 3× vs 20 VUs (p95 1.04 s).
- 100 VUs: No failed requests, but overall p95 2.57 s crosses the 2 s load threshold — degradation without hard errors.
- Inflection point: Between 50 and 100 steady VUs on single-process Uvicorn + default SQLAlchemy pool (30 + 40 overflow).
- Commands used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=load-20 k6 run scripts/load/api-load.js
K6_SCENARIO=load-50 k6 run scripts/load/api-load.js
K6_SCENARIO=load-100 k6 run scripts/load/api-load.js
Payload tests (payload-upload.js, payload-read.js)¶
Generate fixtures (once per machine)¶
Upload tests need binary files at 1 MiB, 6 MiB, and 12 MiB (the app limit is 12 MiB per file in app/services/upload_storage.py):
python3 scripts/load/generate_fixtures.py
On Windows, use python scripts/load/generate_fixtures.py if python3 is not on your PATH.
Prerequisites¶
- Everything required for concurrency tests, plus:
- Fixtures generated under
scripts/load/fixtures/ - If the test user has no factory/machine yet,
setup()creates one automatically (override withK6_FACTORY_ID/K6_MACHINE_ID)
Upload script (payload-upload.js)¶
Exercises POST /api/machines/{id}/images with multipart JPEG uploads.
K6_SCENARIO |
Profile | File size |
|---|---|---|
upload-smoke |
1 VU, 30s | 1 MiB (quick wiring check only) |
upload-1m |
2 VUs, ~3 min | 1 MiB each |
upload-6m |
2 VUs, ~3 min | 6 MiB each |
upload-12m |
1→2 VUs, ~3 min | 12 MiB each (at app limit) |
upload-mix |
3 VUs, ~3 min | random 1 / 6 / 12 MiB |
Override file size on any scenario: K6_UPLOAD_SIZE=6mb
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
python3 scripts/load/generate_fixtures.py
K6_SCENARIO=upload-smoke k6 run scripts/load/payload-upload.js
Read / download script (payload-read.js)¶
Exercises large outbound data paths:
| Endpoint | Share | Notes |
|---|---|---|
GET /api/factories?limit=300 |
~35% | default limit via K6_FACTORY_LIMIT |
GET /api/machines?limit=100 |
~25% | default via K6_MACHINE_LIMIT |
GET /api/tasks?limit=100 |
~20% | default via K6_TASK_LIMIT |
GET /api/uploads/... (machine image) |
~15% | downloads an uploaded image |
GET /api/factories/{id}/mindmap |
~5% | only in read-mindmap scenario; needs Graphviz |
K6_SCENARIO |
Profile | Purpose |
|---|---|---|
read-smoke |
5 VUs, 30s | quick fat-read check (12 MiB image download when seeded) |
read-fat |
5 VUs, ~3 min | sustained large list + download mix |
read-mindmap |
1→2 VUs, ~1.5 min | Graphviz PDF generation stress |
Skip mindmap if Graphviz is not installed: K6_SKIP_MINDMAP=1
K6_SCENARIO=read-smoke k6 run scripts/load/payload-read.js
K6_SCENARIO=read-fat k6 run scripts/load/payload-read.js
What to watch on payload runs¶
http_req_sending— upload bandwidth / client-side send timehttp_req_receiving— download bandwidth for lists and imageshttp_req_waiting— server DB, disk, or PDF workresponse_bytesandupload_bytescustom trends- Disk usage on
uploads/after upload scenarios - 413 responses if a file exceeds 12 MiB
Payload results log¶
| # | Date | Test type | Scenario | Max VUs | Run duration | Payload | Environment | p95 (all) | Error rate | Outcome | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|
| P1 | 2026-06-24 | upload | upload-smoke | 1 | ~31 s | 1 MiB JPEG × 10 | Local dev (Mac) | 111 ms | 0.00% | Pass | Wiring check only — not a real payload test |
| P2 | 2026-06-24 | read | read-smoke | 5 | ~31 s | limit=300/100 + 1 MiB download | Local dev (Mac) | 36 ms | 0.00% | Pass | Quick read check on small DB |
| P3 | 2026-06-24 | upload | upload-6m | 2 | ~3m 05s | 6 MiB JPEG × 40 | Local dev (Mac) | 31 ms | 0.00% | Pass | ~252 MB sent; upload p95 ~31 ms |
| P4 | 2026-06-24 | upload | upload-12m | 2 | ~3m 01s | 12 MiB JPEG × 20 (app max) | Local dev (Mac) | 46 ms | 0.00% | Pass | ~252 MB sent at max allowed size |
| P5 | 2026-06-24 | upload | upload-mix | 3 | ~3m 05s | random 1 / 6 / 12 MiB × 60 | Local dev (Mac) | 45 ms | 0.00% | Pass | ~331 MB sent; mixed sizes |
| P6 | 2026-06-24 | read | read-fat | 5 | ~3m 02s | limit=300/100 + 12 MiB download | Local dev (Mac) | 43 ms | 0.00% | Pass | 382 iterations; ~643 MB received |
Payload size ladder (what to run)¶
| Scenario | File / response size | Role |
|---|---|---|
upload-smoke / read-smoke |
1 MiB | Fast wiring check only |
upload-6m |
6 MiB | Mid-size upload stress |
upload-12m |
12 MiB | App limit (MAX_UPLOAD_SIZE_BYTES) |
upload-mix |
1 / 6 / 12 MiB random | Realistic mixed traffic |
read-fat |
Lists + 12 MiB image download | Outbound bandwidth stress |
Payload run P1–P6 summary¶
- P1–P2 (smoke): 1 MiB is intentionally small — confirms scripts and auth work.
- P3–P5 (upload ladder): 6 MiB and 12 MiB (max allowed) with 2–3 concurrent VUs — all passed, p95 31–46 ms on localhost. ~835 MB total uploaded across P3–P5.
- P6 (read-fat): 5 VUs for 3 min with 12 MiB image downloads — 643 MB received, p95 43 ms, 0% errors.
- Note: localhost loopback hides real network latency; re-run on staging for WAN/mobile-like numbers.
- Commands used:
python3 scripts/load/generate_fixtures.py
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=upload-6m k6 run scripts/load/payload-upload.js
K6_SCENARIO=upload-12m k6 run scripts/load/payload-upload.js
K6_SCENARIO=upload-mix k6 run scripts/load/payload-upload.js
K6_SCENARIO=read-fat k6 run scripts/load/payload-read.js
Payload template (copy for next run)¶
| # | Date | upload / read | scenario | Max VUs | Duration | Size / limits | p95 | Errors | Outcome | Notes |
|---|---|---|---|---|---|---|---|---|---|---|
Findings — concurrency (local dev, 2026-06-24)¶
All five concurrency runs used single-process Uvicorn against PostgreSQL on a Mac dev machine with the loadtest user and 1 s think time between requests.
Capacity snapshot¶
| Steady VUs | Outcome | p95 (all) | Error rate | Notes |
|---|---|---|---|---|
| ≤20 | Pass | 369 ms | 0% | Comfortable headroom |
| ~50 | Pass | 1.04 s | 0% | Meets load thresholds (p95 < 2 s, errors < 1%) |
| ~100 | Fail (latency) | 2.57 s | 0% | No HTTP failures; p95 exceeds 2 s target |
| ramp → 200 | Fail | 2.78 s | 6.05% | 334 request timeouts; API became unresponsive until restart |
Practical local limit (current setup): plan for ~50 concurrent users per single Uvicorn worker. Degradation starts between 50 and 100 VUs; 200+ causes timeouts and DB pool exhaustion.
Does the server break or just get slow?¶
Our measured runs distinguish three behaviors. High load usually means extreme slowness, not a mass crash. A true breakdown (timeouts, pool exhaustion, API unresponsive until restart) appeared only under the most aggressive concurrency ramp.
| Behavior | What the user sees | HTTP / server | When we saw it |
|---|---|---|---|
| Healthy | Snappy responses | 0% errors, low p95 | ≤20 VUs on read mix; 1 user on most matrix routes |
| Slow but working | Long waits; app still usable | 0% HTTP failures; p95 rises (1–70+ s on worst routes) | ~50–100 VUs read mix; API matrix at 10–50 users hammering writes/auth |
| Partial breakdown | Spinners time out; new requests may hang | Non-zero error rate (timeouts); DB pool exhausted; restart required | Default load ramp to 200 VUs (~6% errors, 334 timeouts) |
Important clarifications:
- Slow ≠ broken. At 100 steady VUs on the mobile read mix, every request still returned 2xx — p95 was ~2.57 s with 0% errors. The API stayed up; it was just too slow for the 2 s target.
- Matrix at 50 users (2026-06-28, pre-fix): p95 on
users.createreached ~67 s, yet the run logged 0 HTTP errors — requests queued on bcrypt, the DB pool, and one worker; they completed eventually rather than failing en masse. - Signup benchmark (2026-06-30, post-fix): focused
POST /api/usersharness with uvicorn 4 workers + QR worker — p95 193 ms at 1 user, 1.25 s at 50 concurrent signups, 0 errors. Seedocs/performance-results/run-20260630-signup-benchmark/. Run:venv/bin/python scripts/performance/run_signup_benchmark.py. - Breakdown is real but narrower. Run #2 (ramp 50→200 VUs) hit
QueuePool limit of size 30 overflow 40 reached; the API stopped accepting work until restart. That is capacity exhaustion, not random 500s. - Business errors ≠ server crash. The API matrix at 30 users recorded one 409 on
admin.users.delete(disposable user still had a notification) — a validation/conflict under load, not infrastructure failure.
Bottlenecks observed¶
- Single-process Uvicorn — one event loop handles all concurrent requests; auth (bcrypt) and list endpoints queue under load.
- SQLAlchemy connection pool —
pool_size=30,max_overflow=40inapp/database.py. Run #2 hitQueuePool limit of size 30 overflow 40 reachedand the API stopped accepting requests until restart. - Auth endpoint — password hashing keeps login among the slowest paths at higher VU counts (expected).
- Heavy list endpoints —
factories?limit=300,tasks?limit=100show the highest p95 at 50–100 VUs.
Recommended next steps¶
Before more heavy k6 runs:
- Increase Uvicorn workers — e.g.
uvicorn main:app --host 0.0.0.0 --port 8000 --workers 4(ensurepool_size × workers ≤ PostgreSQL max_connections). - Review DB pool sizing — align
pool_size/max_overflowinapp/database.pywith worker count and Postgresmax_connections. - Re-test steady load — rerun
K6_SCENARIO=load-100after config changes; target p95 < 2 s with 0% errors.
Optional follow-up runs (log as Run #6+ or Payload P1+):
| Priority | Scenario | Script | Purpose |
|---|---|---|---|
| After tuning | load-100 |
api-load.js |
Confirm 100 VUs pass thresholds |
| Baseline | stress |
api-load.js |
Ramp 100 → 500 → 0 over 7 min |
| Payload | upload-smoke → upload-12m |
payload-upload.js |
Upload size ladder (1 / 6 / 12 MiB) |
| Payload | read-smoke → read-fat |
payload-read.js |
Large list + image download |
| Payload | read-mindmap |
payload-read.js |
Factory PDF generation (needs Graphviz) |
| Pre-prod | load-50 or upload-mix on staging |
either | Realistic data volume and hardware |
Also capture on heavier runs (see checklist below): CPU/memory on the API host, PostgreSQL active connections, and whether list payloads reflect production row counts.
Template (copy for next run)¶
| # | Date | YYYY-MM-DD | smoke / load / stress | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
What to record on heavier runs¶
When running load or stress, also note (in Notes):
- CPU / memory on the API host
- PostgreSQL connection count or pool exhaustion
- First VU count where p95 exceeds 2 s or errors appear
- Whether data volume (factories/machines/tasks rows) was realistic
Overall backend capacity (slow vs breakdown)¶
This section combines all performance test types logged for this repo. Read it when you need a single answer to “What happens when many people use the app?”
Three test types — three questions¶
| Test | Question | Tooling | Breakdown vs slow |
|---|---|---|---|
| Concurrency load | How many users browsing lists at once? | scripts/load/api-load.js |
Slow up to ~100 VUs (0% errors); partial breakdown at 200 VU ramp |
| Volume / cardinality | How do lists behave as tables grow (one user)? | scripts/volume/ |
Always slow-only at 1 client — 0% errors at T1; no server crash |
| API matrix | What if many users hit every route? | scripts/performance/ |
Mostly extremely slow (0% HTTP errors at r100/vu-50); SLO failures, not outage |
Do not mix them: a passing 50-VU load test on a small database does not mean 50 users can run the full app on production-sized data.
Combined capacity picture (local dev, single Uvicorn worker)¶
| Scenario | Approx. load | Server behavior | Errors | Practical verdict |
|---|---|---|---|---|
| Single user, modest data | 1 client | Healthy | 0% | Fine |
| Single user, 6-month data | 1 user | Tasks 437 ms → 23 ms; factories 345 ms → 24 ms; machines 351 ms → 320 ms | 0% errors | Task, factory, and machine lists fixed |
| ~20 users, mobile read mix | 20 VUs | Healthy (~370 ms p95) | 0% | Comfortable |
| ~50 users, mobile read mix only | 50 VUs | Slow (~1 s p95) | 0% | Works; at threshold |
| ~100 users, mobile read mix only | 100 VUs | Very slow (~2.6 s p95) | 0% HTTP failures | Up but unacceptable latency |
| Ramp to 200 VUs, read mix | 200 VU peak | Partial breakdown | ~6% timeouts | Pool exhausted; restart needed |
| 10+ users, full API surface (matrix) | 10–50 concurrent route runners | Extremely slow on writes/auth | 0% HTTP (one 409 at 30 users) | Not production-ready for heavy workflows |
| Large uploads / fat reads (payload) | 1–5 VUs | Healthy on localhost | 0% | Payload size not the main limit |
What this means for traffic handling¶
- Read-heavy internal use (~20 users, list screens): the backend generally stays up and responds; latency is acceptable on current hardware.
- ~50 users doing only what the mobile load script simulates: slow but functional — no measured HTTP failure rate on steady
load-50. - ~50 users doing real writes, registration, login, notifications, quotes: does not meet SLOs — the API matrix shows multi-second to minute-long tails, but requests still mostly complete rather than crashing the process.
- Aggressive spikes (200+ VUs or exhausted DB pool): the server can enter a partial breakdown — requests time out, new connections stall, and manual restart was required in Run #2.
Raw matrix artifacts: docs/performance-results/run-20260628-api-matrix/. Volume detail: Volume Testing.
Related docs¶
- Analysis Report — project changelog and checklist
- Volume Testing — row-count / cardinality (1 client)
- Local Deployment — running the API locally