Skip to content

API Load Testing

This page records backend load-test runs for the Sword API (Gezen Backend). Add a new row to the results table after each run.


Quick Navigation


Executive TL;DR & Key Takeaways

1. Practical Concurrency Limit: Our current single-process setup comfortably handles up to 50 concurrent virtual users (VUs). Beyond 50 VUs, latency degrades significantly, and at 200 VUs, the system experiences partial breakdown due to database connection pool exhaustion.

2. Slow vs. Broken (A Crucial Distinction): High load on our API causes extreme slowness (queuing on password hashing and database connections) rather than a mass crash of the application. Even at 100 VUs, the error rate remains 0%, but latency increases to ~2.5 seconds p95 (exceeding our 2-second target).

3. Concurrency vs. Payload: Concurrency tests (many users, small requests) and Payload tests (few users, heavy uploads/downloads) answer different architectural questions. Do not mix them in a single run.

4. Non-Technical View: For a simpler, high-level explanation of these results designed for business stakeholders, please refer to the General Load Testing Report.


Stack

Layer Technology Description
Load generator k6 (Grafana) High-performance, developer-centric load testing tool
Concurrency script scripts/load/api-load.js Simulates typical mobile/web application traffic
Payload scripts scripts/load/payload-upload.js, payload-read.js Simulates heavy file uploads and outbound data transfers
Shared helpers scripts/load/lib/common.js Common authentication and request utilities
Upload fixtures scripts/load/fixtures/ Generated via python3 scripts/load/generate_fixtures.py
API Framework FastAPI (main.py) High-performance Python web framework
ASGI server Uvicorn Single-process ASGI server under test
Database PostgreSQL Relational database (via SQLAlchemy ORM)
Auth under test OAuth2 Password Flow POST /api/auth/token (JWT bearer on subsequent requests)

Concurrency testing vs payload testing

These are two different questions. Run both before production, but use separate scripts and separate result rows — mixing them in one run makes failures hard to interpret.

Concurrency testing Payload testing
Question How many users can use the app at the same time? What happens when each request is large or expensive?
Script api-load.js payload-upload.js, payload-read.js
What you vary Virtual users (VUs), ramp rate, think time Request/response size, upload file size, list limit, PDF generation
Typical VUs 20–200 1–5 (few users, heavy actions)
Typical request size Small JSON (lists, profile) 1–12 MiB uploads, large list responses, file downloads
Main bottlenecks Event loop queue, DB pool, bcrypt login Disk I/O, memory, serialization, Graphviz/PDF CPU
Key metrics p95 latency, error rate, req/s at N VUs http_req_sending, http_req_receiving, response_bytes, upload duration
Prerequisite Test user exists Test user + at least one factory/machine; fixtures generated for uploads

See Glossary — terms and metrics for definitions.

Concurrency answers capacity: “Can 50 people browse factories and tasks at once?”

Payload answers volume and heaviness: “Can 3 people upload 12 MiB images while others pull limit=300 lists?”

Glossary — terms and metrics

Read this section first if the tables or k6 output look unfamiliar.

Core concepts

Term Meaning
VU (Virtual User) One simulated client running the script in a loop. 50 VUs ≈ 50 people using the app at the same time (each doing one action, then waiting).
Iteration One complete pass through the main script for one VU (one HTTP request + think time). 2,782 iterations with 20 VUs means those users collectively performed 2,782 actions during the run.
Think time Pause between actions, like a real user reading a screen. Concurrency tests use 1 s; payload read tests use 1–5 s depending on scenario. Lower think time = more pressure on the API.
Ramp / stages How VUs increase or decrease over time. load-50 ramps up to 50 VUs over 30 s, holds for 2 min, then ramps down — not an instant spike.
K6_SCENARIO Named test profile (e.g. smoke, load-50, upload-12m). Sets VU count, duration, and pass/fail thresholds.
Setup Runs once before the test (login, create factory/machine, seed a 12 MiB image). Not counted as load — prepares data for the main loop.

Endpoint percentages (traffic mix)

Tables like “GET /api/factories — 25%” describe how often each endpoint is chosen, not success rate or server load share.

Each virtual user picks a random number and hits one endpoint:

Range Endpoint Share
0–10% POST /api/auth/token 10%
10–35% GET /api/factories?limit=300 25%
35–60% GET /api/machines?limit=100 25%
60–85% GET /api/tasks?limit=100 25%
85–100% GET /api/users/me 15%

So over a long run, roughly 25% of requests go to factories — mimicking a mobile app where list screens dominate. Payload scripts use the same idea with different weights (e.g. 35% factories, 15% image download in read-fat).

Latency and throughput

Metric What it measures How to read it
p95 (95th percentile) Request duration — 95% of requests were faster than this value; 5% were slower. Primary latency number in our logs. p95 1.04 s means most requests were under a second, but the slow tail reached ~1 s.
p95 (all) p95 across every HTTP request in the run. Overall API responsiveness under that scenario.
p95 auth / factories / … p95 broken out per endpoint (custom k6 trends). Shows which route is slow — e.g. login (bcrypt) vs list endpoints.
Req/s HTTP requests completed per second (throughput). Higher with more VUs and shorter think time. Not the same as “users per second” — each user waits between actions.
Run duration Wall-clock time of the k6 run (includes ramp up/down). Slightly longer than the “hold” time in the scenario name.

Errors, checks, and pass/fail

Metric What it measures How to read it
Error rate Share of requests that failed (non-2xx HTTP status, timeouts, or failed script checks). 0.00% = no failures. 6.05% = about 1 in 16 requests failed — often timeouts under heavy load, not always HTTP 500.
Thresholds Pass/fail rules defined in the script, e.g. p95 < 2 s and errors < 1%. k6 prints ✓ or ✗ at the end. Our Outcome column matches this.
Outcome (Pass / Fail) Did the run meet its thresholds? Fail (latency) = requests succeeded but were too slow. Fail with errors = timeouts or HTTP errors.
Checks passed (e.g. 2783/2783) Script assertions (status 2xx, login returns token, etc.). All checks can pass while thresholds still fail if latency is too high.
Request timeout k6 gave up waiting (default 60 s). Shows up as errors without a 5xx body — common when the server is overloaded or the DB pool is exhausted.

HTTP timing breakdown (k6 defaults)

Useful on payload runs where bodies are large:

Metric Phase
http_req_sending Time sending the request body to the server (uploads).
http_req_waiting Time waiting for the server (DB queries, disk, PDF generation).
http_req_receiving Time downloading the response body (large lists, image files).
http_req_duration Sum of sending + waiting + receiving (same basis as p95).

Payload-specific terms

Term Meaning
MiB Mebibyte (1 MiB = 1,048,576 bytes). Upload fixtures are 1 / 6 / 12 MiB; 12 MiB is the app’s max upload size.
limit=300 Query parameter on list endpoints — returns up to 300 rows. Larger limits = bigger JSON responses.
response_bytes / upload_bytes Custom k6 trends tracking bytes sent or received per request.
413 HTTP “Payload Too Large” — file exceeded the 12 MiB cap.

Results table columns (concurrency log)

Column Meaning
Max VUs Peak concurrent virtual users during the run.
Total requests All HTTP requests k6 completed (including login in setup and per-iteration calls).
Environment / Server / DB Where and how the API ran — results are not comparable across different setups without noting these.
Thresholds The pass criteria for that scenario (copied from the script).
Notes Free text — timeouts, restarts, config changes, anything that explains outliers.

Results table columns (payload log)

Column Meaning
Test type upload or read.
Payload What was heavy — file size, list limits, or download size.
p95 (all) Overall latency; on localhost, upload p95 can look deceptively low because loopback has no real network.

Reading k6 terminal output (quick map)

After a run, k6 prints a summary. Common lines:

  • http_req_duration — overall latency stats; p(95)=… is what we log as p95 (all).
  • http_req_failed — fraction of failed requests (our error rate).
  • iterations — total loop count across all VUs.
  • vus — current / max virtual users.
  • ✓ / ✗ next to threshold names — pass or fail for that rule.

Concurrency tests (api-load.js)

The k6 script simulates typical mobile/web app traffic:

  1. Setup (once): logs in with K6_USERNAME / K6_PASSWORD and obtains a JWT.
  2. Main loop (per virtual user): picks a weighted random endpoint, sends one authenticated request, then sleeps 1 second (think time between actions).

Endpoint Approx. share = traffic mix (how often each route is hit), not success rate. See Glossary.

Endpoint Approx. share
POST /api/auth/token 10%
GET /api/factories?limit=300 25%
GET /api/machines?limit=100 25%
GET /api/tasks?limit=100 25%
GET /api/users/me 15%

Scenarios

Set with K6_SCENARIO:

Scenario Profile Default thresholds
smoke 5 VUs for 30s p95 < 3s, errors < 5%
load (default) ramp 50 → 200 → 0 over 4 min p95 < 2s, errors < 1%
load-20 steady 20 VUs (30s ramp, 2m hold, 30s down) p95 < 2s, errors < 1%
load-50 steady 50 VUs (30s ramp, 2m hold, 30s down) p95 < 2s, errors < 1%
load-100 steady 100 VUs (30s ramp, 2m hold, 30s down) p95 < 2s, errors < 1%
stress ramp 100 → 500 → 0 over 7 min p95 < 5s, errors < 5%

How to run

Install k6

macOS (Homebrew)

brew install k6

Linux (Debian / Ubuntu)

curl -fsSL https://dl.k6.io/key.gpg | sudo gpg --dearmor -o /usr/share/keyrings/k6-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/k6-archive-keyring.gpg] https://dl.k6.io/deb stable main" | sudo tee /etc/apt/sources.list.d/k6.list
sudo apt-get update
sudo apt-get install k6

Linux (Fedora / CentOS)

sudo dnf install https://dl.k6.io/rpm/repo.rpm
sudo dnf install k6

Windows (winget)

winget install k6

Windows (Chocolatey)

choco install k6

Verify: k6 version. Other options (Docker, standalone binary): k6 install docs.

Start the API

In one terminal, from the repo root:

macOS / Linux (bash or zsh)

cd /path/to/gezen-backend
source venv/bin/activate
uvicorn main:app --host 0.0.0.0 --port 8000

Windows (PowerShell)

cd C:\path\to\gezen-backend
.\venv\Scripts\Activate.ps1
uvicorn main:app --host 0.0.0.0 --port 8000

Set variables and run (smoke test)

macOS / Linux (bash or zsh)

export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=smoke k6 run scripts/load/api-load.js

Windows (PowerShell)

$env:BASE_URL = "http://127.0.0.1:8000"
$env:K6_USERNAME = "loadtest"
$env:K6_PASSWORD = "loadtest123"
$env:K6_SCENARIO = "smoke"
k6 run scripts/load/api-load.js

Windows (Command Prompt)

set BASE_URL=http://127.0.0.1:8000
set K6_USERNAME=loadtest
set K6_PASSWORD=loadtest123
set K6_SCENARIO=smoke
k6 run scripts/load/api-load.js

Use a dedicated test user (for example loadtest via POST /api/users). Do not run high load against production without approval.

Command examples below use bash / zsh syntax (export, VAR=value cmd). On Windows, use the PowerShell or Command Prompt patterns above for env vars and scenario selection.

Quick reference: scripts/load/README.md

Concurrency results log

Fill in one row per run. Copy the template row at the bottom when adding a new entry.

# Date Scenario Max VUs Run duration Environment Base URL Server DB Test user Total requests Req/s Error rate p95 (all) p95 auth p95 factories p95 machines p95 tasks p95 users/me Thresholds Outcome Notes
1 2026-06-24 smoke 5 ~31 s Local dev (Mac) http://127.0.0.1:8000 Uvicorn, single process PostgreSQL loadtest 75 2.41 0.00% 203 ms 206 ms 24 ms 12 ms 15 ms 10 ms Pass (p95 < 3s, errors < 5%) Pass 5 concurrent users handled easily; 74 iterations, 76/76 checks passed, no 4xx/5xx
2 2026-06-24 load 200 ~4m 30s Local dev (Mac) http://127.0.0.1:8000 Uvicorn, single process PostgreSQL loadtest 5513 20.43 6.05% 2.78 s 3.23 s 2.95 s 1.27 s 3.01 s 1.28 s Fail (p95 < 2s, errors < 1%) Fail Ramp 50 → 200 → 0 over 4 min; 5509 iterations, 5180/5514 checks passed; 334 request timeouts at peak load
3 2026-06-24 load-20 20 ~3m 01s Local dev (Mac) http://127.0.0.1:8000 Uvicorn, single process PostgreSQL loadtest 2782 15.38 0.00% 369 ms 389 ms 383 ms 199 ms 378 ms 194 ms Pass (p95 < 2s, errors < 1%) Pass Steady 20 VUs; 2781 iterations, 2783/2783 checks passed
4 2026-06-24 load-50 50 ~3m 00s Local dev (Mac) http://127.0.0.1:8000 Uvicorn, single process PostgreSQL loadtest 5600 31.05 0.00% 1.04 s 1.13 s 1.17 s 616 ms 1.18 s 593 ms Pass (p95 < 2s, errors < 1%) Pass Steady 50 VUs; 5599 iterations, 5601/5601 checks passed
5 2026-06-24 load-100 100 ~3m 01s Local dev (Mac) http://127.0.0.1:8000 Uvicorn, single process PostgreSQL loadtest 7398 40.83 0.00% 2.57 s 2.41 s 2.87 s 1.39 s 2.80 s 1.39 s Fail (p95 < 2s, errors < 1%) Fail Steady 100 VUs; 7397 iterations, 7399/7399 checks passed; 0 HTTP failures but p95 exceeded 2 s

Run #1 summary

  • Conclusion: 5 virtual users for 30 seconds was handled without errors; overall p95 203 ms, well under the 3 s smoke threshold.
  • Slowest path: login (~200 ms avg) due to password hashing — expected.
  • List endpoints: factories, machines, tasks, and users/me mostly < 25 ms p95 on local hardware.
  • Command used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=smoke k6 run scripts/load/api-load.js

Run #2 summary

  • Conclusion: Default load profile (50 → 200 VUs over 4 minutes) failed thresholds — overall p95 2.78 s (target < 2 s) and 6.05% errors (target < 1%).
  • Throughput: 5513 requests at ~20.4 req/s with 1 s think time; 5509 completed iterations.
  • Failure mode: Mostly request timeouts (60 s k6 default) under ~200 concurrent VUs on single-process Uvicorn — not 5xx responses.
  • Slowest paths at p95: auth 3.23 s, tasks 3.01 s, factories 2.95 s; machines and users/me stayed lower (1.27–1.28 s).
  • Command used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=load k6 run scripts/load/api-load.js

Runs #3–#5 summary (steady-state: 20 / 50 / 100 VUs)

Separate fixed-load runs to measure capacity at each level (30s ramp → 2m hold → 30s ramp down). API restarted between Run #2 and these runs after DB pool exhaustion.

VUs Outcome p95 (all) Req/s Errors
20 Pass 369 ms 15.4 0%
50 Pass 1.04 s 31.1 0%
100 Fail (latency) 2.57 s 40.8 0%
  • 20 VUs: Comfortable headroom — all endpoint p95 under 400 ms.
  • 50 VUs: Still within thresholds; latency roughly 3× vs 20 VUs (p95 1.04 s).
  • 100 VUs: No failed requests, but overall p95 2.57 s crosses the 2 s load threshold — degradation without hard errors.
  • Inflection point: Between 50 and 100 steady VUs on single-process Uvicorn + default SQLAlchemy pool (30 + 40 overflow).
  • Commands used:
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=load-20 k6 run scripts/load/api-load.js
K6_SCENARIO=load-50 k6 run scripts/load/api-load.js
K6_SCENARIO=load-100 k6 run scripts/load/api-load.js

Payload tests (payload-upload.js, payload-read.js)

Generate fixtures (once per machine)

Upload tests need binary files at 1 MiB, 6 MiB, and 12 MiB (the app limit is 12 MiB per file in app/services/upload_storage.py):

python3 scripts/load/generate_fixtures.py

On Windows, use python scripts/load/generate_fixtures.py if python3 is not on your PATH.

Prerequisites

  • Everything required for concurrency tests, plus:
  • Fixtures generated under scripts/load/fixtures/
  • If the test user has no factory/machine yet, setup() creates one automatically (override with K6_FACTORY_ID / K6_MACHINE_ID)

Upload script (payload-upload.js)

Exercises POST /api/machines/{id}/images with multipart JPEG uploads.

K6_SCENARIO Profile File size
upload-smoke 1 VU, 30s 1 MiB (quick wiring check only)
upload-1m 2 VUs, ~3 min 1 MiB each
upload-6m 2 VUs, ~3 min 6 MiB each
upload-12m 1→2 VUs, ~3 min 12 MiB each (at app limit)
upload-mix 3 VUs, ~3 min random 1 / 6 / 12 MiB

Override file size on any scenario: K6_UPLOAD_SIZE=6mb

export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
python3 scripts/load/generate_fixtures.py
K6_SCENARIO=upload-smoke k6 run scripts/load/payload-upload.js

Read / download script (payload-read.js)

Exercises large outbound data paths:

Endpoint Share Notes
GET /api/factories?limit=300 ~35% default limit via K6_FACTORY_LIMIT
GET /api/machines?limit=100 ~25% default via K6_MACHINE_LIMIT
GET /api/tasks?limit=100 ~20% default via K6_TASK_LIMIT
GET /api/uploads/... (machine image) ~15% downloads an uploaded image
GET /api/factories/{id}/mindmap ~5% only in read-mindmap scenario; needs Graphviz
K6_SCENARIO Profile Purpose
read-smoke 5 VUs, 30s quick fat-read check (12 MiB image download when seeded)
read-fat 5 VUs, ~3 min sustained large list + download mix
read-mindmap 1→2 VUs, ~1.5 min Graphviz PDF generation stress

Skip mindmap if Graphviz is not installed: K6_SKIP_MINDMAP=1

K6_SCENARIO=read-smoke k6 run scripts/load/payload-read.js
K6_SCENARIO=read-fat k6 run scripts/load/payload-read.js

What to watch on payload runs

  • http_req_sending — upload bandwidth / client-side send time
  • http_req_receiving — download bandwidth for lists and images
  • http_req_waiting — server DB, disk, or PDF work
  • response_bytes and upload_bytes custom trends
  • Disk usage on uploads/ after upload scenarios
  • 413 responses if a file exceeds 12 MiB

Payload results log

# Date Test type Scenario Max VUs Run duration Payload Environment p95 (all) Error rate Outcome Notes
P1 2026-06-24 upload upload-smoke 1 ~31 s 1 MiB JPEG × 10 Local dev (Mac) 111 ms 0.00% Pass Wiring check only — not a real payload test
P2 2026-06-24 read read-smoke 5 ~31 s limit=300/100 + 1 MiB download Local dev (Mac) 36 ms 0.00% Pass Quick read check on small DB
P3 2026-06-24 upload upload-6m 2 ~3m 05s 6 MiB JPEG × 40 Local dev (Mac) 31 ms 0.00% Pass ~252 MB sent; upload p95 ~31 ms
P4 2026-06-24 upload upload-12m 2 ~3m 01s 12 MiB JPEG × 20 (app max) Local dev (Mac) 46 ms 0.00% Pass ~252 MB sent at max allowed size
P5 2026-06-24 upload upload-mix 3 ~3m 05s random 1 / 6 / 12 MiB × 60 Local dev (Mac) 45 ms 0.00% Pass ~331 MB sent; mixed sizes
P6 2026-06-24 read read-fat 5 ~3m 02s limit=300/100 + 12 MiB download Local dev (Mac) 43 ms 0.00% Pass 382 iterations; ~643 MB received

Payload size ladder (what to run)

Scenario File / response size Role
upload-smoke / read-smoke 1 MiB Fast wiring check only
upload-6m 6 MiB Mid-size upload stress
upload-12m 12 MiB App limit (MAX_UPLOAD_SIZE_BYTES)
upload-mix 1 / 6 / 12 MiB random Realistic mixed traffic
read-fat Lists + 12 MiB image download Outbound bandwidth stress

Payload run P1–P6 summary

  • P1–P2 (smoke): 1 MiB is intentionally small — confirms scripts and auth work.
  • P3–P5 (upload ladder): 6 MiB and 12 MiB (max allowed) with 2–3 concurrent VUs — all passed, p95 31–46 ms on localhost. ~835 MB total uploaded across P3–P5.
  • P6 (read-fat): 5 VUs for 3 min with 12 MiB image downloads — 643 MB received, p95 43 ms, 0% errors.
  • Note: localhost loopback hides real network latency; re-run on staging for WAN/mobile-like numbers.
  • Commands used:
python3 scripts/load/generate_fixtures.py
export BASE_URL=http://127.0.0.1:8000
export K6_USERNAME=loadtest
export K6_PASSWORD=loadtest123
K6_SCENARIO=upload-6m k6 run scripts/load/payload-upload.js
K6_SCENARIO=upload-12m k6 run scripts/load/payload-upload.js
K6_SCENARIO=upload-mix k6 run scripts/load/payload-upload.js
K6_SCENARIO=read-fat k6 run scripts/load/payload-read.js

Payload template (copy for next run)

# Date upload / read scenario Max VUs Duration Size / limits p95 Errors Outcome Notes

Findings — concurrency (local dev, 2026-06-24)

All five concurrency runs used single-process Uvicorn against PostgreSQL on a Mac dev machine with the loadtest user and 1 s think time between requests.

Capacity snapshot

Steady VUs Outcome p95 (all) Error rate Notes
≤20 Pass 369 ms 0% Comfortable headroom
~50 Pass 1.04 s 0% Meets load thresholds (p95 < 2 s, errors < 1%)
~100 Fail (latency) 2.57 s 0% No HTTP failures; p95 exceeds 2 s target
ramp → 200 Fail 2.78 s 6.05% 334 request timeouts; API became unresponsive until restart

Practical local limit (current setup): plan for ~50 concurrent users per single Uvicorn worker. Degradation starts between 50 and 100 VUs; 200+ causes timeouts and DB pool exhaustion.

Does the server break or just get slow?

Our measured runs distinguish three behaviors. High load usually means extreme slowness, not a mass crash. A true breakdown (timeouts, pool exhaustion, API unresponsive until restart) appeared only under the most aggressive concurrency ramp.

Behavior What the user sees HTTP / server When we saw it
Healthy Snappy responses 0% errors, low p95 ≤20 VUs on read mix; 1 user on most matrix routes
Slow but working Long waits; app still usable 0% HTTP failures; p95 rises (1–70+ s on worst routes) ~50–100 VUs read mix; API matrix at 10–50 users hammering writes/auth
Partial breakdown Spinners time out; new requests may hang Non-zero error rate (timeouts); DB pool exhausted; restart required Default load ramp to 200 VUs (~6% errors, 334 timeouts)

Important clarifications:

  • Slow ≠ broken. At 100 steady VUs on the mobile read mix, every request still returned 2xx — p95 was ~2.57 s with 0% errors. The API stayed up; it was just too slow for the 2 s target.
  • Matrix at 50 users (2026-06-28, pre-fix): p95 on users.create reached ~67 s, yet the run logged 0 HTTP errors — requests queued on bcrypt, the DB pool, and one worker; they completed eventually rather than failing en masse.
  • Signup benchmark (2026-06-30, post-fix): focused POST /api/users harness with uvicorn 4 workers + QR worker — p95 193 ms at 1 user, 1.25 s at 50 concurrent signups, 0 errors. See docs/performance-results/run-20260630-signup-benchmark/. Run: venv/bin/python scripts/performance/run_signup_benchmark.py.
  • Breakdown is real but narrower. Run #2 (ramp 50→200 VUs) hit QueuePool limit of size 30 overflow 40 reached; the API stopped accepting work until restart. That is capacity exhaustion, not random 500s.
  • Business errors ≠ server crash. The API matrix at 30 users recorded one 409 on admin.users.delete (disposable user still had a notification) — a validation/conflict under load, not infrastructure failure.

Bottlenecks observed

  1. Single-process Uvicorn — one event loop handles all concurrent requests; auth (bcrypt) and list endpoints queue under load.
  2. SQLAlchemy connection pool — pool_size=30, max_overflow=40 in app/database.py. Run #2 hit QueuePool limit of size 30 overflow 40 reached and the API stopped accepting requests until restart.
  3. Auth endpoint — password hashing keeps login among the slowest paths at higher VU counts (expected).
  4. Heavy list endpoints — factories?limit=300, tasks?limit=100 show the highest p95 at 50–100 VUs.

Before more heavy k6 runs:

  1. Increase Uvicorn workers — e.g. uvicorn main:app --host 0.0.0.0 --port 8000 --workers 4 (ensure pool_size × workers ≤ PostgreSQL max_connections).
  2. Review DB pool sizing — align pool_size / max_overflow in app/database.py with worker count and Postgres max_connections.
  3. Re-test steady load — rerun K6_SCENARIO=load-100 after config changes; target p95 < 2 s with 0% errors.

Optional follow-up runs (log as Run #6+ or Payload P1+):

Priority Scenario Script Purpose
After tuning load-100 api-load.js Confirm 100 VUs pass thresholds
Baseline stress api-load.js Ramp 100 → 500 → 0 over 7 min
Payload upload-smoke → upload-12m payload-upload.js Upload size ladder (1 / 6 / 12 MiB)
Payload read-smoke → read-fat payload-read.js Large list + image download
Payload read-mindmap payload-read.js Factory PDF generation (needs Graphviz)
Pre-prod load-50 or upload-mix on staging either Realistic data volume and hardware

Also capture on heavier runs (see checklist below): CPU/memory on the API host, PostgreSQL active connections, and whether list payloads reflect production row counts.

Template (copy for next run)

# Date YYYY-MM-DD smoke / load / stress

What to record on heavier runs

When running load or stress, also note (in Notes):

  • CPU / memory on the API host
  • PostgreSQL connection count or pool exhaustion
  • First VU count where p95 exceeds 2 s or errors appear
  • Whether data volume (factories/machines/tasks rows) was realistic

Overall backend capacity (slow vs breakdown)

This section combines all performance test types logged for this repo. Read it when you need a single answer to “What happens when many people use the app?”

Three test types — three questions

Test Question Tooling Breakdown vs slow
Concurrency load How many users browsing lists at once? scripts/load/api-load.js Slow up to ~100 VUs (0% errors); partial breakdown at 200 VU ramp
Volume / cardinality How do lists behave as tables grow (one user)? scripts/volume/ Always slow-only at 1 client — 0% errors at T1; no server crash
API matrix What if many users hit every route? scripts/performance/ Mostly extremely slow (0% HTTP errors at r100/vu-50); SLO failures, not outage

Do not mix them: a passing 50-VU load test on a small database does not mean 50 users can run the full app on production-sized data.

Combined capacity picture (local dev, single Uvicorn worker)

Scenario Approx. load Server behavior Errors Practical verdict
Single user, modest data 1 client Healthy 0% Fine
Single user, 6-month data 1 user Tasks 437 ms → 23 ms; factories 345 ms → 24 ms; machines 351 ms → 320 ms 0% errors Task, factory, and machine lists fixed
~20 users, mobile read mix 20 VUs Healthy (~370 ms p95) 0% Comfortable
~50 users, mobile read mix only 50 VUs Slow (~1 s p95) 0% Works; at threshold
~100 users, mobile read mix only 100 VUs Very slow (~2.6 s p95) 0% HTTP failures Up but unacceptable latency
Ramp to 200 VUs, read mix 200 VU peak Partial breakdown ~6% timeouts Pool exhausted; restart needed
10+ users, full API surface (matrix) 10–50 concurrent route runners Extremely slow on writes/auth 0% HTTP (one 409 at 30 users) Not production-ready for heavy workflows
Large uploads / fat reads (payload) 1–5 VUs Healthy on localhost 0% Payload size not the main limit

What this means for traffic handling

  1. Read-heavy internal use (~20 users, list screens): the backend generally stays up and responds; latency is acceptable on current hardware.
  2. ~50 users doing only what the mobile load script simulates: slow but functional — no measured HTTP failure rate on steady load-50.
  3. ~50 users doing real writes, registration, login, notifications, quotes: does not meet SLOs — the API matrix shows multi-second to minute-long tails, but requests still mostly complete rather than crashing the process.
  4. Aggressive spikes (200+ VUs or exhausted DB pool): the server can enter a partial breakdown — requests time out, new connections stall, and manual restart was required in Run #2.

Raw matrix artifacts: docs/performance-results/run-20260628-api-matrix/. Volume detail: Volume Testing.