Sword Backend - Deep Code Analysis Report¶
1. Executive Summary¶
Following the initial deep architectural and security sweep, several critical defects have been corrected: the hardcoded OpenAI API key fallback was removed, OpenAI model selection was centralized with gpt-4.1 primary and gpt-4o fallback, the broken authentication token cache and later DB-backed auth lookup cache were removed, duplicate API route blocks were removed and protected by route inventory coverage, the remaining major main.py endpoint families were extracted into domain routers, machine file uploads now validate extension, MIME type, and size, unused heavy development dependencies were pruned, core Factory, Machine, Quote, and Device deletes now use deleted_at, production reference delete guards now use the correct association-table columns, production print() calls in app modules were replaced with module-level logging, the duplicate soft-delete migration in the legacy Alembic tree was removed, digital competency answer/analysis persistence was added to match the API contract, focused pytest coverage now exists for auth, token security, RBAC, push notifications, registration, login, OpenAI fallback, file uploads, device CRUD, catalog/reference routes, small domain routes, production references, route inventory, active Alembic migration behavior, soft-delete behavior, and calendar event query efficiency, and the GitHub Actions test workflow now runs the suite in CI with a focused Python 3.9-compatible test requirements file.
The codebase is closer to being maintainable, but still not production-ready. The main remaining risks are architectural and operational: many extracted routers still contain copied business logic that should move into services over time, app/crud.py remains broad, some non-core admin/reference deletes still use hard deletes, several lower-priority query paths still need batching, and production observability still needs a consistent request and business-event logging policy. Local disk file storage is acceptable for the current internal single-server usage as long as uploads/ is persistent and backed up; S3-compatible storage is a future scaling option, not an urgent next step.
2. In-Depth Technical Review¶
2.1 Code Quality & Architecture¶
- Application Entrypoint (
main.py): The former monolithic controller has been reduced to app bootstrap, static mounts, explicit file-serving routes, router inclusion, and a temporary OpenAI test compatibility shim. Domain endpoints now live underapp/routers/, including auth, users, notifications, factories, catalog/reference data, production PDF, digital competency, quote documents/workflows, solution center, tasks, calendar events, sectors, ROI, diagnostics, devices, machines, production references, quotes, and machine images. - Route Inventory Protection: Duplicate API route registrations were removed and
tests/test_route_inventory.pynow guards against reintroducing duplicate(method, path)pairs. - Service Layer Still Thin: The OpenAI client fallback now lives in
app/services/openai_client.py, and upload storage remains inapp/services/upload_storage.py. Most business workflows are now in routers instead ofmain.py, but task/calendar/solution/PDF/email logic still needs gradual service extraction. - Structured Logging: Production
print()calls in app modules have been replaced with module-levellogging.getLogger(__name__)loggers. Remaining console output is limited to standalone maintenance scripts; broader request, error, and business-event logging conventions still need to be defined.
2.2 Security Risks & Vulnerabilities¶
- Exposed Secrets: The hardcoded OpenAI key fallback has been removed from
main.py, OpenAI now readsOPENAI_API_KEYfrom the environment, and the previously exposed key has been revoked externally. - File Uploads: Machine image uploads now enforce an extension whitelist, MIME type matching, and a 12 MiB application-level size limit (
MAX_UPLOAD_SIZE_BYTES = 12 * 1024 * 1024). The validation, local file write, cleanup, and machine image URL helpers now live inapp/services/upload_storage.py, whilemain.pykeeps the existing endpoint and database behavior. Focused tests cover allowed types, exact-at-limit acceptance, unsupported extensions, content-type mismatches, oversized files, empty filename rejection, failed multi-file uploads, image list/delete lifecycle, mobile image upload/list routes, local service persistence, partial-file cleanup, and both current URL contracts (/uploads/...and/api/uploads/...). Current Nginx deployment templates setclient_max_body_size 100M, so the application remains the limiting layer. For internal single-server use, local disk storage is acceptable ifuploads/survives deploys, is backed up, and disk usage is monitored. - Authentication Cache: The broken in-memory token cache and the later
lru_cachewrapper around DB-backed user lookup have been removed. Login and bearer-token validation now perform normal database reads, preserving immediate visibility of user status and credential changes. - User Deactivation and Deletion (2026-07-08): Inactive users are now blocked at login with HTTP 403 before any access token is issued. Admin user deletion no longer returns 409 when business records exist; transient personal rows (notifications, assignees, notes, calendar assignees, QR jobs, user-created calendar events) are removed, and critical business records keep history with
created_byset toNULL.
2.3 Performance & Scalability¶
- List screen speed: Task list 437 ms → 23 ms (fixed). Factory list 523 ms → 24 ms (fixed). Machine list 351 ms → 320 ms (fixed 2026-07-01). See optimization.md.
- Factory mindmap event-loop safety (2026-07-01):
GET /api/factories/{id}/mindmapno longer performs Graphviz rendering directly on the event loop. The endpoint uses a bounded async semaphore (MINDMAP_RENDER_CONCURRENCY, default1) and runs token validation, synchronous SQLAlchemy reads, and Graphviz PDF rendering through the threadpool only after a render slot is available, so queued mindmap requests do not hold database sessions. Mindmap graph construction moved toapp/services/factory_mindmap.py, render output now uses a per-request temporary directory instead of sharedfactory_{id}_mindmap.pdffiles, and response cleanup removes the directory afterFileResponsecompletes. The Pythongraphviz==0.20.3package is now pinned inrequirements/runtime.txt; the server image or host still must provide the Graphvizdotbinary. Local verification withdot15.0.0: one real mindmap request returned HTTP 200 in 165 ms; six concurrent real mindmap requests all returned HTTP 200 with max 790 ms while/api/openapi.jsonreturned in 4 ms and/api/factories?limit=1returned in 17 ms. Focused tests cover the async wrapper/threadpool contract, isolated output directory cleanup, consultant access rejection, and graph service rendering behavior. - User signup async QR (2026-06-30): Signup no longer renders QR images in the request path.
POST /api/usersnow returnsqr_status="pending"withqr_code=null, and theusersrow plus the database-backeduser_qr_jobsrow are committed in the same transaction so a failed signup cannot leave behind a hidden persisted account. In the sync fallback path (SIGNUP_QR_ASYNC=false), QR generation also stays inside that transaction, so QR render failures roll the signup back cleanly. QR generation moved toapp/services/user_qr.pywith a separate worker (scripts/workers/qr_worker.py). Clients can pollGET /api/users/me/qrfor readiness. Duplicate email/username checks now rely on unique indexes plusIntegrityErrormapping instead of race-prone pre-insertSELECTs. PostgreSQL pool sizing is environment-backed (DB_POOL_SIZE,DB_MAX_OVERFLOW,DB_POOL_TIMEOUT; defaults5/5/20per process). Focused registration tests now cover enqueue failure and sync QR failure rollback behavior. Signup benchmark (2026-06-30, uvicorn 4 workers + QR worker): p95 193 ms at 1 user (target < 1 s), 1.25 s at 50 concurrent signups (target < 2 s), 0 HTTP errors — down from 67 s p95 before the fix. Artifacts:docs/performance-results/run-20260630-signup-benchmark/. Harness:scripts/performance/run_signup_benchmark.py. - API Matrix Benchmark (
r100): The real harness underscripts/performance/completed the fullr100user band (1/10/20/30/50); outputs live underdocs/performance-results/run-20260628-api-matrix/. At1user onlyusers.create(p95 ~1404 ms) andfactories.create(p95 ~2694 ms) missed the 1 second target;factories.mindmapwas honestly skipped without Graphviz. By10users 11 routes, by20users 28 routes, by30users 32 routes, and by50users all 50 routes crossed 1 second, with notable p95s onusers.create,auth.token,notifications.system,solutions.create, andfactories.create. Ther500,r1000, andr5000bands are not finished; higher tiers should wait untilr100SLO failures are addressed. - Load behavior — slow vs breakdown: Measured runs show extreme slowness before outage. Up to ~100 concurrent readers on the mobile mix returned 0% HTTP failures (slow p95). The API matrix at 50 users logged 0% HTTP errors with p95 up to ~67 s on
users.create— queueing on bcrypt, one worker, and the DB pool, not process crash. Partial breakdown (timeouts, pool exhaustion, restart required) appeared only on the aggressive 200 VU ramp (~6% errors). Seedocs/load-testing.md#overall-backend-capacity-slow-vs-breakdown(and the general report indocs/general-load-testing.md). - Stateful File Storage: Files are still saved to a local
uploads/directory. This is acceptable for the current internal single-server deployment model if the directory is persistent, backed up, and excluded from destructive deploy steps. Object storage such as S3-compatible storage should remain a future scaling option for multi-server, containerized, or ephemeral-disk deployments. - In-Memory App State: The auth token/user cache risk has been resolved. Any remaining global runtime state should be reviewed during modularization.
2.4 Data Integrity¶
- Soft Deletes for Core Records:
Factory,Machine,Quote, andDevicenow include nullabledeleted_atcolumns and their primary delete paths mark records deleted while preserving dependent history such as quote items, logs, tasks, calendar events, and device consumable links. Normal CRUD helpers and key API list/detail paths hide soft-deleted rows, Solution Center now rejects soft-deleted factories, factory mindmap generation excludes soft-deleted machines, quote update paths reject reassignment to soft-deleted factories, and device reads/updates/deletes ignore soft-deleted devices. Production reference delete guards now correctly detect factory assignments throughproduction_category_idandproduction_area_id. Non-core reference/admin entities still need a separate lifecycle decision before replacing every remaining hard delete. - Device Lifecycle & Schema Hardening: Device create/update payloads now use a typed consumable schema with positive quantity validation. Device create/update also validates that referenced consumables exist as
QuoteItemrows of typesarf, and device deletes now soft-delete the device while preserving consumable link history. - Alembic Migration Path: The active Alembic path is
alembic_clean. Runtime dependencies now include Alembic, and the clean migration graph includes idempotent soft-delete revisions for existing baseline-stamped databases:0002_add_soft_delete_columnsforFactory,Machine, andQuote,0003_add_device_soft_deleteforDevice,0004_add_digital_competency_answer_scorefor digital competency answer scores and analysis persistence,0005_add_user_qr_statusfor async signup QR state plus theuser_qr_jobsqueue table, and0006_add_notification_indexesfor notification list/unread-count query indexes. Upgrade/downgrade idempotence is covered by focused migration tests. The duplicate soft-delete revision that had been added to the legacyalembic/versionstree has been removed; that tree remains legacy and should not receive new migrations unless the project intentionally switches back to it.
2.5 Third-Party Integrations¶
- OpenAI Model Handling: Model selection is centralized in
app/services/openai_client.py; the approved order isgpt-4.1first withgpt-4ofallback. - Bloated Requirements: The unused heavy development dependencies
google-cloud-aiplatformandgoogle-cloud-bigqueryhave been removed fromrequirements/dev.txt. CI no longer installs the full local development freeze; it installsrequirements/runtime.txtplus the focusedrequirements/test.txtfile so Python 3.9 jobs do not resolve incompatible dev-only pins such aspytest==9.0.3.
3. Estimated Adaptation Time & Adoptability¶
- Can the internal team take over this code? With high difficulty.
- Adaptation Time: It will take a mid/senior developer 2-3 weeks just to trace the tangled logic and resolve the immediate crashing bugs in
main.pybefore they can comfortably add new features. - Refactor Effort: Estimated at 140 - 180 hours. Breaking
main.pyinto distinct domains (Users, Machines, Factories, Quotes) and establishing a clean Service Layer is a massive but unavoidable undertaking.
4. Top 10 Priority Action Items¶
- Extract Services From Routers: Move task, calendar, solution center, PDF/proforma email, and ROI business logic out of routers into focused service modules.
- Broaden Observability: Define consistent request, error, and business-event logging on top of the module-level loggers now in production app code.
- Batch Remaining N+1 Loops: Refactor factory create/update relationship assignment loops in
app/crud.pywithin_()queries. - Expand Integration Tests: Add tests for core creation/update/delete flows and representative workflow endpoints beyond the existing soft-delete and router extraction coverage.
- Harden Configuration: Move remaining magic values and integration settings into typed configuration.
- Pydantic V2 Cleanup: Migrate deprecated class-based Pydantic config and
.dict()usage to current APIs. - Decide Lifecycle for Remaining Reference Data: Review hard deletes for users, catalog/reference entities, tasks, and admin taxonomies.
- Keep S3 as Future Scaling Work: Keep local
uploads/for the internal single-server app, but preserve the storage service boundary so S3-compatible storage can be added later if deployment topology changes. - Keep Logging Noise Low: Keep noisy diagnostic data at debug level and avoid logging sensitive payloads.
- Keep CI Green: Expand the test workflow as new service extraction slices add coverage.
5. Python Refactoring Principles & Best Practices¶
To guide the upcoming refactoring effort, the following Python-specific principles and best practices should be adhered to:
5.1 Core Principles¶
- DRY (Don't Repeat Yourself): Eliminate the extreme code duplication found in
main.pyby centralizing common logic into reusable functions or service classes. - KISS (Keep It Simple, Stupid): Replace complex, nested logic with flatter, more readable Pythonic patterns.
- Single Responsibility Principle (SRP): Each module, class, and function must have one clear purpose.
main.pyis now mostly bootstrap and router wiring; remaining oversized workflow logic should continue moving from routers into focused services. - Separation of Concerns: Business logic currently scattered in routers and CRUD functions must be moved to a dedicated service layer.
5.2 Python-Specific Best Practices¶
- Leverage Pythonic Idioms: Use list comprehensions,
enumerate(), andzip()to replace verboseforloops where appropriate. - Type Hinting: Implement robust type hints (
typingmodule) across the codebase to improve IDE support and catch data structure mismatches. - Guard Clauses: Simplify nested
if-elseblocks (like those in the auth flow) by using early returns for error/edge cases. - Symbolic Constants: Replace hardcoded magic numbers and strings (e.g., model names, file paths) with constants or configuration objects.
- Functional Decomposition: Break down the multi-hundred line functions in
main.pyinto smaller, independently testable units. - Dependency Pruning: Remove unused imports and libraries from
requirements/dev.txtto minimize the application footprint and security surface area. - Standardized Logging: Replace all
print()statements with a structured logging configuration using the standardlogginglibrary. - Statelessness: Ensure all logic is stateless (avoiding in-memory dictionaries for cache) to support horizontal scaling.
6. Final Decision Expectation¶
Technical Decision: Continue by refactoring.
Justification: The foundational database schema (SQLAlchemy models), Alembic migration setup, and test structure are usable, and several critical security defects have already been fixed. The project should not be rewritten from scratch, but it still needs disciplined refactoring around routing, services, data lifecycle, query performance, logging, and storage before it can be considered production-ready.
7. Next Architecture Slice¶
The first router extraction wave is now in place: duplicate API route blocks were removed, machine image routes live in app/routers/machine_images.py, core machine routes live in app/routers/machines.py, core quote/quote-item routes live in app/routers/quotes.py, device CRUD routes live in app/routers/devices.py, production reference routes live in app/routers/production_references.py, and catalog/reference routes now live in domain routers with focused coverage. The next narrow refactor should move toward side-effect-heavy workflow routes while preserving behavior with focused tests and service boundaries.
Recommended shape:
- Keep
tests/test_route_inventory.pygreen so duplicate routes cannot return. - Extract one remaining workflow family at a time with a clear service boundary.
- Stabilize side-effect-heavy PDF/proforma/email flows with tests before moving logic.
- Keep full-suite CI and focused local tests passing after each slice.
8. Brief Evaluation Table¶
| Category | Status | Risk | Note / Action |
|---|---|---|---|
| Code Quality | Fair | Medium | main.py is now lean bootstrap/router wiring and duplicate routes are guarded by tests; remaining debt is mostly service extraction and broad CRUD cleanup. |
| Security | Good | Medium | Critical issues are fixed: exposed key handling, auth cache removal, upload validation, and safer deletes. Production config, storage operations, and logging policy still need hardening. |
| Architecture | Fair | Medium | Domain routers are in place across the main API surface; the next risk is business logic still living in routers instead of focused services. |
| Logging | Fair | Medium | Production app modules use module loggers instead of print(); request, error, and business-event logging conventions still need definition. |
| Database | Fair | Medium | Soft delete and active Alembic migration coverage are much stronger; remaining work is lifecycle decisions for non-core entities and a few batching candidates. |
| API/Integration | Fair | Medium | OpenAI fallback, upload contracts, and digital competency persistence are covered; broader integration settings still need typed configuration. |
| Testing/Deployment | Fair | Medium | Focused pytest coverage and GitHub Actions CI are in place; coverage should expand around end-to-end business workflows. |
| Adoptability | Fair | Medium | The project is no longer blocked by a giant main.py; onboarding is still slowed by business logic spread across routers and broad CRUD helpers. |
9. Live Refactoring Checklist¶
Updated as changes are applied. Do not delete completed items — they serve as an audit trail.
Critical / Security¶
- [x] 1. Remove hardcoded OpenAI API key fallback from
main.py—main.py:120 - [x] 2. Revoke the old exposed key on platform.openai.com — External / manual, confirmed by project owner
- [x] 3. Fix broken
token_cache— memory leak + guaranteedNameErrorcrash on cache hit —main.py:130-159 - [x] Remove process-local auth user/token caching and DB-backed
lru_cachelookup —app/auth.py,tests/test_token_security.py
Bugs / Crashes¶
- [x] 4. Centralize approved OpenAI model selection —
gpt-4.1primary,gpt-4ofallback —main.py - [x] 5. Remove 2 duplicate
upload_machine_images_mobileendpoint definitions —main.py:~3503, ~4545 - [x] Fix SQLAlchemy overlapping relationship warning in
QuoteItemDetail(conflict betweendeviceandquote_item) —app/models.py - [x] Stabilize soft-delete API leak paths for Solution Center, factory mindmap, and quote factory reassignment —
main.py,tests/test_soft_deletes.py - [x] Cover OpenAI model fallback and final failure behavior without network calls —
tests/test_openai_fallback.py
Code Quality / Low Risk¶
- [x] 6. Remove duplicate ReportLab import block (copy-pasted twice in header) —
main.py:46-52 - [x] 7. Untrack
__pycache__,sword.db,temp/PDFs from git index - [x] 8. Update
.gitignoreto explicitly coveruploads/andsword.db - [x] 9. Prune unused heavy deps from
requirements/dev.txt(google-cloud-aiplatform,google-cloud-bigquery) - [x] Fix MkDocs Material Turkish language switch by enabling i18n Material alternates —
mkdocs.yml - [x] Remove one repeated machine brand/model route block from
main.py - [x] Remove duplicate API route blocks and add route inventory coverage —
main.py,tests/test_route_inventory.py - [x] Create comprehensive installation and deployment guides in English (
work.md) and Turkish (work-tr.md) for Windows & Linux, focusing on IP-only, non-SSL setups
Architecture / Higher Risk (Deferred)¶
- [x] 10. Secure file uploads — MIME validation, 12 MiB size limit, extension whitelist, and upload storage helper extraction —
app/services/upload_storage.py,main.py - [x] 11. Implement soft deletes (
deleted_at) onFactory,Machine,Quote—app/models.py+ Alembic migration - [x] Verify and reconcile Alembic tooling for the soft-delete migration — active path is
alembic_clean, with idempotent revision0002_add_soft_delete_columns - [x] Remove duplicate soft-delete migration from the legacy Alembic tree —
alembic/versions - [x] Add focused migration idempotence coverage for active soft-delete revisions —
tests/test_migrations.py - [x] 12. Fix primary calendar event N+1 query path in
app/crud.pywith eager loading and query-count regression coverage —tests/test_crud_query_efficiency.py - [x] 13. Start modularizing
main.pyintoapp/routers/+app/services/—app/routers/machine_images.py,app/routers/machines.py,app/routers/quotes.py,app/routers/devices.py,app/routers/production_references.py - [x] Extract production reference routes and fix association delete guards —
app/routers/production_references.py,app/crud.py,tests/test_production_references.py - [x] Harden device lifecycle and request schema — soft delete, typed consumable payloads,
sarfconsumable validation, and focused API/migration tests —app/routers/devices.py,app/schemas.py,app/models.py,alembic_clean/versions/0003_add_device_soft_delete.py,tests/test_devices.py,tests/test_migrations.py - [x] 14. Replace production app
print()statements with structuredlogging—main.py,app/crud.py,app/auth.py,app/routers/machine_images.py - [ ] 15. Keep local
uploads/for internal single-server use; revisit S3-compatible storage only if scaling or deployment topology requires it - [x] 16. Add focused pytest tests for auth, token security, registration, login, RBAC, push notifications, device CRUD, and file upload validation
- [x] Add GitHub Actions test workflow for pytest —
.github/workflows/tests.yml - [x] Fix CI dependency resolution by replacing the full dev requirements install with focused test requirements —
.github/workflows/tests.yml,requirements/test.txt - [x] Make application database engine creation compatible with SQLite CI URLs while preserving PostgreSQL pool settings —
app/database.py - [x] Add focused pytest coverage for core soft-delete behavior and quote delete API visibility —
tests/test_soft_deletes.py - [x] Expand machine image upload lifecycle coverage for web and mobile routes —
tests/test_file_uploads.py - [x] Add direct upload storage service coverage for local persistence, oversize cleanup, and URL contract preservation —
tests/test_upload_storage_service.py - [x] Decide device lifecycle hardening — soft delete, historical consumable link preservation, and typed consumable request schemas
- [ ] 17. Expand pytest coverage for remaining core creation/update/delete business flows
- [x] Add k6 API load test script for auth, factories, machines, tasks, and users/me —
scripts/load/api-load.js,scripts/load/README.md - [x] Add load-test results documentation with methodology and first smoke run (5 VUs, ~31s, 0% errors, p95 203ms) —
docs/load-testing.md,docs/load-testing.tr.md,mkdocs.yml - [x] Log default load scenario run #2 (50→200 VUs, ~4m30s, 6.05% errors, p95 2.78s, thresholds failed) —
docs/load-testing.md,docs/load-testing.tr.md - [x] Add steady-state load scenarios (
load-20,load-50,load-100) and log runs #3–#5 —scripts/load/api-load.js,docs/load-testing.md,docs/load-testing.tr.md - [x] Document load-test findings and recommended next steps (capacity ~50 VUs/worker, pool/Uvicorn bottlenecks, re-test plan) —
docs/load-testing.md,docs/load-testing.tr.md - [x] Add k6 payload load scripts for upload size (1/6/12 MiB) and fat read/download paths —
scripts/load/payload-upload.js,scripts/load/payload-read.js,scripts/load/lib/common.js,scripts/load/generate_fixtures.py - [x] Document concurrency vs payload testing methodology and payload runbook —
docs/load-testing.md,docs/load-testing.tr.md,scripts/load/README.md - [x] Log payload smoke runs P1–P2 (upload-smoke 1 MiB, read-smoke fat lists + download, 0% errors) —
docs/load-testing.md,docs/load-testing.tr.md - [x] Log payload ladder runs P3–P6 (6/12 MiB uploads, upload-mix, read-fat with 12 MiB downloads, ~643 MB received) —
docs/load-testing.md,docs/load-testing.tr.md,payload-read.js12 MiB seed - [x] Add load-test glossary (VU, p95, endpoint traffic mix %, error rate, thresholds, k6 output) —
docs/load-testing.md,docs/load-testing.tr.md - [x] Add cross-platform k6 install and run instructions (macOS, Linux, Windows) —
docs/load-testing.md,docs/load-testing.tr.md,scripts/load/README.md - [x] Add PostgreSQL volume/cardinality test harness (seed tiers, 1-client probe, k6 volume-baseline, query-shape tests) —
scripts/volume/,tests/test_volume_query_efficiency.py,docs/volume-testing.md,docs/volume-testing.tr.md - [x] Log volume test run #1 T1 (200/4k/5k rows, factories p95 345ms, tasks p95 437ms, 0% errors at 1 VU) —
docs/volume-results/run-20260626-051613/,docs/volume-testing.md - [x] Log volume test run #2 T1 post task-list fix —
docs/volume-results/run-20260629-171059/,docs/optimization.md - [x] Add non-technical, results-focused general load and volume testing documentation in English and Turkish —
docs/general-load-testing.md,docs/general-volume-testing.md,docs/general-load-testing.tr.md,docs/general-volume-testing.tr.md - [x] Add full API matrix benchmark harness for seeded row tiers (100/500/1000/5000) and user tiers (1/10/20/30/50), with exact reseeding, canonical route coverage, and workflow-safe execution order —
scripts/performance/,tests/test_api_matrix_catalog.py - [x] Fix benchmark blockers found during the matrix build: cleanup idempotence for seeded and workflow-created rows,
broadcast-notificationlist response shape, and public user registration double-hash plus admin-escalation behavior —scripts/performance/seed_api_matrix_db.py,scripts/performance/workflow.py,app/routers/admin.py,app/routers/users.py,app/auth.py - [x] Block inactive users at login and implement transactional admin user deletion cleanup —
app/schemas.py,app/auth.py,app/routers/auth.py,app/routers/admin.py - [x] Validate the patched 100-row / 1-user matrix baseline on a fresh server process: all runnable routes pass except
users.createat 1332.91 ms;factories.mindmapremains an honest skip without Graphviz — local benchmark run on 2026-06-28 - [x] Complete the full
r100user-concurrency band (1,10,20,30,50) and persist scenario outputs underdocs/performance-results/run-20260628-api-matrix/;users.createremains above target even at1user, and broad SLO failure begins by10users — local benchmark run on 2026-06-28 - [ ] Run the remaining seeded row tiers (
r500,r1000,r5000) across1/10/20/30/50users and publish the aggregated findings; fixr100SLO failures first —r100/vu-50alone took several minutes, so upper tiers need a longer uninterrupted run - [ ] Reduce
users.createbelow the 1 second target in the matrix baseline without weakening password hashing or QR generation semantics