HEAVYHAUL AGENT — PRE-LAUNCH PRODUCTION AUDIT AND IMPLEMENTATION PLAN Date: 2026-09-16 Scope: full read of src/, supabase/migrations/, scripts/, deploy files, .env (values redacted), plus real runs of the test suite, type-check, lint, and a migration dry-run on an empty PostgreSQL 17 database. Ratings BLOCKER will cause data loss, a security breach, or total failure SERIOUS will cause visible bugs or support load LATER worth fixing, not launch-critical ===================================================================================== PART A — HOW THE SYSTEM IS WIRED (context for everything below) ===================================================================================== - No Supabase Auth. Accounts are a JSON array in AUTH_USERS in .env, passwords in plaintext (src/lib/auth/env-users.ts:34-42, 96-105). - The database is reached ONLY through the service-role key, which bypasses RLS (src/lib/supabase/admin.ts:10-19; supabase/migrations/0002_env_auth.sql:50-53). - The proxy protects page prefixes only, never /api/* (src/proxy.ts:7-15). - Therefore every authorization decision lives in application code, with no second layer behind it. ===================================================================================== PART B — FINDINGS BY SECTION ===================================================================================== ------------------------------------------------------------------------------------- 1. AUTH & AUTHORIZATION ------------------------------------------------------------------------------------- All 36 API routes were read. No route lets user A read user B's trip by changing an id. - trips/[id]/* routes all use requireParticipant (src/lib/api-guard.ts:13-26). - Admin routes check role === 'admin' or requireInternal. - Company routes re-check approved membership via getCompanyContext. - List pages (alerts, history, billing, dashboards) all use the shared visibility loader (src/lib/data/trips.ts:42-84). - Public routes (intentional, token/secret gated): early-access, intake, pilot/claim, pilot/unsubscribe, company-claims/[token], pilot/email-webhook, run-day (cron header). BLOCKER No way to onboard a real user. Signup page only posts to the early-access list (src/app/(auth)/signup/page.tsx:18-37). Adding a customer = editing .env on the server with a plaintext password + PM2 restart. The checked .env holds 4 test accounts sharing one password. No forgot-password, no email verification, no lockout on the login action (src/app/(auth)/actions.ts:26-44). BLOCKER AUTH_ROLE_SWITCH=true in the checked .env. Correctly limited to admins (src/app/(auth)/actions.ts:63-71), but it gives the admin a real session as the FIRST configured account of each role (src/lib/auth/env-users.ts:163-168). With real users in AUTH_USERS that is customer impersonation with write access. The code comment at env-users.ts:141 says to turn it off before customers. SERIOUS Anyone holding an invite link can join as that participant. invite/accept checks only token exists + not expired, never that the signed-in user's email matches (src/app/api/invite/accept/route.ts:11-33). Links are shown on the inviter's screen and forwarded by hand because no email is sent (src/app/api/trips/[id]/invite/route.ts:112-116). SERIOUS Email-based linking depends on a field nobody fills. Users without `email` in AUTH_USERS get username@users.local (src/lib/auth/env-users.ts:76), so invitations to their real address never auto-link (src/app/(auth)/actions.ts:103-110). SERIOUS Imported trips can never be claimed. claimImportedTrips must be called from an email-verification flow that does not exist (src/lib/data/imported-claims.ts:16-19). Imported Synchron participants sign in and see nothing. LATER Webhook secret accepted as ?key= in the URL, so it lands in access logs (src/app/api/pilot/email-webhook/route.ts:24). Webhook and cron secrets use === rather than a constant-time compare (run-day/route.ts:17). LATER User ids are a hash of the username (env-users.ts:57-67). Renaming a username orphans every row that person owns. ------------------------------------------------------------------------------------- 2. SECRETS & CONFIG ------------------------------------------------------------------------------------- OK .env never committed (git history clean); .gitignore correct; no key patterns in tracked source. The pickaxe hit on the first commit is the eyJ... placeholder in .env.example. SERIOUS Checked .env would break production links: NEXT_PUBLIC_APP_URL=http://localhost:3000 is what every invite/intake link is built from (src/lib/app-url.ts:23-26). Verify the server's copy. SERIOUS Two production domains in the code: SITE_URL = heavyhaulagent.com (src/lib/app-url.ts:14); webhook notes and PM2 app name use heavyhaulgbt.com (email-webhook/route.ts:10, deploy.sh:18). LATER Seven .env.backup-* / .env.bak-* copies in the project root (ignored, but live secrets on disk). LATER No security headers in next.config.ts; poweredByHeader on. CORS fine (no permissive headers). Dev preview correctly 404s outside `next dev`. LATER scripts/clear-test-data.mjs deletes every trip on whatever DB .env points at; requires --yes but nothing distinguishes prod from test. ------------------------------------------------------------------------------------- 3. DATA LAYER ------------------------------------------------------------------------------------- Migration dry-run, empty PostgreSQL 17 with stubbed auth/storage schemas: Apply 0001..0018 in order 18/18 succeeded, 41 tables Re-apply all (idempotency) 8 fail: 0001, 0003-0009 lack IF NOT EXISTS SERIOUS Silent schema fallbacks can drop data. isMissingColumn matches any error message containing the column name (src/lib/db-compat.ts:13-20). On a match, routes retry without the new fields: paid_with / payer / replaces_permit_id on a permit re-order are discarded with a success response (src/app/api/trips/[id]/requests/route.ts:95-98). No migration tracking table exists, so nothing says which of the 18 are applied in prod. SERIOUS Trip creation is not atomic and follow-up inserts are unchecked. Public intake: trip inserted, then broker participant inserted with the error ignored (src/app/api/intake/route.ts:139-202). If that fails the broker cannot see the trip (visibility is participant-row only). Same in src/app/trips/new/actions.ts:43-70 and the invite route where the invitation insert error is ignored (invite/route.ts:88-98). SERIOUS Permit/route orders succeed locally even when the vendor call fails. service_requests row written first, then createSynchronOrder, response ok:true regardless (requests/route.ts:167-193; intake/route.ts:322-340). No retry, no flag. LATER Missing indexes (verified): profiles(lower(email)) for three ilike lookups, service_requests(requested_by), warnings(resolved), company_memberships(company_id), permits(document_id). Small tables today. LATER Cascades wider than the UI implies: deleting a trip removes chat, moderation tickets and ticket history (0009:24-25); deleting a profile removes intake page, contacts, memberships. Only scripts delete today. LATER Ref codes from Math.random with no collision retry (src/lib/domain/refcode.ts:4-10). ------------------------------------------------------------------------------------- 4. ERROR HANDLING ------------------------------------------------------------------------------------- SERIOUS No error.tsx / global-error.tsx / not-found.tsx / loading.tsx anywhere in src/app. Any thrown server error shows Next's unbranded "Application error". SERIOUS Expired session mid-work is a dead end. Chat poll stops silently on non-OK (src/components/app/chat-panel.tsx:183); send shows "Sign in first." with no redirect (chat-panel.tsx:276). All workspace actions behave the same. No shared fetch helper, no 401 handler. SERIOUS Permit upload can outlive the reverse proxy. Route allows 120s and extraction waits up to 120s (documents/route.ts:10-11; src/lib/adapters/flask.ts:108). Typical cPanel/Apache timeout is 60s: user sees "upload failed" while the server finishes and stores the file. Retry = duplicates (content_hash never computed). LATER One route echoes a raw storage error (company/logo/route.ts:42). LATER Vendor failure text persisted as a shared chat message (chat/route.ts:127). LATER req.formData() on non-multipart body throws unguarded in 4 routes -> bare 500. OK No stack traces leak; no empty one-line catch blocks; the 39 multi-line catches all return a fallback. ------------------------------------------------------------------------------------- 5. INPUT VALIDATION ------------------------------------------------------------------------------------- OK Every JSON route validates with zod server-side; ids checked as UUIDs; transitions enforced; all DB access via PostgREST parameters; ilike escaped (src/lib/like.ts:15-22); `next` redirect restricted (actions.ts:157-160). SERIOUS Anonymous unbounded upload: public intake accepts any number of 25 MB permit files (intake/route.ts:102-104); rate limit is per-process keyed on x-forwarded-for (intake/route.ts:49-57). Each submission creates a trip on the broker's dashboard. SERIOUS No login throttling. LATER Trip document upload has no MIME allowlist (documents/route.ts:29-37) while intake has one (intake/route.ts:23-28). state_code not checked against a list. Blog HTML injected from repo content via marked (trusted input today). ------------------------------------------------------------------------------------- 6. FAILURE AND EDGE STATES ------------------------------------------------------------------------------------- OK First-time user: profile row + broker intake page created at sign-in (actions.ts:86-101); every dashboard has an empty state. SERIOUS Admin trip page loads every trip and every participant row on the platform to compute completions (src/lib/data/trip-page.ts:49-70; trips.ts:150-153). Cost grows with total trips. No loading states anywhere. SERIOUS Expired session: page loads redirect; in-page actions don't; server actions redirect('/login') and drop the form. LATER Double submit: login/intake/invite protected; the two trip creation forms have no idempotency key (fast double click = two trips). OK Back button: wizard state is client-only; invite accept and company decisions idempotent. ------------------------------------------------------------------------------------- 7. OPS ------------------------------------------------------------------------------------- SERIOUS Could not debug a 2am incident: 5 console.error + 1 console.warn in the whole server. No request ids; no logging of failed logins, vendor timeouts, storage failures, or fallback triggers. Audit tables record who/what, never why it failed. PM2 stdout is the only log; no rotation configured. SERIOUS No error tracking (no Sentry or equivalent). BLOCKER Backup without restore. scripts/backup-db.mjs dumps tables to JSON from a laptop; last run 2026-09-11. No restore script; reloading JSON needs FK ordering and enum handling nobody has written. Supabase PITR status unknown. SERIOUS No rollback path: deploy.sh pulls, installs, builds, restarts in place. A failed build leaves the old process running but the tree advanced; next reboot starts a mismatched app. Migrations applied by hand, no down scripts. LATER CI runs tests on every push (good) but does not build; lint advisory at 330 errors; no health endpoint; no cron configured for run-day. ------------------------------------------------------------------------------------- 8. WHAT IS MISSING THAT SHOULD BE ADDED ------------------------------------------------------------------------------------- - Email delivery for everything except pilot campaigns. ZeptoMail adapter works but only two call sites use it (src/lib/data/pilot-network.ts:381, template test sends). NOT sent today: trip invitations, intake dispatcher invitations, company approval links (company/claims/route.ts:258), contact invitations (contacts/route.ts:124), welcome / password emails (trips/new/actions.ts:229-231), permit-handling follow-ups, the `notify` list on service requests, every Synchron order email (placeholder address at src/lib/email-templates.ts:105). - Self-serve auth: signup, email verification, forgot password, invite-to-create-account. Also unblocks imported-trip claiming. - Notifications: none in-app, push, or SMS. Chat updates only by 4s polling while open. - Flask backend: extraction, AI chat, Synchron order creation all route through it. Without it chat says so, extraction stays pending, orders are silently not placed. - FMCSA lookup stub (src/lib/adapters/company-lookup.ts:44). Voice input stub. Route credits / billing are demo data (src/lib/demo/credits.ts:127). - Edge rate limiting (login, intake, early-access). Health endpoint, error tracking, log shipping, cron for run-day, migration ledger, document dedupe by hash, S3 branch the document resolver anticipates (src/lib/data/document-urls.ts:9-16). ===================================================================================== PART C — TEST RESULTS (real output) ===================================================================================== vitest run Test Files 18 passed (18) Tests 203 passed (203) Duration 906ms tsc --noEmit exit 0, 0 errors eslint 7390 problems (330 errors, 7060 warnings) — advisory in CI migrations, empty DB 18/18 applied; 8 not re-runnable Coverage honesty: Tests import 17 pure modules + 5 render components. No test touches any API route, server action, api-guard.ts, session.ts, accounts.ts, getSessionUser, the chat or document loaders, or any adapter. 12 critical paths identified: login+revocation, trip access guard, 3 trip creation paths, invite+accept, upload->warnings->status, chat+privacy, service requests-> Synchron, company claim approval, admin password/role changes, pilot campaign sends, public intake. 5 have tests for the underlying pure rule (status machine, warnings, permissions, participant access, Synchron mapping). 7 of 12 (~60%) have no test at all. Handler-level coverage (the code that runs in production): 0%. ===================================================================================== PART D — BLOCKERS, ORDERED BY WHAT TO FIX FIRST ===================================================================================== 1. Provision real accounts safely: separate passwords set through the admin console (scrypt-hashed in auth_accounts), real `email` on every AUTH_USERS entry, login throttling. (env-users.ts:96-105; actions.ts:26-44) 2. Set AUTH_ROLE_SWITCH=false in the production .env. (env-users.ts:163-168) 3. Connect the Flask backend OR make permit/route requests refuse honestly. Today they return success and place nothing. (requests/route.ts:167-193; intake/route.ts:322-340) 4. Confirm Supabase point-in-time recovery is enabled and rehearse one restore. A JSON dump with no loader is not a backup. (scripts/backup-db.mjs) ===================================================================================== PART E — NEXT STEPS: IMPLEMENTATION PLAN ===================================================================================== Ordered so that each phase can ship on its own. Phase 0 is the minimum before any real user signs in. Estimates are working days for one engineer. ------------------------------------------------------------------------------------- PHASE 0 — LAUNCH GATE (do all of these before the first real login) ~2-3 days ------------------------------------------------------------------------------------- 0.1 Production .env hygiene - Set AUTH_ROLE_SWITCH=false. - Set NEXT_PUBLIC_APP_URL to the real public domain (decide heavyhaulagent.com vs heavyhaulgbt.com once; update email-webhook/route.ts:10 comment and deploy.sh PM2 name to match). - Delete the seven .env.backup-*/.env.bak-* files from the project root. - Generate a fresh AUTH_SECRET (openssl rand -hex 32) if the current one has ever been in a backup file that left the machine. 0.2 Real accounts without plaintext passwords - For each real user: add an AUTH_USERS entry with username, role, company, real email, and a throwaway password; then immediately use Admin -> Users -> password_reset (src/app/api/admin/users/route.ts) so the effective password is a scrypt hash in auth_accounts and the env password is dead. - Remove Test_Dispatcher / Test_Driver / Test_Admin from the production AUTH_USERS (or give each a unique strong password). Never share a password across accounts. - Add login throttling: in src/app/(auth)/actions.ts signIn, keep an in-memory map keyed by lower(username) + ip with a 5-attempt / 15-minute lockout, and log every failed attempt with console.error (structured JSON, see 0.5). Edge rate limiting comes in Phase 2. 0.3 Make vendor failures honest - src/app/api/trips/[id]/requests/route.ts: after createSynchronOrder, if !order.ok -> update service_requests set status='failed' (add the enum value or use a text column via a small migration 0019), log the error, and return { ok:false, error:'Order could not be sent to Synchron. Nothing was ordered.' } with status 502. Same in intake/route.ts synchron branch. - If FLASK_API_BASE_URL is not configured in prod, hide the "Order permits / Order route" buttons behind isFlaskConfigured() and show "Ordering is not connected yet" instead of a working button. 0.4 Backups and restore - In Supabase dashboard: confirm daily backups / PITR are enabled for the project. - Write scripts/restore-db.mjs that reads db-backups// and reloads tables in FK order (profiles, broker_pages, companies, trips, trip_participants, trip_invitations, documents, permits, warnings, service_requests, chat_messages, trip_events, trip_units, review_tickets, ticket_events, company_*, pilot_*, auth_accounts, email_templates, email_template_versions, moderation_settings, carrier_contacts, early_access_signups, import_runs, import_records) using upsert on id, then re-uploads storage objects. Test it once against a scratch Supabase project. Add a cron on the server: node scripts/backup-db.mjs nightly. - Guard scripts/clear-test-data.mjs: refuse unless NEXT_PUBLIC_SUPABASE_URL matches an explicit ALLOW_CLEAR_PROJECT_REF env var. 0.5 Minimum observability - Add src/lib/log.ts exporting log.info/warn/error that writes one JSON line ({ts, level, msg, ...fields}) to stdout. Replace the 6 console.* calls. - Log (at minimum): failed sign-in, requireParticipant 403s, every isMissingColumn fallback that triggers (db-compat.ts:13-20 — add a log line there), every Flask / ZeptoMail / storage failure, every unhandled route exception. - Install pm2-logrotate on the server (pm2 install pm2-logrotate). - Add src/app/api/health/route.ts: GET returns { ok, db: