Skip to content

Gateway ops — deploy, migrate, seed ​

Live

The hosted gateway at api.tenkeybridge.com went live 2026-07-09 (Fly.io sjc + Neon aws-us-west-2), proven end-to-end against a real QuickBooks Enterprise company file. This page documents the deploy flow that stood it up.

The gateway is a single Node/TypeScript process (Fly.io + Neon Postgres) that terminates the QBO-compatible REST API, brokers the agent-plane WebSocket, and runs the OAuth2 code-flow clone. This page covers standing one up: first deploy, DNS + TLS, database migrations, and seeding the first tenant.

One-time: create the Fly app ​

From the repo root:

bash
cd apps/gateway
fly launch --no-deploy

This reads apps/gateway/fly.toml (app name tenkeybridge-gateway, region sjc — the nearest live region to the Neon database in aws-us-west-2; Fly deprecated den and sea in mid-2026 — min_machines_running = 1, auto_stop_machines = "off" — the agent's persistent WebSocket connection needs a machine that's always warm) and creates the Fly app without deploying yet.

Secrets ​

DATABASE_URL is never committed to fly.toml — it's set as a Fly secret so it isn't visible in the build config or flyctl config show:

bash
fly secrets set DATABASE_URL="postgres://<user>:<pass>@<neon-host>/<db>?sslmode=require"

Use the pooled Neon connection string (PgBouncer, port 6543) — the gateway holds one long-lived pg.Pool per process, and Fly may run more than one machine.

Auth secrets (#91 P1) ​

BETTER_AUTH_SECRET signs session cookies and tokens. Generate 32+ bytes and set it once — rotating it invalidates every session, so treat it as permanent unless it leaks.

bash
fly secrets set --app tenkeybridge-gateway \
  BETTER_AUTH_SECRET="$(openssl rand -base64 32)" \
  RESEND_API_KEY="<your Resend API key>"

RESEND_API_KEY sends the magic-link email; src/config.ts requires both of these (and AUTH_EMAIL_FROM, below) — the gateway refuses to boot without them, the same as it already refuses to boot without DATABASE_URL.

Google and GitHub sign-in are optional — set both halves of a pair or set neither:

bash
fly secrets set --app tenkeybridge-gateway \
  GOOGLE_CLIENT_ID="..." GOOGLE_CLIENT_SECRET="..." \
  GITHUB_CLIENT_ID="..." GITHUB_CLIENT_SECRET="..."

src/config.ts's socialProvider() treats an id with no matching secret (or vice versa) as absent rather than erroring, so a half-configured pair just means the portal doesn't offer that sign-in button — the gateway still boots and the other provider (or magic-link email) still works. Nobody is blocked on setting these up; add them whenever the Google/GitHub OAuth apps exist.

Apple sign-in (#215) ​

Apple is optional the same way, but its secret is minted, not copied: Apple gives you a signing key (AuthKey_<KEYID>.p8, downloadable exactly once), and the client secret is a JWT you sign with that key. Apple refuses a secret older than six months, so this is a recurring chore, not a one-off.

You need four values, all from the Apple Developer account: the Team ID (top right of the account page), the Key ID of the Sign in with Apple key (Certificates, Identifiers & Profiles → Keys), the Services ID (ours is com.tenkeybridge.portal; its return URL must be https://api.tenkeybridge.com/admin/v1/auth/callback/apple and its domain api.tenkeybridge.com), and the .p8 file. Keep the key outside every repo — ~/.config/tenkeybridge/ is fine, *.p8 is gitignored as a backstop — and never on Fly: the gateway only ever sees the minted JWT.

bash
export APPLE_TEAM_ID="XXXXXXXXXX" APPLE_KEY_ID="YYYYYYYYYY" \
  APPLE_SERVICES_ID="com.tenkeybridge.portal" \
  APPLE_PRIVATE_KEY_PATH="$HOME/.config/tenkeybridge/AuthKey_YYYYYYYYYY.p8"
APPLE_CLIENT_SECRET="$(pnpm --silent gateway:apple-secret)" || exit 1
[ -n "$APPLE_CLIENT_SECRET" ] || { echo "mint failed — see the error above"; exit 1; }
fly secrets set --app tenkeybridge-gateway \
  APPLE_CLIENT_ID="$APPLE_SERVICES_ID" \
  APPLE_CLIENT_SECRET="$APPLE_CLIENT_SECRET"

pnpm --silent gateway:apple-secret (apps/gateway/src/auth/appleSecretCli.ts) prints the JWT alone on stdout and its expiry on stderr — the --silent matters, or pnpm's own banner lands in the captured value; it mints a 180-day secret. The APPLE_CLIENT_ID / APPLE_CLIENT_SECRET pair is gated by socialProvider() like the others — half-set means no button, and the gateway still boots.

Rotation. When you set the secret, put a calendar reminder about five months out; when it fires, re-run the block above. An expired secret does not take the button down — Apple rejects the token exchange, so "Continue with Apple" fails with a provider error while every other sign-in method keeps working. If the .p8 is ever lost, revoke that key in the Apple portal, create a new one (new Key ID), and mint again.

Non-secret values (PORTAL_BASE_URL, AUTH_COOKIE_DOMAIN, AUTH_EMAIL_FROM, GATEWAY_BASE_URL) live in fly.toml [env], not as secrets — they're already there as of #91.

Admin gate secret (#91 final review) ​

ADMIN_GATE_USER / ADMIN_GATE_PASSWORD close the entire /admin/v1 surface (sign-in, sign-up, org CRUD, the api-key endpoints, and the read endpoints) behind HTTP Basic auth — the human partner's call, made because that surface has no self-serve gate of its own until the P4 portal ships. Both are optional as a pair: unset entirely, src/admin/adminGate.ts skips the gate (this is what keeps local dev, CI, and the test suite green); set only one and loadConfig() throws at boot, since a half-configured gate is a deployment mistake, not a safely-degraded state.

bash
fly secrets set --app tenkeybridge-gateway \
  ADMIN_GATE_USER="..." \
  ADMIN_GATE_PASSWORD="$(openssl rand -base64 24)"

The password is a Fly secret, never a fly.toml [env] entry — that's exactly the mistake this gate exists to prevent one layer up. Remove both once the P4 portal ships its own account-signup/auth flow — leaving them set past that point just adds a second, forgotten front door.

A valid x-api-key bypasses this gate on its own — no -u Basic flag needed — everywhere on /admin/v1 except in front of /admin/v1/auth/* (sign-in, sign-up, OAuth callbacks, magic links), which stays Basic-gated unconditionally: an API key must never be a way into the sign-up surface this gate exists to hide. Once you hold an org-scoped key, every other admin call looks like:

bash
curl -s https://api.tenkeybridge.com/admin/v1/realms?orgId=$ORG_ID \
  -H "x-api-key: $TKB_API_KEY"

See Admin API for the full endpoint list and response shapes.

OAuth provider callback URLs — register these with each provider (Better Auth is mounted at /admin/v1/auth, see src/auth/auth.ts's basePath):

ProviderAuthorized redirect URI
Googlehttps://api.tenkeybridge.com/admin/v1/auth/callback/google
GitHubhttps://api.tenkeybridge.com/admin/v1/auth/callback/github

AUTH_COOKIE_DOMAIN=.tenkeybridge.com is what lets portal.tenkeybridge.com read a session cookie set by api.tenkeybridge.com. Leave it unset locally — browsers reject a Domain attribute on localhost.

Running db:migrate or seed from a shell (not on Fly) needs the same required vars in that shell's environment, not just DATABASE_URL — loadConfig() validates the full config eagerly before either command touches the database. For a one-off migration or seed run, a throwaway 32+ byte BETTER_AUTH_SECRET and any non-empty RESEND_API_KEY / AUTH_EMAIL_FROM are enough; neither command sends email or opens a session.

Billing secrets (#92) ​

Billing (Stripe subscriptions, metered per company file — see the billing guide) is opt-in: src/config.ts's loadBilling() treats STRIPE_SECRET_KEY, STRIPE_WEBHOOK_SECRET, and STRIPE_PRICE_PRODUCTION as all-or-nothing. Leave all three unset and the gateway boots exactly as before #92 (no plan enforcement, no /billing/v1 route). Set only one or two and loadConfig() throws at boot — a half-configured Stripe integration is a deployment mistake, the same idiom as the admin gate pair above.

bash
fly secrets set --app tenkeybridge-gateway \
  STRIPE_SECRET_KEY="sk_test_..." \
  STRIPE_WEBHOOK_SECRET="whsec_..." \
  STRIPE_PRICE_PRODUCTION="price_..."

Use Stripe test-mode keys (sk_test_...) until launch. Erik sets the live (sk_live_...) keys when the product actually goes live — do not swap in live keys unprompted.

BILLING_ENFORCE controls whether a non-entitled org's requests are actually blocked:

bash
fly secrets set --app tenkeybridge-gateway BILLING_ENFORCE="false"
  • Unset (or empty) defaults to true once the three Stripe vars above are all set — enforcement is on by default, not an opt-in on top of opt-in.
  • Only the exact lowercase literals "true" and "false" are recognized. Any other value ("0", "False", "no", …) is silently treated as false — it does not error, and it does not mean "true" — so always set it to exactly one of those two strings.
  • BILLING_ENFORCE is ignored (effectively false) whenever billing itself isn't configured — there's nothing to enforce.
  • The daily trial-email sweep (startTrialNotifierJob, #243 lane 2 — the day-23 warning and day-30 "trial has ended" emails) only starts when billing is both enabled and enforced. It stays off in shadow mode (BILLING_ENFORCE=false) too, since its day-30 email says "API access is paused," which would be false while enforcement is off.

Register the webhook endpoint in the Stripe Dashboard (or stripe listen/stripe trigger for local testing) pointed at:

https://api.tenkeybridge.com/billing/v1/webhook

Subscribe it to exactly these six event types — src/billing/webhook.ts only handles these (everything else is a no-op "ignored" branch):

  • checkout.session.completed
  • customer.subscription.created — a fresh checkout emits this (not .updated), and it is what stores the subscription item id the quantity sync needs; without it, realm creates/deletes don't adjust the Stripe quantity until the first renewal
  • customer.subscription.updated
  • customer.subscription.deleted
  • invoice.payment_failed
  • invoice.paid

Copy the endpoint's signing secret (whsec_...) from the Stripe Dashboard into STRIPE_WEBHOOK_SECRET above — that's what stripe.webhooks.constructEvent verifies deliveries against.

Rollout: deploy with BILLING_ENFORCE=false first. In shadow mode every non-entitled request still succeeds (200, not 402), but the entitlement check still runs and logs a warn — { orgId, realmId, reason }, message "billing: organization is not entitled" — for every one it would have blocked. Watch that line in Axiom for a few days to confirm it's only firing for orgs you expect (test/trial accounts, not real customers), then flip:

bash
fly secrets set --app tenkeybridge-gateway BILLING_ENFORCE="true"

Staff console secrets (#243) ​

The staff console (console.tenkeybridge.com) is a separate, allow-listed surface — see Staff console below for how it's gated end to end. Two Fly secrets control it:

bash
fly secrets set --app tenkeybridge-gateway \
  SUPERADMIN_EMAILS="erik@empireinnovators.com" \
  CONSOLE_HOST="console.tenkeybridge.com"
  • SUPERADMIN_EMAILS — comma-separated, case-insensitive (config.ts's loadStaff lower-cases, trims, dedupes, and sorts the list). Empty or unset closes the console to everyone — this is the fail-safe default, not a degraded state, so it must be set deliberately.
  • CONSOLE_HOST — the hostname server.ts's host router matches to route a request to the console app instead of the customer app/portal. In dev, leave this unset and set CONSOLE_DEV=1 instead (not a secret — a plain env var). CONSOLE_DEV=1 makes console.<host-of-GATEWAY_BASE_URL> (e.g. console.localhost:8080) the dev console host, so pnpm gateway:dev serves the console at its own console.localhost:<port> while plain localhost:<port> keeps routing to the customer app/portal. The two are mutually exclusive — setting both CONSOLE_HOST and CONSOLE_DEV is a config error the gateway refuses to boot with.
  • STAFF_SESSION_MAX_AGE_MS — optional, default 12 hours (43200000). A staff session expires this long after sign-in, with no sliding refresh: the staff Better Auth instance's own session lifetime is set to the same value with refresh disabled, and requireStaff also checks the session's createdAt against it — so the console re-prompts sign-in more often than the portal on purpose, and activity never extends it.
  • AXIOM_EXPLORER_BASE_URL — optional, not a secret and not a config.ts knob (a raw process.env read in server.ts, same convention as PORTAL_VITE). When set, each request-log row's axiomUrl becomes an "Open in Axiom" deep link (${AXIOM_EXPLORER_BASE_URL}?q=requestId == "<id>"); when unset every row's axiomUrl is null rather than guessing at a URL shape.
  • API_REQUEST_LOG_RETENTION_DAYS — Lane 1's knob (api_requests retention); see Request log — documented there, not repeated here, since it's the data both the Requests page (console UI, Lane 4) and /staff/v1/requests (below) read.

Web Connector ​

QuickBooks Web Connector (QBWC) is the second transport — see the Web Connector guide for what it is from a customer's side. Everything below is what changes operationally on the gateway.

Knobs ​

All optional, all in apps/gateway/src/config.ts, none secret — set as plain fly.toml [env] entries or fly secrets set if you'd rather not have them visible in the repo (they carry no sensitive value either way):

Env varDefaultMeaning
QBWC_BASE_URLGATEWAY_BASE_URLOrigin written into every generated .qwc (AppURL, AppSupport) and advertised in the WSDL. Production sets https://qbwc.tenkeybridge.com — the CloudFront edge whose certificate chains to a root baked into Windows itself (see Web Connector edge, #231). api.tenkeybridge.com/qbwc stays mounted for .qwc files issued before the edge existed. Must be https.
QBWC_MIN_RUN_EVERY_S10Poll interval written into every .qwc (RunEveryNSeconds). It is not echoed back in the authenticate reply any more: the Web Connector that ships with QuickBooks 2024 (34.0.10010.76) reads authRet[3] as whole minutes, rounds 10 s to 0 and clamps the row to a 1-minute poll on every successful authenticate (seen on Rightworks 2026-09-21), so the .qwc value alone sets the cadence. The config loader refuses a value above QBWC_REQUEST_TIMEOUT_MS / 3 (80 s at the default 240 s Web Connector budget) — a slower poll than that can't finish inside the timeout even in the best case.
QBWC_AUTH_HOLD_MS10000How long authenticate may park a poll waiting for queued work before answering "nothing to do." Set to 0 for plain polling with no hold.
QBWC_SESSION_IDLE_MS10000How long receiveResponseXML may park after a response, waiting for the next queued request before ending the session. Set to 0 for plain polling with no hold.
QBWC_MAX_NOOPS2How many NoOp keep-alives a work-less sendRequestXML sends before giving up and letting the session close.
QBWC_OFFLINE_AFTER_MS90000No authenticate call for this long ⇒ the realm's Web Connector is reported offline (portal badge, isOnline()) — unless a session is still live. Web Connector only authenticates to open a session, and a client that keeps queueing work inside QBWC_SESSION_IDLE_MS keeps one session open for the whole run, so a realm also counts as online while it has a session whose last call is within QBWC_SESSION_TTL_MS (#239).
QBWC_SESSION_TTL_MS300000An idle session ticket older than this is swept; any request it had claimed is marked lost rather than re-run.
QBWC_REQUEST_TIMEOUT_MS240000How long a REST call served by Web Connector waits for QuickBooks before 504 AGENT_TIMEOUT (#190). Deliberately longer than AGENT_REQUEST_TIMEOUT_MS (60 s): under the "Yes, always" unattended grant the first call after QuickBooks was closed makes Web Connector cold-launch it — measured 26–27 s end to end on EMPIRE when nothing is in the way (2026-09-16 §4.5 proof), but a launch that stalls behind a dialog or a Windows UAC prompt on the host takes as long as the prompt goes unanswered — and that call should complete, not time out. The loader refuses a value below AGENT_REQUEST_TIMEOUT_MS. A realm with a live agent socket never sees this budget — the router prefers the agent. Ceiling 290000 (below the 300 s Fly idle_timeout and Node's default 300 s requestTimeout); must also stay below QBWC_SESSION_TTL_MS.
TRANSPORT_JOURNAL_RETENTION_DAYS30How long a terminal row (done/failed/expired/lost) stays in transport_requests before the daily retention sweep deletes it.
PORTAL_AGENT_UIunsetExactly true shows the Windows-agent section and agent-token UI on the portal's realm page; anything else hides them so the portal offers only Web Connector (the launch posture). Presentation only — the Admin API's agent-token endpoints and pnpm gateway:seed work either way.

Fly proxy idle timeout must exceed the Web Connector budget. Fly Proxy closes an HTTP connection with no bytes flowing after idle_timeout (60 s by default), which would cut a 4-minute Web Connector wait before the gateway answered. apps/gateway/fly.toml sets [http_service.http_options] idle_timeout = 300; if you ever raise QBWC_REQUEST_TIMEOUT_MS, raise idle_timeout with it so that idle_timeout > QBWC_REQUEST_TIMEOUT_MS / 1000 always holds. Node's http.Server has no response timeout of its own, so nothing else in the stack needs adjusting.

When to set a hold to 0: QBWC_AUTH_HOLD_MS and QBWC_SESSION_IDLE_MS are what turn a 10-second poll into something close to push — a queued request wakes a parked call immediately instead of waiting for the next scheduled poll. They're not part of the published SOAP contract, so if a particular Web Connector build ever turns out to time out on a held call (shows up as QBWC1012/QBWC1041 errors that clear the moment you shorten the hold), set both to 0. The transport still works correctly at 0 — every request just waits for the next full poll interval instead of being picked up mid-interval, which only costs latency, not correctness.

Web Connector edge (CloudFront) ​

Why (#231). Hosted Windows desktops (Rightworks images) don't receive automatic root-certificate updates, so Fly's Let's Encrypt chain (ECDSA, four certs through two 2026 cross-signs, no OCSP, hourly CRLs) fails Windows chain validation inside the Web Connector process even after a user installs ISRG Root X1 by hand — every fresh TLS handshake fails, and a row only looks healthy while a handshake that IE happened to prime is kept alive (proven 2026-09-21 on a Rightworks desktop). Web Connector therefore dials https://qbwc.tenkeybridge.com/qbwc, an AWS CloudFront distribution whose ACM certificate chains to Amazon Root CA 1 (Starfield Services Root Certificate Authority - G2, present in every Windows image since XP), proxying to api.tenkeybridge.com with caching disabled. REST, OAuth, the agent socket and the portal stay on api..

As built (2026-09-22).

ItemValue
AWS account488914879355 (Empire Innovations) — Builder ID sign-in (social login, MFA on the Builder ID), no IAM user and no long-lived key; CLI access is aws login --profile tkb (browser flow, temporary creds + refresh token, role AccountFullAccessRole, region us-east-1) — re-run it in a real Terminal if the session expires
ACM certificate (us-east-1)arn:aws:acm:us-east-1:488914879355:certificate/bca2d608-119b-4fdd-b0fd-9327e59779b8 — RSA_2048, DNS-validated, ISSUED 2026-09-22 03:25Z, NotAfter 2027-04-07; keep the validation CNAME (below) forever, ACM renews through it
DistributionE3N8DN7EDC9XR9 → d345dlnahbgejo.cloudfront.net, alias qbwc.tenkeybridge.com, PriceClass_100, TLSv1.2_2021, sni-only, http2, IPv6 on
BehaviourCachingDisabled + AllViewerExceptHostHeader, all 7 methods, origin api.tenkeybridge.com https-only (TLSv1.2), read timeout 60 s, keepalive 5 s
Path guardCloudFront Function tkb-qbwc-path-guard (arn:aws:cloudfront::488914879355:function/tkb-qbwc-path-guard, cloudfront-js-2.0, viewer-request): only /qbwc, /qbwc/*, /healthz; everything else 404s with "qbwc.tenkeybridge.com serves QuickBooks Web Connector only. API: https://api.tenkeybridge.com"
DNS (Vercel)qbwc CNAME d345dlnahbgejo.cloudfront.net (TTL 60); validation _7eeaf5e38b4e8ccb45740e14f0cd9bed.qbwc CNAME _7f8041f034b132e74bdfcf6cb2c171b5.wzccmgtwzk.acm-validations.aws (keep forever); apex CAA 0 issue "amazon.com"
GatewayQBWC_BASE_URL = https://qbwc.tenkeybridge.com in fly.toml [env]
MonitorBetter Stack 4959902 — https://qbwc.tenkeybridge.com/qbwc/support (status page section "QB Web Connector")
Web-filter ratingFortiGuard category Information Technology for qbwc.tenkeybridge.com — submitted 2026-09-22 13:50Z at https://globalurl.fortinet.net/rate/submit.php, approved 14:09Z, effective within the hour; escalation link is in the FortiGuard confirmation email (Erik's inbox)
Cutover proofA customer's hosted (Rightworks) desktop, 2026-09-22: .qwc re-added 14:22Z with the Amazon RSA 2048 M01 chain shown in the Authorize dialog; Web Connector exited 14:26Z, relaunched 14:38:54Z → authenticate = none at 14:39:05Z on a fresh handshake, first work 14:40:54Z, first customer reads 14:41Z (three REST GETs through the edge, all 200: 47.4 s cold, 0.9 s, 3.5 s)

Web filters on hosted desktops (found 2026-09-22). Rightworks desktops sit behind a FortiGate with FortiGuard web filtering. A hostname FortiGuard has never rated is "Unrated": the filter answers with its own block page over a certificate it issues itself (IE shows Issued by: FG…, the appliance serial), and Web Connector reports that as QBWC1048 at add time — indistinguishable from a chain problem until someone opens the URL in the desktop's IE and reads the certificate issuer. qbwc.tenkeybridge.com was Unrated on its first customer day, so the CloudFront cutover looked like it had failed (api.tenkeybridge.com, rated since July, was never affected). Rule: any new customer-facing hostname is submitted for a FortiGuard rating the day it is created, before the first hosted-desktop customer is pointed at it — treat the rating as part of provisioning, alongside DNS and the certificate. The form is https://globalurl.fortinet.net/rate/submit.php (category Information Technology, an email address for the confirmation); the 09-22 request was approved in about 20 minutes and took effect within the hour. If a customer on another hosting provider reports QBWC1048 on a current .qwc, ask for the IE certificate screenshot first and look up which filter vendor the issuer belongs to before touching certificates.

Operating notes. tenkeybridge.com already carried Vercel-managed CAA records (letsencrypt.org, pki.goog, sectigo.com), so the first two ACM certificate requests failed with CAA_ERROR; the fix was an apex CAA record 0 issue "amazon.com" on Vercel DNS — a CAA record directly on the qbwc name isn't possible (a name that's a CNAME can't carry other record types), and CAA lookups for a CNAME fall back to the apex anyway, so the apex record is also what lets ACM's auto-renewal keep working going forward. The zone's negative-cache TTL is 600 s: wait at least 10 minutes after a CAA change before re-requesting a certificate. OCSP stapling is not in play here — openssl s_client -status against every CloudFront edge IP reports "no responses sent." The chain carries OCSP and CRL URLs on amazontrust.com that a client fetches itself if it checks revocation; Web Connector does not by default, and the EMPIRE .NET check with CheckCertificateRevocationList = $true passed regardless. Fly-Client-IP on /qbwc behind the edge is now a CloudFront address; the viewer's IP is in X-Forwarded-For. /qbwc carries no per-IP rate limiter, so nothing keys on it. .qwc files issued before the edge still point at api.tenkeybridge.com/qbwc, which stays mounted; customers on hosted desktops should be moved to a fresh .qwc (new connection in the portal, remove + re-add in Web Connector). To rotate or rebuild: aws cloudfront get-distribution-config --id E3N8DN7EDC9XR9 → edit → update-distribution --if-match. Cost: CloudFront's permanent free tier (1 TB, 10 M requests/month); ACM public certificates are free.

/qbwc and the rate limiter ​

/qbwc is mounted outside the REST app (src/server.ts) and is not covered by the REST rate limiter, OAuth, or entitlement checks — those apply to your API callers, not to the Web Connector client polling on its own schedule. There's no separate rate limit on /qbwc either; a per-username failure counter locks a username out for 60 seconds after five failed authenticate attempts in a minute, which is what actually guards the endpoint against password guessing. An authenticate carrying an empty or absurdly long username is refused with nvu before it reaches either the lockout map or the database, and the map itself is pruned every 60 seconds and capped, so a flood of junk usernames can't grow it without bound.

Every call also emits its own qbwc call log line — requestId, method, the session ticket, the realmId behind that ticket, and a one-word outcome (nvu/none/work for authenticate; qbxml/NoOp/empty for sendRequestXML; the returned integer for receiveResponseXML) — on top of the usual per-request line. Filter Axiom by realmId to follow one customer's connector, or by outcome: "nvu" to spot a client guessing passwords. The line deliberately carries no username, no password, no HCP response and no qbXML.

transport_requests retention ​

Every qbXML exchange that goes through Web Connector — and, later, the agent path once it adopts the same journal — is a row in transport_requests: queued → inflight → done/failed, or expired (deadline passed before a poll claimed it) / lost (claimed, then the session died without a response).

The journal stores no qbXML. A row carries metadata only: the status above, the timings (enqueued_at, deadline, claimed_at, completed_at), the size of the request and of the response in bytes (request_bytes, response_bytes), a SHA-256 digest of the request (request_sha256), the error code and message plus the raw COM HRESULT when QuickBooks failed it (error_code, error_message, qbwc_hresult), and the session ticket that claimed it (claimed_by). The request and response bodies live only in the gateway process, in memory, for as long as the request is in flight — nothing from your company file is ever written to the database, which is what the Privacy Policy §3 commits to. The digest is enough to confirm that two support reports describe the same request without the request itself being recoverable from it.

A daily job, the same pattern as the trial reaper, deletes terminal rows older than TRANSPORT_JOURNAL_RETENTION_DAYS (default 30 days) in bounded batches of 5,000. Two sweeps run every 60 seconds alongside it: queued rows past their deadline become expired, and rows a crashed process left inflight well past their deadline (5 minutes' grace) become lost so they are never re-served. The journal is per-org data — it's included in the cascade when a realm is deleted, same as agent tokens and Web Connector credentials.

Reading Web Connector presence ​

web_connector_credentials carries each connector's live status: lastPollAt, lastSessionAt, lastError, clientVersion, plus the pin (companyFile, companyName, pinnedAt). The normal way to read it is GET /admin/v1/realms/:realmId, which folds the newest non-revoked credential's presence into webConnector: { online, lastPollAt, lastSessionAt, companyFile, companyName, lastError } and lists every credential (revoked included) under webConnectorCredentials. For a direct look — support, or debugging a stuck connector — against the database itself:

sql
SELECT username, revoked, company_file, company_name, last_poll_at, last_error
FROM web_connector_credentials
WHERE realm_id = '<realm-id>'
ORDER BY created_at DESC;

A healthy connector's last_poll_at should be within QBWC_OFFLINE_AFTER_MS (90 s by default) of now; anything older than that is what flips the portal badge to offline even though the row still exists.

Developer sandbox ​

Sandbox realms (realms.kind = 'sandbox') are answered by the in-gateway qbXML emulator, never by an agent or Web Connector. Two knobs bound what one sandbox can cost (sandbox spec §6.2):

Env varDefaultWhat it bounds
SANDBOX_DAILY_CALLS2000REST calls per sandbox per UTC day → 429 SANDBOX_DAILY_LIMIT
SANDBOX_MAX_RECORDS10000Records per sandbox (lines don't count) → 400 SANDBOX_RECORD_LIMIT

Both are plain env, not secrets — set them in fly.toml's [env] block (or with fly secrets set, either works) to change the defaults. Abuse beyond the caps: the staff console's Disable user action (POST /staff/v1/users/:id/disable) blocks that developer's sign-in, or seed delete-org removes the whole org and its realms. Iterators live in process memory for 10 minutes, which the single-machine constraint makes safe; a restart only forces the REST layer's stateless re-walk.

Sandbox parity ​

Plan L3-R2/R3/R6. The emulator (above) is only as good as its agreement with real QuickBooks Desktop — pnpm sandbox:parity runs the same scripted scenarios (apps/gateway/src/sandbox/parity/scenarios.ts, ALL_SCENARIOS) against both sides and diffs the normalized responses. There is no scheduled/CI workflow for this (L3-R3): it needs EMPIRE's real QuickBooks and a human at the console to keep it open on the pinned file, so it runs by hand, on demand — see the PR template's Sandbox parity checklist.

One-time EMPIRE setup ​

Company file. Erik, in the QuickBooks UI:

  • A new company "ExampleCo Landscaping", created with Express Start, industry General Service-based Business (not Lawn Care / Landscaping — its extra accounts and items would collide with the seed's API-layer creates), with no sample data.
  • Preferences:
    • Sales tax on. QuickBooks refuses to save this until a "most common sales tax item" exists, so create the seed's item right there: Sales Tax Item ExampleCo County Tax, description County sales tax, rate 7.5%, tax agency = new vendor Example County Treasurer, and pick it as the most common item. The base layer's ensure-by-name then reuses it, so these values must match apps/gateway/src/sandbox/seed/exampleco.ts exactly. Codes stay the defaults Tax / Non.
    • Inventory and purchase orders on.
    • Estimates on.
    • Class tracking on.
    • Time tracking on.
    • Account numbers off.
    • Payments ▸ Company Preferences ▸ "Use Undeposited Funds as a default deposit to account" — turn this on. QuickBooks has no "Undeposited Funds" account on a fresh file; it creates its own special account (SpecialAccountType UndepositedFunds) lazily, the first time a payment needs it — this preference is what makes that happen before --apply-seed runs. Recording one payment through the UI instead works too. Without it, the seed's payments land in an ordinary look-alike account instead of QuickBooks' real one, and every deposit that links a payment fails Desktop 3180 ("The given record number is not in the Payments to Deposit list") — the real-QB finding that shaped the seed runner's special account handling (seed/runner.ts).
  • Open the file single-user, non-elevated.
  • Company-info screens (legal name, LLC type, EIN, address, admin email) take fictional values — ExampleCo Landscaping LLC, single-member LLC, EIN 12-3456789, 100 Example Way, Boise, ID 83702, admin@example.com, fiscal year starting January. Parity allow-lists CompanyInfo, so they don't matter beyond staying fictional.

(Done 2026-09-24 — this file already exists on EMPIRE with the tax item and treasurer vendor in place.)

Parity agent. A second, separate copy of the Windows agent, so it never fights tkb-agent (the production dev-realm agent) over the same company file — QuickBooks serves one company file per session, so stop tkb-agent first. Same binary (tenkeybridge-agent.exe, same one tkb-agent runs); it's configured by environment variables in a small launcher script, not by its own appsettings.json (the agent reads appsettings.json next to the exe, then TENKEYBRIDGE_-prefixed env vars, which win — AgentConfig.cs/agent-install.md's Configuration reference has the full key list):

  • C:\Users\erik\tkb\parity\run-parity-agent.cmd:

    bat
    set TENKEYBRIDGE_GATEWAYURL=ws://127.0.0.1:8787/agent
    set TENKEYBRIDGE_AGENTTOKEN=<the Mac's ~/.config/tenkeybridge/parity/agent-token>
    set TENKEYBRIDGE_COMPANYFILE=C:\Users\Public\Documents\Intuit\QuickBooks\Company Files\ExampleCo Landscaping.qbw
    set TENKEYBRIDGE_AGENTID=parity
    cd /d C:\Users\erik\tkb\parity
    ..\tenkeybridge-agent.exe >> parity-agent.log 2>&1

    TENKEYBRIDGE_GATEWAYURL is the reverse tunnel's local end (see "Each run" below; use a different port than 8787 only if you also pass --port to sandbox:parity). TENKEYBRIDGE_AGENTTOKEN is the token sandbox:parity creates at ~/.config/tenkeybridge/parity/agent-token on the Mac — printed once, the first time any mode runs against EMPIRE (an operator can also create that file by hand on first setup). >> parity-agent.log keeps its own log, separate from tkb-agent's.

  • Scheduled task tkb-parity-agent, on-demand only (unlike tkb-agent's /SC ONLOGON, this one should never auto-start and fight tkb-agent for the company file):

    schtasks /Create /TN tkb-parity-agent /TR C:\Users\erik\tkb\parity\run-parity-agent.cmd /SC ONCE /ST 00:00 /SD 01/01/2030 /IT /RL LIMITED

    (/SC ONCE /ST 00:00 /SD 01/01/2030 is a schtasks idiom for "never fires on its own" — the task exists only to be triggered by hand with schtasks /Run below. /IT /RL LIMITED runs it non-elevated on the interactive desktop session, which COM requires; see auto-memory windows-box-access.)

Each run ​

  1. On the Mac: ssh -N -R 8787:127.0.0.1:8787 empire in its own terminal — this reverse-tunnels EMPIRE's port 8787 back to the Mac's own 8787, where sandbox:parity --real's throwaway local gateway (startRealHost, apps/gateway/src/sandbox/parity/realHost.ts) listens. Leave it running for the duration.
  2. On EMPIRE (over ssh empire): stop tkb-agent if it's running, then schtasks /Run /TN tkb-parity-agent.
  3. On the Mac: pnpm sandbox:parity -- --real.
  4. When done: stop tkb-parity-agent (schtasks /End /TN tkb-parity-agent), restart tkb-agent if the dev realm needs it, and Ctrl-C the tunnel.

sandbox:parity's own --port/--out flags default to 8787 and ~/.config/tenkeybridge/parity/reports; pass --port to both the CLI and the ssh -R command together if 8787 is taken locally. --only <regex> narrows the suite to scenarios whose id matches, for a triage rerun of just the handful a fix touches (valid with --self/--real).

--dump-xml <dir> (valid with --real only, off by default) writes every qbXML request/response pair the run makes against EMPIRE's real agent to <dir>, one NNNN-<scenario>.<step>.request.xml / …response.xml pair per round trip (realHost.ts's dumpingTransport, wrapping the same Bridge that carries production traffic — nothing in the normal, undumped path changes). Use it when a diff's ROOT CAUSE needs the actual wire qbXML, not just the REST JSON the parity report already shows — e.g. settling whether real Desktop's InvoiceQueryRs genuinely omits LinkedTxnRet for an applied payment (Task 10 triage-3 bucket D) or a BillPayment's SetCredit link (bucket E):

pnpm sandbox:parity -- --real --only '^invoice\.paged$' --dump-xml ~/tmp/tkb-dump

First run only ​

Before the first --real run against a freshly created company file:

  1. pnpm sandbox:parity -- --apply-seed — seeds ExampleCo (apps/gateway/src/sandbox/seed/exampleco.ts, the same EXAMPLECO definition the emulator's own sample template is built from) into the real file over the REST pipeline, exactly as a real customer's API calls would. Writes ~/.config/tenkeybridge/parity/exampleco-ids.json as it goes (one write per record), so a run interrupted partway resumes from where it left off rather than re-creating already-seeded records (ensureExisting: true also re-finds anything the base layer expects to already exist by name).
  2. pnpm sandbox:parity -- --capture-reports — the 5 report types (P&L, Balance Sheet, Trial Balance, AR aging, AP aging) and the Budget entity's report fan-out have no emulator implementation to diff against (Desktop's own report engine is the only source of truth), so this captures real responses once and commits them; apps/gateway/src/sandbox/reportOps.ts replays them verbatim for sandbox realms. Writes apps/gateway/src/sandbox/seed/reports.json.
  3. Commit reports.json (and re-run --capture-reports — never hand-edit it — if the seed or the report registry ever changes what it should contain).

Reading the report ​

pnpm sandbox:parity -- --real prints a markdown summary (formatSummary) and writes the same thing plus the full JSON diff list to <out>/parity-<run>.{json,md} (default ~/.config/tenkeybridge/parity/reports/). It exits 1 when unexplained > 0.

Every diff falls into one of two buckets:

  • A real emulator bug. Fix apps/gateway/src/sandbox/ (a handler, the seed, or a translation-core mapping) so the emulator agrees with Desktop, then re-run.
  • A difference that can't be closed (Desktop non-determinism, a field the emulator deliberately doesn't model, timing). Add an entry to apps/gateway/src/sandbox/parity-allowlist.ts (PARITY_ALLOWLIST) with a stable id (A1, A2, …) and a reason — spec §5.3's rule is every allow-listed diff is published, with its reason, on the Sandbox docs page, so "reason" means an explanation a developer reading the docs would accept, not just "known issue".

TaxCode run-hash collision ​

TaxCode's Name is capped at 3 characters on Desktop, too short to carry the full {run} token every other list entity's uniqueness field uses, so the TaxCode scenarios' names are <1-char verb code><2-char base-36 hash of the run token> (scenarios.ts's taxCodeName/hashBase36) — only 1,296 possible values per verb, and unlike every other parity-created record, TaxCodes are never deleted (Desktop has no TaxCode delete), so names from every past --real run persist in the company file. After roughly 40 runs, a same-verb collision with an earlier run's leftover TaxCode becomes likely by the birthday bound. It surfaces as a visible diff (a create that should expect: "ok" instead faults a duplicate-name error on the real side only, since the emulator's realm resets between runs and never accumulates old names). Remedy: in the QuickBooks UI, make the colliding TaxCodes inactive, or rename them, freeing up that verb's name space; there is no in-emulator fix, since the collision is a property of the real company file having history the emulator side doesn't share.

Why no scheduled workflow (L3-R3) ​

--real needs EMPIRE's real QuickBooks reachable over a live reverse SSH tunnel and a human to keep the company file open and pinned — there is no unattended path to it (unlike the Web Connector transport, which is exactly what lets customers avoid running anything at all). So parity is a PR gate by convention (the pull request template's checklist), not a CI job: a PR that touches apps/gateway/src/sandbox/ names its --real result (or N/A) before merge, same spirit as the compat-page regen check but run by hand.

Query kill switches ​

Two gateway env vars restore the pre-v2 query behavior without a deploy of new code (fly secrets set / fly config are Erik's hands; both default to on):

  • TKB_QUERY_NAME_PUSHDOWN=off — stop pushing name =/LIKE into a QuickBooks NameFilter (full scan + exact gateway-side check).
  • TKB_QUERY_TWO_PHASE=off — stop the two-phase slim scan (Lane 4): queries whose WHERE / ORDERBY can't be pushed down go back to one wide full-record fetch. Use it if a query misbehaves on a particular QuickBooks version (e.g. a Desktop that rejects IncludeRetElement answers automatically with the single-phase scan, but the extra request shows in the request log). Shape: a slim *QueryRq with IncludeRetElement, then a ListID/TxnID fetch of just the page.

Request log ​

Every /v3/company/* REST call — success or failure, whichever transport served it — is a row in api_requests, written only when the URL's :realmId segment is a numeric string (/^\d{1,32}$/, matching QBO's own realmId shape — store/hash.ts's newRealmId). This gate runs before the bearer token is checked, so without it an unauthenticated request could put arbitrary junk in that URL segment and have it land straight in the table.

An unauthenticated/401 flood against a numeric realm URL still costs one fire-and-forget insert per request (#246) — that's inherent to logging every attempt, not just successes — but it can no longer run unbounded: an IP-keyed limiter (RATE_LIMIT_REST_IP_PER_MIN, rest/routes.ts) is mounted AHEAD of bearer auth on /v3/company/*, so a flood from one address is capped before it can reach either the token DB lookup or this insert. It's a separate rule from the per-realm RATE_LIMIT_REST_PER_MIN limiter below it (#94) — that one only ever sees requests that already carry a valid token — sharing the same in-memory store under its own restip: key prefix.

This is metadata only, same commitment as transport_requests above: no query text, no request/response body, no headers. A row carries the org and realm, the OAuth client, method and route — the matched route template, exactly as Hono reports it (c.req.routePath), e.g. /v3/company/:realmId/query, or /v3/company/:realmId/* for a request an auth/entitlement/rate-limit check rejected before it reached the inner route — never the concrete path, which would carry the realm id and, for other routes, a record id. It also carries the entity (when the route has one) and a coarse operation (read/query/create/update/delete/ cdc/other), the HTTP status and fault code on a failure, timing and byte counts, and which transport served it — agent/qbwc/none — with a link to the transport_requests row for a qbwc-served call.

Byte counts come from the Content-Length header when a handler set one. Most JSON responses don't — Hono's c.json() never sets it — so responseBytes falls back to measuring the actual serialized body's byte length (never its content) from a Response.clone() taken before the logging middleware returns; the read happens inside the same fire-and-forget write, so it never delays the response the client receives.

The write happens fire-and-forget, from the same requestLogger middleware that builds the "gateway http request" Axiom line, right after the response is built — an insert failure is logged as api_request_log_failed and never touches the response. orgId is resolved for every /v3/company/* call now (not just when billing is configured), so both this table and the Axiom line always carry it.

A daily job — the same pattern as transport_requests' own retention sweep — deletes rows older than API_REQUEST_LOG_RETENTION_DAYS (default 90 days, longer than TRANSPORT_JOURNAL_RETENTION_DAYS's 30 because this table holds no payload and the planned staff console wants a wider window) in bounded batches of 5,000.

This table has no reader API yet — it exists so a later staff console (the super-admin console, #243) and a later customer-facing Requests tab can both read from the same data without a second migration. See docs/superpowers/specs/2026-09-22-superadmin-console-design.md §3.1.

Env varDefaultMeaning
API_REQUEST_LOG_RETENTION_DAYS90How long an api_requests row stays before the daily retention sweep deletes it.
RATE_LIMIT_REST_PER_MIN600Per-realm REST limit, mounted after bearer auth (#94) — only authenticated calls count against it.
RATE_LIMIT_REST_IP_PER_MIN1200Per-client-IP REST limit, mounted AHEAD of bearer auth (#246) — caps an unauthenticated flood before it reaches the token lookup or the api_requests insert above. 2x the per-realm default: one IP can legitimately front several realms (a shared backend calling multiple company files), so it needs headroom RATE_LIMIT_REST_PER_MIN doesn't. 0 disables it, same convention as every other RATE_LIMIT_* knob.

Pre-deploy checklist (#91 P1) — completed 2026-08-04 ​

The P1 cutover is done: secrets set, migration rehearsed on a Neon branch, 0002 and seed backfill-org run against production, gateway deployed (v29), and 0003 (#135, org_id NOT NULL) shipped behind it. Kept below as the record of what a schema-plus-data cutover on this gateway involves.

  1. Set secrets — BETTER_AUTH_SECRET, RESEND_API_KEY (and ADMIN_GATE_USER/ADMIN_GATE_PASSWORD if the gate isn't already configured) — see §Secrets above.
  2. Rehearse the migration on a Neon branch — §Required pre-deploy step, below. Do not skip straight to production.
  3. db:migrate against production — §Run migrations against Neon.
  4. seed backfill-org against production — §The P1 ownership backfill.
  5. Deploy — fly deploy, below.

The ordering mattered because P1's SELECTs (verifyAdminKey, getOAuthClient, verifyClientSecret) name org_id, and the gateway refuses to boot without BETTER_AUTH_SECRET / RESEND_API_KEY — deploying before migrating would have taken the live realm's OAuth surface down, or crash-looped the whole service including the agent WebSocket. The same rule holds for any future migration that adds a column existing code reads: migrate first, deploy second.

For ordinary deploys, fly.toml's release_command = "pnpm db:migrate" (runs after build, before the new image serves traffic) handles migrations on its own; re-running an applied migration is a no-op. It never runs a backfill or any other seed command — those stay by hand.

Deploy ​

bash
fly deploy

This builds apps/gateway/Dockerfile from the repo root (multi-stage, node:24-slim). One stage compiles @tenkeybridge/portal (Vite) and copies apps/portal/dist into the runtime image, and — since #243 — another stage compiles @tenkeybridge/console (the staff console SPA) and copies apps/console/dist the same way; the gateway stage is pnpm install --prod and runs via tsx directly (no compile step). Then release_command (pnpm db:migrate, [deploy] in fly.toml) runs against the production database, and only then is the new image rolled out to traffic.

Locally, pnpm gateway:dev is the one command that yields API + portal on the same origin (Vite middleware on the gateway port). Do not run the portal Vite dev server on :5173. pnpm console:dev (root) is the console counterpart — pnpm --filter @tenkeybridge/gateway console:dev under the hood, its own gateway script (CONSOLE_DEV=1 CONSOLE_VITE=1 tsx watch src/server.ts, distinct from dev, which hardcodes PORTAL_VITE=1) — attaching the console's own Vite HMR server instead of the portal's. The two are mutually exclusive locally (only one of PORTAL_VITE/CONSOLE_VITE is ever set per boot), so run pnpm gateway:dev or pnpm console:dev, not both at once.

Health & readiness ​

The gateway exposes two check endpoints, both excluded from the request logging described in Logs below (Fly polls them continuously, and logging every poll would drown real traffic):

  • GET /healthz — liveness only: "the HTTP listener is up." Always 200 { ok: true } as long as the process is running and accepting connections. It does not touch the database.
  • GET /readyz (#94) — readiness: "the gateway can actually serve traffic." Runs SELECT 1 against the database with a 1.5s timeout and returns 200 { ok: true, db: "ok" } on success or 503 { ok: false, db: "error" } if the query fails or times out. The timeout races the query with a timer via Promise.race — the timer is always cleared and is unref()d, so a hung DB connection can never hang this endpoint or leave a dangling handle behind. On failure only a short, static reason is logged at debug level — the underlying driver error can carry connection-string material, so it's never logged at info/warn or included in the response body.

Sustained outage: what /readyz does not cover

The 1.5s timeout guarantees the response to the poller never hangs, but it only gives up client-side — it does not cancel the in-flight SELECT 1 or the connection attempt underneath it. During a genuine sustained Postgres outage, each 30s poll that lands on a hung query or a stuck connection attempt can leave that query/connection abandoned rather than freed, and over enough polls this can pressure the gateway's own connection pool (the same pool real REST traffic uses) even though every individual /readyz response still comes back in ~1.5s. A per-query statement_timeout was investigated as a fix (#94 review) but isn't a cheap addition here: the gateway's Db handle is deliberately driver-agnostic (real Postgres in production, PGlite in tests, both behind the same execute() call — see src/store/db.ts), and PGlite's driver neither supports the multi-statement SET LOCAL statement_timeout; SELECT 1 trick node-postgres allows nor actually enforces statement_timeout cancellation when driven the way that would require. Making it real-Postgres-only would mean a second, untested code path solely for production. Follow-up, not yet built: a small, separate connection (or 1-connection pool) dedicated to /readyz, with its own connectionTimeoutMillis/statement_timeout set at construction time. That configuration would live only on this separate connection — the shared app pool real REST/OAuth traffic uses stays untouched — so a hung readiness check could never accumulate against production traffic's own connections either.

fly.toml wires both into [http_service.checks]:

toml
[[http_service.checks]]
  method = "get"
  path = "/healthz"
  interval = "15s"
  timeout = "2s"

[[http_service.checks]]
  method = "get"
  path = "/readyz"
  interval = "30s"
  timeout = "3s"
  grace_period = "10s"

/readyz is checked less often and given a longer timeout and a startup grace period than /healthz — a DB blip shouldn't trip a restart as eagerly as the process being fully unresponsive would.

Paired with the checks is a restart policy:

toml
[[restart]]
  policy = "always"
  max_retries = 10

A machine that fails its checks (crashes, or reports not-ready repeatedly) is restarted automatically rather than left down until someone notices fly logs or an alert. max_retries = 10 caps the restart loop so a consistently-broken deploy (e.g. bad DATABASE_URL) fails loud instead of churning forever.

Single-machine constraint ​

The gateway currently runs as exactly one Fly machine (min_machines_running = 1, auto_stop_machines = "off"), and that is a hard constraint, not a cost-saving default: AgentHub (apps/gateway/src/hub/agentHub.ts) keeps every connected edge agent's WebSocket in an in-memory map, keyed by realm, that exists only on the machine that accepted that agent's upgrade. If a second machine were added, a REST request for a realm whose agent dialed into the other machine would 503 AGENT_OFFLINE even though that agent is online — just on the wrong box. So min_machines_running must never be raised above 1 as it stands today. (A restart under the policy above is safe — it replaces the one machine, it doesn't add a second one alongside it.)

Upgrade path, if the gateway ever needs to scale past one machine:

  1. fly-replay by realm — keep a realm→machine-id lookup (e.g. in Postgres or Fly's own machine metadata) and have any machine that receives a REST request for a realm it doesn't hold an agent socket for respond with a Fly-Replay header pointing at the machine that does. Fly re-routes the request there. Lowest-effort option; keeps AgentHub's in-memory design as-is.
  2. A shared broker — move agent socket state (or the requests/responses themselves) out of process into something all machines can reach, e.g. Redis pub/sub or a similar message bus, so any machine can serve any realm regardless of which one holds the actual WebSocket. More work, but removes the single-machine constraint entirely rather than routing around it.

Neither is built. This section exists so a future scale-up doesn't rediscover the constraint the hard way by silently raising min_machines_running and shipping intermittent AGENT_OFFLINEs.

DNS + TLS ​

Point api.tenkeybridge.com at the Fly app with a CNAME in Vercel DNS:

api.tenkeybridge.com.  CNAME  tenkeybridge-gateway.fly.dev.

Then request the certificate on the Fly side:

bash
fly certs add api.tenkeybridge.com
fly certs show api.tenkeybridge.com   # poll until status is "Ready"

Fly's edge proxy forwards the WebSocket upgrade for /agent over the same internal_port as the REST traffic — no separate listener or extra Fly config is needed for the agent plane.

qbwc.tenkeybridge.com is a CNAME to CloudFront, not Fly — see Web Connector edge.

Run migrations against Neon ​

Migrations are drizzle-kit generated SQL, checked into apps/gateway/drizzle/. Apply them with the migrate CLI, pointed at the real (non-pooled, for DDL) Neon URL:

bash
DATABASE_URL="postgres://<user>:<pass>@<neon-host>/<db>?sslmode=require" \
  pnpm --filter @tenkeybridge/gateway db:migrate

To regenerate migrations after a schema change:

bash
cd apps/gateway && pnpm db:generate

schema.ts (Drizzle table definitions) and applySchema in src/store/db.ts (hand-written idempotent DDL used by the test suite's PGlite instances) are the same schema expressed twice — keep them in agreement and re-run the gateway test suite after any change to either.

Required pre-deploy step: rehearse on a Neon branch first ​

db:migrate runs drizzle-orm/node-postgres/migrator, which wraps every pending migration in a single transaction. That path has only ever been exercised against PGlite in tests — PGlite doesn't go through the real Postgres migrator, so nothing in the test suite proves this migrator works against real Postgres. Before running db:migrate against production, branch Neon's production database (§Staging, below), point DATABASE_URL at the branch, and run the full sequence — db:migrate then seed backfill-org — there first. Confirm the branch ends with organization/user/member tables present, realms.org_id/oauth_clients.org_id populated for the existing rows, and org_id still nullable on both, then tear the branch down. Do not treat this as optional or skip straight to production — it is the only place the production data path gets exercised before it runs for real.

The gateway test suite cannot be pointed at the branch — every test config hardcodes pglite://memory, which is the whole reason this rehearsal exists. Run pnpm --filter @tenkeybridge/gateway test as a normal regression check, but to exercise the migrated database, boot the gateway itself against the branch (DATABASE_URL=<branch url> pnpm --filter @tenkeybridge/gateway dev) and drive the OAuth flow end to end: POST /oauth2/v1/authorize → POST /oauth2/v1/tokens → an authenticated /v3/company/... call. That covers getOAuthClient and verifyClientSecret — the two SELECTs this branch changed to name org_id. A 503 AGENT_OFFLINE on the REST call is the expected pass: it means auth cleared and only the agent is absent.

A gateway booted this way against a Neon branch — which carries real owner emails copied from production — only sends trial-expiry emails (startTrialNotifierJob, #243 lane 2) when billing is both enabled (all three STRIPE_* vars set) and enforced (BILLING_ENFORCE unset/"true"); leave the STRIPE_* vars unset for a rehearsal boot, and set RESEND_API_KEY=log regardless, so any send that does fire logs instead of reaching a real inbox.

Rehearsed 2026-08-03 on branch p1-migrator-rehearsal (a copy of production: 1 realm, 3 oauth_clients). 0002 applied cleanly in one transaction, adding Better Auth's eight tables and a nullable org_id; backfill-org attached 1 realm + 3 clients and was a verified no-op on a second run. Credential columns (admin_key_hash — since dropped by 0006 — secret_hash, redirect_uris) hashed identically to production before and after, and the full OAuth flow succeeded against the migrated branch. Branch torn down.

One-time: the P1 ownership backfill (#91) ​

realms and oauth_clients gained an owning organization. 0002 adds Better Auth's eight tables plus a nullable org_id column on both — it does not, and as shipped on this branch cannot, make org_id NOT NULL in the same migration run. drizzle-orm's Postgres migrator applies all pending migrations in one transaction (node_modules/drizzle-orm/pg-core/dialect.js), so a SET NOT NULL migrated in the same batch as the ADD COLUMN would fail on production's existing ownerless row and roll the ADD COLUMN back with it — leaving no org_id column at all and an unrunnable backfill. Tightening org_id to NOT NULL was therefore deferred to its own later migration and deploy, tracked in #135.

Both halves are now done. This section is kept as the record of a sequence that is finished, not a runbook to re-run: the backfill ran against production on 2026-08-04 (1 realm, 3 oauth_clients, 0 rows left NULL), and 0003 — the two-statement SET NOT NULL — shipped after it. A fresh database gets both from db:migrate in order and needs no backfill at all, because every write path has required org_id since P1.

So there were two steps here, not three, and no schema change happened between them:

bash
export NEON_URL="postgres://...neon.tech/tenkeybridge?sslmode=require"   # DIRECT url, not pooled

# 1. Better Auth's eight tables + a nullable org_id on realms/oauth_clients
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway db:migrate

# 2. Create the owning org + owner account and attach every ownerless row.
#    Idempotent: safe to re-run.
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway seed backfill-org \
  --org-name "ExampleCo" --org-slug exampleco \
  --owner-email owner@exampleco.example --owner-name "Pat Owner"

The backfill (src/store/backfill.ts) never recreates a realm and never touches admin_key_hash (dropped by 0006), secret_hash, or redirect_uris — the credentials that work before it work after it. It matches the org on slug and the user on email, and only attaches rows where org_id is still NULL, so a realm that already belongs to someone is never moved — re-running it after a partial failure, or after it has already succeeded, is a no-op on anything already attached.

The owner account it creates has a verified email and no linked provider. Signing in with Google or GitHub at that same address links to it rather than forking a second user, so the first portal login lands on the org that owns the existing realm.

Seed the first tenant ​

Use the seed CLI (apps/gateway/src/seed.ts, exposed as pnpm gateway:seed from the repo root — there is no root-level pnpm seed, that name only exists as a package script inside apps/gateway) against the same DATABASE_URL.

Realms and OAuth clients are owned by an organization as of #91 P1 — create-realm and create-client both now require --org <orgId>. If no organization exists yet, seed backfill-org (above) creates one and prints its orgId; otherwise look the id up in the organization table.

bash
# 1. Create a realm (tenant) under an org — prints realmId
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway seed create-realm \
  --name "Acme Co" --org "$ORG_ID"

# 2. Issue an agent token for that realm — the edge agent's appsettings.json needs this
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway seed issue-agent-token --realm $REALM_ID

# 3. Register an OAuth client for whatever app will call the REST API, also org-owned
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway seed create-client \
  --name "My App" --org "$ORG_ID" --redirect https://myapp.example.com/oauth/callback

Each command prints its secret (agentToken, client_secret) once — store it immediately, it is not retrievable later (only the hash is persisted).

create-realm no longer takes or prints an admin key — the adminKey/consent-form credential (P1) was fully retired in #91 P2/P3. There is nothing to pass and nothing to store beyond realmId.

gateway:seed is now the internal / break-glass path only — the self-serve Admin API (/admin/v1, P2) does everything above (create a realm, issue an agent token, register a client) for a normal customer, gated by org membership instead of shell access to this database. Use seed when you need to provision something before an organization/owner account exists to call the admin API with, or when debugging directly against the database. seed backfill-org remains the only way to create an organization from scratch until the portal (P4) ships its own signup flow.

Deleting an organization (break-glass) ​

seed delete-org is the mechanism behind the Privacy Policy's account/organization deletion commitment — there is no self-serve "delete my org" button in the portal, so an operator runs this by hand against the org's orgId.

bash
DATABASE_URL=$NEON_URL pnpm --filter @tenkeybridge/gateway seed delete-org \
  --org "$ORG_ID" --yes --reason "customer requested deletion, ticket #482"

It deletes, in one transaction: every realm the org owns and that realm's oauth_tokens/oauth_codes/agent_tokens/transport_requests/ web_connector_credentials; the org's oauth_clients; its usage_daily, billing_accounts, and api_requests (#243 Lane 1's request log — no FK, deleted explicitly for privacy completeness) rows; and the Better Auth apikey/invitation/member rows scoped to the org, then the organization row itself.

It never calls Stripe — and refuses outright if a subscription is live. delete-org and set-plan both call into apps/gateway/src/staff/actions.ts (the same functions the staff console's write endpoints use — see Staff console), whose deleteOrg/setPlan check the org's billing account before touching anything: if its status is active, trialing, or past_due, the command exits with an error instead of deleting — cancel it in the Stripe Dashboard first. This used to be a printed warning that didn't block the delete; it is now a hard refusal.

--reason <text> is required whenever --yes is given (a dry run needs none). It's written into staff_audit_log alongside the delete, with actor_email set to "seed-cli" — the same append-only log every console-driven action writes to, so a CLI break-glass delete shows up next to console actions with a reason recorded either way.

Flags:

  • No --yes: dry run. Prints a summary (org name/slug, realm/client/API-key/ member counts, billing plan+status, the Stripe warning if applicable) and exits 2 with refusing without --yes — nothing is touched. --reason is not required for a dry run.
  • --yes --reason <text>: performs the delete, or — if the org has a live Stripe subscription per above — exits non-zero with seed delete-org: org "<orgId>" has a live Stripe subscription — cancel it in the Stripe Dashboard first. on stderr and touches nothing. On success, prints a count of what was removed from each table, including deleted api_requests rows:.
  • --with-owner: also deletes any member user left with no other organization after this one is gone (their session/account/apikey/ verification rows, then the user row itself). A member who still belongs to another org is left alone and reported as skipped, never deleted.

Run the OAuth2 code-flow (authorize with realm_id → consent → redirect with code → token exchange) against /oauth2/v1/authorize and /oauth2/v1/tokens to get an ACCESS_TOKEN scoped to that realm — see Authentication.

Setting an org's plan (break-glass) ​

seed set-plan moves an org straight to the free enterprise plan, or back to trial, without a raw DB edit. It exists for two cases that never go through Checkout: the founder's own org (which must never be charged), and the "email sales@tenkeybridge.com for volume pricing" path (#92) — before this, closing that gap meant hand-editing billing_accounts.

Pull DATABASE_URL from the running Fly machine the same way as any other prod DB op, then run from the repo root:

bash
DATABASE_URL=$(fly ssh console -a tenkeybridge-gateway -C 'printenv DATABASE_URL') \
  pnpm gateway:seed set-plan --org "$ORG_ID" --plan enterprise --yes \
  --reason "founder org, never billed"

(pnpm gateway:seed needs only DATABASE_URL, same as delete-org since #212 — no other gateway secret required.)

It sets plan to the value given, resets status to "none", and clears stripeCustomerId/stripeSubscriptionId/stripeSubscriptionItemId/ currentPeriodEnd/graceUntil on the org's billing_accounts row — creating the row first if the org didn't have one yet. Neither trial nor enterprise carries a Stripe subscription, so nothing Stripe-linked is left behind to go stale. Moving to trial fills a null trialEndsAt with now (an already-set one is kept) so this can never produce a trial that never expires.

It never calls Stripe — and refuses outright if a subscription is live. set-plan calls the same apps/gateway/src/staff/actions.ts function (setPlan) the staff console's write endpoint uses — see Staff console. If the org's billing account status is active, trialing, or past_due, the command exits with an error instead of writing: cancel it in the Stripe Dashboard first. This used to be a printed warning that let the write through anyway; it is now a hard refusal — cancel the subscription in the Stripe Dashboard, then re-run.

--reason <text> is required whenever --yes is given (a dry run needs none), written into staff_audit_log with actor_email set to "seed-cli" — the same append-only log every console-driven plan change writes to.

Flags:

  • --org <orgId> (required) and --plan trial|enterprise (required) — any other --plan value is a validation error, before the DB pool even opens.
  • No --yes: dry run. Prints the same org summary as delete-org (name/slug, realm/client/API-key/member counts, current billing plan+status, the Stripe warning if applicable) plus the plan change it would make, and exits 2 with refusing without --yes — nothing is written. --reason is not required for a dry run.
  • --yes --reason <text>: performs the update, or — if the org has a live Stripe subscription per above — exits non-zero with seed set-plan: org "<orgId>" has a live Stripe subscription — cancel it in the Stripe Dashboard first. on stderr and writes nothing. On success, prints the org id and the plan it was set to.

Backfilling sandboxes ​

seed sandbox-backfill gives every organization its sandbox realm (<Org> Sandbox, ExampleCo sample data). New orgs get one automatically; run this once after the deploy that introduced sandboxes, for orgs that predate it, and any time an org is missing its sandbox (the on-create step never blocks signup, so a failure there is healed here).

bash
DATABASE_URL=… pnpm gateway:seed sandbox-backfill            # dry run, exits 2
DATABASE_URL=… pnpm gateway:seed sandbox-backfill --yes      # apply
DATABASE_URL=… pnpm gateway:seed sandbox-backfill --org <orgId> --yes

Idempotent; each line prints exists, would-create, created or healed. Rehearse on a Neon branch first (production is production).

Staff console ​

Operator-only control plane for #243 — every org and user, billing/plan state, request logs, and connection health, plus the break-glass actions that used to require the seed CLI and a raw DATABASE_URL. It replaces none of the CLI's commands — seed set-plan/seed delete-org (above) still work, and as of this lane go through the exact same code path the console's write endpoints do (apps/gateway/src/staff/actions.ts) — but it's the preferred way to do these things day to day.

Host and gating. Reachable only at https://console.tenkeybridge.com (CONSOLE_HOST — see Staff console secrets above). server.ts's host router sends every request whose Host header matches CONSOLE_HOST to a completely separate Hono app (staff/routes.ts's makeStaffApp); every other host serves the customer app/portal as before, and never has a /staff/* route at all. The console host itself 404s every customer route — /v3, /admin/v1, /qbwc, /portal, /agent — before falling through to the console SPA, so a customer path typed against the console host can't accidentally reach the SPA's fallback either.

Signing in requires an account whose email is in SUPERADMIN_EMAILS, but the allowlist is enforced at the API layer, not the auth layer: the staff Better Auth instance's own sign-in endpoint isn't allowlist-gated, so any existing account (e.g. a portal user who set a password via onboarding) can obtain a valid staff-cookie session. What is allowlist-gated is every email the instance would otherwise send — magic link, password reset, email verification — so a non-staff address can never receive a working sign-in link in the first place. Either way, every /staff/v1/* call re-checks the caller's (lower-cased) email against SUPERADMIN_EMAILS on every request, in addition to requiring emailVerified, a fresh-enough session, and user.disabled_at IS NULL. Any of the five checks failing returns a bare 404 {} — on purpose, no distinguishing detail: a wrong email, an unknown route, and an expired session are all indistinguishable from each other and from "no such route." Only a write endpoint's validation 400s (bad reason, bad body) carry a message — those are past the gate.

Sign-in methods. Email + password and magic link only (no Google/ GitHub/Apple, no organization plugin) — and no sign-up: POST /staff/v1/auth/sign-up/* is 404'd before the Better Auth passthrough even runs, and both plugins also set disableSignUp: true at the option level. A staff member's user row must already exist (created via the portal, or by seed backfill-org) before they can ever sign into the console.

Session isolation. The console has its own Better Auth instance (staff/staffAuth.ts), sharing the same user/session tables as the portal's instance but with its own cookie name (tkb-staff.session_token vs. the portal's) and, critically, its own signing secret — an HMAC-SHA256 of config.auth.secret derived with a "tkb-staff" label, not the raw secret the portal instance signs with. A portal session cookie renamed to the staff cookie name fails the staff instance's signature check outright; the console is also host-only (no crossSubDomainCookies), so the portal's cookie is never even sent to console.tenkeybridge.com in the first place. Signing into the portal never signs you into the console, and vice versa.

Staff sessions expire at STAFF_SESSION_MAX_AGE_MS (12h default) after sign-in, and never slide: the staff instance sets Better Auth's session.expiresIn to the same value with disableSessionRefresh, so GET /staff/v1/auth/get-session returns null at exactly the moment requireStaff (which also checks the session's createdAt) starts 404ing — the console can tell it has been signed out. The console re-prompts sign-in more often than the portal by design. The staff cookie is also SameSite=Strict (not the portal's Lax), HttpOnly, Secure on https, and host-only.

Staff writes need the console Origin and a JSON body. Every other *.tenkeybridge.com host is same-site with the console, so the cookie alone can't tell a console request from one a sibling subdomain's page makes (a text/plain POST with credentials: "include" is a CORS simple request — no preflight). So every /staff/v1/* request other than GET/ HEAD (outside /staff/v1/auth/*) must carry both an Origin header exactly equal to the console's base URL (https://<CONSOLE_HOST>; a trailing slash on the configured value is ignored) andcontent-type: application/json (a ; charset=… parameter is fine). Anything else — wrong or missing Origin, text/plain, a form body — is refused with 403 {"error":"forbidden_origin", ...} before any route runs. The check runs after the session gate, so an unauthenticated write still gets the bare 404 {}. Scripting a write with curl therefore needs -H "Origin: https://console.tenkeybridge.com" -H "content-type: application/json" as well as the cookie.

Disabling a user. POST /staff/v1/users/:id/disable sets user.disabled_at and deletes every one of that user's session rows in the same transaction — an already-open browser session is killed immediately, not just future sign-ins refused. A staff member cannot disable their own account — the API refuses with 400 {"error":"invalid_request","message":"You cannot disable your own staff account."}. enable clears disabled_at; revoke-sessions deletes session rows without touching disabled_at — a lighter "kick them out, but they can still sign back in" action.

Writes are audited, never call Stripe. All seven mutations (set_plan, extend_trial, set_realm_limit, delete_org, disable_user, enable_user, revoke_sessions) require a non-empty reason (≤500 characters) and write an append-only row to staff_audit_log (GET /staff/v1/audit, filterable by ?target=/?actor=) in the same transaction as the mutation — an audit-insert failure rolls the mutation back too. None of them ever call Stripe. Every refusal (staff/actions.ts's ActionError, mapped explicitly to HTTP by staff/writeHttp.ts's mapActionError — no catch-all branch) comes back as one of these, never the internal (uppercase) code name:

RefusalHTTPWire body
Missing/too-long reason400{"error":"invalid_request","message":"reason is required and must be at most 500 characters."}
Target org/user not found404{} — indistinguishable from requireStaff's own gate, on purpose
Org has a live Stripe subscription (set_plan, delete_org; status active/trialing/past_due)409{"error":"stripe_subscription_active","message":"This organization has a live Stripe subscription — cancel it in Stripe first."}
delete_org's confirm_slug doesn't match400{"error":"invalid_request","message":"confirm_slug does not match this organization's slug."}
disable_user targeting yourself400{"error":"invalid_request","message":"You cannot disable your own staff account."}
set_realm_limit with a non-null limit on a production/enterprise org409{"error":"not_trial_plan","message":"A realm-limit override only applies to trial-plan organizations."}

Cancel the subscription in the Stripe Dashboard first, then retry, for the stripe_subscription_active row above. Clearing a realm-limit override to null is always allowed regardless of plan. extend_trial is allowed on any plan — it's inert on a non-trial org, so there's nothing to protect, and it has no refusal of its own beyond the shared reason/not-found ones.

delete_org's audit row stores ids, not emails. staff_audit_log is append-only and never purged, so the persisted after for a delete_org keeps the deletion counts, the deleted realm ids, and the deleted/skipped user ids plus their counts (userIdsDeleted, userIdsSkipped, usersDeletedCount, usersSkippedCount) — never the deleted users' email addresses. The org's own name and slug stay in before. The emails still appear in the transient output (the API response's result, the seed delete-org summary), just not in the log. After the delete, every deleted realm's live agent connection is closed, same as a portal realm delete.

The seed CLI and the console share one implementation. seed set-plan/seed delete-org (see Deleting an organization and Setting an org's plan above) call the exact same apps/gateway/src/staff/actions.ts functions the console's write endpoints do — same Stripe-active refusal, same audit trail, just with actor_email = "seed-cli" instead of the signed-in staff email. Both CLI commands now require --reason <text> for the same reason the console does: every break-glass action gets a reason in the audit log, whether it was clicked or typed.

Request log. GET /staff/v1/requests reads the api_requests table (#243 Lane 1 — see Request log for what's in a row and its retention). Each row carries a link to its transport_requests row (status, QBWC HRESULT) when it came in over Web Connector, and — only when AXIOM_EXPLORER_BASE_URL is configured (see Staff console secrets) — an "Open in Axiom" deep link scoped to that request's x-request-id, for the full qbXML detail that never lands in Postgres.

Realm health. GET /staff/v1/realms returns every realm across every org (a bare array, not paginated — the realm count is small enough to return whole), sorted so a realm that's online right now (agent WebSocket or Web Connector) sorts last, and among offline realms, one that has never made contact sorts first, then oldest contact first. ?health=offline / ?health=error narrow it to realms with neither transport currently online, or with a Web Connector lastError on record, respectively — the same online/offline definitions the portal's own realm page uses (agent WebSocket presence + Web Connector's last-poll window).

Rate limits. The console's sign-in surface (/staff/v1/auth/* POSTs) shares the same RATE_LIMIT_AUTH_PER_MIN knob the portal's own auth surface uses (default 20/min), keyed per-IP under a distinct staffauth: prefix so the two buckets never collide. Every other /staff/v1/* call is additionally capped at a fixed 120 requests/minute per staff session (not configurable via env — STAFF_API_RATE_LIMIT_PER_MIN in staff/routes.ts), keyed on the session id so it only applies once a caller is past requireStaff.

Per-address auth email cap (#219). RATE_LIMIT_AUTH_PER_MIN above is per source IP across all of /admin/v1/auth/* — it does nothing to stop one IP from email-bombing many different target addresses, or many IPs each hitting the same one. RATE_LIMIT_AUTH_EMAIL_PER_ADDRESS (default 3) and RATE_LIMIT_AUTH_EMAIL_WINDOW_MS (default 600000, 10 minutes) add a second, address-keyed cap shared across the four unauthenticated triggers that can mail an arbitrary address: sign-in/magic-link, request-password-reset, send-verification-email, and sign-up/email's existing-account notice. It's enforced at the email-send layer (auth/emailThrottle.ts wraps the AuthEmail both the customer and staff Better Auth instances share — see server.ts), not in route middleware, so it covers every trigger uniformly without parsing a request body. The key is a sha256 hash of the trimmed, lower-cased address, not the address itself, in the same shared MemoryRateLimitStore the other rate limits use (a distinct authEmail: prefix keeps its buckets from colliding with theirs). A capped send is silent by design (#219's no-enumeration requirement): the route still returns its normal 200 / "check your email" response, the send simply never reaches Resend, and a warn-level log line records the suppression by hash + domain only, never the raw address. 0 disables the cap entirely, same convention as the other three RATE_LIMIT_* knobs.

Entitlement cache lag. A staff-console write to plan, trial end, or realm-limit override invalidates the gateway's 60s EntitlementCache for that org immediately, in the same request — the effect is visible on the very next REST call. seed set-plan/seed delete-org run in a separate process with no access to that in-memory cache, so a CLI write against a gateway that's already running can lag up to 60 seconds before it takes effect on live traffic (a fresh gateway boot, or a machine restart, has no such lag).

Health checks and uptime monitors must target the API host, not the console host. /healthz and /readyz are routes on the customer app (server.ts) — on console.tenkeybridge.com they aren't special-cased by staff/routes.ts, so they fall through to the console SPA's catch-all and answer with the SPA, not a health check. Keep monitoring https://api.tenkeybridge.com/healthz and /readyz (see Uptime & status); do not add a console-host monitor for either path.

Console SPA. staff/routes.ts serves static files from apps/console/dist (server.ts resolves it relative to the built gateway). The gateway image's console-build stage (see Deploy above) compiles @tenkeybridge/console and copies its dist/ into the runtime image, so a normal fly deploy ships the console along with the portal and the API — no separate build step. If that stage is ever skipped (a stale image, or a local dist/ deleted), every console page — including /staff/v1/auth/*'s own UI, which the SPA would normally host — answers 503 { error: "console_dist_missing" }. The /staff/v1/* JSON API works the same regardless; only the browser UI is missing.

Local dev. Set CONSOLE_DEV=1 (see Staff console secrets) instead of CONSOLE_HOST, then reach the console at http://console.localhost:<port> — the same port pnpm gateway:dev prints for everything else, with console. prefixed onto the host so the gateway's host router can tell it apart from the customer app/portal at plain localhost:<port>.

Knobs not already covered above:

Env varDefaultNotes
CONSOLE_HOSTunset (console unreachable)Fly secret in prod; mutually exclusive with CONSOLE_DEV
CONSOLE_DEVunset"1" or "true" in dev — console served at console.<host-of-GATEWAY_BASE_URL>
SUPERADMIN_EMAILSunset (console closed to everyone)Fly secret, comma-separated, case-insensitive
STAFF_SESSION_MAX_AGE_MS43200000 (12h)Fixed lifetime from sign-in — Better Auth expiresIn + requireStaff's createdAt check; no sliding refresh
AXIOM_EXPLORER_BASE_URLunset (axiomUrl is null on every request row)Not a secret — a display convenience, read directly from process.env rather than config.ts
API_REQUEST_LOG_RETENTION_DAYS90See Request log — the data /staff/v1/requests reads

Milestone proof ​

With an edge agent connected for the realm and a valid access token, this is the end-to-end proof the gateway is live and bridging real QuickBooks data:

bash
curl -s -H "Authorization: Bearer $ACCESS_TOKEN" \
  https://api.tenkeybridge.com/v3/company/$REALM_ID/customer/$CUSTOMER_ID | jq .

Logs ​

Every HTTP request — REST or OAuth, success or failure, even an early auth-middleware rejection — produces exactly one structured pino line (/healthz and /readyz are deliberately excluded: Fly polls both continuously and the noise would drown real traffic — see Health & readiness). This is what "check the Fly logs for that request" should actually show:

json
{"level":40,"method":"GET","path":"/v3/company/:realmId/:entity/:id","status":404,"durationMs":3,"realmId":"4562...","entity":"widget","msg":"gateway http request"}
json
{"level":40,"method":"GET","path":"/v3/company/:realmId/query","status":400,"durationMs":2,"realmId":"4562...","query":"SELECT * FROM Customer MAXRESULTS -5","errorCode":"UNSUPPORTED_QUERY","msg":"gateway http request"}
json
{"level":30,"method":"POST","path":"/oauth2/v1/tokens","status":200,"durationMs":41,"clientId":"tkbcl_...","grantType":"refresh_token","msg":"gateway http request"}

Fields: method, path (the route template, e.g. :entity/:id, not the raw URL — so log lines group cleanly), status, durationMs, and realmId/entity when the route carries them. A REST query request additionally logs the raw query text (the thing you need to diagnose an UNSUPPORTED_QUERY 422/400 without guessing), and any 4xx/5xx logs an errorCode pulled from the response body's fault/error code. OAuth requests log grantType/responseType/clientId when known. Level is info for 2xx/3xx, warn for 4xx, error for 5xx.

Hard constraint: never a request/response body, an Authorization header, a token, an auth code, or a client secret — fly logs is safe to paste into an issue.

Durable sink: Axiom ​

fly logs is lossy — Fly's NATS-based log pipeline silently drops lines under normal operation (observed live, #53), and keeps no history. The source of truth is Axiom: when the AXIOM_TOKEN + AXIOM_DATASET secrets are set, the gateway ships every log line directly to Axiom over HTTPS (via @axiomhq/pino), bypassing Fly's pipeline. stdout stays wired in parallel, so fly logs still works for casual live-tailing — just never treat a missing line there as evidence the request didn't happen; query Axiom instead.

  • Where: app.axiom.co, org empire-innovations, dataset tkb-gateway (free tier: 30-day retention, 500 GB/mo — orders of magnitude above gateway volume).

  • Setup: create an ingest-only API token scoped to the dataset (Axiom → Settings → API tokens), then:

    bash
    fly secrets import -a tenkeybridge-gateway
    # paste, then Ctrl-D:
    #   AXIOM_TOKEN=xaat-...
    #   AXIOM_DATASET=tkb-gateway
  • Rotation: revoke the token in Axiom, create a new one, re-run the import. With either var unset the gateway logs to stdout only (local dev, tests, and CI never touch Axiom).

  • Querying: the msg field is gateway http request; filter on path, status, realmId, or errorCode. Example APL:

    ['tkb-gateway'] | where path contains "companyinfo" | sort by _time desc

Alerting ​

Five Axiom monitors are specified — APL, thresholds, and an email notifier to erik@empireinnovators.com — covering 5xx rate, agent-offline spikes, unhandled errors, gateway silence, and rate-limit storms. Full spec, exact queries, field-name verification against the pino call sites, and the MCP calls to create them: apps/gateway/ops/axiom-monitors.md (repo path, not a published docs page). Not yet created — creation is blocked on Axiom MCP re-authentication; that file has a "How to create" section a follow-up session can run verbatim once /mcp is re-authed.

Staging ​

Staging E2E runs against a Neon branch database seeded with throwaway tenant data — not a second Fly app. Branch Neon's main database, point a local or preview gateway process at the branch's connection string via DATABASE_URL, run migrations and seed as above, and tear the branch down when done.

Backups & restore ​

Neon backup posture (point-in-time recovery) ​

Neon protects the database continuously via PITR (it retains the write-ahead log for a window, not periodic snapshots) — any point inside that window is restorable, not just a nightly checkpoint. The gateway project's configured retention window is UNVERIFIED as of 2026-08-28 — this session had no signed-in Neon console session and no neonctl/API key available on the build machine, so the actual value was not read. Check it before relying on it: Neon console → project tenkeybridge → Settings → History retention (neon projects get once authenticated shows the same value).

Neon's plan defaults, at time of writing:

PlanHistory retention
Free6 hours
Launch7 days

If the actual configured value is under 24 hours, upgrading the project to Launch is a launch prerequisite — a bad migration or an unqualified DELETE discovered the next morning would otherwise be unrecoverable. This is Erik's call, not an autonomous purchase — do not upgrade the plan without his sign-off; just report what the console shows.

Point-in-time branch restore ​

The normal way to inspect, or recover, a past state without touching production: branch the database as of a timestamp.

Console: project tenkeybridge → Branches → Create branch → parent main → Time → pick the point in time → Create. Connect with psql using the branch's own connection string from the console.

neon (binary neon; neonctl is an alias for it — either name runs the same commands):

bash
neon branches create --project-id <project-id> \
  --name drill-2026-08-28 --parent 2026-08-28T12:00:00Z

--parent accepts a branch name/id, a timestamp, or an LSN — a bare RFC 3339 timestamp defaults to branching off main — see neon branches create --help if this syntax has shifted since.

Then, against the branch's own connection string (never the production one):

sql
SELECT count(*) FROM realms;
SELECT max(created_at) FROM "session";

Both should come back sane — a realm count matching production, a session row timestamped near the point you picked. If either errors or looks wrong, the chosen point predates a write you needed. Tear the branch down when done:

bash
neon branches delete --project-id <project-id> <branch-id>

See apps/gateway/ops/restore-drill.md for the dated record of drills run against this recipe.

Manual dump/restore (belt-and-suspenders, off-Neon) ​

PITR covers "restore to a point in time" as long as the Neon project itself exists; a manual dump is the fallback if the project is ever lost, misconfigured, or you want an offline copy outside Neon entirely.

bash
pg_dump "$NEON_URL" --no-owner --format=custom -f tkb-$(date -u +%F).dump

Restore into a fresh Neon branch — never into main — to inspect or verify before doing anything destructive with it:

bash
pg_restore --no-owner -d "$BRANCH_URL" tkb-2026-08-28.dump

Use the direct (non-pooled) Neon URL for both, same as migrations (§Run migrations against Neon). pg_dump/pg_restore are not installed on the Mac by default — brew install libpq gets a matching client (it's keg-only, so add it to PATH), or run the dump from the Fly gateway machine (fly ssh console -a tenkeybridge-gateway) instead — the runtime image doesn't ship Postgres client tools either, so that's the same install either place.

Where dumps live: Erik's encrypted local disk, and nowhere else. Never commit a dump to the repo, attach one to a GitHub issue/PR/Actions artifact, or upload it anywhere off Erik's own machine — it's a full copy of every tenant's data.

Uptime & status ​

CREATED 2026-09-01 — Erik opened the Better Stack account (erik@empireinnovators.com, phone alerts configured during onboarding) and the four monitors + status page were created via the Uptime API (https://uptime.betterstack.com/api/v2, Bearer token from the account's API tokens page):

Monitor idNameURL
4885495Gateway liveness (healthz)https://api.tenkeybridge.com/healthz
4885514Gateway readiness (readyz)https://api.tenkeybridge.com/readyz
4885515Docs sitehttps://docs.tenkeybridge.com/
4885516Marketing sitehttps://tenkeybridge.com/
4959902Web Connector edge (qbwc) — added 2026-09-22 (#231)https://qbwc.tenkeybridge.com/qbwc/support
  • Interval: created at 30s (check_frequency: 30); as of 2026-09-22 all five run every 3 minutes (the plan's current allowance).
  • Alerts: email + SMS + call on every monitor (email/sms/call: true).
  • Status page: a public status page covering all four monitors. The default Better Stack subdomain is fine to start. A custom status.tenkeybridge.com is optional — if set up, add the CNAME Better Stack provides as a Vercel DNS record (same pattern as the api CNAME in §DNS + TLS above) and note it here once done.

Status page URL: https://tenkeybridge.betteruptime.com (status page id 261369, sections: API gateway, API gateway — database, Documentation, Website, QB Web Connector). status.tenkeybridge.com CNAME not set up (optional, see above).

Ops checklist ​

  • [ ] Monthly restore drill — run the point-in-time branch restore above (or follow the recipe in apps/gateway/ops/restore-drill.md directly), verify both queries look sane, delete the branch, and record the result as a new dated entry in apps/gateway/ops/restore-drill.md.

The public legal pages (Terms, Privacy, DPA, Subprocessors) live in apps/marketing, driven by apps/marketing/lib/legal.ts. To publish a change:

  1. Bump LEGAL.effectiveDate in apps/marketing/lib/legal.ts.
  2. Add a dated entry to apps/marketing/content/legal/CHANGELOG.md (newest first).
  3. Merge to main — marketing redeploys automatically (Vercel Git integration).

The trademark notice and legal-page links are also hardcoded in apps/portal/src/components/Trademark.tsx, LegalConsent.tsx, and apps/docs/.vitepress/config.ts — if the wording itself changes, update those alongside legal.ts.

Launch day ​

An ordered checklist the project owner executes by hand (spec docs/superpowers/specs/2026-09-15-qbwc-first-launch-design.md §6.5). Nothing here is automated, and nothing here is thrown by anyone else.

  1. Preconditions. The three QBWC-first lanes are merged and deployed: gateway via fly deploy -c apps/gateway/fly.toml from the repo root; docs and marketing deploy from Git on merge. The Web Connector cold-launch proof numbers are recorded in the Web Connector guide. Better Stack and Axiom monitors are green. Neon retention is still 7 days.
  2. Re-enable Axiom monitor 4 (gateway silence). It was disabled 2026-09-02 for flapping pre-launch (apps/gateway/ops/axiom-monitors.md); it is meaningful once there is traffic.
  3. Flip billing. Confirm the three live STRIPE_* secrets are set and the shadow-mode log line count is what you expect (see Billing secrets), then fly secrets set --stage --app tenkeybridge-gateway BILLING_ENFORCE="true". --stage writes the secret without restarting; step 4's command performs the single restart that applies both changes.
  4. The switch. fly secrets unset --app tenkeybridge-gateway ADMIN_GATE_USER ADMIN_GATE_PASSWORD — one command, so the pair never half-exists (see Admin gate secret). Watch /readyz and the status page https://tenkeybridge.betteruptime.com through the rolling restart.
  5. Merge launch/open-signup-copy (the prepared draft PR) and confirm the docs deploy — the docs must never advertise a signup the gate blocks, so this happens in the same hour as step 4, after it.
  6. Retire the "prod is dev" posture. Update the auto-memory prod-is-dev-until-launch.md: from here deploys go through PRs, DB ops through a Neon branch rehearsal, migrations only via release_command.
  7. First customer (hosted QuickBooks). Send the Rightworks guide. They create the org, the realm, and the Web Connector connection themselves in the portal. Ask them to keep QuickBooks open on the pinned file for responsiveness (or accept the ~30-second first call after a quiet period — longer if a dialog or a Windows UAC prompt stalls the launch on their host). Confirm the portal badge reads Web Connector online before their first API call.

Rollback ​

One command re-sets the gate pair and turns enforcement off:

bash
fly secrets set --app tenkeybridge-gateway \
  ADMIN_GATE_USER="<user>" ADMIN_GATE_PASSWORD="<password>" BILLING_ENFORCE="false"

The switches are reversible; what happened while they were open is not: organizations, realms, and QuickBooks writes made by a customer stay, and the merged signup copy (step 5) needs a revert PR to go back.

TenkeyBridge is an independent product, not affiliated with, endorsed by, or sponsored by Intuit Inc. QuickBooks, QuickBooks Online, and QuickBooks Desktop are trademarks of Intuit Inc., used only to describe compatibility.