Skip to content

Incident guides

Check cookie domain, BETTER_AUTH_URL, BETTER_AUTH_URL_API, CORS_ORIGINS, trusted proxy settings, and browser Origin behavior. For mobile, confirm the app points at the same API host used by Better Auth.

First checks: VALKEY_URL, TLS settings, QUEUE_SUFFIX (must match the tier — -dev / -stg / empty in PRD), and whether the worker process is running at all.

Bull Board exists on DEV/STG only — PRD has no dashboard (a Cloud Run worker pool has no URL). Queue inspection and failed-job recovery in every tier go through apps/worker/scripts/queue-admin.ts:

  • DEV/STG: a terminal in the worker container — bun run queue:admin <cmd>; or locally, cd apps/worker && bun --env-file=../../.env.local run queue:admin <cmd>.
  • PRD: a one-off Cloud Run job using the worker image with a command override (Memorystore is VPC-private, so the job must run inside the VPC). The exact gcloud run jobs execute invocation lands with the provisioning work.

Safety model: every mutating command is a dry run by default — re-run with --execute to apply. The script prints an env banner (Valkey host, QUEUE_SUFFIX, resolved queue names) before acting; read it before executing against a shared Valkey. Lead payloads contain PII (name/email/phone/message) — the list commands (list-failed, list-dlq) print them only with --json; inspect always prints the full job, payload included.

Start with queue:admin counts for waiting/active/delayed/failed totals and repeatable-scheduler health across all seven queues, then list-failed <queue> / inspect <queue> <jobId> for detail.

QueueFailure recordRecovery
leads (lead-ingestion)failed_lead_ingestions DLQ row + admin Knock alert + admin UI (/dashboard/admin/failed-leads)list-dlq --status PENDING → retry-dlq --id <row> --execute. Bad payload: fix it in SQL (jsonb_set) first — the script skips rows the worker’s type guard would reject, leaving them PENDING. The script is the authoritative replay path: it clears the stale lead-<dedupeKey> job that otherwise makes a re-enqueue silently no-op (the admin-UI replay has that limitation). Purge semantics: the daily 03:00 dlq-purge deletes terminal-status rows immediately and PENDING rows after DLQ_RETENTION_DAYS (default 90) — the script therefore enqueues, verifies, and only then marks REPLAYED.
alerts (price-drop-notify, status-change-notify)Valkey retained-failed set only (no DLQ table) — genuinely lossy once removeOnFail: 500 evictslist-failed alerts → retry-failed alerts --all --execute, promptly. saved-search-scan self-heals on its next tick.
search-indexingValkey retained-failed set onlyretry-failed search-indexing for individual jobs; for wholesale drift use the existing backfills — bun run backfill:properties / backfill:search.
billingRepeatable scans only (trial-scan, stripe-reconcile)Self-healing — each run re-derives state. Verify the schedulers exist via counts billing.
ai (generate-description)ai_jobs row (durable), but no automatic redispatch — the 30s reconciler only covers ROOM_STAGINGretry-failed ai for retained failures; redispatch-descriptions for lost dispatches / stuck rows (rebuilds the payload from the ai_jobs row; never writes ai_jobs.status).
cdn-purge (purge-property-images)Valkey retained-failed set onlyretry-failed cdn-purge --all --execute, or do nothing: a lost purge is bounded by the 24h s-maxage clamp on /api/images/*, so de-listed photos expire on their own. A terminal failure (bad CLOUDFLARE_API_TOKEN, wrong zone) does not retry at all and needs the env fixed. A once-per-process Sentry warning cdn-purge-unsupported means the zone plan has no tag purge — expected on a non-Enterprise Cloudflare zone, not an incident.
ai-staging (stage-room)Fully protected: Postgres dispatch outbox + 30s reconcilerRead-only inspection only — the script refuses to mutate this queue. The manual lever is the ai_jobs row in Postgres; see AI room staging.
storage (drain-deletion-outbox, orphan-sweep)The storage_deletion_outbox row itself — status / attempts / lastError / nextAttemptAt. No DLQ table on purpose: a job failure only means “this tick did not finish”, and the next tick re-claims the same rows, so a DLQ row would be a weaker second copy of a record that already exists. A row still failing after 10 attempts raises a Sentry event tagged storage-deletion-outbox.Self-healing — the drain runs every minute and re-claims PENDING/FAILED/lease-expired PROCESSING rows with exponential backoff capped at 60 min. Nothing is lost while the queue is down; deletions simply defer, and the bucket keeps the bytes. Triage with SELECT "status", count(*) FROM "storage_deletion_outbox" GROUP BY 1 and inspect lastError on FAILED rows — Object is still referenced by <table> is not a fault, it is the drain correctly deferring a key another live row points at for 24h. Rising non-referenced FAILED rows means the worker cannot reach S3. orphan-sweep is report-only unless STORAGE_ORPHAN_SWEEP_EXECUTE=true, and even then it enqueues into this outbox rather than deleting, so the drain’s reference re-check stays the only deletion path.
image-sanitize (sanitize-property-image)The property_images row itself (processingStatus / processingError) is the primary record; transient exhaustion also writes a failed_image_sanitize row + a Sentry event. A permanent decode failure writes no DLQ row — the FAILED tile is the signal, and the agent can Retry it.Self-healing for anything left PENDING: the repeatable image-sanitize-sweep (*/5 * * * *) re-enqueues rows older than 10 minutes, deduped on a job id derived from processingUpdatedAt. Nothing is lost while the queue is down — the raw bytes are in S3 under pending/properties/… and nothing public changed at upload time. Watch for a rising count of FAILED('transient') rows, which means the worker cannot reach S3 or Postgres; the agent-side fix is the per-photo Retry endpoint. State machine, races and queries: apps/worker/src/image-sanitize.FEATURE.md.

Rehearsal (run quarterly, and after changing queue code)

Section titled “Rehearsal (run quarterly, and after changing queue code)”

On DEV, with the worker running (cd apps/worker && bun --env-file=../../.env.local run dev); all queue:admin commands via bun --env-file=../../.env.local run queue:admin …:

  1. queue:admin counts — banner shows the -dev suffix and seven queues.
  2. queue:admin inject-test-failure --execute — enqueues a lead with an invalid propertyId; the FK violation exhausts 3 attempts in ~30s. (The command refuses to run when QUEUE_SUFFIX is empty, i.e. production-style names.)
  3. queue:admin list-dlq --status PENDING — the row appears; admins receive the lead-ingestion-failed Knock alert.
  4. queue:admin retry-dlq --id <rowId> (dry run) — the printed plan includes removing the stale failed job.
  5. Fix the payload to a real property: UPDATE failed_lead_ingestions SET payload = jsonb_set(payload, '{propertyId}', to_jsonb('<realPropertyId>'::text)) WHERE id = '<rowId>'; then re-run with --execute — the row flips to REPLAYED and the script reports the created inquiry.
  6. Loss-safety check: inject a second failure and replay it without fixing the payload — a fresh PENDING row appears after it fails again, proving a REPLAYED-then-purged row loses nothing. discard-dlq it.
  7. Clean up: DELETE FROM inquiries WHERE "dedupeKey" = '<dedupeKey>'; (lead_events cascade).

Check workflow payload shape, recipient id, channel ids, signing key, device registration, and generated workflow docs for trigger sites.

Check S3 credentials, bucket endpoint, upload route ownership checks, imgproxy key/salt, URL TTL, and CSP/image host allowlists.

Check whether the source of truth changed in Postgres, the worker indexed the update, and Valkey/Meilisearch keys match the expected index.

Use the repo reset script documented in Conventions → Monorepo. Avoid ad hoc force-reset commands because pre/post SQL ordering matters.