Development runs in Railway project kottia.com, environment development, fronted by
Cloudflare DNS (proxied only where the zone certificate covers the multi-label DEV hosts — see the
TLS note in Deployment). Only web, apps/api, and imgproxy have public domains; docs and data
services use Railway private networking at <service>.railway.internal. Object storage remains
Cloudflare R2 bucket uploads-dev. Staging and production target GCP us-central1 (Iowa) —
decided, not provisioned; see Production landscape.
flowchart TB
subgraph Public["Public internet"]
Browser["Browser"]
Mobile["Mobile (native)"]
R2[("Cloudflare R2 (S3)")]
end
subgraph Railway["Railway: kottia.com / development"]
SvelteKit["apps/svelte-web (Bun + adapter-node)"]
API["apps/api (Bun + Hono)"]
Worker["apps/worker (Bun + BullMQ)"]
DB[("PostgreSQL + PostGIS 16")]
VK[("Valkey")]
MS[("Meilisearch v1.9")]
IP["imgproxy v3.30"]
end
Browser -->|HTTPS| SvelteKit
Browser -->|HTTPS| API
Browser -->|HTTPS, signed| IP
Mobile -->|HTTPS| API
SvelteKit --> DB
SvelteKit --> VK
SvelteKit --> MS
SvelteKit --> R2
API --> DB
API --> VK
Worker --> DB
Worker --> VK
Worker --> MS
R2 --> IP
| Service | Internal URL pattern | Public URL |
|---|
| PostgreSQL | postgres.railway.internal:5432 (Railway service reference) | — |
| Valkey | valkey.railway.internal:6379 (Railway service reference) | — |
| Meilisearch | http://meilisearch.railway.internal:7700 | — |
| Cloudflare R2 | https://<ACCOUNT_ID>.r2.cloudflarestorage.com (S3 API, off-box) | — |
| SvelteKit web | — | https://dev.useast.kottia.com |
| apps/api | — | https://api.dev.useast.kottia.com |
| imgproxy | — | https://assets.dev.useast.kottia.com → container port 8080 |
| docs | http://docs.railway.internal:8080 | — (docs.dev.useast.kottia.com reserved for DEV, Access-gated) |
| Bull Board | — | Internal by default; optional bull.dev.useast.kottia.com only behind Basic auth. DEV/STG only — PRD uses bun run queue:admin (see Incident guides) |
PostgreSQL, Valkey, and Meilisearch have Railway volumes mounted at
/var/lib/postgresql/data, /data, and /meili_data, respectively. Their services have no
public domains. The docs service has no public domain yet; docs.dev.useast.kottia.com is
reserved for DEV and is added only once Cloudflare Access/SSO is verified: an anonymous request must return 401/403 or a 30x whose
Location points to the configured Cloudflare Access login host; a 200 response containing the
docs body or any X-Railway-Edge header fails the access-boundary check.
| Service | Why it exists |
|---|
| PostgreSQL + PostGIS | System of record. 45 models, 85+ RLS policies. PostGIS for property ST_MakePoint coordinates. |
| Valkey | (a) Property search via @repo/redis-search GEOSEARCH; (b) BullMQ queue backend (db1); (c) rate-limit buckets; (d) area price funnel cache. |
| Meilisearch | Locations (SEPOMEX-sourced municipalities, cities, localities), agents, agent service areas. Fuzzy + filterable text search. |
| Cloudflare R2 | All uploaded user content: property images, floor plans, avatars. S3-compatible, zero egress fees. |
| imgproxy | On-demand resize + WebP conversion. HMAC-signed URLs — attackers can’t manipulate dimensions or request arbitrary resources. |
| apps/svelte-web | Public site + dashboards. SSR + form actions. Hosts its own Better Auth handler for the web flow. |
| apps/api | Standalone tRPC + Better Auth host for mobile (and eventually browser). Same DB + auth tables, interchangeable sessions. |
| apps/worker | BullMQ consumer. Search indexing, lead ingestion, alerts, AI jobs (Gemini, Replicate). |
| Queue | Concurrency | Notes |
|---|
search-indexing | default | sync-property (2s debounce, 3 retries) → Valkey; sync-agent, sync-team → Meilisearch |
leads | default | lead-ingestion (reliable, DLQ to failed_lead_ingestions); sla-monitor (repeatable */15 * * * *) |
alerts | default | price-drop-notify, status-change-notify, saved-search-scan (repeatable */30 * * * *) |
ai | 1 | limiter: { max: 6/min } to stay under Replicate’s per-account cap at $0 credit |
Every queue name is suffixed with process.env.QUEUE_SUFFIX. Railway DEV and local dev use
-dev; production will leave it empty. Producers and workers in a tier must use the same suffix.
Bull Board registration is suffix-aware too.
| Tool | Coverage |
|---|
| Sentry | Error tracking + session replay, one project per deployable: apps/svelte-web (server + client), apps/api, apps/worker, and the native mobile apps (ios-user / android-user / ios-agent / android-agent, org na-v1f). The non-mobile inits run a beforeSend scrubber (redacts cookies, authorization, CSRF, `password |
| PostHog | Consent-gated product analytics. |
| Bull Board | BullMQ dashboard — internal by default, or exposed through the optional Basic-auth-protected DEV domain. It grants queue control and exposes lead PII. DEV/STG only; PRD uses bun run queue:admin (Incident guides). |
| Knock dashboard | Inspect in-app feed deliveries + email channel state per workflow. |
| Scenario | Mitigation |
|---|
| Lost Postgres password | Update the Railway postgres service credential and its service references. Reset requires re-initializing the volume or ALTER USER from a deployed container. |
| Stale Valkey state | redis-cli -u "$VALKEY_URL" FLUSHDB then restart the worker so re-indexing fires. |
| Stale Meilisearch index | Restart the worker to restore locations / service_areas from R2 search-seeds/*.ndjson.gz; agents/teams re-index through the worker’s search-indexing queue. |
| imgproxy SSRF | Already mitigated: URLs are HMAC-signed server-side. IMGPROXY_KEY / IMGPROXY_SALT rotation invalidates all in-flight URLs. |
| Replicate traffic or spend spike | Set AI_ROOM_STAGING_ENABLED=false on web, API, and worker to stop new predictions; do not delete persisted jobs. Inspect the dedicated ai-staging queue and ai_provider_attempts, then adjust AI_STAGING_RATE_LIMIT_PER_MINUTE only against the current account limits. Existing prediction IDs remain resumable/cancelable. |
| DB schema drift | cd packages/database && bun run db:reset — drops the schema, then re-runs the Kysely migration chain to latest (0001 ID-generator functions → 0002 baseline schema → 0003 triggers, then the incremental migrations after it) so generated-ID columns and triggers come back intact. See Architecture → Database. |