Operations Guide: Setup, Deployment, Runtime Internals & Troubleshooting

This is the guide for actually running orvo-abha — as a developer setting it up locally, or as whoever is responsible for it in staging/production. It complements the other guides rather than repeating them: Architecture Overview explains what the system is and why it's shaped this way; the API Reference documents endpoints; this document covers running the thing — local setup, environment variables, build/deploy, the runtime patterns that make the async ABDM protocol tractable in practice, secret rotation, and where to look when something goes wrong.

Everything below is verified against the actual repo state (Dockerfile, docker-compose.yml, bitbucket-pipelines.yml, entrypoint.sh, package.json, .env.example) as of September 2026, including a couple of places where the code doesn't match what an older doc (the root README.md) claims — those are called out explicitly rather than silently repeated.


Prerequisites

Local development setup

npm install
npx prisma generate
npx prisma migrate deploy   # or `npx prisma migrate dev` if you're actively changing schema.prisma
npm run dev                 # nodemon, restarts on file change

npm run dev alone assumes DATABASE_URL in your .env already points at a reachable Postgres. If you're developing against the real staging database rather than a local Postgres, npm run tunnel opens an SSH tunnel to the RDS instance through the abha.orvo.app bastion host (the exact command, including the key path, is in package.json — it's a local dev convenience, not something to hardcode elsewhere), and npm run dev:local runs the tunnel and the dev server together via concurrently.

Once running, / serves a small live status page (environment, uptime, DB connectivity, version — see the Design & operational discrepancies section below for why this is not a real health check), /docs serves the Swagger/OpenAPI UI, and /docs/guide serves this documentation set itself (rendered by docs-site.routes.ts from the docs/guide/ markdown files — the same files you're reading now, browsable in a live deployment).

Environment variables

.env.example at the repo root is the canonical, actively-maintained variable reference — every variable is commented in place with what it's for, sandbox-vs-production differences, and known gotchas (e.g. why sslmode=disable must never reach the production DATABASE_URL, or why ABDM_PHR_ABHA_URL includes a path prefix while ABDM_BASE_URL doesn't). Copy it to .env and fill in real values rather than treating any table in a guide as the source of truth — a duplicated table drifts, .env.example doesn't.

A few things worth knowing that aren't obvious just from reading the file top to bottom:

Testing

npm test                # vitest run — the full suite
npm run test:watch      # interactive
npm run test:coverage

Per-module scripts exist for the areas that have historically needed isolated, fast iteration: test:middleware, test:log, and the test:abdm:* family (link-token, hip-linking, gateway, phr-login, phr-profile) — npm run test:abdm runs all five together. Test files sit next to the code they test, in __tests__/ directories per module (e.g. src/modules/abdm/shared/__tests__/consent-scope.subscription-event.test.ts) — there's no separate top-level test tree to keep in sync.

There is no sandbox integration-test suite that actually calls ABDM's servers — the tests mock the gateway client. Validating a change against real ABDM behavior means running it against the sandbox manually (or via one of the dry-run-*/configure-* scripts in package.json) and watching GatewayTransaction/AbdmCallbackEvent rows, or the realtime events, as described in Troubleshooting below.

Build & deployment

Build (npm run build): prisma generate → wipe dist/ and .cache/typescript → tsc → tsc-alias (rewrites @/... path aliases to relative imports, since Node can't resolve tsconfig.json path mappings at runtime) → copies the PDF-generation font assets into dist/ by hand, because tsc only emits .ts output and doesn't copy static files.

Docker image (Dockerfile, multi-stage):

  1. Builder (node:22-alpine): npm ci, then npm run build with NODE_OPTIONS=--max-old-space-size=2048 set — memory-constrained EC2 build hosts otherwise OOM partway through tsc. A dummy DATABASE_URL is set at build time only so prisma generate can run without a real database being reachable during the image build.
  2. Runtime (node:22-alpine): copies node_modules, dist, prisma, docs (so /docs/guide has something to render), package*.json, and global-bundle.pem (Amazon RDS's root CA bundle — required because RDS's TLS cert chain isn't in Node's default trust store, and the pg pool enables TLS verification whenever NODE_ENV=production). Runs as the image's non-root default; installs dumb-init and curl for signal handling and the container healthcheck.

Local Compose run: docker compose up -d --build. docker-compose.yml defines two services: abha-api (always), and db — a Postgres 16 container gated behind COMPOSE_PROFILES=staging, used only for the staging deployment (production points at RDS directly and must leave COMPOSE_PROFILES empty, or a second, unused database starts alongside RDS for nothing).

CI/CD (bitbucket-pipelines.yml): on a push to main (production) or the staging branch, the pipeline SSHes into the target EC2 host, git reset --hards to the pushed commit, regenerates .env from Bitbucket's own pipeline/deployment variables (every variable in .env.example has a corresponding echo "VAR=${VAR}" >> .env line here — if you add a new env var to the app, it needs a matching line here too, or it will silently be empty in every deployed environment), then runs docker compose up -d --build.

Design & operational discrepancies worth knowing about

Two things the code actually does differently from what an older doc (README.md) describes — worth knowing before you rely on either:

Runtime internals worth understanding before debugging a production issue

These are patterns that live in src/modules/abdm/core/ and apply across every ABDM-facing module — knowing them is usually the difference between diagnosing an ABDM issue in minutes versus re-deriving it from scratch each time.

Dual session management. Two independent token caches exist because they serve genuinely different credential shapes:

Per-facility gateway client. Every facility gets its own FacilityGatewayClient (via GatewayClientRegistry, created lazily and cached in memory — evict(hipId)/evictAll() force a reload from BridgeConfig). It automatically injects the ABDM-required headers (Authorization, X-HIP-ID, X-CM-ID, REQUEST-ID, TIMESTAMP), retries a 401 exactly once after invalidating the cached bridge session, retries transient 5xx/429 with 1s/2s backoff (never retries a 4xx — those are calling-side errors, retrying them just repeats the same mistake), and logs every outbound call to GatewayTransaction. If a facility's calls are behaving strangely, GatewayTransaction rows for that hipId are the first place to look.

ABDM error mapping. abdm-errors.ts normalizes ABDM's error phrases/codes into HTTP status codes, user-facing messages, and a retryable flag — see the ABDM Error Matrix and Error Coverage Checklist for the full phrase-by-phrase mapping. Only two specific codes (900902, ABDM-1023) trigger an automatic session-token refresh-and-retry; every other 401 passes straight through, on the theory that most 401s are a genuine auth-config problem a blind retry won't fix.

Idempotency on inbound callbacks. Because ABDM's gateway redelivers callbacks (a documented, observed behavior, not a hypothetical), every inbound handler that mutates state checks for an existing row first: patient-share callbacks key on requestId, health-information requests key on transactionId, consent notifications upsert on consentId. The subscription LINK-event handler goes further with an explicit dedup check (hasExistingLinkConsent) keyed on the exact set of care-context references, added after a production incident where unbounded redelivery of one notification produced 1,297 duplicate Consent rows for a single patient (cleaned up down to 326 genuine ones) — see the callback-notify sections of the Consent & Health Data doc for the current state of that logic.

Async persistence pattern. After enrollment/login, the patient's profile and linking tokens are written to the database in the background (fire-and-forget) rather than being awaited before the response — database latency never adds to the latency the patient's app sees. Failures are logged, not surfaced to the caller. This is a deliberate availability-over-consistency tradeoff for this specific path; it is not the pattern used for the subscription LINK/DATA handlers described above, which now await their consent/health-info side effects specifically so a failure is visible on the durable callback row instead of silently lost (see the Consent & Health Data doc's subscription-notify section for why that changed).

BridgeConfig secret key rotation

BridgeConfig.clientSecret is encrypted at rest with AES-256-GCM, in the format v<N>:<base64>, where N is a key version from SECRET_ENCRYPTION_KEYS (1:<hex>,2:<hex>,...). The first key in the list encrypts every new write; every key in the list can decrypt existing rows — that's what makes rotation zero-downtime.

Planned rotation:

  1. Generate a new key: node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
  2. Add it as a new version, keeping the old one: SECRET_ENCRYPTION_KEYS=1:<old-hex>,2:<new-hex> — deploy this first. At this point the app can decrypt both versions but hasn't re-encrypted anything yet.
  3. Run the rotation script: npx ts-node scripts/rotate-secrets.ts (add --dry-run to preview without writing). It reports [ROTATED]/[SKIP]/[ERROR] per BridgeConfig row.
  4. Do not proceed if any row reports [ERROR] — the old key must stay in SECRET_ENCRYPTION_KEYS until every row is confirmed re-encrypted.
  5. Once all rows are on the new version, remove the old key and redeploy: SECRET_ENCRYPTION_KEYS=2:<new-hex>.

Emergency rotation (suspected key compromise), in order: (1) revoke the affected ABDM credentials in the ABDM portal immediately — this is the step that actually cuts off an attacker, independent of anything below; (2) rotate DATABASE_URL credentials too, since the encryption key alone is useless to an attacker without database access; (3) run the planned-rotation steps above, keeping the compromised key in the keyset only as long as it takes to re-encrypt every row; (4) issue and store fresh ABDM credentials under the new key once rotation completes.

To check current key-version state at any time without changing anything: npx ts-node scripts/rotate-secrets.ts --dry-run.

Troubleshooting & where to look

A call to ABDM returns 202 and then nothing ever happens (no on-init/callback arrives): this is the single most common ABDM integration failure shape, and it's almost always one of: a trailing space in a hipId/hiuId (ABDM trims nothing and silently drops the request — see the HIP Identity doc's facility-settings section), a "blind" request pinning a HIP/care-context ABDM doesn't yet know about (see the consent-create gotcha in the Consent & Health Data doc), a subscription request missing hips (empirically required despite ABDM's own docs marking it optional), or sandbox/production credential mismatch (see Environment Variables above). Check GatewayTransaction for the outbound call and AbdmCallbackEvent for whether anything ever came back on that requestId — a request with a GatewayTransaction row but no matching AbdmCallbackEvent row confirms the drop happened on ABDM's side, not in this codebase's handling of a callback that arrived.

An inbound callback is rejected with a 422 before any handler code runs: check the relevant *.schema.ts Zod schema — ABDM's actual sandbox payloads sometimes diverge from its own published spec (an omitted "required" field, an unexpected extra one, a category value not yet in an enum — see the recent category: z.string().min(1) relaxation in the subscription-notify schema for exactly this class of fix). The fix is almost always to loosen the schema to accept what ABDM actually sends and let the service layer decide what's a real processing error, rather than rejecting at the validation boundary where the payload can't even be logged.

A specific ABDM error phrase/code and what to do about it: the ABDM Error Matrix and Error Coverage Checklist list every known phrase pulled from ABDM's Swagger/Postman artifacts, whether this codebase's normalizer already maps it, and whether it's retryable.

Diagnosing in production: prod Postgres is only reachable through the SSH tunnel via the abha.orvo.app bastion (npm run tunnel); the Log table carries full ABDM request/response payloads for a given transaction, and is usually the fastest way to see exactly what was sent and what came back, byte for byte, without needing to reproduce anything.

Realtime events not arriving on the frontend: confirm PUSHER_* env vars are set (they're optional — an empty set disables realtime entirely, silently, which looks identical to "events aren't firing" from the frontend's perspective) and use GET /v1/realtime/events?channel=...&since=... (durable replay, ~1 hour retention) to check whether the event was actually published, independent of whether the frontend's Pusher subscription received it.

Feature flags

GET /v1/feature-flags?app=orvo-hub / the /v1/admin/feature-flags pair let orvo-phr/orvo-web/orvo-hub toggle features per app per ABDM environment without a redeploy — see the Feature Flags section of the Platform API doc for the full request/response shapes. Operationally relevant here: there's no local override for the environment a flag applies to — it's always resolved server-side from ABDM_ENV, so testing a flag's staging behavior means toggling it against the deployment that's actually running in that ABDM_ENV, not passing a query parameter.