Operations Guide: Setup, Deployment, Runtime Internals & Troubleshooting
This is the guide for actually running orvo-abha — as a developer setting it up locally, or as whoever is responsible for it in staging/production. It complements the other guides rather than repeating them: Architecture Overview explains what the system is and why it's shaped this way; the API Reference documents endpoints; this document covers running the thing — local setup, environment variables, build/deploy, the runtime patterns that make the async ABDM protocol tractable in practice, secret rotation, and where to look when something goes wrong.
Everything below is verified against the actual repo state (Dockerfile, docker-compose.yml, bitbucket-pipelines.yml, entrypoint.sh, package.json, .env.example) as of September 2026, including a couple of places where the code doesn't match what an older doc (the root README.md) claims — those are called out explicitly rather than silently repeated.
Prerequisites
- Node.js 20+ (the Docker image builds on Node 22)
- PostgreSQL 14+ (production runs 14 on RDS; the staging Compose profile runs Postgres 16)
- ABDM sandbox or production credentials (from the ABDM Integrator Portal — see the Environment Variables section below for which client id/secret pairs you need)
- Docker & Docker Compose, if you want to run it containerized rather than with
npm run dev
Local development setup
npm install
npx prisma generate
npx prisma migrate deploy # or `npx prisma migrate dev` if you're actively changing schema.prisma
npm run dev # nodemon, restarts on file change
npm run dev alone assumes DATABASE_URL in your .env already points at a reachable Postgres. If you're developing against the real staging database rather than a local Postgres, npm run tunnel opens an SSH tunnel to the RDS instance through the abha.orvo.app bastion host (the exact command, including the key path, is in package.json — it's a local dev convenience, not something to hardcode elsewhere), and npm run dev:local runs the tunnel and the dev server together via concurrently.
Once running, / serves a small live status page (environment, uptime, DB connectivity, version — see the Design & operational discrepancies section below for why this is not a real health check), /docs serves the Swagger/OpenAPI UI, and /docs/guide serves this documentation set itself (rendered by docs-site.routes.ts from the docs/guide/ markdown files — the same files you're reading now, browsable in a live deployment).
Environment variables
.env.example at the repo root is the canonical, actively-maintained variable reference — every variable is commented in place with what it's for, sandbox-vs-production differences, and known gotchas (e.g. why sslmode=disable must never reach the production DATABASE_URL, or why ABDM_PHR_ABHA_URL includes a path prefix while ABDM_BASE_URL doesn't). Copy it to .env and fill in real values rather than treating any table in a guide as the source of truth — a duplicated table drifts, .env.example doesn't.
A few things worth knowing that aren't obvious just from reading the file top to bottom:
- Credentials are split by ABDM registration group, not by module.
ABDM_PHR_CLIENT_ID/SECRETis for the PHR-app registration only;ABDM_M234_CLIENT_ID/SECRETis shared across everything else (identity/ABHA, consent, subscription, HPR). Using the wrong pair for a call produces an ABDM auth failure that looks like a bad credential when it's actually the right credential for the wrong API family. - One deployment = one ABDM environment.
ABDM_ENV(sbxorabdm) has to match every base-URL variable and every credential pair you set — there's no per-request override. Mixing a sandboxhipIdwith production credentials (or vice versa) is a documented source of a call that returns202 Acceptedand then silently never produces a callback, which looks identical to a real integration bug (see Integration Contract §5). SECRET_ENCRYPTION_KEYSis required, not optional, despite not being marked required in older docs —BridgeConfig.clientSecretrows can't be written or read without it. See Secret rotation below for the format and how to change it safely.ABDM_CALLBACK_IP_CHECKshould befalsein sandbox/dev (ABDM's sandbox gateway IPs aren't worth chasing) andtruein production withABDM_CALLBACK_ALLOWED_IPSpopulated from the portal — this is the only thing standing between the inbound callback routes and the open internet, since those routes have no login by design (ABDM can't present a bearer token — see the Architecture doc's security model).ORVO_CORE_SERVICE_SECRETgatesPOST /v1/orvo/consultationsinbound, and is also sent outbound asX-Abha-Secreton the Rx-template fetch toapi.orvo.app— one secret, two different header names depending on direction. LeavingORVO_CORE_BASE_URLempty disables the outbound calls entirely rather than erroring.
Testing
npm test # vitest run — the full suite
npm run test:watch # interactive
npm run test:coverage
Per-module scripts exist for the areas that have historically needed isolated, fast iteration: test:middleware, test:log, and the test:abdm:* family (link-token, hip-linking, gateway, phr-login, phr-profile) — npm run test:abdm runs all five together. Test files sit next to the code they test, in __tests__/ directories per module (e.g. src/modules/abdm/shared/__tests__/consent-scope.subscription-event.test.ts) — there's no separate top-level test tree to keep in sync.
There is no sandbox integration-test suite that actually calls ABDM's servers — the tests mock the gateway client. Validating a change against real ABDM behavior means running it against the sandbox manually (or via one of the dry-run-*/configure-* scripts in package.json) and watching GatewayTransaction/AbdmCallbackEvent rows, or the realtime events, as described in Troubleshooting below.
Build & deployment
Build (npm run build): prisma generate → wipe dist/ and .cache/typescript → tsc → tsc-alias (rewrites @/... path aliases to relative imports, since Node can't resolve tsconfig.json path mappings at runtime) → copies the PDF-generation font assets into dist/ by hand, because tsc only emits .ts output and doesn't copy static files.
Docker image (Dockerfile, multi-stage):
- Builder (
node:22-alpine):npm ci, thennpm run buildwithNODE_OPTIONS=--max-old-space-size=2048set — memory-constrained EC2 build hosts otherwise OOM partway throughtsc. A dummyDATABASE_URLis set at build time only soprisma generatecan run without a real database being reachable during the image build. - Runtime (
node:22-alpine): copiesnode_modules,dist,prisma,docs(so/docs/guidehas something to render),package*.json, andglobal-bundle.pem(Amazon RDS's root CA bundle — required because RDS's TLS cert chain isn't in Node's default trust store, and thepgpool enables TLS verification wheneverNODE_ENV=production). Runs as the image's non-root default; installsdumb-initandcurlfor signal handling and the container healthcheck.
Local Compose run: docker compose up -d --build. docker-compose.yml defines two services: abha-api (always), and db — a Postgres 16 container gated behind COMPOSE_PROFILES=staging, used only for the staging deployment (production points at RDS directly and must leave COMPOSE_PROFILES empty, or a second, unused database starts alongside RDS for nothing).
CI/CD (bitbucket-pipelines.yml): on a push to main (production) or the staging branch, the pipeline SSHes into the target EC2 host, git reset --hards to the pushed commit, regenerates .env from Bitbucket's own pipeline/deployment variables (every variable in .env.example has a corresponding echo "VAR=${VAR}" >> .env line here — if you add a new env var to the app, it needs a matching line here too, or it will silently be empty in every deployed environment), then runs docker compose up -d --build.
Design & operational discrepancies worth knowing about
Two things the code actually does differently from what an older doc (README.md) describes — worth knowing before you rely on either:
- There is no
GET /healthorGET /health/readyendpoint. Both are documented inREADME.md's Getting Started section, but neither route exists insrc/app.ts. What exists instead isGET /(status-page.routes.ts) — a small HTML status page (environment, uptime, DB round-trip latency, version) intended for a human hitting the bare URL. It always returns200, even when the DB check inside it fails, by deliberate design (so a transient DB blip doesn't trip container-restart logic) — which means it cannot function as a true liveness/readiness probe. The DockerHEALTHCHECKand the (currently commented-out) health-check line inbitbucket-pipelines.ymlboth point at this same always-200 route, so today neither one can actually detect "the app is up but the database is unreachable." If you need a real readiness probe, it would need to be a new route that reflects DB status in its HTTP status code, not just its body. entrypoint.sh's retry-loop migration logic is not wired into the image. The script (bounded retries, 5s apart, warns and starts the server anyway after 12 failed attempts rather than hanging forever) is a genuinely well-designed pattern, anddocker-compose.ymlhas a comment attributing startup-ordering behavior to it — but theDockerfileneverCOPYs it in, and itsCMDrunsprisma migrate deploy && node dist/server.jsdirectly, once, with no retry. In practice this means a transient database-unavailability at container start (an RDS failover mid-boot, a stagingdbcontainer that's a few seconds slow to accept connections) currently fails the container outright rather than retrying — the resilience the script describes doesn't apply until it's actually referenced from theDockerfile(e.g.COPY entrypoint.sh ./+ENTRYPOINT ["./entrypoint.sh"]ahead of the existingCMD).
Runtime internals worth understanding before debugging a production issue
These are patterns that live in src/modules/abdm/core/ and apply across every ABDM-facing module — knowing them is usually the difference between diagnosing an ABDM issue in minutes versus re-deriving it from scratch each time.
Dual session management. Two independent token caches exist because they serve genuinely different credential shapes:
AbhaSessionManager— a single in-memory + DB-persisted (AbhaSessiontable) token for the M1/identity family, one set of env credentials for the whole deployment. On cold start, the last-cached token is restored from the DB and reused if still valid, so a rolling deploy doesn't force a fresh token fetch for every request that lands right after restart.BridgeSessionManager— per-facility tokens backed byBridgeConfig, since each facility can have its own bridge/client credentials. Concurrent refresh requests for the samebridgeId:environmentare deduplicated behind aMap<string, Promise>guard, so a burst of requests arriving the instant a token expires triggers exactly one refresh call to ABDM, not one per request.- Both include a 10-second failure cooldown — a token-fetch failure doesn't get retried on every subsequent request for 10 seconds, so an ABDM outage degrades gracefully instead of hammering the gateway.
Per-facility gateway client. Every facility gets its own FacilityGatewayClient (via GatewayClientRegistry, created lazily and cached in memory — evict(hipId)/evictAll() force a reload from BridgeConfig). It automatically injects the ABDM-required headers (Authorization, X-HIP-ID, X-CM-ID, REQUEST-ID, TIMESTAMP), retries a 401 exactly once after invalidating the cached bridge session, retries transient 5xx/429 with 1s/2s backoff (never retries a 4xx — those are calling-side errors, retrying them just repeats the same mistake), and logs every outbound call to GatewayTransaction. If a facility's calls are behaving strangely, GatewayTransaction rows for that hipId are the first place to look.
ABDM error mapping. abdm-errors.ts normalizes ABDM's error phrases/codes into HTTP status codes, user-facing messages, and a retryable flag — see the ABDM Error Matrix and Error Coverage Checklist for the full phrase-by-phrase mapping. Only two specific codes (900902, ABDM-1023) trigger an automatic session-token refresh-and-retry; every other 401 passes straight through, on the theory that most 401s are a genuine auth-config problem a blind retry won't fix.
Idempotency on inbound callbacks. Because ABDM's gateway redelivers callbacks (a documented, observed behavior, not a hypothetical), every inbound handler that mutates state checks for an existing row first: patient-share callbacks key on requestId, health-information requests key on transactionId, consent notifications upsert on consentId. The subscription LINK-event handler goes further with an explicit dedup check (hasExistingLinkConsent) keyed on the exact set of care-context references, added after a production incident where unbounded redelivery of one notification produced 1,297 duplicate Consent rows for a single patient (cleaned up down to 326 genuine ones) — see the callback-notify sections of the Consent & Health Data doc for the current state of that logic.
Async persistence pattern. After enrollment/login, the patient's profile and linking tokens are written to the database in the background (fire-and-forget) rather than being awaited before the response — database latency never adds to the latency the patient's app sees. Failures are logged, not surfaced to the caller. This is a deliberate availability-over-consistency tradeoff for this specific path; it is not the pattern used for the subscription LINK/DATA handlers described above, which now await their consent/health-info side effects specifically so a failure is visible on the durable callback row instead of silently lost (see the Consent & Health Data doc's subscription-notify section for why that changed).
BridgeConfig secret key rotation
BridgeConfig.clientSecret is encrypted at rest with AES-256-GCM, in the format v<N>:<base64>, where N is a key version from SECRET_ENCRYPTION_KEYS (1:<hex>,2:<hex>,...). The first key in the list encrypts every new write; every key in the list can decrypt existing rows — that's what makes rotation zero-downtime.
Planned rotation:
- Generate a new key:
node -e "console.log(require('crypto').randomBytes(32).toString('hex'))" - Add it as a new version, keeping the old one:
SECRET_ENCRYPTION_KEYS=1:<old-hex>,2:<new-hex>— deploy this first. At this point the app can decrypt both versions but hasn't re-encrypted anything yet. - Run the rotation script:
npx ts-node scripts/rotate-secrets.ts(add--dry-runto preview without writing). It reports[ROTATED]/[SKIP]/[ERROR]perBridgeConfigrow. - Do not proceed if any row reports
[ERROR]— the old key must stay inSECRET_ENCRYPTION_KEYSuntil every row is confirmed re-encrypted. - Once all rows are on the new version, remove the old key and redeploy:
SECRET_ENCRYPTION_KEYS=2:<new-hex>.
Emergency rotation (suspected key compromise), in order: (1) revoke the affected ABDM credentials in the ABDM portal immediately — this is the step that actually cuts off an attacker, independent of anything below; (2) rotate DATABASE_URL credentials too, since the encryption key alone is useless to an attacker without database access; (3) run the planned-rotation steps above, keeping the compromised key in the keyset only as long as it takes to re-encrypt every row; (4) issue and store fresh ABDM credentials under the new key once rotation completes.
To check current key-version state at any time without changing anything: npx ts-node scripts/rotate-secrets.ts --dry-run.
Troubleshooting & where to look
A call to ABDM returns 202 and then nothing ever happens (no on-init/callback arrives): this is the single most common ABDM integration failure shape, and it's almost always one of: a trailing space in a hipId/hiuId (ABDM trims nothing and silently drops the request — see the HIP Identity doc's facility-settings section), a "blind" request pinning a HIP/care-context ABDM doesn't yet know about (see the consent-create gotcha in the Consent & Health Data doc), a subscription request missing hips (empirically required despite ABDM's own docs marking it optional), or sandbox/production credential mismatch (see Environment Variables above). Check GatewayTransaction for the outbound call and AbdmCallbackEvent for whether anything ever came back on that requestId — a request with a GatewayTransaction row but no matching AbdmCallbackEvent row confirms the drop happened on ABDM's side, not in this codebase's handling of a callback that arrived.
An inbound callback is rejected with a 422 before any handler code runs: check the relevant *.schema.ts Zod schema — ABDM's actual sandbox payloads sometimes diverge from its own published spec (an omitted "required" field, an unexpected extra one, a category value not yet in an enum — see the recent category: z.string().min(1) relaxation in the subscription-notify schema for exactly this class of fix). The fix is almost always to loosen the schema to accept what ABDM actually sends and let the service layer decide what's a real processing error, rather than rejecting at the validation boundary where the payload can't even be logged.
A specific ABDM error phrase/code and what to do about it: the ABDM Error Matrix and Error Coverage Checklist list every known phrase pulled from ABDM's Swagger/Postman artifacts, whether this codebase's normalizer already maps it, and whether it's retryable.
Diagnosing in production: prod Postgres is only reachable through the SSH tunnel via the abha.orvo.app bastion (npm run tunnel); the Log table carries full ABDM request/response payloads for a given transaction, and is usually the fastest way to see exactly what was sent and what came back, byte for byte, without needing to reproduce anything.
Realtime events not arriving on the frontend: confirm PUSHER_* env vars are set (they're optional — an empty set disables realtime entirely, silently, which looks identical to "events aren't firing" from the frontend's perspective) and use GET /v1/realtime/events?channel=...&since=... (durable replay, ~1 hour retention) to check whether the event was actually published, independent of whether the frontend's Pusher subscription received it.
Feature flags
GET /v1/feature-flags?app=orvo-hub / the /v1/admin/feature-flags pair let orvo-phr/orvo-web/orvo-hub toggle features per app per ABDM environment without a redeploy — see the Feature Flags section of the Platform API doc for the full request/response shapes. Operationally relevant here: there's no local override for the environment a flag applies to — it's always resolved server-side from ABDM_ENV, so testing a flag's staging behavior means toggling it against the deployment that's actually running in that ABDM_ENV, not passing a query parameter.