High availability (v0.60)¶
From v0.60 a Network Authority (NA) can run as several instances on one
PostgreSQL database (see High Availability). Usually a load balancer sits in front; the SDK then
needs only baseUrl pointing at it. Without a load balancer, give the SDK
every instance and it fails over itself.
import { GenesisMeshClient, isRetryableConflict } from 'genesis-mesh-sdk';
const gm = new GenesisMeshClient({
baseUrls: ['https://na-a.example', 'https://na-b.example'],
signer: operatorSigner,
retry: { attempts: 2, baseDelayMs: 200 },
});
Failover rules¶
Requests go to one instance at a time.
Idempotent requests (every GET, public verification and evidence submission) move to the next instance on a transport failure or 502/503/504, then use
retrywith backoff.Non-idempotent requests (evaluate, issue, publish, revoke) move to the next instance only when the connection could not be made (refused, unresolvable, unreachable), because the request never reached the NA. After any other failure they are not replayed: a request that may have run must not run twice. The error reaches the caller, and the next request starts at another instance.
The client stays on the instance that last answered.
NetworkError.connectFailed says whether a failure happened before the
request was sent.
Readiness¶
const r = await gm.health.readiness();
// { ready, status, instance, ha_mode, rate_limiter,
// database: { backend, writable, schema_version, expected_schema_version, error? },
// signing_key: { key_id, provider, fingerprint } }
for (const e of await gm.health.endpoints()) {
console.log(e.base_url, e.reachable, e.ready, e.readiness?.signing_key.fingerprint);
}
readiness() returns a not-ready NA as ready: false with the failing checks
instead of throwing. endpoints() probes every configured instance directly.
It is meant for monitoring and for spotting an instance with a different key:
every instance of one NA reports the same fingerprint. liveness() and
health() wrap /healthz and /health.
Conflicts between instances¶
Some operations are decided by the database across instances. The request
that loses a race gets 409 and changed nothing, so it can be retried:
Code |
Meaning |
|---|---|
|
another version of the policy was activated at the same moment |
|
concurrent publishes of one policy kept taking the next version |
|
concurrent revocations kept taking the next CRL sequence |
|
evidence retention is already running on another instance |
try {
await gm.policy.activate(policyId, version);
} catch (err) {
if (isRetryableConflict(err)) await gm.policy.activate(policyId, version);
else throw err;
}
Evidence submission is already safe to retry: an identical record returns
status: 'duplicate'.
Testing against a cluster¶
In the sdk-typescript repository, npm run test:e2e also runs tests/e2e-ha.test.ts when
GM_E2E_HA_DATABASE_URL names an empty PostgreSQL database (C collation)
and nginx is installed. It starts two NA instances behind nginx with the
core’s cluster harness. It kills one mid-run, then checks that a secret’s
lifecycle continues through the load balancer and through baseUrls
failover, with one unbroken, verified history.