Demo: three-node gossip cluster
Location: contrib/demo/cluster/
Runs three ahdapa instances on loopback ports 8080, 8081, and 8082 behind a shared self-signed TLS certificate. The script verifies that CRDT state converges across all three nodes, cross-node token issuance works, and a token issued on one node is introspected successfully on another.
What it shows
- CRDT convergence — a public OAuth2 client registered on node1 appears on
node2 and node3 within milliseconds (the CRDT write wakes the gossip loop
immediately via
Notify; the 2 s polling interval is only a fallback). - Tombstone propagation — deleting the client on node3 propagates back to node1 within the same convergence window.
- Any-node token issuance — a confidential OAuth2 client registered on
node1 can obtain a
client_credentialsaccess token from any of the three nodes once the CRDT has converged. - Cross-node token introspection — a token issued by node1 is accepted by
node2’s
/introspectendpoint; all nodes share the same signing keys via gossip. - Cross-node session revocation — enabled by
distributed_mode = "eventual"in each node config. Logouts written into therevoked_sessionsLwwMap propagate to all peers within one gossip round. - Scope definition replication — custom scope-to-claim mappings created on any node propagate to all peers.
- Sustained multi-client traffic — three confidential clients (created on different nodes) are exercised across all nodes in rotation for 35 s; per-node latency and gossip push statistics are printed at the end.
- SPNEGO key self-registration (Kerberos mode) — each node authenticates to its
two peers’
POST /api/gossip/register-kemendpoint over SPNEGO and self-registers its own ML-KEM-768 and gossip signing keys. - Stale-key self-heal (Kerberos mode) — a node that restarts with a wiped database regenerates its ML-KEM-768 key; on the next gossip round its peers’ envelopes fail to open, the node detects the stale stored key and re-registers its fresh keys with each peer over SPNEGO — no admin API call is involved.
Prerequisites
openssl(1)— for generating the demo TLS PKI (present on all modern Linux systems).python3— for JSON parsing in convergence checks.- Ports 8080, 8081, and 8082 free. The script evicts any stale processes bound to those ports before starting.
- The
ahdapabinary — the script looks in$PATHfirst (resolved to its full absolute path), thentarget/release/ahdapa, thentarget/debug/ahdapa, and falls back tocargo buildif none is found. - Kerberos tools (
krb5kdc,kdb5_util,kadmin.local,kinit) — optional. When present, the demo runs in Kerberos mode: it starts an ephemeral KDC (realmCLUSTER.LOCAL) with a single-principal keytab per node and bootstraps the gossip keys via SPNEGO self-registration, and additionally exercises the stale-key self-heal. Without them it falls back to the admin-seed bootstrap and skips the self-heal step; the demo still passes.
The node IDs are nodeN.localhost, which glibc resolves to loopback
automatically — no /etc/hosts entries are required.
Running
# Non-interactive: runs all steps, prints PASS/FAIL, exits.
contrib/demo/cluster/run.sh
# Interactive: same steps, then keeps nodes running until Ctrl-C.
contrib/demo/cluster/run.sh --interactive
What the script does
- Generates a P-256 CA root and a server certificate with
subjectAltName=IP:127.0.0.1,DNS:node1.localhost,DNS:node2.localhost, DNS:node3.localhost, shared by all three nodes. Valid for 1 day. - Starts an ephemeral KDC with per-node keytabs
(
HTTP/nodeN.localhost@CLUSTER.LOCAL) — Kerberos mode only. - Starts all three nodes with fresh SQLite databases under
/tmp/; in Kerberos mode each node runs with its own keytab and credential cache. - Logs in as
aliceon each node, then generates a random 32-byte cluster wrapping key and pushes it to all three viaPUT /api/admin/keys/clusterso that every node can decrypt the others’ gossiped secrets and converge on the same key material. Admin sessions are node-local, so each cross-node admin call authenticates on the target node. - Bootstraps the gossip KEM and signing keys: in Kerberos mode each node
self-registers with its two peers via
POST /api/gossip/register-kem(SPNEGO), in fallback mode it fetchesGET /api/gossip/kem-infoand seeds the other two viaPOST /api/admin/nodes/seed. - Creates a public OAuth2 client on node1 and polls node2 and node3 until it appears (convergence check).
- Deletes the client on node3 and confirms the deletion propagates to node1.
- Creates a confidential client on node1 (
client_secret_post), waits for convergence, then requests aclient_credentialstoken from each node and verifies HTTP 200 withaccess_tokenon each. - Introspects the node1 token via node2’s
/introspectendpoint and assertsactive: true, then exercises a DPoP-bound token request (RFC 9449). - Creates two more clients (one on node2, one on node3), waits for convergence, then runs 35 s of rotating token traffic across all three nodes and clients.
- Stale-key self-heal — Kerberos mode only: wipes node2’s database and
restarts the node so it regenerates its ML-KEM-768 key; the stale keys make
the peers reject node2’s gossip with 401 (and/or node2 fails to open
envelopes addressed to its old key,
NoMatchingRecipient), which triggers node2 to re-register its fresh keys with node3 over SPNEGO; node1 then receives node2’s fresh keys when node3’s updated CRDT state is relayed to it (a silent gossip merge), and a client created on node1 after the heal converges to both peers. - Prints per-node token latency, per-client success rates, and gossip push
statistics, then prints
PASSorFAIL.
Example output (abbreviated)
Generating TLS PKI (CA + server certificate)...
CA cert: .../tls/ca.pem
server cert: .../tls/cert.pem
Starting node1 (:8080)...
Starting node2 (:8081)...
Starting node3 (:8082)...
Waiting for all nodes to be ready...
node1 ready.
node2 ready.
node3 ready.
Synchronizing cluster wrapping key across all nodes...
generated cluster key: dGVzdGtleXRl... (truncated)
node1: cluster key set.
node2: cluster key set.
node3: cluster key set.
re-authenticated on node1 with shared key.
Bootstrap — SPNEGO self-registration of gossip KEM keys
node1 → https://node2.localhost:8081: registered (HTTP 200)
node1 → https://node3.localhost:8082: registered (HTTP 200)
node2 → https://node1.localhost:8080: registered (HTTP 200)
...
node3 → https://node2.localhost:8081: registered (HTTP 200)
Waiting 5 s for first gossip rounds...
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Step 1 — create an OAuth2 client on node1
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Created client: 3f8a1b2c-...
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Step 2 — wait for gossip to replicate the client to node2 and node3
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
node2: client appeared after 3s
node3: client appeared after 3s
...
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Statistics — 35s window, 11 batches × 3 clients
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Token endpoint (by node):
Node Requests Success Avg latency
──────── ──────── ─────── ───────────
node1 8 8/8 6ms
node2 8 8/8 5ms
node3 8 8/8 5ms
total 24 24/24
Gossip (outbound pushes, traffic window only):
Node Pushes Total bytes Avg/push Skips
──────── ────── ─────────── ──────── ─────
node1 3 22500 7500 8
node2 2 15000 7500 9
node3 2 15000 7500 9
total 7 52500 — 26
Generation-skip rate: 26/33 push attempts skipped (78%)
cluster converged — majority of gossip rounds were no-ops
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
PASS — all cluster demo steps succeeded.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Interactive exploration
With --interactive, the script keeps all three nodes running after the
verification steps. The browser will warn about the self-signed certificate;
add contrib/demo/cluster/tls/cert.pem to your browser’s trust store or
click through the warning for local testing.
node1 https://node1.localhost:8080/ui/
node2 https://node2.localhost:8081/ui/
node3 https://node3.localhost:8082/ui/
Log in as alice (password printed at startup) on any node. The admin panel reflects the same
client list on all three nodes; changes made on any node appear on the others
within milliseconds (gossip wakes immediately on CRDT writes).
Configuration notes
| File | Description |
|---|---|
node1.toml | Port 8080, gossip peers: 8081 and 8082, distributed_mode = "eventual" |
node2.toml | Port 8081, gossip peers: 8080 and 8082, distributed_mode = "eventual" |
node3.toml | Port 8082, gossip peers: 8080 and 8081, distributed_mode = "eventual" |
users.toml.in | Users template shared by all nodes; password substituted at runtime |
Key settings in each node config:
[gssapi]
service = "HTTP"
keytab = "__KEYTAB__" # substituted at runtime: per-node keytab in Kerberos
# mode, /nonexistent/keytab in fallback mode
# IPA GSSAPI must be enabled so the node acquires its GSSAPI initiator
# credential (used for SPNEGO self-registration); the empty uri keeps
# LDAP/IPA lookups disabled in the demo.
[ipa]
gssapi = true
uri = ""
[gossip]
peers = ["https://node2.localhost:8081", "https://node3.localhost:8082"]
interval_secs = 2
allowed_node_ids = ["node1.localhost", "node2.localhost", "node3.localhost"]
[cluster]
distributed_mode = "eventual"
In Kerberos mode each node authenticates with its own single-principal keytab
(HTTP/nodeN.localhost@CLUSTER.LOCAL from the ephemeral KDC) so its GSSAPI
initiator can negotiate against peers; alice still authenticates with a
password from the static users file for the admin API. The [ipa] table with
gssapi = true is required for the initiator credential acquisition (an
absent [ipa] table defaults gssapi to false); its empty uri keeps
LDAP/IPA lookups disabled. In fallback mode the
keytab path is /nonexistent/keytab, so GSSAPI stays disabled. The [tls]
section is not present in the static config files; the script appends it at
runtime using the paths of the generated certificate and key.
CI usage
The script is also run as a CI integration test in the cluster-demo job of
.github/workflows/ci.yml. It exits 0 on pass and 1 on any assertion failure.
See also
- Multi-node Cluster — full cluster configuration reference.
- Gossip Protocol — protocol internals.