Gossip Protocol
The gossip protocol synchronises IdpCrdt state across cluster nodes. The wire protocol,
CMS envelope construction (crates/ahdapa-cms/), and CRDT semantics are unchanged from
ahdapa’s original design, but the engine that drives the push-pull loop is not: it’s now
agirru_engine::GossipEngine, a generic
gossip engine extracted from this same code. Ahdapa supplies the application-specific
pieces under src/routes/gossip/:
| File | Implements | Responsibility |
|---|---|---|
transport.rs | agirru_engine::Transport<String> | Resolve a peer node_id to a URL + KEM key, sign_and_seal, POST, verify_and_open the response. Classifies HTTP 401 as SendOutcome::AuthRejected (drives on_auth_rejected below). |
membership.rs | agirru_engine::Membership<String> | The node_id of every configured/discovered peer URL that resolves to a live cluster_nodes entry with a KEM key. |
storage.rs | agirru_engine::Storage<IdpCrdt> | Wraps IdpCrdt::load_from_db / AppState::persist_crdt — see the note on bootstrap_with_state under Background loop for why load() is not actually called at startup. |
hooks.rs | agirru_engine::GossipHooks<String, IdpCrdt> | Admission filtering, ABAC policy rebuild, signing-key cache eviction, wrapping-key pull, 401 re-registration. See Admission filters. |
codec.rs | agirru_engine::Codec<Envelope<IdpCrdt>> | CBOR encode/decode of the engine’s own Envelope<IdpCrdt> type, folding in stale-envelope rejection. |
runtime.rs | agirru_engine::GossipRuntime<IdpCrdt, String> | Bundles the five pieces above into AhdapaGossipRuntime, and defines the AhdapaGossipEngine type alias. |
maintenance.rs | — | Periodic tombstone/expired-entry cleanup, decoupled from gossip rounds (see Background loop). |
mod.rs | — | Axum handlers (gossip_sync, gossip_kem_info, gossip_wrapping_key, gossip_register_kem, gossip_node_stats, gossip_await_client) and run(), which bootstraps and drives the engine. |
Design
The protocol uses delta-based exchange by default, falling back to full-state on first
contact or after an error. After a successful round, a node sends only the CRDT entries
that changed since the last successful exchange with each peer (a sparse IdpCrdt delta),
rather than the full state every time. Peers exchange generation counters in each envelope
to coordinate what they already have. One round-trip still brings both nodes to the same
merged state.
This delta/full-state decision, and the per-peer generation bookkeeping behind it, is now
entirely internal to agirru_engine::GossipEngine’s PeerBook — ahdapa no longer owns
peer_last_gen/peer_response_gen maps directly (see CRDT_GENERATION counter).
Full-state pushes occur in the following cases:
- First contact with a peer (no prior recorded generation for it).
- After any connection error or non-2xx response (the engine clears its bookkeeping for that peer).
- When a
pull_wrapping_keyfailure occurs (AhdapaGossipHooks::on_mergedincrementswrapping_key_pull_errors, but does not itself force a full-state resync — the next round’s delta computation is unaffected by this failure since it’s keyed off successful sync generations, not wrapping-key state).
Full-state exchange is the safe baseline because CRDT merge is always additive: a delta
merged into a full state, or a full state merged into a delta, produces the same result
as two full states merged. Old nodes that do not understand the is_delta field decode
it as false (full state) and merge safely. See Payload size and bandwidth for the full bandwidth model.
Protocol
Endpoint
POST /api/gossip/sync
Content-Type: application/pkcs7-mime
X-Ahdapa-Node-Id: <sender's node_id>
<DER SignedData wrapping EnvelopedData>
The gossip_sync axum handler verifies and decrypts the message (message-origin
authentication — see Receiver logic — this part stays in the handler,
outside the engine), then delegates everything else to
engine.handle_inbound(&sender_node_id, plaintext): CBOR decode + staleness check
(AhdapaCodec), admission filtering (AhdapaGossipHooks::filter_inbound), merge +
persist, and building the response — either a delta (when request_delta_since is set in
the inbound envelope) or the full CRDT — which the handler then seals in the same CMS
format before replying.
CMS wire format
Gossip messages use a two-layer CMS structure:
OUTER: SignedData {
eContentType = id-envelopedData
eContent = <inner EnvelopedData DER>
certificates = { sender_self_signed_cert } // carries sender's P-256 public key
signerInfos = { SignerInfo {
signatureAlgorithm = id-ecPublicKey (P-256)
signature = ECDSA-P256-Sign(sha256(eContent))
} }
}
INNER: EnvelopedData {
recipientInfos = SET OF OtherRecipientInfo {
oriType = id-ori-kem
oriValue = KEMRecipientInfo {
kem = id-alg-ml-kem-768
kemct = <ML-KEM-768 encapsulated ciphertext>
wrap = id-aes256-wrap
encryptedKey = AES-256-KeyWrap(kek, CEK) // 40 bytes
}
} // one ORI per recipient
encryptedContentInfo {
contentEncryptionAlgorithm = id-aes256-gcm + GcmParameters(nonce)
encryptedContent = AES-256-GCM(CEK, nonce, CBOR) ‖ tag[16]
}
}
Key derivation: kek = HKDF-SHA256(ml_kem_shared_secret, info="ahdapa-cms-kek", 32)
Key material
Each node has two key pairs stored in the local node_keys table (never gossiped):
| Key | Type | PKCS#8 DER column | SPKI DER column | Published in CRDT |
|---|---|---|---|---|
| KEM encryption | ML-KEM-768 | private_key_der | public_key_der | NodeEntry.kem_public_key_der |
| Gossip signing | ECDSA P-256 | signing_private_key_der | signing_public_key_der | NodeEntry.gossip_signing_pub_key_der |
| JWT signing | Configured by jwt_signing_algorithm (default: ES256) | jwt_signing_priv_der | (derived from jwt_signing_priv_der) | SigningKeyEntry.public_key_der (public only) |
A minimal self-signed X.509 certificate for the P-256 key is also stored
(signing_certificate_der) and embedded in every outbound SignedData so the
receiver can extract the sender’s public key for SPKI comparison without needing a CA.
Keys are generated on first start by bootstrap_node_kem_key() in src/routes/mod.rs
and reused across restarts.
Sender logic
Split between agirru_engine::GossipEngine::run_round (peer-agnostic: snapshotting,
delta computation, envelope construction, encoding) and AhdapaTransport::send_recv
(ahdapa-specific: KEM key resolution, CMS sealing, the actual HTTP call):
# GossipEngine::run_round, per peer (PeerBook tracks last_sent_gen/last_reported_peer_gen)
1. if peer_book.is_backed_off(peer): skip
2. select payload:
- if PeerBook has last_sent_gen for peer: crdt.delta_since(last_sent_gen) → is_delta=true
- otherwise (first contact or after error): crdt.clone() → is_delta=false
3. AhdapaGossipHooks::filter_outbound(peer, &mut payload) // strips peer's own hbac_rules replica
4. if payload == IdpCrdt::default(): skip (nothing new to send)
5. wrap in agirru_engine::Envelope {
state: payload, is_delta, sender_gen: current_gen,
request_delta_since: PeerBook's last_reported_peer_gen for peer,
}
6. AhdapaCodec::encode(&envelope) → CBOR bytes
# AhdapaTransport::send_recv(peer_node_id, request_bytes), spawned concurrently per peer
7. resolve_peer_by_node_id(peer_node_id) → url; look up peer's kem_public_key_der in
local CRDT → NoKemKey error if either is missing
8. ahdapa_cms::sign_and_seal(request_bytes, [peer_kem_spki],
own_signing_priv_pkcs8, own_signing_cert_der)
a. seal(plaintext, recipients):
i. generate random CEK (256-bit) and nonce (96-bit)
ii. AES-256-GCM(CEK, nonce, plaintext) → ciphertext ‖ tag
iii. for each recipient: ML-KEM-768 encapsulate → (kemct, ss)
kek = HKDF-SHA256(ss, "ahdapa-cms-kek", 32)
AES-256 key-wrap(kek, CEK) → encryptedKey
encode KEMRecipientInfo + OtherRecipientInfo
iv. EnvelopedDataBuilder.build() → enveloped_der
b. CmsContentInfo::sign(enveloped_der, own_cert, own_priv_key) → signed_der
9. POST /api/gossip/sync with Content-Type: application/pkcs7-mime
10. on HTTP 401: classify_error → SendOutcome::AuthRejected, driving on_auth_rejected
11. on success: verify_and_open the response, return plaintext bytes to the engine, which
decodes, filter_inbound + merges it (same merge_and_persist path as the passive side)
Receiver logic
Split between the gossip_sync axum handler (message-origin authentication — unchanged,
stays outside the engine on purpose) and engine.handle_inbound:
# gossip_sync handler
1. read X-Ahdapa-Node-Id header → sender_node_id; reject 400 if absent
2. look up sender's gossip_signing_pub_key_der in local CRDT
→ None (no pinned key) → reject 401; TOFU is no longer accepted
3. ahdapa_cms::verify_and_open(body, own_kem_priv_pkcs8, sender_signing_pub_spki)
a. cms.certs()[0] → embedded signer cert → extract SPKI
b. compare embedded SPKI against pinned sender_signing_pub_spki; mismatch → 401
c. cms.verify(NO_SIGNER_CERT_VERIFY) → validates ECDSA signature; → 401 on failure
d. open(enveloped_der, own_kem_priv_pkcs8):
i. find OtherRecipientInfo with oriType = id-ori-kem
ii. ML-KEM-768 decapsulate(own_priv, kemct) → ss
kek = HKDF-SHA256(ss, "ahdapa-cms-kek", 32)
AES-256 key-unwrap(kek, encryptedKey) → CEK
iii. parse GcmParameters → nonce
iv. AES-256-GCM decrypt(CEK, nonce, encryptedContent) → CBOR bytes (plaintext)
4. engine.handle_inbound(&sender_node_id, plaintext) → response_plaintext (steps 5-9 below);
empty Vec means "reject" (malformed or stale envelope) → 400
# GossipEngine::handle_inbound (steps 5-9 happen here)
5. AhdapaCodec::decode(plaintext) → Envelope<IdpCrdt>; folds in the staleness check:
reject if issued_at < now - tombstone_ttl_secs (default 7 days) → replay prevention
6. AhdapaGossipHooks::filter_inbound(sender, &mut envelope.state) — see Admission filters
7. merge_and_persist: merge into engine's IdpCrdt (the same Arc<RwLock<IdpCrdt>> as
AppState.crdt — see Background loop); if changed, call on_merged, then persist
8. determine response payload (anti-echo: excludes anything local_gen > pre_merge_gen,
i.e. entries just merged in from this same sender):
- if envelope.request_delta_since is Some(since): state.delta_range(since, pre_merge_gen) → is_delta=true
- otherwise: state.clone() → is_delta=false
wrap in Envelope { state: response_state, is_delta, sender_gen: post_merge_gen, request_delta_since: None }
AhdapaCodec::encode → response bytes
# back in gossip_sync
9. look up sender_node_id's kem_public_key_der in local CRDT (now updated by the merge —
may have just been introduced by this very request, if this is the sender's first
contact) → reject 400 if still absent
sign_and_seal(response_plaintext, [sender_kem_key], own_signing_priv, own_signing_cert) → response
Admission filters
AhdapaGossipHooks::filter_inbound applies two filters before merging, both operating on
envelope.state (the inbound IdpCrdt) in place:
Node allowlist (gossip.allowed_node_ids): the combined static and topology-derived
allowlist is always enforced. An empty union of both lists admits nobody (fail-closed).
Any new NodeEntry whose node_id is not in the combined allowlist is dropped
(cluster_nodes.retain_keys(...)) before merge. Protects against rogue nodes
self-registering and obtaining the cluster wrapping key.
HBAC replica-ownership filter (agirru_crdt::ReplicaScoped): hbac_rules and
rar_rules (RFC 9396 rules, same causal model) are the only fields where IdpCrdt
implements ReplicaScoped. filter_inbound calls
retain_replicas(&|r| r == replica_id_from_str(sender)), keeping only causal tags
attributed to the authenticated sender’s own replica and dropping anything claiming to be
a different replica. filter_outbound does the mirror-image thing on the way out
(retain_replicas(&|r| r != replica_id_from_str(peer))) — stripping the peer’s own
entries before echoing them back, a bandwidth optimization rather than a security
control. This is new: the pre-migration code had no per-field replica-ownership check for
hbac_rules at all.
There used to be a second node-allowlist layer here — “a sender may only add its own
NodeEntry via gossip,” defending against a forged X-Ahdapa-Node-Id header — removed
during the agirru_engine migration. It is safe to drop because, by the time sender
reaches this hook, the identity is cryptographically corroborated rather than a bare
header claim: gossip_sync looks up the pinned signing key for the specific node_id
the header claims, and verify_and_open only succeeds if the envelope was signed by
that key. A forged header naming a different node_id would need to hold that node’s
private key to pass verification. This depends on that exact
lookup-then-verify-against-the-one-pinned-key sequence being preserved in gossip_sync —
if that ordering ever changes (e.g. verifying against any known peer’s key instead of
specifically the claimed sender’s), this layer must be reinstated.
KEM self-registration and signing-key pinning
A node whose KEM key is not yet in the CRDT cannot receive encrypted gossip — the
sender skips peers with no known KEM key. POST /api/gossip/register-kem seeds both
the ML-KEM-768 public key and the ECDSA P-256 gossip signing key before the first
gossip exchange.
There are two independent triggers for calling register_self_with_peer(): proactively
after each topology refresh (below), and reactively via AhdapaGossipHooks::on_auth_rejected
whenever a peer responds 401 to a gossip push — GossipEngine’s own PeerBook throttles
repeat calls to at most once per auth_reject_notify_cooldown (default 60s) per peer, so a
persistently-rejecting peer doesn’t trigger Kerberos churn every round.
In IPA deployments, after each topology refresh, the local node calls
register_self_with_peer() for every newly discovered peer that does not yet have
this node’s KEM key. That function:
- Acquires a Kerberos service ticket for
HTTP@<peer_host>using the local machine credential (gss_initiator). - POSTs this node’s ML-KEM-768 public key and its ECDSA P-256 gossip signing
public key to
<peer_url>/api/gossip/register-kemwithAuthorization: Negotiate <AP-REQ>. - The peer verifies the AP-REQ, extracts the authenticated principal
(
HTTP/<hostname>@<REALM>viaServicePrincipal::parse), and stores both keys in theNodeEntryunder<hostname>— provided that hostname matches thenode_idin the request body, is in the allowlist, and the principal’s realm matches the expected realm (gossip.kerberos_realmif set, otherwise the local[server] realm).
The insert uses a three-case match: insert-fresh (neither key known), upsert-signing-
key-only (KEM key known but signing key absent), or no-op (both keys already set). Once
the signing key is pinned, gossip_sync rejects any message from that sender whose
embedded ECDSA key does not match the pinned value — there is no TOFU fallback. This
requires gssapi.initiator_principal to be set so that AppState::ipa.gss_initiator
is Some; the mechanism is a no-op when it is absent.
Background loop
routes::gossip::run(state) is spawned from main.rs after AppState is constructed.
When ipa_topology = true, a separate task (topology::run_topology_refresh) is also
spawned; it populates AppState::dynamic_peers and AppState::dynamic_allowed_nodes
before the first gossip round.
#![allow(unused)]
fn main() {
tokio::spawn(routes::gossip::run(state.clone()));
// When ipa_topology = true:
tokio::spawn(topology::run_topology_refresh(state.clone()));
}
run(state) returns immediately (no gossip task at all) if no peers are configured and
ipa_topology is disabled. Otherwise it:
- Builds
AhdapaGossipRuntime::new(state.clone()), bundlingAhdapaTransport/AhdapaMembership/AhdapaStorage/AhdapaGossipHooks/AhdapaCodec. - Bootstraps the engine via
AhdapaGossipEngine::bootstrap_with_state, not the more obviousbootstrap:bootstrapwould callStorage::load()and wrap the result in its own brand-new, engine-privateArc<RwLock<IdpCrdt>>— completely disconnected fromAppState.crdt, theArcevery other part of ahdapa (admin handlers, token issuance, HBAC) reads and writes directly. Gossip merges would land in a copy nothing else ever reads, and local writes would never reach the engine’s outbound push.bootstrap_with_statetakesArc::clone(&state.crdt)directly instead, so gossip and every other code path share exactly one lock. (This was a real bug caught bycluster-demoduring the migration — node2/node3 never saw a client node1 created, because the engine was gossiping its own private copy.) - Stores the running engine in
state.gossip_engine(Arc<RwLock<Option<Arc<AhdapaGossipEngine>>>>) sogossip_sync,gossip_await_client, and every CRDT-mutating admin handler can reach it. - Spawns the periodic maintenance task (
maintenance::spawn_maintenance_task) — see below — as an independent task, not tied to gossip rounds. - Calls
engine.run_with(callback), which loops forever: a debounced periodic tick (round_interval, fromgossip.interval_secs) plus an immediate off-cycle round wheneverengine.notify_local_change()is called (the replacement for the oldgossip_notifyNotify— every CRDT-mutating admin handler calls this after writing, same trigger-an-immediate-push behavior as before). The callback updatesgossip_stats.rounds_completed/last_round_at, but only when the round’speers_attemptedlist was non-empty — an idle round with no configured/discovered peers at all doesn’t count as a completed round, matching the old loop’s semantics.
Each round (GossipEngine::run_round, inside the engine, generic over the replicated
state type) still has the same three phases as before, just no longer ahdapa-owned code:
// Phase 1: snapshot + build per-peer payloads (single read lock on the shared IdpCrdt)
peers = AhdapaMembership::peers() // node_ids with a resolvable URL + KEM key
snapshot = state.read().clone(); current_gen = CRDT_GENERATION.current()
for each peer (via internal PeerBook, tracking last_sent_gen/last_reported_peer_gen):
if peer_book.is_backed_off(peer): skip
payload = last_sent_gen.map(|g| snapshot.delta_since(g)).unwrap_or(snapshot.clone())
AhdapaGossipHooks::filter_outbound(peer, &mut payload)
if payload == IdpCrdt::default(): skip (nothing new for this peer)
encode Envelope{state: payload, is_delta, sender_gen: current_gen, request_delta_since}
// Phase 2: send to every peer concurrently (JoinSet)
for (peer, bytes) in outbound: spawn AhdapaTransport::send_recv(peer, bytes)
// Phase 3: process responses one at a time as they arrive
for (peer, result) in completed sends:
on error: classify via AhdapaTransport::classify_error;
if AuthRejected (peer's 401) and not in cooldown: on_auth_rejected(peer)
on success: AhdapaCodec::decode → Envelope; filter_inbound; merge_and_persist
(same path as the passive side — see Receiver logic above)
merge_and_persist (used by both the active round above and the passive
handle_inbound) is the one place a merge actually lands: it merges into the shared
IdpCrdt, and only if that merge actually changed something does it call
AhdapaGossipHooks::on_merged, persist via AhdapaStorage::persist_full, and fire the
changed notification gossip_await_client waits on. A genuinely redundant delivery
(e.g. the same data learned from a different peer in a full-mesh cluster) triggers none
of that.
AhdapaGossipHooks::on_merged does three things whenever a merge changes state, on
whichever side (inbound or the active side’s post-response merge) triggered it:
- Rebuilds the ABAC policies (
AppState::rebuild_abac_policies) — but only wheninbound.hbac_rulesactually contains live rules, not unconditionally every round like the old loop did. Likewise rebuilds the RFC 9396 policy (AppState::rebuild_rar_policy) when the inbound state carries any RAR rule content (including a deletion) or RAR types, and persists the HBAC/RAR snapshot when either rule set arrived. - Evicts any cached signing key whose
kidwas tombstoned by this merge. - If the peer’s
wrapping_key_iddiffers from the local one, pulls the new key from them (pull_wrapping_key) — the same logic the old loop ran, but now triggered fromon_mergedon either side rather than only the active side’s response handling.
If a peer is unreachable, the error is recorded in the engine’s internal PeerBook and
the round continues to the next peer. GossipConfig::backoff controls per-peer retry
backoff (BackoffPolicy::None by default, matching the old loop’s “retry every
interval, no backoff” behavior).
Periodic maintenance task
maintenance::spawn_maintenance_task (src/routes/gossip/maintenance.rs) runs
independently of gossip rounds, on the same cadence the old loop’s “hourly” block used
(interval_secs × rounds_per_hour, so effectively every ~1 hour at the default 5s
interval):
loop:
sleep(period)
purge_expired_families(now) // moved here from every gossip round --
// see the note under CRDT primitives → LwwMap
purge_old_tombstones / purge_old_revocations / purge_expired_access_token_revocations /
purge_expired_upstream_tokens / saml2_sp_entries / saml2_idp_entries /
federation_policies .purge_old_tombstones(tombstone_cutoff)
cleanup_expired_families(db, now) // DB-level purge, same six cleanup_* fns
cleanup_old_tombstones(db, tombstone_cutoff)
cleanup_old_crdt_revocations / cleanup_expired_at_revocations /
cleanup_expired_upstream_tokens(db, ...)
cleanup_expired_pending_saml2 / cleanup_expired_pending_saml2_sp /
cleanup_expired_saml2_artifacts(db, now)
jti_cache.retain(|_, exp| exp > now)
The topology refresh task runs an initial fetch immediately on startup (before the first
gossip sleep) and then sleeps for ipa_topology_interval_secs (minimum 30 s, default
300 s). On LDAP error, the previous peer list is kept unchanged and a warning is logged.
After each successful topology fetch, if gss_initiator is available, the topology task
also calls register_self_with_peer() for each newly-discovered peer whose KEM key is
not yet in the CRDT. This pre-seeds the key via POST /api/gossip/register-kem with a
Kerberos AP-REQ so that the legitimate node wins the OR-Map first-write-wins race before
the first gossip round fires.
Cluster wrapping key
The 32-byte cluster wrapping key (used for session cookies) is stored node-locally
in node_keys.wrapping_key_cms_der as a CMS EnvelopedData blob sealed to the node’s
own ML-KEM-768 public key. It is never gossiped in plaintext or as a multi-recipient
blob.
Only a short UUID string (wrapping_key_id) is gossiped in the CRDT. When a node
observes a different UUID after a gossip merge, it fetches the actual key on demand:
GET /api/gossip/wrapping-key
X-Ahdapa-Node-Id: <requester's node_id>
Response: 200 OK
Content-Type: application/octet-stream
X-Ahdapa-Node-Id: <responder's node_id>
Body: SignedData(EnvelopedData) DER
The response is a full SignedData(EnvelopedData) blob produced by sign_and_seal(),
sealed to exactly one recipient (the requester’s ML-KEM-768 public key) and signed
with the responder’s ECDSA P-256 gossip signing key. The requester looks up the
responder’s pinned signing key in the CRDT (from the X-Ahdapa-Node-Id header) and
calls verify_and_open(). A response from a node with no pinned signing key is
rejected. Confidentiality is ensured by the inner ML-KEM-768 encryption; integrity and
sender authentication are ensured by the outer ECDSA P-256 signature.
Node statistics endpoint
GET /api/gossip/stats
Unauthenticated. Intentionally unauthenticated — like /api/gossip/kem-info — because the
admin web UI fetches it before an admin session is established. What is exposed is aggregate
counts and gossip health indicators; no key material, user data, or token content is returned.
Response body (JSON):
{
"node_id": "ipa1.example.com",
"crdt_generation": 42,
"counts": {
"clients": 3,
"signing_keys": 2,
"cluster_nodes": 3,
"refresh_families": 7,
"revoked_sessions": 1,
"scope_definitions": 8,
"ipa_idp_overrides": 0
},
"peers": ["https://ipa2.example.com/idp", "https://ipa3.example.com/idp"],
"active_signing_kid": "abc123",
"kem_enrolled": true,
"gossip_signing_enrolled": true,
"gossip": {
"started_at": 1716000000,
"rounds_completed": 12,
"last_round_at": 1716000060,
"peer_last_sync": { "ipa2.example.com": 1716000058, "ipa3.example.com": 1716000059 },
"persist_errors": 0,
"wrapping_key_pull_errors": 0
}
}
Field notes:
crdt_generation— current value of theCRDT_GENERATIONatomic counter.counts.*— live (non-tombstoned) entry counts for each CRDT collection.peers— union of configuredgossip.peersand topology-discovered peers.active_signing_kid— thekidof the currently active JWT signing key.kem_enrolled/gossip_signing_enrolled— whether both cryptographic identities are registered in the CRDT.gossip.started_at— Unix timestamp when the gossip background task started.gossip.rounds_completed— number of gossip rounds in which at least one peer was successfully synced. Idle rounds (CRDT unchanged, all pushes skipped) and rounds where all peers fail do not increment this counter.gossip.last_round_at— Unix timestamp of the most recent round that synced at least one peer.nulluntil the first successful sync.gossip.peer_last_sync— Unix timestamp of the most recent successful inbound sync from each peer (recorded by the/api/gossip/syncreceiver).gossip.persist_errors— cumulative DB persist failures since startup (incremented after both inbound sync and outbound merge failures).gossip.wrapping_key_pull_errors— cumulative failures to pull the cluster wrapping key from a peer after detecting a UUID change.
This endpoint is used by the admin web UI Cluster Nodes page to display per-node runtime gossip health alongside the static CRDT node entries from GET /api/admin/nodes.
Client convergence endpoint
GET /api/gossip/await-client?client_id=<id>&timeout_ms=<ms>
Unauthenticated. Long-poll endpoint that blocks until client_id appears in the
local CRDT or timeout_ms elapses (default: 30000 ms). Uses
GossipEngine::wait_for_change() to detect convergence without polling — it fires
after any merge that actually changed state, on either the active or passive side (see
Background loop); this replaced the old, ahdapa-owned crdt_changed
Notify. When gossip is disabled or the engine hasn’t bootstrapped yet
(state.gossip_engine is None), there is no change-notification source to await, so
the endpoint falls back to a 200ms poll interval instead of busy-looping. Returns
200 OK when the client is found, 408 Request Timeout on timeout.
Primarily used by the ahdapa-bench converge scenario to measure gossip
propagation time with sub-millisecond precision.
Kerberos KEM self-registration endpoint
POST /api/gossip/register-kem
Authorization: Negotiate <base64-AP-REQ>
Content-Type: application/json
{
"node_id": "<hostname>",
"kem_public_key_der": "<base64url-ML-KEM-768-SPKI-DER>",
"gossip_signing_pub_key_der": "<base64url-ECDSA-P256-SPKI-DER>"
}
Used by topology-discovered peers to seed both their ML-KEM-768 public key and their
ECDSA P-256 gossip signing key before the first gossip round. All three fields are
required; missing or empty fields return 400 Bad Request. The server:
- Returns
503 Service Unavailableif the GSSAPI server credential is unavailable (state.gss_credisNone— indicates a configuration or keytab problem). - Calls
try_spnego()to accept the Kerberos AP-REQ. Returns401 Negotiateif absent,401if the token is invalid. - Calls
ServicePrincipal::parse()on the authenticated principal. Rejects with403if the principal is notHTTP/<host>@<REALM>(user principals and non-HTTP service types are excluded). - Rejects with
403if the principal’s realm does not match the expected realm —gossip.kerberos_realmwhen set, otherwise the local[server] realm. The check is always enforced, preventing cross-realm trust escalation. - Checks that
req.node_id.to_lowercase() == authed_host. Rejects with403if they differ — a machine can only register its own identity. - Checks that
authed_hostis admitted: it must appear in the staticallowed_node_idslist, in the topology-derived allowlist, or match the hostname of a staticgossip.peersURL (static peers are implicitly trusted for KEM registration). Rejects with403if not admitted. - Applies a three-case match on the existing CRDT entry for this
node_id:- Insert-fresh: neither key known → insert
NodeEntrywith both keys. - Upsert-signing-key-only: KEM key present but
gossip_signing_pub_key_derempty → update the entry to add the signing key. - No-op: both keys already present → return
200 OKimmediately (idempotent).
- Insert-fresh: neither key known → insert
- Returns
200 OK, optionally with aWWW-Authenticate: Negotiate <mutual-auth-token>header if GSSAPI produced a mutual-authentication output token.
At startup, bootstrap_wrapping_key() reads node_keys.wrapping_key_cms_der. If
present, it decrypts the blob to recover the 32-byte key. If absent (first start), it
generates a fresh key, seals it to the node’s own KEM key, and stores the result in
node_keys. A UUID is generated and published to the CRDT as wrapping_key_id with
timestamp=1 so that the established cluster’s UUID wins the LWW merge on the first
gossip round.
When the cluster wrapping key is rotated via PUT /api/admin/keys/cluster, the node
re-seals the new key to its own KEM key, stores it in node_keys, and updates
crdt.wrapping_key_id to a new UUID. Peers detect the UUID change via gossip and pull
the new key via the on-demand endpoint.
Convergence
| Scenario | Convergence |
|---|---|
| Single node | Instant (no peers) |
| Two-node cluster (KEM keys known) | After 1 gossip round (≤ interval_secs seconds) |
| Three-node cluster, all connected | After 1–2 gossip rounds |
| Partition healed after T seconds | After ≤ 2 gossip rounds from partition heal |
| New node joining (static peers) | After 2 gossip rounds (learn KEM key → pull wrapping key via on-demand endpoint) |
New node joining (IPA topology, gss_initiator set) | After 1 gossip round — both the KEM key and the gossip signing key are pre-seeded via Kerberos register-kem before first gossip push; wrapping key pulled on first exchange. Requires both nodes to complete their mutual register-kem calls before the first gossip interval fires; this holds in practice because the topology refresh runs immediately on startup, before the first gossip sleep. |
New signing key propagation: a key added on node A is available on node B after at most 1 gossip round from A to B. Resource servers should cache JWKS with a short TTL (≤ interval_secs × 2) to avoid key-not-found errors during propagation.
Security considerations
| Property | Value |
|---|---|
| Confidentiality | AES-256-GCM per-recipient (inner EnvelopedData) |
| Integrity | AES-256-GCM auth tag + ECDSA P-256 signature |
| Sender authentication | ECDSA P-256 over eContent (outer SignedData); signing key pinned via register-kem before first gossip |
| Node admission control | allowed_node_ids allowlist (fail-closed on empty) + ReplicaScoped filtering restricting hbac_rules / rar_rules tags to the authenticated sender’s own replica |
| Replay prevention | Envelope.issued_at checked against now - tombstone_ttl_secs (AhdapaCodec::decode) |
| Post-quantum | ML-KEM-768 for key encapsulation (FIPS 203) |
/api/gossip/syncSHOULD be firewalled to the cluster’s subnet as defense-in-depth. CMS encryption ensures confidentiality even if traffic is captured, but network isolation prevents unauthorized nodes from attempting to self-register.- The allowlist is fail-closed. When both the static
allowed_node_idslist and the topology-derived allowlist are empty, no node can self-register via gossip or the wrapping-key endpoint. This is intentional: operators must either configure an explicit allowlist or enableipa_topologyso that hostnames are discovered automatically. - Gossip envelopes carry a timestamp (
issued_at). Envelopes older thantombstone_ttl_secs(default 7 days) are rejected, preventing an attacker from replaying a captured gossip message after its tombstones have been GC-purged. - The gossip
reqwest::Clienthas a 10-second request timeout. Slow peers do not block the gossip loop. - ECDSA P-256, not Ed25519, is used for gossip signing. OpenSSL’s
CMS_sign()API requires a key type that has a default digest algorithm; Ed25519 (PureEdDSA) does not satisfy this requirement. P-256 provides the same 128-bit security level. - The JWT signing algorithm is configurable, not gossip signing. Each node generates
its own JWT signing key pair (algorithm set by
[server] jwt_signing_algorithm, default: ES256) stored innode_keys.jwt_signing_priv_der. The private key never leaves the node; only the public key is gossiped inSigningKeyEntry. This is distinct from the ECDSA P-256 gossip signing key.
Payload size and bandwidth
Raw field sizes
The binary fields that dominate gossip payload size (measured from a three-node demo cluster):
| Field | Bytes | Gossiped |
|---|---|---|
ML-KEM-768 public key SPKI (NodeEntry.kem_public_key_der) | 1,206 | Yes |
ECDSA P-256 gossip signing pub key SPKI (NodeEntry.gossip_signing_pub_key_der) | 91 | Yes |
JWT signing private key DER (SigningKeyEntry.private_key_der) | varies by algorithm | No — #[serde(skip_serializing)]; stays in node_keys |
JWT signing public key SPKI DER (SigningKeyEntry.public_key_der) | varies by algorithm | Yes |
ECDSA P-256 gossip signing certificate (node_keys.signing_certificate_der) | 291 | No — local only |
ML-KEM-768 private key PKCS#8 (node_keys.private_key_der) | 2,498 | No — local only |
The ML-KEM-768 public key is the dominant field by a factor of ~13× over the next largest gossiped value.
Per-entity CBOR contribution
The CRDT is serialised as CBOR (ciborium). CBOR stores binary fields as raw bytes (no base64 overhead). Approximate CBOR size per entry:
| Entity | ~CBOR bytes | Dominant field |
|---|---|---|
NodeEntry (one cluster node) | ~1,530 B | ML-KEM-768 pub key (1,206 B raw) |
SigningKeyEntry (one JWT signing key) | ~100 B (ES256) – ~2,600 B (ML-DSA-87) | JWT public key (size varies by algorithm); private key not gossiped |
ClientEntry (typical OAuth2 client) | ~150 B | UUIDs + scopes; short serde field names (2 chars) keep the CBOR compact |
RefreshFamilyState (one active session) | ~100 B | UUIDs + counters |
CMS envelope overhead
Each gossip message is sent to exactly one peer, so the CMS overhead is constant regardless of cluster size:
| Layer | Bytes |
|---|---|
Outer SignedData headers + ECDSA P-256 signature (64 B) + signer cert (291 B) | ~555 B |
Inner EnvelopedData KEMRecipientInfo: ML-KEM-768 ct (1,088 B) + wrapped CEK (40 B) + headers | ~1,233 B |
| AEAD overhead (12 B nonce + 16 B GCM tag) | 28 B |
| Fixed CMS overhead per gossip message | ~1,816 B |
Wire size per gossip push
Total CMS-encrypted wire size for one outbound push (one recipient). The cluster wrapping key blob is no longer in the gossip body; only its UUID is gossiped:
| Scenario | Wire bytes |
|---|---|
| 3 nodes, 3 signing keys, 4 clients, 0 sessions (demo measured) | ~7,454 B |
| 3 nodes, 5 clients, 50 sessions | ~8 KB |
| 5 nodes, 5 clients, 50 sessions | ~11 KB |
| 10 nodes, 10 clients, 100 sessions | ~19 KB |
Bandwidth per gossip cycle
The topology is full-mesh: each node pushes to every configured peer and receives a response. Total cluster bandwidth per cycle (worst case, full-state) = N × (N−1) × 2 × wire_bytes. In practice, delta exchange reduces per-push payload to the size of changed entries only.
These figures are theoretical maximums — they assume every gossip round produces an actual push. In practice, the generation-skip optimisation suppresses pushes when the CRDT has not changed since the last successful round. In the demo cluster (active token issuance, no schema mutations), 93% of rounds were skipped, reducing steady-state bandwidth to near zero. Pushes happen only when the CRDT actually changes (client creation/deletion, key rotation, node join/leave).
At interval_secs = 2 (demo default; config default is 5 s):
| Nodes | Wire/msg | Per round (2 s) | Per hour (theoretical max) |
|---|---|---|---|
| 2 | 7.5 KB | 60 KB | 108 MB |
| 3 | 7.5 KB | 90 KB | 162 MB |
| 5 | 10.5 KB | 420 KB | 756 MB |
| 10 | 19 KB | 3.4 MB | 6.1 GB |
The O(N²) topology is practical for the expected deployment range of 2–5 nodes. Above ~10 nodes the bandwidth cost becomes significant and a partial-mesh peer configuration (each node lists only a subset of peers) should be considered.
Marginal cost per added entity
Measured at a three-node baseline:
| Change | Extra bytes per gossip message |
|---|---|
| +1 cluster node | ~+1,530 B (NodeEntry: ML-KEM-768 pub key 1,206 B + other fields; CBOR-encoded, no wrapping key blob) |
| +1 removed node (tombstone) | +~220 B (tombstone metadata; key fields absent) |
| +1 OAuth2 client | +~150 B |
| +1 active refresh token family (session) | +~100 B |
| +1 JWT signing key (rotation) | +~100 B (public key only; private key not gossiped) |
Sessions and clients are cheap. Nodes are the dominant cost because every node contributes 1,206 B of ML-KEM-768 public key material (gossiped in NodeEntry). The previous per-node cost of +1,644 B from the cluster wrapping key CMS blob is eliminated — the wrapping key is no longer gossiped.
Tombstone accumulation and GC
Removed nodes, deleted clients, and revoked signing keys leave OrMap tombstones. Each
tombstone adds ~220 B to gossip messages until it is garbage-collected. Tombstones older
than gossip.tombstone_ttl_secs (default: 7 days) are purged from both the in-memory
CRDT and the database approximately once per hour. This bounds tombstone growth even in
high-churn deployments.
The TTL must exceed the longest expected node downtime: a node that is offline longer than the TTL may re-gossip entries that were since deleted (those entries would be re-merged on reconnection). The default 7-day TTL is conservative and suitable for most deployments.
Configuration
| Key | Default | Description |
|---|---|---|
peers | [] | Peer node base URLs. Gossip is disabled when this list is empty and ipa_topology is false. |
interval_secs | 5 | Push interval in seconds. |
allowed_node_ids | [] | Allowlist of node_id values permitted to self-register. The allowlist is always enforced — an empty union of both the static list and the topology-derived list admits nobody (fail-closed). When ipa_topology = true, discovered replica hostnames are appended automatically. |
tombstone_ttl_secs | 604800 | Seconds to retain OR-Map tombstones before GC. Also the maximum age of accepted gossip envelopes (issued_at window). Must exceed the longest expected node downtime. Default: 7 days. |
ipa_topology | false | When true, a background task (src/topology.rs) queries cn=topology,cn=ipa,cn=etc,<suffix> for ipaReplTopoSegment entries and derives gossip peer URLs of the form https://<hostname><base_path>. The peer list is stored in AppState::dynamic_peers and the allowlist in AppState::dynamic_allowed_nodes; both are merged with any statically configured peers and allowed_node_ids at each gossip round. |
ipa_topology_interval_secs | 300 | How often (seconds) to re-query the IPA topology. Only used when ipa_topology = true. Minimum: 30 s. |
kerberos_realm | — | Expected Kerberos realm for register-kem callers (e.g. "IPA.EXAMPLE.COM"). When set, principals whose realm does not match are rejected with 403, preventing cross-realm trust escalation. When unset, the local [server] realm is used as the expected realm, so cross-realm principals are always rejected. |
[gossip]
peers = ["https://node2.example.com:8080", "https://node3.example.com:8080"]
interval_secs = 5
allowed_node_ids = ["node1.example.com", "node2.example.com", "node3.example.com"]
tombstone_ttl_secs = 604800 # 7 days
For IPA-integrated deployments, the static peers list can be omitted entirely when
ipa_topology = true. A single-node deployment requires no gossip configuration at all.