Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Gossip Protocol

The gossip protocol synchronises IdpCrdt state across cluster nodes. The wire protocol, CMS envelope construction (crates/ahdapa-cms/), and CRDT semantics are unchanged from ahdapa’s original design, but the engine that drives the push-pull loop is not: it’s now agirru_engine::GossipEngine, a generic gossip engine extracted from this same code. Ahdapa supplies the application-specific pieces under src/routes/gossip/:

FileImplementsResponsibility
transport.rsagirru_engine::Transport<String>Resolve a peer node_id to a URL + KEM key, sign_and_seal, POST, verify_and_open the response. Classifies HTTP 401 as SendOutcome::AuthRejected (drives on_auth_rejected below).
membership.rsagirru_engine::Membership<String>The node_id of every configured/discovered peer URL that resolves to a live cluster_nodes entry with a KEM key.
storage.rsagirru_engine::Storage<IdpCrdt>Wraps IdpCrdt::load_from_db / AppState::persist_crdt — see the note on bootstrap_with_state under Background loop for why load() is not actually called at startup.
hooks.rsagirru_engine::GossipHooks<String, IdpCrdt>Admission filtering, ABAC policy rebuild, signing-key cache eviction, wrapping-key pull, 401 re-registration. See Admission filters.
codec.rsagirru_engine::Codec<Envelope<IdpCrdt>>CBOR encode/decode of the engine’s own Envelope<IdpCrdt> type, folding in stale-envelope rejection.
runtime.rsagirru_engine::GossipRuntime<IdpCrdt, String>Bundles the five pieces above into AhdapaGossipRuntime, and defines the AhdapaGossipEngine type alias.
maintenance.rs—Periodic tombstone/expired-entry cleanup, decoupled from gossip rounds (see Background loop).
mod.rs—Axum handlers (gossip_sync, gossip_kem_info, gossip_wrapping_key, gossip_register_kem, gossip_node_stats, gossip_await_client) and run(), which bootstraps and drives the engine.

Design

The protocol uses delta-based exchange by default, falling back to full-state on first contact or after an error. After a successful round, a node sends only the CRDT entries that changed since the last successful exchange with each peer (a sparse IdpCrdt delta), rather than the full state every time. Peers exchange generation counters in each envelope to coordinate what they already have. One round-trip still brings both nodes to the same merged state.

This delta/full-state decision, and the per-peer generation bookkeeping behind it, is now entirely internal to agirru_engine::GossipEngine’s PeerBook — ahdapa no longer owns peer_last_gen/peer_response_gen maps directly (see CRDT_GENERATION counter). Full-state pushes occur in the following cases:

  • First contact with a peer (no prior recorded generation for it).
  • After any connection error or non-2xx response (the engine clears its bookkeeping for that peer).
  • When a pull_wrapping_key failure occurs (AhdapaGossipHooks::on_merged increments wrapping_key_pull_errors, but does not itself force a full-state resync — the next round’s delta computation is unaffected by this failure since it’s keyed off successful sync generations, not wrapping-key state).

Full-state exchange is the safe baseline because CRDT merge is always additive: a delta merged into a full state, or a full state merged into a delta, produces the same result as two full states merged. Old nodes that do not understand the is_delta field decode it as false (full state) and merge safely. See Payload size and bandwidth for the full bandwidth model.

Protocol

Endpoint

POST /api/gossip/sync
Content-Type: application/pkcs7-mime
X-Ahdapa-Node-Id: <sender's node_id>

<DER SignedData wrapping EnvelopedData>

The gossip_sync axum handler verifies and decrypts the message (message-origin authentication — see Receiver logic — this part stays in the handler, outside the engine), then delegates everything else to engine.handle_inbound(&sender_node_id, plaintext): CBOR decode + staleness check (AhdapaCodec), admission filtering (AhdapaGossipHooks::filter_inbound), merge + persist, and building the response — either a delta (when request_delta_since is set in the inbound envelope) or the full CRDT — which the handler then seals in the same CMS format before replying.

CMS wire format

Gossip messages use a two-layer CMS structure:

OUTER: SignedData {
  eContentType = id-envelopedData
  eContent     = <inner EnvelopedData DER>
  certificates = { sender_self_signed_cert }   // carries sender's P-256 public key
  signerInfos  = { SignerInfo {
    signatureAlgorithm = id-ecPublicKey (P-256)
    signature = ECDSA-P256-Sign(sha256(eContent))
  } }
}

INNER: EnvelopedData {
  recipientInfos = SET OF OtherRecipientInfo {
    oriType  = id-ori-kem
    oriValue = KEMRecipientInfo {
      kem           = id-alg-ml-kem-768
      kemct         = <ML-KEM-768 encapsulated ciphertext>
      wrap          = id-aes256-wrap
      encryptedKey  = AES-256-KeyWrap(kek, CEK)   // 40 bytes
    }
  }                                                // one ORI per recipient
  encryptedContentInfo {
    contentEncryptionAlgorithm = id-aes256-gcm + GcmParameters(nonce)
    encryptedContent           = AES-256-GCM(CEK, nonce, CBOR) ‖ tag[16]
  }
}

Key derivation: kek = HKDF-SHA256(ml_kem_shared_secret, info="ahdapa-cms-kek", 32)

Key material

Each node has two key pairs stored in the local node_keys table (never gossiped):

KeyTypePKCS#8 DER columnSPKI DER columnPublished in CRDT
KEM encryptionML-KEM-768private_key_derpublic_key_derNodeEntry.kem_public_key_der
Gossip signingECDSA P-256signing_private_key_dersigning_public_key_derNodeEntry.gossip_signing_pub_key_der
JWT signingConfigured by jwt_signing_algorithm (default: ES256)jwt_signing_priv_der(derived from jwt_signing_priv_der)SigningKeyEntry.public_key_der (public only)

A minimal self-signed X.509 certificate for the P-256 key is also stored (signing_certificate_der) and embedded in every outbound SignedData so the receiver can extract the sender’s public key for SPKI comparison without needing a CA.

Keys are generated on first start by bootstrap_node_kem_key() in src/routes/mod.rs and reused across restarts.

Sender logic

Split between agirru_engine::GossipEngine::run_round (peer-agnostic: snapshotting, delta computation, envelope construction, encoding) and AhdapaTransport::send_recv (ahdapa-specific: KEM key resolution, CMS sealing, the actual HTTP call):

# GossipEngine::run_round, per peer (PeerBook tracks last_sent_gen/last_reported_peer_gen)
1. if peer_book.is_backed_off(peer): skip
2. select payload:
   - if PeerBook has last_sent_gen for peer: crdt.delta_since(last_sent_gen) → is_delta=true
   - otherwise (first contact or after error): crdt.clone() → is_delta=false
3. AhdapaGossipHooks::filter_outbound(peer, &mut payload)  // strips peer's own hbac_rules replica
4. if payload == IdpCrdt::default(): skip (nothing new to send)
5. wrap in agirru_engine::Envelope {
       state: payload, is_delta, sender_gen: current_gen,
       request_delta_since: PeerBook's last_reported_peer_gen for peer,
   }
6. AhdapaCodec::encode(&envelope) → CBOR bytes

# AhdapaTransport::send_recv(peer_node_id, request_bytes), spawned concurrently per peer
7. resolve_peer_by_node_id(peer_node_id) → url; look up peer's kem_public_key_der in
   local CRDT → NoKemKey error if either is missing
8. ahdapa_cms::sign_and_seal(request_bytes, [peer_kem_spki],
                              own_signing_priv_pkcs8, own_signing_cert_der)
   a. seal(plaintext, recipients):
      i.   generate random CEK (256-bit) and nonce (96-bit)
      ii.  AES-256-GCM(CEK, nonce, plaintext) → ciphertext ‖ tag
      iii. for each recipient: ML-KEM-768 encapsulate → (kemct, ss)
           kek = HKDF-SHA256(ss, "ahdapa-cms-kek", 32)
           AES-256 key-wrap(kek, CEK) → encryptedKey
           encode KEMRecipientInfo + OtherRecipientInfo
      iv.  EnvelopedDataBuilder.build() → enveloped_der
   b. CmsContentInfo::sign(enveloped_der, own_cert, own_priv_key) → signed_der
9. POST /api/gossip/sync with Content-Type: application/pkcs7-mime
10. on HTTP 401: classify_error → SendOutcome::AuthRejected, driving on_auth_rejected
11. on success: verify_and_open the response, return plaintext bytes to the engine, which
    decodes, filter_inbound + merges it (same merge_and_persist path as the passive side)

Receiver logic

Split between the gossip_sync axum handler (message-origin authentication — unchanged, stays outside the engine on purpose) and engine.handle_inbound:

# gossip_sync handler
1. read X-Ahdapa-Node-Id header → sender_node_id; reject 400 if absent
2. look up sender's gossip_signing_pub_key_der in local CRDT
   → None (no pinned key) → reject 401; TOFU is no longer accepted
3. ahdapa_cms::verify_and_open(body, own_kem_priv_pkcs8, sender_signing_pub_spki)
   a. cms.certs()[0] → embedded signer cert → extract SPKI
   b. compare embedded SPKI against pinned sender_signing_pub_spki; mismatch → 401
   c. cms.verify(NO_SIGNER_CERT_VERIFY) → validates ECDSA signature; → 401 on failure
   d. open(enveloped_der, own_kem_priv_pkcs8):
      i.   find OtherRecipientInfo with oriType = id-ori-kem
      ii.  ML-KEM-768 decapsulate(own_priv, kemct) → ss
           kek = HKDF-SHA256(ss, "ahdapa-cms-kek", 32)
           AES-256 key-unwrap(kek, encryptedKey) → CEK
      iii. parse GcmParameters → nonce
      iv.  AES-256-GCM decrypt(CEK, nonce, encryptedContent) → CBOR bytes (plaintext)
4. engine.handle_inbound(&sender_node_id, plaintext) → response_plaintext (steps 5-9 below);
   empty Vec means "reject" (malformed or stale envelope) → 400

# GossipEngine::handle_inbound (steps 5-9 happen here)
5. AhdapaCodec::decode(plaintext) → Envelope<IdpCrdt>; folds in the staleness check:
   reject if issued_at < now - tombstone_ttl_secs (default 7 days) → replay prevention
6. AhdapaGossipHooks::filter_inbound(sender, &mut envelope.state) — see Admission filters
7. merge_and_persist: merge into engine's IdpCrdt (the same Arc<RwLock<IdpCrdt>> as
   AppState.crdt — see Background loop); if changed, call on_merged, then persist
8. determine response payload (anti-echo: excludes anything local_gen > pre_merge_gen,
   i.e. entries just merged in from this same sender):
   - if envelope.request_delta_since is Some(since): state.delta_range(since, pre_merge_gen) → is_delta=true
   - otherwise: state.clone() → is_delta=false
   wrap in Envelope { state: response_state, is_delta, sender_gen: post_merge_gen, request_delta_since: None }
   AhdapaCodec::encode → response bytes

# back in gossip_sync
9. look up sender_node_id's kem_public_key_der in local CRDT (now updated by the merge —
   may have just been introduced by this very request, if this is the sender's first
   contact) → reject 400 if still absent
   sign_and_seal(response_plaintext, [sender_kem_key], own_signing_priv, own_signing_cert) → response

Admission filters

AhdapaGossipHooks::filter_inbound applies two filters before merging, both operating on envelope.state (the inbound IdpCrdt) in place:

Node allowlist (gossip.allowed_node_ids): the combined static and topology-derived allowlist is always enforced. An empty union of both lists admits nobody (fail-closed). Any new NodeEntry whose node_id is not in the combined allowlist is dropped (cluster_nodes.retain_keys(...)) before merge. Protects against rogue nodes self-registering and obtaining the cluster wrapping key.

HBAC replica-ownership filter (agirru_crdt::ReplicaScoped): hbac_rules and rar_rules (RFC 9396 rules, same causal model) are the only fields where IdpCrdt implements ReplicaScoped. filter_inbound calls retain_replicas(&|r| r == replica_id_from_str(sender)), keeping only causal tags attributed to the authenticated sender’s own replica and dropping anything claiming to be a different replica. filter_outbound does the mirror-image thing on the way out (retain_replicas(&|r| r != replica_id_from_str(peer))) — stripping the peer’s own entries before echoing them back, a bandwidth optimization rather than a security control. This is new: the pre-migration code had no per-field replica-ownership check for hbac_rules at all.

There used to be a second node-allowlist layer here — “a sender may only add its own NodeEntry via gossip,” defending against a forged X-Ahdapa-Node-Id header — removed during the agirru_engine migration. It is safe to drop because, by the time sender reaches this hook, the identity is cryptographically corroborated rather than a bare header claim: gossip_sync looks up the pinned signing key for the specific node_id the header claims, and verify_and_open only succeeds if the envelope was signed by that key. A forged header naming a different node_id would need to hold that node’s private key to pass verification. This depends on that exact lookup-then-verify-against-the-one-pinned-key sequence being preserved in gossip_sync — if that ordering ever changes (e.g. verifying against any known peer’s key instead of specifically the claimed sender’s), this layer must be reinstated.

KEM self-registration and signing-key pinning

A node whose KEM key is not yet in the CRDT cannot receive encrypted gossip — the sender skips peers with no known KEM key. POST /api/gossip/register-kem seeds both the ML-KEM-768 public key and the ECDSA P-256 gossip signing key before the first gossip exchange.

There are two independent triggers for calling register_self_with_peer(): proactively after each topology refresh (below), and reactively via AhdapaGossipHooks::on_auth_rejected whenever a peer responds 401 to a gossip push — GossipEngine’s own PeerBook throttles repeat calls to at most once per auth_reject_notify_cooldown (default 60s) per peer, so a persistently-rejecting peer doesn’t trigger Kerberos churn every round.

In IPA deployments, after each topology refresh, the local node calls register_self_with_peer() for every newly discovered peer that does not yet have this node’s KEM key. That function:

  1. Acquires a Kerberos service ticket for HTTP@<peer_host> using the local machine credential (gss_initiator).
  2. POSTs this node’s ML-KEM-768 public key and its ECDSA P-256 gossip signing public key to <peer_url>/api/gossip/register-kem with Authorization: Negotiate <AP-REQ>.
  3. The peer verifies the AP-REQ, extracts the authenticated principal (HTTP/<hostname>@<REALM> via ServicePrincipal::parse), and stores both keys in the NodeEntry under <hostname> — provided that hostname matches the node_id in the request body, is in the allowlist, and the principal’s realm matches the expected realm (gossip.kerberos_realm if set, otherwise the local [server] realm).

The insert uses a three-case match: insert-fresh (neither key known), upsert-signing- key-only (KEM key known but signing key absent), or no-op (both keys already set). Once the signing key is pinned, gossip_sync rejects any message from that sender whose embedded ECDSA key does not match the pinned value — there is no TOFU fallback. This requires gssapi.initiator_principal to be set so that AppState::ipa.gss_initiator is Some; the mechanism is a no-op when it is absent.

Background loop

routes::gossip::run(state) is spawned from main.rs after AppState is constructed. When ipa_topology = true, a separate task (topology::run_topology_refresh) is also spawned; it populates AppState::dynamic_peers and AppState::dynamic_allowed_nodes before the first gossip round.

#![allow(unused)]
fn main() {
tokio::spawn(routes::gossip::run(state.clone()));
// When ipa_topology = true:
tokio::spawn(topology::run_topology_refresh(state.clone()));
}

run(state) returns immediately (no gossip task at all) if no peers are configured and ipa_topology is disabled. Otherwise it:

  1. Builds AhdapaGossipRuntime::new(state.clone()), bundling AhdapaTransport/ AhdapaMembership/AhdapaStorage/AhdapaGossipHooks/AhdapaCodec.
  2. Bootstraps the engine via AhdapaGossipEngine::bootstrap_with_state, not the more obvious bootstrap: bootstrap would call Storage::load() and wrap the result in its own brand-new, engine-private Arc<RwLock<IdpCrdt>> — completely disconnected from AppState.crdt, the Arc every other part of ahdapa (admin handlers, token issuance, HBAC) reads and writes directly. Gossip merges would land in a copy nothing else ever reads, and local writes would never reach the engine’s outbound push. bootstrap_with_state takes Arc::clone(&state.crdt) directly instead, so gossip and every other code path share exactly one lock. (This was a real bug caught by cluster-demo during the migration — node2/node3 never saw a client node1 created, because the engine was gossiping its own private copy.)
  3. Stores the running engine in state.gossip_engine (Arc<RwLock<Option<Arc<AhdapaGossipEngine>>>>) so gossip_sync, gossip_await_client, and every CRDT-mutating admin handler can reach it.
  4. Spawns the periodic maintenance task (maintenance::spawn_maintenance_task) — see below — as an independent task, not tied to gossip rounds.
  5. Calls engine.run_with(callback), which loops forever: a debounced periodic tick (round_interval, from gossip.interval_secs) plus an immediate off-cycle round whenever engine.notify_local_change() is called (the replacement for the old gossip_notify Notify — every CRDT-mutating admin handler calls this after writing, same trigger-an-immediate-push behavior as before). The callback updates gossip_stats.rounds_completed/last_round_at, but only when the round’s peers_attempted list was non-empty — an idle round with no configured/discovered peers at all doesn’t count as a completed round, matching the old loop’s semantics.

Each round (GossipEngine::run_round, inside the engine, generic over the replicated state type) still has the same three phases as before, just no longer ahdapa-owned code:

// Phase 1: snapshot + build per-peer payloads (single read lock on the shared IdpCrdt)
peers = AhdapaMembership::peers()   // node_ids with a resolvable URL + KEM key
snapshot = state.read().clone(); current_gen = CRDT_GENERATION.current()
for each peer (via internal PeerBook, tracking last_sent_gen/last_reported_peer_gen):
    if peer_book.is_backed_off(peer): skip
    payload = last_sent_gen.map(|g| snapshot.delta_since(g)).unwrap_or(snapshot.clone())
    AhdapaGossipHooks::filter_outbound(peer, &mut payload)
    if payload == IdpCrdt::default(): skip (nothing new for this peer)
    encode Envelope{state: payload, is_delta, sender_gen: current_gen, request_delta_since}

// Phase 2: send to every peer concurrently (JoinSet)
for (peer, bytes) in outbound: spawn AhdapaTransport::send_recv(peer, bytes)

// Phase 3: process responses one at a time as they arrive
for (peer, result) in completed sends:
    on error: classify via AhdapaTransport::classify_error;
              if AuthRejected (peer's 401) and not in cooldown: on_auth_rejected(peer)
    on success: AhdapaCodec::decode → Envelope; filter_inbound; merge_and_persist
                (same path as the passive side — see Receiver logic above)

merge_and_persist (used by both the active round above and the passive handle_inbound) is the one place a merge actually lands: it merges into the shared IdpCrdt, and only if that merge actually changed something does it call AhdapaGossipHooks::on_merged, persist via AhdapaStorage::persist_full, and fire the changed notification gossip_await_client waits on. A genuinely redundant delivery (e.g. the same data learned from a different peer in a full-mesh cluster) triggers none of that.

AhdapaGossipHooks::on_merged does three things whenever a merge changes state, on whichever side (inbound or the active side’s post-response merge) triggered it:

  1. Rebuilds the ABAC policies (AppState::rebuild_abac_policies) — but only when inbound.hbac_rules actually contains live rules, not unconditionally every round like the old loop did. Likewise rebuilds the RFC 9396 policy (AppState::rebuild_rar_policy) when the inbound state carries any RAR rule content (including a deletion) or RAR types, and persists the HBAC/RAR snapshot when either rule set arrived.
  2. Evicts any cached signing key whose kid was tombstoned by this merge.
  3. If the peer’s wrapping_key_id differs from the local one, pulls the new key from them (pull_wrapping_key) — the same logic the old loop ran, but now triggered from on_merged on either side rather than only the active side’s response handling.

If a peer is unreachable, the error is recorded in the engine’s internal PeerBook and the round continues to the next peer. GossipConfig::backoff controls per-peer retry backoff (BackoffPolicy::None by default, matching the old loop’s “retry every interval, no backoff” behavior).

Periodic maintenance task

maintenance::spawn_maintenance_task (src/routes/gossip/maintenance.rs) runs independently of gossip rounds, on the same cadence the old loop’s “hourly” block used (interval_secs × rounds_per_hour, so effectively every ~1 hour at the default 5s interval):

loop:
    sleep(period)
    purge_expired_families(now)               // moved here from every gossip round --
                                                // see the note under CRDT primitives → LwwMap
    purge_old_tombstones / purge_old_revocations / purge_expired_access_token_revocations /
        purge_expired_upstream_tokens / saml2_sp_entries / saml2_idp_entries /
        federation_policies .purge_old_tombstones(tombstone_cutoff)
    cleanup_expired_families(db, now)          // DB-level purge, same six cleanup_* fns
    cleanup_old_tombstones(db, tombstone_cutoff)
    cleanup_old_crdt_revocations / cleanup_expired_at_revocations /
        cleanup_expired_upstream_tokens(db, ...)
    cleanup_expired_pending_saml2 / cleanup_expired_pending_saml2_sp /
        cleanup_expired_saml2_artifacts(db, now)
    jti_cache.retain(|_, exp| exp > now)

The topology refresh task runs an initial fetch immediately on startup (before the first gossip sleep) and then sleeps for ipa_topology_interval_secs (minimum 30 s, default 300 s). On LDAP error, the previous peer list is kept unchanged and a warning is logged.

After each successful topology fetch, if gss_initiator is available, the topology task also calls register_self_with_peer() for each newly-discovered peer whose KEM key is not yet in the CRDT. This pre-seeds the key via POST /api/gossip/register-kem with a Kerberos AP-REQ so that the legitimate node wins the OR-Map first-write-wins race before the first gossip round fires.

Cluster wrapping key

The 32-byte cluster wrapping key (used for session cookies) is stored node-locally in node_keys.wrapping_key_cms_der as a CMS EnvelopedData blob sealed to the node’s own ML-KEM-768 public key. It is never gossiped in plaintext or as a multi-recipient blob.

Only a short UUID string (wrapping_key_id) is gossiped in the CRDT. When a node observes a different UUID after a gossip merge, it fetches the actual key on demand:

GET /api/gossip/wrapping-key
X-Ahdapa-Node-Id: <requester's node_id>

Response: 200 OK
Content-Type: application/octet-stream
X-Ahdapa-Node-Id: <responder's node_id>
Body: SignedData(EnvelopedData) DER

The response is a full SignedData(EnvelopedData) blob produced by sign_and_seal(), sealed to exactly one recipient (the requester’s ML-KEM-768 public key) and signed with the responder’s ECDSA P-256 gossip signing key. The requester looks up the responder’s pinned signing key in the CRDT (from the X-Ahdapa-Node-Id header) and calls verify_and_open(). A response from a node with no pinned signing key is rejected. Confidentiality is ensured by the inner ML-KEM-768 encryption; integrity and sender authentication are ensured by the outer ECDSA P-256 signature.

Node statistics endpoint

GET /api/gossip/stats

Unauthenticated. Intentionally unauthenticated — like /api/gossip/kem-info — because the admin web UI fetches it before an admin session is established. What is exposed is aggregate counts and gossip health indicators; no key material, user data, or token content is returned.

Response body (JSON):

{
  "node_id":          "ipa1.example.com",
  "crdt_generation":  42,
  "counts": {
    "clients":          3,
    "signing_keys":     2,
    "cluster_nodes":    3,
    "refresh_families": 7,
    "revoked_sessions": 1,
    "scope_definitions": 8,
    "ipa_idp_overrides": 0
  },
  "peers":               ["https://ipa2.example.com/idp", "https://ipa3.example.com/idp"],
  "active_signing_kid":  "abc123",
  "kem_enrolled":        true,
  "gossip_signing_enrolled": true,
  "gossip": {
    "started_at":      1716000000,
    "rounds_completed": 12,
    "last_round_at":   1716000060,
    "peer_last_sync":  { "ipa2.example.com": 1716000058, "ipa3.example.com": 1716000059 },
    "persist_errors":  0,
    "wrapping_key_pull_errors": 0
  }
}

Field notes:

  • crdt_generation — current value of the CRDT_GENERATION atomic counter.
  • counts.* — live (non-tombstoned) entry counts for each CRDT collection.
  • peers — union of configured gossip.peers and topology-discovered peers.
  • active_signing_kid — the kid of the currently active JWT signing key.
  • kem_enrolled / gossip_signing_enrolled — whether both cryptographic identities are registered in the CRDT.
  • gossip.started_at — Unix timestamp when the gossip background task started.
  • gossip.rounds_completed — number of gossip rounds in which at least one peer was successfully synced. Idle rounds (CRDT unchanged, all pushes skipped) and rounds where all peers fail do not increment this counter.
  • gossip.last_round_at — Unix timestamp of the most recent round that synced at least one peer. null until the first successful sync.
  • gossip.peer_last_sync — Unix timestamp of the most recent successful inbound sync from each peer (recorded by the /api/gossip/sync receiver).
  • gossip.persist_errors — cumulative DB persist failures since startup (incremented after both inbound sync and outbound merge failures).
  • gossip.wrapping_key_pull_errors — cumulative failures to pull the cluster wrapping key from a peer after detecting a UUID change.

This endpoint is used by the admin web UI Cluster Nodes page to display per-node runtime gossip health alongside the static CRDT node entries from GET /api/admin/nodes.

Client convergence endpoint

GET /api/gossip/await-client?client_id=<id>&timeout_ms=<ms>

Unauthenticated. Long-poll endpoint that blocks until client_id appears in the local CRDT or timeout_ms elapses (default: 30000 ms). Uses GossipEngine::wait_for_change() to detect convergence without polling — it fires after any merge that actually changed state, on either the active or passive side (see Background loop); this replaced the old, ahdapa-owned crdt_changed Notify. When gossip is disabled or the engine hasn’t bootstrapped yet (state.gossip_engine is None), there is no change-notification source to await, so the endpoint falls back to a 200ms poll interval instead of busy-looping. Returns 200 OK when the client is found, 408 Request Timeout on timeout.

Primarily used by the ahdapa-bench converge scenario to measure gossip propagation time with sub-millisecond precision.

Kerberos KEM self-registration endpoint

POST /api/gossip/register-kem
Authorization: Negotiate <base64-AP-REQ>
Content-Type: application/json

{
  "node_id":                  "<hostname>",
  "kem_public_key_der":       "<base64url-ML-KEM-768-SPKI-DER>",
  "gossip_signing_pub_key_der": "<base64url-ECDSA-P256-SPKI-DER>"
}

Used by topology-discovered peers to seed both their ML-KEM-768 public key and their ECDSA P-256 gossip signing key before the first gossip round. All three fields are required; missing or empty fields return 400 Bad Request. The server:

  1. Returns 503 Service Unavailable if the GSSAPI server credential is unavailable (state.gss_cred is None — indicates a configuration or keytab problem).
  2. Calls try_spnego() to accept the Kerberos AP-REQ. Returns 401 Negotiate if absent, 401 if the token is invalid.
  3. Calls ServicePrincipal::parse() on the authenticated principal. Rejects with 403 if the principal is not HTTP/<host>@<REALM> (user principals and non-HTTP service types are excluded).
  4. Rejects with 403 if the principal’s realm does not match the expected realm — gossip.kerberos_realm when set, otherwise the local [server] realm. The check is always enforced, preventing cross-realm trust escalation.
  5. Checks that req.node_id.to_lowercase() == authed_host. Rejects with 403 if they differ — a machine can only register its own identity.
  6. Checks that authed_host is admitted: it must appear in the static allowed_node_ids list, in the topology-derived allowlist, or match the hostname of a static gossip.peers URL (static peers are implicitly trusted for KEM registration). Rejects with 403 if not admitted.
  7. Applies a three-case match on the existing CRDT entry for this node_id:
    • Insert-fresh: neither key known → insert NodeEntry with both keys.
    • Upsert-signing-key-only: KEM key present but gossip_signing_pub_key_der empty → update the entry to add the signing key.
    • No-op: both keys already present → return 200 OK immediately (idempotent).
  8. Returns 200 OK, optionally with a WWW-Authenticate: Negotiate <mutual-auth-token> header if GSSAPI produced a mutual-authentication output token.

At startup, bootstrap_wrapping_key() reads node_keys.wrapping_key_cms_der. If present, it decrypts the blob to recover the 32-byte key. If absent (first start), it generates a fresh key, seals it to the node’s own KEM key, and stores the result in node_keys. A UUID is generated and published to the CRDT as wrapping_key_id with timestamp=1 so that the established cluster’s UUID wins the LWW merge on the first gossip round.

When the cluster wrapping key is rotated via PUT /api/admin/keys/cluster, the node re-seals the new key to its own KEM key, stores it in node_keys, and updates crdt.wrapping_key_id to a new UUID. Peers detect the UUID change via gossip and pull the new key via the on-demand endpoint.

Convergence

ScenarioConvergence
Single nodeInstant (no peers)
Two-node cluster (KEM keys known)After 1 gossip round (≤ interval_secs seconds)
Three-node cluster, all connectedAfter 1–2 gossip rounds
Partition healed after T secondsAfter ≤ 2 gossip rounds from partition heal
New node joining (static peers)After 2 gossip rounds (learn KEM key → pull wrapping key via on-demand endpoint)
New node joining (IPA topology, gss_initiator set)After 1 gossip round — both the KEM key and the gossip signing key are pre-seeded via Kerberos register-kem before first gossip push; wrapping key pulled on first exchange. Requires both nodes to complete their mutual register-kem calls before the first gossip interval fires; this holds in practice because the topology refresh runs immediately on startup, before the first gossip sleep.

New signing key propagation: a key added on node A is available on node B after at most 1 gossip round from A to B. Resource servers should cache JWKS with a short TTL (≤ interval_secs × 2) to avoid key-not-found errors during propagation.

Security considerations

PropertyValue
ConfidentialityAES-256-GCM per-recipient (inner EnvelopedData)
IntegrityAES-256-GCM auth tag + ECDSA P-256 signature
Sender authenticationECDSA P-256 over eContent (outer SignedData); signing key pinned via register-kem before first gossip
Node admission controlallowed_node_ids allowlist (fail-closed on empty) + ReplicaScoped filtering restricting hbac_rules / rar_rules tags to the authenticated sender’s own replica
Replay preventionEnvelope.issued_at checked against now - tombstone_ttl_secs (AhdapaCodec::decode)
Post-quantumML-KEM-768 for key encapsulation (FIPS 203)
  • /api/gossip/sync SHOULD be firewalled to the cluster’s subnet as defense-in-depth. CMS encryption ensures confidentiality even if traffic is captured, but network isolation prevents unauthorized nodes from attempting to self-register.
  • The allowlist is fail-closed. When both the static allowed_node_ids list and the topology-derived allowlist are empty, no node can self-register via gossip or the wrapping-key endpoint. This is intentional: operators must either configure an explicit allowlist or enable ipa_topology so that hostnames are discovered automatically.
  • Gossip envelopes carry a timestamp (issued_at). Envelopes older than tombstone_ttl_secs (default 7 days) are rejected, preventing an attacker from replaying a captured gossip message after its tombstones have been GC-purged.
  • The gossip reqwest::Client has a 10-second request timeout. Slow peers do not block the gossip loop.
  • ECDSA P-256, not Ed25519, is used for gossip signing. OpenSSL’s CMS_sign() API requires a key type that has a default digest algorithm; Ed25519 (PureEdDSA) does not satisfy this requirement. P-256 provides the same 128-bit security level.
  • The JWT signing algorithm is configurable, not gossip signing. Each node generates its own JWT signing key pair (algorithm set by [server] jwt_signing_algorithm, default: ES256) stored in node_keys.jwt_signing_priv_der. The private key never leaves the node; only the public key is gossiped in SigningKeyEntry. This is distinct from the ECDSA P-256 gossip signing key.

Payload size and bandwidth

Raw field sizes

The binary fields that dominate gossip payload size (measured from a three-node demo cluster):

FieldBytesGossiped
ML-KEM-768 public key SPKI (NodeEntry.kem_public_key_der)1,206Yes
ECDSA P-256 gossip signing pub key SPKI (NodeEntry.gossip_signing_pub_key_der)91Yes
JWT signing private key DER (SigningKeyEntry.private_key_der)varies by algorithmNo — #[serde(skip_serializing)]; stays in node_keys
JWT signing public key SPKI DER (SigningKeyEntry.public_key_der)varies by algorithmYes
ECDSA P-256 gossip signing certificate (node_keys.signing_certificate_der)291No — local only
ML-KEM-768 private key PKCS#8 (node_keys.private_key_der)2,498No — local only

The ML-KEM-768 public key is the dominant field by a factor of ~13× over the next largest gossiped value.

Per-entity CBOR contribution

The CRDT is serialised as CBOR (ciborium). CBOR stores binary fields as raw bytes (no base64 overhead). Approximate CBOR size per entry:

Entity~CBOR bytesDominant field
NodeEntry (one cluster node)~1,530 BML-KEM-768 pub key (1,206 B raw)
SigningKeyEntry (one JWT signing key)~100 B (ES256) – ~2,600 B (ML-DSA-87)JWT public key (size varies by algorithm); private key not gossiped
ClientEntry (typical OAuth2 client)~150 BUUIDs + scopes; short serde field names (2 chars) keep the CBOR compact
RefreshFamilyState (one active session)~100 BUUIDs + counters

CMS envelope overhead

Each gossip message is sent to exactly one peer, so the CMS overhead is constant regardless of cluster size:

LayerBytes
Outer SignedData headers + ECDSA P-256 signature (64 B) + signer cert (291 B)~555 B
Inner EnvelopedData KEMRecipientInfo: ML-KEM-768 ct (1,088 B) + wrapped CEK (40 B) + headers~1,233 B
AEAD overhead (12 B nonce + 16 B GCM tag)28 B
Fixed CMS overhead per gossip message~1,816 B

Wire size per gossip push

Total CMS-encrypted wire size for one outbound push (one recipient). The cluster wrapping key blob is no longer in the gossip body; only its UUID is gossiped:

ScenarioWire bytes
3 nodes, 3 signing keys, 4 clients, 0 sessions (demo measured)~7,454 B
3 nodes, 5 clients, 50 sessions~8 KB
5 nodes, 5 clients, 50 sessions~11 KB
10 nodes, 10 clients, 100 sessions~19 KB

Bandwidth per gossip cycle

The topology is full-mesh: each node pushes to every configured peer and receives a response. Total cluster bandwidth per cycle (worst case, full-state) = N × (N−1) × 2 × wire_bytes. In practice, delta exchange reduces per-push payload to the size of changed entries only.

These figures are theoretical maximums — they assume every gossip round produces an actual push. In practice, the generation-skip optimisation suppresses pushes when the CRDT has not changed since the last successful round. In the demo cluster (active token issuance, no schema mutations), 93% of rounds were skipped, reducing steady-state bandwidth to near zero. Pushes happen only when the CRDT actually changes (client creation/deletion, key rotation, node join/leave).

At interval_secs = 2 (demo default; config default is 5 s):

NodesWire/msgPer round (2 s)Per hour (theoretical max)
27.5 KB60 KB108 MB
37.5 KB90 KB162 MB
510.5 KB420 KB756 MB
1019 KB3.4 MB6.1 GB

The O(N²) topology is practical for the expected deployment range of 2–5 nodes. Above ~10 nodes the bandwidth cost becomes significant and a partial-mesh peer configuration (each node lists only a subset of peers) should be considered.

Marginal cost per added entity

Measured at a three-node baseline:

ChangeExtra bytes per gossip message
+1 cluster node~+1,530 B (NodeEntry: ML-KEM-768 pub key 1,206 B + other fields; CBOR-encoded, no wrapping key blob)
+1 removed node (tombstone)+~220 B (tombstone metadata; key fields absent)
+1 OAuth2 client+~150 B
+1 active refresh token family (session)+~100 B
+1 JWT signing key (rotation)+~100 B (public key only; private key not gossiped)

Sessions and clients are cheap. Nodes are the dominant cost because every node contributes 1,206 B of ML-KEM-768 public key material (gossiped in NodeEntry). The previous per-node cost of +1,644 B from the cluster wrapping key CMS blob is eliminated — the wrapping key is no longer gossiped.

Tombstone accumulation and GC

Removed nodes, deleted clients, and revoked signing keys leave OrMap tombstones. Each tombstone adds ~220 B to gossip messages until it is garbage-collected. Tombstones older than gossip.tombstone_ttl_secs (default: 7 days) are purged from both the in-memory CRDT and the database approximately once per hour. This bounds tombstone growth even in high-churn deployments.

The TTL must exceed the longest expected node downtime: a node that is offline longer than the TTL may re-gossip entries that were since deleted (those entries would be re-merged on reconnection). The default 7-day TTL is conservative and suitable for most deployments.

Configuration

KeyDefaultDescription
peers[]Peer node base URLs. Gossip is disabled when this list is empty and ipa_topology is false.
interval_secs5Push interval in seconds.
allowed_node_ids[]Allowlist of node_id values permitted to self-register. The allowlist is always enforced — an empty union of both the static list and the topology-derived list admits nobody (fail-closed). When ipa_topology = true, discovered replica hostnames are appended automatically.
tombstone_ttl_secs604800Seconds to retain OR-Map tombstones before GC. Also the maximum age of accepted gossip envelopes (issued_at window). Must exceed the longest expected node downtime. Default: 7 days.
ipa_topologyfalseWhen true, a background task (src/topology.rs) queries cn=topology,cn=ipa,cn=etc,<suffix> for ipaReplTopoSegment entries and derives gossip peer URLs of the form https://<hostname><base_path>. The peer list is stored in AppState::dynamic_peers and the allowlist in AppState::dynamic_allowed_nodes; both are merged with any statically configured peers and allowed_node_ids at each gossip round.
ipa_topology_interval_secs300How often (seconds) to re-query the IPA topology. Only used when ipa_topology = true. Minimum: 30 s.
kerberos_realm—Expected Kerberos realm for register-kem callers (e.g. "IPA.EXAMPLE.COM"). When set, principals whose realm does not match are rejected with 403, preventing cross-realm trust escalation. When unset, the local [server] realm is used as the expected realm, so cross-realm principals are always rejected.
[gossip]
peers              = ["https://node2.example.com:8080", "https://node3.example.com:8080"]
interval_secs      = 5
allowed_node_ids   = ["node1.example.com", "node2.example.com", "node3.example.com"]
tombstone_ttl_secs = 604800  # 7 days

For IPA-integrated deployments, the static peers list can be omitted entirely when ipa_topology = true. A single-node deployment requires no gossip configuration at all.