fleet: make the hub mirror publish non-blocking (bounded hung-hub tail latency) #11
Labels
No labels
correctness
coverage
milestone:M4
polish
security
tech-debt
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
public/warden#11
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
From the fleet whole-branch review (v0.22.0).
hub_mirrorpublishes to the local bus first (unguarded) then best-effort to the hub;NatsConn.publishblocks on.result(timeout=5). If the hub connection is HUNG (not cleanly failed), the hub publish can add up to a bounded ~5s to the TAIL of the transition/notify call before being swallowed. Local publish + local incident persistence run first and are byte-for-byte unaffected (invariant holds), and nats_pub is the last notifier in the composite — so this is bounded + local-safe, not a correctness bug. Follow-up: make the hub mirror fire-and-forget / async (or a shorter hub-publish timeout) so a hung hub adds ~0 to the local notify tail. Non-blocking; watch only if hub-network hangs appear in prod.Fixed in v0.24.1. The fleet hub mirror hop is now fire-and-forget: NatsConn.publish_nowait submits the hub publish to the loop via run_coroutine_threadsafe WITHOUT .result(), so a hung/slow hub can no longer add the ~5s tail to local incident processing; failures log via a done-callback, never raised. The LOCAL publish stays the blocking, ordered path (byte-for-byte unchanged). Whole-branch review: APPROVE.