Coordinate only the invariant
Why unnecessary agreement limits distributed-system throughput and write availability.
Key points
Coordination protects invariants that concurrent local decisions could violate. Raft makes one ordering boundary explicit through a replicated log and majority commitment [1]. That shared path limits horizontal scale and needs enough available participants and links.
Commutative operations can often execute locally and reconcile later. Deterministic merge rules let replicas converge without one execution order, so independent nodes can remain writable across a partition [2].
The claim is bounded: for this independent-increment task, five equal-capacity nodes acknowledge convergent local operations at up to five times the simplified Raft rate. The contracts differ. Raft acknowledges a majority-committed prefix. The others acknowledge divergent local state that can be lost before replication.
Confluence before distribution
A term-rewriting system makes confluence visible. It can choose different valid redexes, but each path remains joinable at one normal form.
The viewer starts with neg(neg(add(a,b))). Path A removes the outer double negation. Path B distributes both negations before eliminating them. Step both paths to verify that each reaches add(a,b).
Loading the interactive explanation. The complete static explanation follows.
| Rule | Rewrite |
|---|---|
| Double negation | neg(neg(x)) → x |
| Distribute negation | neg(add(x,y)) → add(neg(x),neg(y)) |
| Join criterion | Both paths reach the same add(a,b) normal form |
The invariant chooses the mechanism
Start with the observable result, not with a protocol.
The task accepts uniquely identified increments at one origin. After synchronization, every replica must hold the admitted-operation count. It needs no order and excludes decrements, transfers, limits, and exactly-once effects.
That narrow contract permits two coordination-free constructions:
- Operation-ledger reconciliation. Nodes record operation IDs and merge by set union. Union is associative, commutative, and idempotent. Duplicates do not change the result, but metadata grows with history.
- Monotonic G-Counter. A node increments its component. Replicas merge components with
max, then sum them. This compact form assumes exactly-once delivery to one origin and cannot represent a decrement.
The G-Counter is also a convergent system. Monotonicity is a way to obtain confluence, not a competing category [3].
From joinability to an operation ledger
An operation ledger treats each accepted ID as a fact. Replicas can learn facts in different orders. Set union joins those paths and deduplicates repeated IDs. Partition the replicas, append IDs, then heal or exchange them.
Loading the interactive explanation. The complete static explanation follows.
| Action | Replica effect | Merge effect |
|---|---|---|
Append A:1 at A | A records one new identity | B learns it on union |
Retry A:1 | The set remains unchanged | Idempotence prevents a duplicate count |
| Append on both sides while partitioned | Views diverge safely | Union contains both identities after heal |
The ledger retains identity for deduplication. The smaller G-Counter requires a stronger exactly-once origin assumption.
The strongest alternative is Raft. Its leader acknowledges an ordered log entry after majority storage. This is the right boundary for one authoritative order, linearizable reads, or rejection of conflicting visible effects.
How Raft restores one leader
Raft uses terms and majority votes to select one log leader. Isolate that leader, then step the election. The majority advances the term, elects a leader, and receives its heartbeat. After healing, the isolated node recognizes the higher term.
Loading the interactive explanation. The complete static explanation follows.
| Election stage | Majority component | Isolated former leader |
|---|---|---|
| Timeout | Stops receiving a valid heartbeat | Cannot reach a majority |
| Candidate | One node increments the term and requests votes | Retains stale local state |
| Elected | At least three nodes recognize the new leader | Still cannot commit |
| Heal | The new leader sends its higher term | Steps down and catches up |
Production Raft also handles log matching, persisted votes, retries, snapshots, and membership changes.
Why coordination limits scale and write availability
A coordination point is a shared service center. Followers improve fault tolerance, but they do not automatically increase one leader’s ordering rate.
A five-node Raft group needs three communicating members to commit. A minority can preserve state but cannot acknowledge a committed entry. After a leader partition, the majority waits for an election. This safety rule limits availability.
Convergent updates move that boundary. Coordination-avoidance analysis makes the invariant-specific trade explicit [4]. Reachable nodes can acknowledge local durability while partitions create valid intermediate views. Eventual communication and finite load are needed for convergence.
Interactive five-node model
The viewer has five nodes, A through E. Each models 60 local operations per second against 300 distributed increments per second. A tick is 100 milliseconds, and the chart keeps 60 ticks. The ledger compacts contiguous per-origin IDs into prefixes while preserving counts and duplicate rejection.
The operation ledger and G-Counter use one origin slot per local acknowledgment. Raft routes proposals through one leader, limiting admission to 60 entries per modeled second. Majority commit appears after two ticks. An isolated minority leader triggers a deterministic majority election after ten ticks.
The shared topology control applies one partition to all mechanisms. PARTITION / AB | CDE leaves the initial Raft leader in the minority; local acknowledgment continues in both convergent components while CDE elects a Raft leader.
Loading the interactive comparison. The exact static assumptions and comparison follow.
Reading the comparison
The healthy default has a simple upper bound:
| Mechanism | Acknowledgment | Modeled upper bound |
|---|---|---|
| Operation ledger | Durable local operation ID | 5 × 60 = 300 operations/s |
| G-Counter | Durable local origin component | 5 × 60 = 300 increments/s |
| Simplified Raft | Majority-committed ordered entry | 1 × 60 = 60 increments/s |
The five-to-one ratio follows from this model, not a universal estimate. At 60 operations per second or less, all three keep up after initial latency. Raft batching can increase throughput. A workload that needs ordering changes the decision even if that path is slower.
The acknowledgments are not equivalent. Local state can be lost with its only node before synchronization; a majority commit has a stronger durability boundary.
Convergence is a whole-system property
Convergence is a whole-system property, including all systems it impacts. Merged storage is not equivalent if replicas sent different emails, duplicated a payment, or issued conflicting actuator commands.
The merge rule must cover the externally observable outcome. Use coordination, escrowed rights, an idempotency boundary, or compensation when an affected system cannot merge the effect safely. Keep the coordinated scope as small as the invariant permits.
When Raft is the right answer
Choose the coordinated log when the system must provide one of these properties:
- one authoritative write order;
- linearizable reads and writes;
- a decision that concurrent actors must not both accept;
- a majority-durable acknowledgment before success is exposed;
- a compact operational model whose stronger semantics are worth the leader and quorum dependency.
Choose a convergent representation only when its merge rule preserves the actual invariant. Define admissible intermediate divergence, retry behavior, deletion behavior, convergence assumptions, and the fate of effects already exposed.
The design question is not “Can consensus solve this?” It usually can. The question is “Which invariant earns the coordination cost?”
Model limitations and references
Publisher: Cabrillo Coast. Revision: 1. Publication date: 2026-09-26. Correction contact: hello@cabrillocoast.com.
The deterministic viewer does not measure this device. It excludes Byzantine faults, correlated loss, clock uncertainty, membership change, snapshots, and log repair. Corrections update this source and revision.
| [1] | D. Ongaro and J. Ousterhout, “In Search of an Understandable Consensus Algorithm,” in Proc. USENIX ATC, 2014, pp. 305–319. [Online]. Available: PDF |
|---|---|
| [2] | M. Shapiro, N. Preguiça, C. Baquero, and M. Zawirski, “Conflict-Free Replicated Data Types,” in Stabilization, Safety, and Security of Distributed Systems, LNCS 6976, 2011, pp. 386–400, doi: 10.1007/978-3-642-24550-3_29. |
| [3] | P. Alvaro, N. Conway, J. M. Hellerstein, and W. R. Marczak, “Consistency Analysis in Bloom: A CALM and Collected Approach,” Univ. California, Berkeley, Tech. Rep. UCB/EECS-2011-117, Oct. 17, 2011. [Online]. Available: PDF |
| [4] | P. Bailis, A. Fekete, J. M. Hellerstein, A. Ghodsi, and I. Stoica, “Coordination Avoidance in Database Systems,” Proc. VLDB Endow., vol. 8, no. 3, pp. 185–196, 2014, doi: 10.14778/2735508.2735509. |