Skip to content

Recover promptly from a cluster failover; fix the configuration channel under RESP3 and with a channel prefix (#3254) - #3255

Merged
mgravell merged 4 commits into
mainfrom
marc/cluster-failover
Oct 1, 2026
Merged

mgravell merged 4 commits into
mainfrom
marc/cluster-failover

Conversation

@mgravell

Copy link
Copy Markdown
Collaborator

Addresses #3254 (the cluster failover outage and the config-channel problems). Three commits, meant to be read separately.

1. Repair primary/replica roles after a cluster failover

After CLUSTER FAILOVER, writes failed with Command cannot be issued to a replica until the next configCheckSeconds tick. Reproduced against a local 6-node OSS cluster (three primaries, three replicas), failing over all three replicas under a write loop:

failure window failed writes
before ~55s ~1080
after none 0

Same result under RESP2 and RESP3. What was happening (measured, not inferred):

  • The first write to each old primary got -MOVED, so the slot map was updated. The resend to the new primary was then refused client-side, because that node was still flagged as a replica.
  • From then on, writes to those slots were refused before being sent. No -MOVED comes back, so nothing triggers a refresh.
  • Reconfigures did run (about every 6s), but they re-read the slot map and trusted the cached roles, so they changed nothing. The role only refreshed from the periodic INFO replication check, which is the configCheckSeconds delay.

Two changes:

  • On -MOVED (not -ASK), a target still flagged as a replica is flipped to primary before the resend. Only a primary can own a slot. Cluster mode only.
  • When a cluster topology is applied, each known node's role is set from it. CLUSTER SLOTS wins where it lists the node, as it does for the slot map, and CLUSTER NODES covers nodes that serve no slots. It never creates a server.

Trade-off: a lagging node's view can briefly flip a role. The next topology reply or the INFO replication check corrects it.

The in-process test server gains Failover(), and READONLY/READWRITE (a promoted replica's connection has to leave read mode).

2. Subscribe to the configuration channel under RESP3

The channel was only subscribed in the subscription connection's handshake, and RESP3 has no such connection, so under RESP3 nothing listened (PUBSUB NUMSUB: 0 on every node). The manual PUBLISH __Booksleeve_MasterChanged pattern therefore did nothing.

It now subscribes on the interactive connection once that connection is known to be RESP3. Deliberately not in the handshake: a connection that fell back to RESP2 must not be put into subscriber mode.

3. Apply the channel prefix when the library fires the channel

The channel is subscribed with ChannelPrefix applied, like any pub/sub channel, and the public PublishReconfigure publishes it that way. ReplicaOf/ReplicaOfAsync and MakePrimaryAsync published it as a raw value instead, so with a prefix their broadcasts reached nobody. They now publish it as a channel.

For comparison I ran 2.9.25 (published 2025-09-29): it also subscribed with the prefix, but its receive side compared the prefixed wire name with the unprefixed setting, so with a prefix no broadcast was ever acted on, however it was sent. Prefixed broadcasts are acted on now.

The prefix is kept, since it is long-standing behaviour, and is now documented in docs/Configuration.md: a manual PUBLISH must use the prefixed name (once per distinct prefix), and an ACL must grant the prefixed channel.

Testing

  • New unit tests for each change, each run under RESP2 and RESP3. I checked that each fails with its own fix reverted.
  • All 1971 unit-style tests pass on net10.0, and Build.csproj -c Release /p:CI=true builds clean.
  • The integration suite has not been run against the shared docker topology.
  • MakePrimaryAsync's broadcast is covered by a manual probe against a real Redis but has no unit test (the fake server has no replication). ReplicaOfAsync is unit tested.

Not covered by this PR

Other points raised in #3254 that this does not change: use of INFO to resolve a changed cluster, the same node being tracked under a DNS name and an IP, and documentation on controlled downtime for maintenance.

After a failover the slot map can be correct while the role flag on each
node is stale. A write routed to a node still believed to be a replica is
refused client-side, so no -MOVED comes back and nothing prompts a refresh;
writes then fail until the next scheduled INFO replication check, i.e. up
to configCheckSeconds. Reconfigures during that window ran but changed
nothing, as they re-read the slot map and trusted the cached roles.

- On -MOVED, a target still flagged as a replica is corrected before the
  resend: only a primary can own a slot, and the resend would otherwise be
  refused as a write to a replica. Cluster mode only, and not for -ASK.
- When cluster topology is applied, set each known node's role from it.
  CLUSTER SLOTS wins where it lists the node, as it does for the slot map,
  with CLUSTER NODES covering nodes that serve no slots. Never creates a
  server.

Measured against a local 6-node OSS cluster, failing over all three
replicas under a write loop: ~55s of failures before, none after, under
both RESP2 and RESP3.

The in-process test server gains Failover(), and READONLY/READWRITE so a
promoted replica's connection can leave read mode.
The channel (__Booksleeve_MasterChanged by default) is how a client is told
by hand that the topology has moved: a PUBLISH to it makes every subscribed
client refresh. It was only ever subscribed in the subscription connection's
handshake, and RESP3 has no such connection, so under RESP3 nothing was
listening and the broadcast reached nobody (PUBSUB NUMSUB: 0 on every node).

Subscribe on the interactive connection once it is known to be RESP3. That
is deliberately not part of the handshake: a connection that fell back to
RESP2 must not be put into subscriber mode.

The channelPrefix behaviour of this channel is unchanged.
…nnel (#3254)

The configuration channel is subscribed with ChannelPrefix applied, as any
pub/sub channel is, and the public PublishReconfigure publishes it that
way too. The library's own broadcasts - after ReplicaOf, and from
MakePrimary - passed the name as a raw value instead, so with a prefix they
were published to a name nobody was subscribed to and reached no one.

Publish them as a channel so the prefix applies consistently, and document
that it does: a manual PUBLISH has to use the prefixed name, and an ACL
needs to grant it.

Measured against 2.9.25 (published 2025-09-29) for comparison: it also
subscribed with the prefix, and its receive side compared the prefixed wire
name with the unprefixed setting, so with a prefix no broadcast was ever
acted on, whichever way it was sent. Prefixed broadcasts are acted on now.
Under RESP3 the interactive connection now holds the configuration-channel
subscription, and the server counts any subscribed connection as a pubsub
client rather than a normal one. ClientKill filtered on TYPE normal, so under
RESP3 it matched nothing and killed nothing.

The test now looks the target up in CLIENT LIST and asserts both its protocol
and the client type that follows from it (pubsub under RESP3, normal under
RESP2), then kills by that type.
@mgravell
mgravell force-pushed the marc/cluster-failover branch from cad204c to 9332fb9 Compare October 1, 2026 13:51
@mgravell
mgravell merged commit 64e2d15 into main Oct 1, 2026
7 of 8 checks passed
@mgravell
mgravell deleted the marc/cluster-failover branch October 1, 2026 15:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant