Conversation
|
|
||
| assertEquals(1, datanodeDetails.size()); | ||
| // After both racks fail, placement falls back to rack2, the only rack with usable nodes. | ||
| assertTrue(cluster.isSameParent(datanodes.get(10), datanodeDetails.get(0))); |
There was a problem hiding this comment.
Could we also assert metrics.getDatanodeChooseFallbackCount() is 1 here and 0 in the preceding test, since trying another affinity rack should not count as a fallback?
There was a problem hiding this comment.
Thanks @peterxcli for working on this.
I have a question regarding the fallback behavior when both affinityNodes and excludedNodes are not null, with fallback = true, see following:
In the method entry, we have the logic:
// When affinity node is null, in this case new node to be selected
// should be in different rack than used nodes rack.
// Exclude nodes should be just excluded from topology node selection,
// which is filled in excludedNodesForCapacity
// Used node rack should not be part of rack selection
// which is filled in excludedNodes
if (affinityNodes == null && excludedNodes != null) {
for (DatanodeDetails node : excludedNodes) {
excludedNodesForCapacity.add(node.getNetworkFullPath());
}
excludedNodes = usedNodes;
}Since affinityNodes != null initially, this block is skipped, leaving excludedNodes as the original excluded list .
When all tries in Step 1 fail, we proceed to Step 2:
steps.add(new Step(null, RACK_LEVEL, affinityNodes != null));
Here step.affinityNode is null and step.ancestorGen is RACK_LEVEL. In this step:
- networkTopology.chooseRandom will exclude entire racks of the nodes in excludedNodes.
- However, because excludedNodes was never swapped with usedNodes, the topology will exclude the racks containing the original excludedNodes, rather than excluding the racks of usedNodes.
Should we perform the same swapping/handling between excludedNodes and usedNodes when transitioning to the fallback step?
What changes were proposed in this pull request?
When a container needs a new replica,
SCMContainerPlacementRackAwareoften tries to put it on the same rack as one of the existing replicas. It picks a few random nodes from that rack, and if none of them can be used after 3 tries, it gives up on the whole placement.The catch is that decommissioned and in-maintenance datanodes are still part of the network topology. So if the first rack it tries has only such nodes left, all 3 tries are wasted there, and SCM never even looks at the racks of the other replicas or at the fallback. The container stays under-replicated, and a decommission waiting for it never finishes. We found this with the SCM simulation (HDDS-16627).
With this change:
chooseNodeis rewritten as a short list of places to try, in order, instead of one loop that juggled several counters. This also makes the rule above easy to see.However, with the fix above, placement reaches the fallback much more often, and that exposed an older problem in it (found by @chungen0126 in review). The fallback stayed off the racks of the excluded nodes instead of the racks of the existing replicas. So an excluded node, usually a replica being decommissioned, ruled out its whole rack, while the racks that had just failed were tried again. The fallback now stays off the racks of the existing replicas.
What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16628
How was this patch tested?
Three new tests in
TestSCMContainerPlacementRackAware: