Conversation
# Conflicts: # hadoop-hdds/server-scm/src/main/java/org/apache/hadoop/hdds/scm/container/placement/algorithms/SCMContainerPlacementRackAware.java
# Conflicts: # hadoop-hdds/server-scm/src/main/java/org/apache/hadoop/hdds/scm/container/placement/algorithms/SCMContainerPlacementRackAware.java
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
This builds on #11403 (HDDS-16628) and includes its commits. To see only what this PR adds, use this comparison.
Take a cluster with two racks and decommission one whole rack. Every container with a replica on that rack needs a new copy, and
SCMContainerPlacementRackAwarewants that copy on a different rack from the existing replicas. The only other rack is the one going away, so placement fails, the containers stay under-replicated, and the decommission never finishes.This PR makes two changes:
MisReplicationHandlernow checks whether copying the container to the chosen node would actually improve its placement. A copy to a rack that already has a replica doesn't: the container becomes over-replicated, the extra copy is deleted, it is mis-replicated again, and the cycle starts over. In that case the handler now skips the copy.We found this with the SCM simulation (HDDS-16627).
What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16629
How was this patch tested?
TestSCMContainerPlacementRackAware: when the other rack is being decommissioned, the new replica goes to the same rack, and without fallback, placement fails.TestRatisMisReplicationHandler: no copy is sent when it would not improve the placement.