Repository navigation
HDDS-16723. Drop queued ICRs for re-registering endpoint to avoid oversized heartbeat. - #11423
Draft
hani-fouladgar wants to merge 1 commit into
Draft
hani-fouladgar wants to merge 1 commit into
hani-fouladgar wants to merge 1 commit into
Conversation
…rsized heartbeat.
hani-fouladgar
marked this pull request as draft
October 6, 2026 21:43
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
When an SCM sends a re-register command to a datanode,
HeartbeatEndpointTask.processReregisterCommand()only transitions the endpoint state (HEARTBEAT→GETVERSION); it does not clear the accumulated incremental container reports (ICRs) for that endpoint. Those ICRs remain inStateContext.incrementalReportsQueue, keyed per endpoint. On the next successful heartbeat,StateContext.getAllAvailableReports()returns the refreshed full container report together with the entire ICR backlog, producing an oversized heartbeat message.Re-registration (
RegisterEndpointTask) already resends a fresh full container report reflecting the datanode's complete current state, so the ICRs queued for the re-registering endpoint are redundant and only serve to bloat the message.This PR:
StateContext.clearIncrementalContainerReports(HostAndPort endpoint)— a scoped, synchronized method that drops the queued ICRs for a single endpoint.processReregisterCommand()for the re-registering endpoint only.Design notes:
IncrementalContainerReportProtoentries are removed. The incremental queue also holdsCommandStatusReportsProto(and pipeline reports). Registration regenerates the full container/pipeline/node state, but it does not regenerate command status reports, so those are intentionally left in place. This mirrors the existinggetFullContainerReportDiscardPendingICR()precedent, which also filters by type rather than clearing the whole queue.What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-16723
If you do not have an ASF Jira account yet, please follow the first-time contributor
instructions in the Jira guideline.
(Please replace this section with the link to the Apache JIRA)
How was this patch tested?
TestStateContext#testClearIncrementalContainerReports: queues a mix of ICRs and command-status reports on two endpoints, clears one endpoint, and asserts that endpoint retains only its command-status reports (ICRs dropped) while the other endpoint keeps both — verifying both the type filtering and the per-endpoint scoping.TestStateContextsuite.