Skip to content

HDDS-16723. Drop queued ICRs for re-registering endpoint to avoid oversized heartbeat. - #11423

Draft
hani-fouladgar wants to merge 1 commit into
apache:masterfrom
hani-fouladgar:HDDS-16723
Draft

hani-fouladgar wants to merge 1 commit into
apache:masterfrom
hani-fouladgar:HDDS-16723

Conversation

@hani-fouladgar

@hani-fouladgar hani-fouladgar commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this pull request?

When an SCM sends a re-register command to a datanode, HeartbeatEndpointTask.processReregisterCommand() only transitions the endpoint state (HEARTBEAT → GETVERSION); it does not clear the accumulated incremental container reports (ICRs) for that endpoint. Those ICRs remain in StateContext.incrementalReportsQueue, keyed per endpoint. On the next successful heartbeat, StateContext.getAllAvailableReports() returns the refreshed full container report together with the entire ICR backlog, producing an oversized heartbeat message.

Re-registration (RegisterEndpointTask) already resends a fresh full container report reflecting the datanode's complete current state, so the ICRs queued for the re-registering endpoint are redundant and only serve to bloat the message.

This PR:

  • Adds StateContext.clearIncrementalContainerReports(HostAndPort endpoint) — a scoped, synchronized method that drops the queued ICRs for a single endpoint.
  • Calls it from processReregisterCommand() for the re-registering endpoint only.

Design notes:

  • Scoped to the re-registering endpoint. ICRs are queued and drained per endpoint, so clearing all endpoints would discard reports still owed to other, healthy SCMs. The clear touches only the endpoint that asked to re-register.
  • Only IncrementalContainerReportProto entries are removed. The incremental queue also holds CommandStatusReportsProto (and pipeline reports). Registration regenerates the full container/pipeline/node state, but it does not regenerate command status reports, so those are intentionally left in place. This mirrors the existing getFullContainerReportDiscardPendingICR() precedent, which also filters by type rather than clearing the whole queue.
  • The full report state is left untouched, since registration regenerat

What is the link to the Apache JIRA

https://issues.apache.org/jira/browse/HDDS-16723

If you do not have an ASF Jira account yet, please follow the first-time contributor
instructions in the Jira guideline.

(Please replace this section with the link to the Apache JIRA)

How was this patch tested?

  • Added TestStateContext#testClearIncrementalContainerReports: queues a mix of ICRs and command-status reports on two endpoints, clears one endpoint, and asserts that endpoint retains only its command-status reports (ICRs dropped) while the other endpoint keeps both — verifying both the type filtering and the per-endpoint scoping.
  • Ran the full TestStateContext suite.

@hani-fouladgar
hani-fouladgar marked this pull request as draft October 6, 2026 21:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant