Skip to content

CAS: existsDirectory on a part file and listDirectory on a table dir issue an S3 LIST; a restart makes 77k LISTs and gets 503 Slow Down #2439

Description

@filimonov

Describe the situation

On a CAS disk, two directory operations of the generic MergeTree code are answered with an S3 LIST of the namespace's table-level files prefix. At startup this produces one LIST per part file, S3 throttles, and startup takes over two minutes. In steady state the same path costs one LIST per table per minute.

Environment

  • Version: 26.6.4.20001.altinityantalya
  • 2 replicas, CAS disk cas with a cache disk on top, used as the default disk
  • 26 tables, 288 active and 1,384 outdated parts on the CAS disk
  • Backend: AWS S3

Observed

Restart on 2026-09-25 15:38 (the restart of 2026-09-24 16:41 shows the same 61k-LIST burst):

minute S3ListObjects CASRootList 503 throttling events
15:40 19,020 18,964 54
15:41 39,945 39,704 337
15:42 18,223 18,225 69

537 AWSClient: Response status: 503, Slow Down lines. 139 uploads failed with Please reduce your request rate, 108 of them ref-lane writes. Ready for connections 2 min 11 s after Starting ClickHouse.

Steady state: CASRootList is 155 per 10 minutes at all times, 22k per day, with 26 tables.

Call chains (from system.trace_log, trace_type = 'Real')

Startup, 111 of 116 sampled LIST stacks:

IMergeTreeDataPart::loadColumnsChecksumsIndexes
  -> IMergeTreeDataPart::checkConsistency
  -> MergeTreeDataPartChecksum::checkSize            // calls existsDirectory(name) for every checksum entry
  -> DataPartStorageOnDiskFull::existsDirectory
  -> ContentAddressedMetadataStorage::existsDirectory
  -> CasPlainObjects::listNamespaceFiles             // S3 LIST of roots/<ns>/files/
  -> ObjectStorageBackend::list

Steady state, all non-GC LIST stacks:

MergeTreeData::clearOldTemporaryDirectories
  -> ContentAddressedMetadataStorage::iterateDirectory(<table dir>)
  -> ContentAddressedMetadataStorage::listDirectory   // TableDir shape
  -> CasPlainObjects::listNamespaceFiles              // S3 LIST of roots/<ns>/files/

Root cause

  1. classifyDirectory (ContentAddressedMetadataStorage.cpp) has no shape for a file inside a part. For <table>/<part>/<file> the part branch returns only when the file is a projection directory. Otherwise the path falls through to parseTableFilePath, which matches on the table uuid, and the path is classified as TableSubdir. The TableSubdir branch of existsDirectory calls listNamespaceFiles and scans the result. Upstream checkSize asks existsDirectory for every checksum entry, so every part file at load costs one LIST. 1,672 parts × ~40 files ≈ 67k LISTs.

  2. listDirectory for a table directory merges two sources: part names from listRefs(ns) (the in-memory ref table, no request) and table-level verbatim files (format_version.txt, mutation_*.txt, deduplication_logs/...). The verbatim files are plain objects under roots/<ns>/files/<name> with no index in manifests or in _ckpt, so the only way to enumerate them is a LIST. clearOldTemporaryDirectories calls this once per table per minute.

Proposal

  1. A PartFile shape in classifyDirectory. A path <table>/<part>/<file> that is not a projection directory gets its own shape. existsDirectory answers it from the part-folder view with view->hasDirectory(file), the same call the ProjectionDir branch already uses. No LIST, one cached manifest read per part. This removes the startup burst.

  2. Cache the names of table-level files that live outside the ref table. Keep the file-name set per namespace life in memory: one LIST on first use after start, then write-through invalidation in putNamespaceFile and removeNamespaceFile. This is sound because the namespace includes server_root_id, so only this node writes these files. This removes the steady-state LISTs from clearOldTemporaryDirectories and any other listDirectory on a table dir.

  3. Consider moving table-level files into the ref table. Today they are the only per-table objects with no index. If their names (or the files themselves) were recorded in the ref table next to the part refs, a cold start would need no LIST either and listDirectory would be a single in-memory merge. This is a format change and is listed here as a design question, not as part of the fix.

Acceptance: a restart of a node with ~1,700 parts issues on the order of one LIST per table, not per part file, and gets no 503; CASRootList in steady state is zero with no CREATE/ALTER activity.

Reproduce

  1. CAS disk as the default disk, a few tables with a few hundred parts each.
  2. Restart the server.
  3. SELECT sum(ProfileEvent_S3ListObjects), sum(ProfileEvent_CASRootList) FROM system.metric_log WHERE event_time > <start> during the first three minutes.
  4. Sample system.trace_log with trace_type = 'Real' and filter symbols for listNamespaceFiles.

Related: #2429

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions