Skip to content

skill(apm-integrations): database-category rule sharpening (SPI-first, eager connect metadata, muzzle) - #12114

Open
jordan-wong wants to merge 2 commits into
masterfrom
skill/database-category-rules-20260730
Open

skill(apm-integrations): database-category rule sharpening (SPI-first, eager connect metadata, muzzle)#12114
jordan-wong wants to merge 2 commits into
masterfrom
skill/database-category-rules-20260730

Conversation

@jordan-wong

@jordan-wong jordan-wong commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

Update — 2026-08-05 autonomous category gap sweep (commit bb7c329778)

An oracle-anchored database category sweep (static diff-vs-master + generation-transcript root-cause across Cassandra/PostgreSQL/R2DBC) re-derived every prior finding and confirmed dedup against this PR's existing commit. Net-new result landing here:

  • R-DB-1 is category-wide, not a Cassandra quirk (proven by evidence: surviving datastax-cassandra-3.0/-3.8 already declare super("cassandra"), so any new top-level cassandra module collides regardless of the deleted 4.0). The existing "modify in place, don't create a parallel module" rule did not cover the case actually hit: the eval slug (cassandra) differs from the family directory (datastax-cassandra/), AND under the blind protocol the same-version module was deleted so "modify in place" had no target — yet the name was still taken by siblings. Added a module-placement rule to instrumenter-module.md: grep for the intended super() name before creating a module; a taken name dictates the directory (join the family), never a new top-level slug. Failure mode it prevents: duplicate @AutoService registration → silent tracing outage (zero spans, no build error).

Everything else the sweep surfaced was already covered here (SPI-miss R-DB-2, eager-connect R-DB-3, muzzle R-DB-4) or routed off-PR (N-DBM-2 silent-downscope → toolkit prompt; db-no-parallel-module grep → toolkit reviewer_check), consistent with the "Not included yet" section below.


What

Sharpens the apm-integrations agent skill with lessons from the database eval cycle (reference PRs #11996 Cassandra / #11997 PostgreSQL / #12032 R2DBC). Draft — a living container that will be refined as feedback comes in from those PR reviews. Only the strongest, review-independent clarifications are in here now; library-specific rules await reviewer validation.

Key insight

Most database findings turned out to be adherence gaps against rules that already exist, not missing rules:

  • SPI-first (ForTypeHierarchy for JDBC's java.sql.*) was already in instrumenter-module.md, but its wording only triggered on "interface-only spec jar" — so an agent given a concrete driver (org.postgresql:postgresql) didn't apply it.
  • assertInverse guidance was already in muzzle.md but was violated.
  • "read the existing super(...) verbatim / scan for an existing module" was already present but violated.

So this PR sharpens the existing rules with the concrete failure modes rather than adding redundant prose, plus adds one genuinely-new idiom.

Changes

references/instrumenter-module.md

  • SPI exception now explicitly covers being handed a concrete implementation of a JDK SPI, not just an interface-only jar. Spells out the two harms of hooking the concrete class: (a) covers only one vendor, and (b) collides at runtime with the existing SPI module on the shared CallDepthThreadLocalMap.incrementCallDepth(<SpiType>.class) guard → mutual span suppression. (Root cause of [reference] eval: blind regeneration of PostgreSQL JDBC driver (toolkit output) #11997's spring-boot test_sql_traces failures.)
  • New: database clients must populate connection metadata eagerly at connect/factory time, not lazily per query — JDBC DriverInstrumentation and R2DBC ConnectionFactoryOptions as the reference points; explains why R2DBC's ConnectionMetadata cannot supply host/port/db/user (root cause of [reference] R2DBC net-new instrumentation (io.r2dbc:r2dbc-spi 1.0.0, toolkit output) #12032's dominant finding).

references/muzzle.md

  • assertInverse rule reinforced with the concrete-driver failure mode: a pinned dependency version is not an API-shape boundary (the PostgreSQL [42.0.0,) + assertInverse example that failed on six pre-42 releases).

Not included yet (deliberately)

Research provenance: docs/eval-research/hypotheses/{cassandra,postgresql,r2dbc}.md on the toolkit repo (R-DB-1 through R-DB-4, N-DBM-2).

🤖 Generated with Claude Code

…gory, add eager-connect idiom

From the database eval cycle (reference PRs #11996/#11997/#12032). Most
database findings turned out to be adherence gaps against rules that
already exist, not missing rules — so this sharpens the existing rules
with the concrete failure modes, plus adds one genuinely-new idiom.

instrumenter-module.md:
- SPI rule: the ForTypeHierarchy exception now explicitly covers being
  handed a CONCRETE driver that implements a JDK SPI (e.g.
  org.postgresql.jdbc.PgStatement implements java.sql.Statement), not
  just interface-only spec jars. The old wording only triggered on
  "interface-only jar", so an agent given a single concrete driver
  didn't apply it — the PostgreSQL regen (R-DB-2) fell into exactly
  this trap and shipped a concrete-class module that also collides at
  runtime with the existing jdbc/ SPI module.
- New: database clients must populate connection metadata eagerly at
  connect/factory time (JDBC DriverInstrumentation; R2DBC
  ConnectionFactoryOptions), not lazily per query (R-DB-3).

muzzle.md:
- assertInverse rule reinforced with the concrete-driver failure mode
  (R-DB-4): a pinned dependency version is not an API-shape boundary.

Draft — will be refined as feedback comes in from the database
reference PR reviews.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@dd-octo-sts

dd-octo-sts Bot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🟡 Java Benchmark SLOs — Performance SLO warning (near threshold)

Suite Status
Startup 🟡 warning

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.82 s 14.68 s [+0.0%; +1.9%] (maybe worse)
startup:insecure-bank:tracing:Agent 13.57 s 13.68 s [-1.7%; -0.0%] (maybe better)
startup:petclinic:appsec:Agent 17.35 s 16.55 s [+0.4%; +9.3%] (maybe worse)
startup:petclinic:iast:Agent 17.40 s 17.52 s [-1.7%; +0.4%] (no difference)
startup:petclinic:profiling:Agent 17.40 s 17.40 s [-1.3%; +1.3%] (no difference)
startup:petclinic:sca:Agent 17.35 s 17.36 s [-1.1%; +1.0%] (no difference)
startup:petclinic:tracing:Agent 16.64 s 16.22 s [-1.7%; +6.8%] (no difference)

Commit: bb7c3297 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

…ame (R-DB-1)

Database category gap sweep (2026-08-05) confirmed R-DB-1 is category-wide,
not a Cassandra quirk: any library with a version-sibling family directory
whose name differs from the integration slug will collide. The existing
"modify in place, don't create a parallel module" rule (line 47) doesn't
cover the case the Cassandra regen actually hit:

- eval slug `cassandra` != family dir `datastax-cassandra/`
- surviving siblings (datastax-cassandra-3.0/-3.8) already declare
  super("cassandra")
- under the blind protocol the same-version (4.0) module was DELETED, so
  "modify it in place" had no target — but the name was still taken

The agent created a new top-level instrumentation/cassandra/ module with a
duplicate super("cassandra") registration -> silent tracing outage (advice
never applied, zero spans, tests timed out, no build error).

Fix: grep the tree for the intended super() name BEFORE creating a module;
if any module (including untouched version-siblings) holds it, join that
family directory rather than minting a new top-level slug. Placement and
name are one decision: a taken name dictates the directory. If there is no
collision-free home, STOP and surface it.

Verified against #12114's existing commit (37e661c): the SPI-collision
case (R-DB-2) and eager-connect (R-DB-3) are already covered there; this is
the distinct version-sibling-family placement case they don't address.
Other sweep findings routed elsewhere (N-DBM-2 silent-downscope -> toolkit
prompt; reviewer-check candidates -> toolkit repo), not this skill PR.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@datadog-datadog-us1-prod

This comment has been minimized.

@jordan-wong
jordan-wong marked this pull request as ready for review August 7, 2026 20:37
@jordan-wong
jordan-wong requested a review from a team as a code owner August 7, 2026 20:37
@jordan-wong
jordan-wong requested a review from dougqh August 7, 2026 20:37
@dd-octo-sts dd-octo-sts Bot added the tag: ai generated Largely based on code generated by an AI or LLM label Aug 7, 2026
@jordan-wong
jordan-wong requested review from ValentinZakharov and removed request for dougqh August 7, 2026 20:37
@dd-octo-sts

dd-octo-sts Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Hi! 👋 Thanks for your pull request! 🎉

To help us review it, please make sure to:

  • Add at least one type, and one component or instrumentation label to the pull request

If you need help, please check our contributing guidelines.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bb7c329778

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

- Implement the **narrowest** `Instrumenter` interface possible:
- Prefer `ForSingleType` > `ForKnownTypes` > `ForTypeHierarchy`
- **EXCEPTION — API specification / interface-only libraries**: when the target library is a specification JAR containing only interfaces (no concrete classes), `ForSingleType` does not work because there are no concrete types to instrument directly. You MUST use `ForTypeHierarchy` with `implementsInterface(named("the.interface.Fqn"))`. This is how vendor implementations of the specification (ActiveMQ, IBM MQ, EclipseLink, Hibernate, etc.) get instrumented through the common interface contract.
- **EXCEPTION applies even when you are handed a CONCRETE implementation, not the spec jar.** The trigger is "does this type implement a shared JDK/spec SPI that other vendors also implement?" — NOT "is the coordinate an interface-only jar?" If you are given a single concrete driver (e.g. `org.postgresql:postgresql`, whose `org.postgresql.jdbc.PgStatement` implements `java.sql.Statement`), you MUST still hook the SPI interface via `ForTypeHierarchy` + `implementsInterface(named("java.sql.Statement"))`, NOT the concrete class via `ForSingleType(named("org.postgresql.jdbc.PgStatement"))`. Hooking the concrete class (a) covers only that one vendor while the SPI hook covers all conforming drivers with one module, and (b) collides at runtime with the existing SPI module that already instruments the same interface — both fire on the same object and mutually suppress spans via the shared `CallDepthThreadLocalMap.incrementCallDepth(<SpiType>.class)` guard. Before instrumenting any concrete class, check whether it implements a type already listed below; if so, the existing SPI module already covers it — do not generate a parallel per-vendor module.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Check behavior before rejecting concrete-driver hooks

When an integration needs vendor-specific behavior that the shared SPI advice cannot provide, implementing an SPI does not mean the concrete type is already covered. For example, DBMCompatibleConnectionInstrumentation.java:39-98 deliberately matches concrete PostgreSQL and other JDBC connection classes—even though they implement java.sql.Connection—to add DBM-specific prepare behavior absent from the generic SPI instrumentation. Scope this prohibition to advice that is behaviorally redundant; otherwise the skill will reject valid concrete instrumentation and silently omit requested features.

Useful? React with 👍 / 👎.


If the existing module targets a genuinely different version range (e.g. existing `foo-1.0/` and you're adding `foo-3.0/`), a version-sibling is correct — but confirm by reading the existing module's muzzle range first.

**The integration name you are given may NOT match the existing family directory — and if it doesn't, the directory wins, not the name.** Before creating a module, grep the whole tree for your intended `super(...)` name: `grep -rn 'super("<name>"' dd-java-agent/instrumentation/`. If ANY existing module already declares that name — including version-sibling modules you are not touching — your module MUST join that family's directory as `<existing-family-dir>/<family>-<version>/`; it must NOT become a new top-level module under a different slug. Placement and name are ONE decision: a taken name dictates the directory.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Do not use the super name as a unique family key

When multiple frameworks intentionally share an enablement name, this rule sends a new module to the wrong family or tells the agent to stop: super("jax-rs", ...) currently appears under the independent rs, jersey, and resteasy families. InstrumenterIndex.loadModules() also indexes module class names without deduplicating InstrumenterModule.name(), so equal super(...) names do not themselves cause the claimed registration outage. Determine placement from the framework/package and actual matcher and muzzle overlap rather than treating a config name as a unique directory key.

AGENTS.md reference: AGENTS.md:L61-L62

Useful? React with 👍 / 👎.


### Database clients: populate connection metadata EAGERLY at connect time, not lazily per query

For database-client integrations (`DatabaseClientDecorator` / `DBTypeProcessingDatabaseClientDecorator`), capture connection metadata (host, port, db name, user) at **connection-establishment** time and cache it in a `ContextStore` keyed on the connection object — not lazily on the first query. The canonical pattern is a dedicated instrumentation on the connect/factory method:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve lazy metadata as a fallback for uncovered connect paths

When a JDBC connection is not created through the instrumented Driver.connect, requiring exclusively eager capture leaves its spans without database metadata. JDBCDecorator.parseDBInfo() explicitly handles this case (JDBCDecorator.java:167-178), and StatementInstrumentation.java:87-90 intentionally invokes that fallback rather than merely reading a pre-populated value. Recommend eager capture as the primary path, but retain lazy extraction where metadata is reachable so DataSource, proxy, and other uncovered creation paths continue to be enriched.

Useful? React with 👍 / 👎.

For database-client integrations (`DatabaseClientDecorator` / `DBTypeProcessingDatabaseClientDecorator`), capture connection metadata (host, port, db name, user) at **connection-establishment** time and cache it in a `ContextStore` keyed on the connection object — not lazily on the first query. The canonical pattern is a dedicated instrumentation on the connect/factory method:

- **JDBC** — `dd-java-agent/instrumentation/jdbc/DriverInstrumentation.java` hooks `Driver.connect(url, props)` and populates `InstrumentationContext.get(Connection.class, DBInfo.class)` at open time. Statement advice then reads the already-cached `DBInfo`.
- **Reactive drivers with an async connect** — the equivalent connect point is the connection FACTORY, not the connection object. For R2DBC, `io.r2dbc.spi.ConnectionFactoryOptions` (available at `ConnectionFactory.create()` / `ConnectionFactories.find(...)`) is the only place host/port/database/user are exposed as structured data; `io.r2dbc.spi.ConnectionMetadata` (on the live `Connection`) exposes ONLY product name/version. Hooking `Connection.createStatement()` + `ConnectionMetadata` therefore CANNOT populate `db.name`/`peer.hostname`/`db.user`/port — you must hook the factory and thread the captured options forward. (OpenTelemetry's R2DBC instrumentation does exactly this; it is a good reference.)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Capture R2DBC options before ConnectionFactory.create

For R2DBC, ConnectionFactory.create() is a zero-argument SPI method returning a Publisher, so ConnectionFactoryOptions is not available at that call as stated. Unless an earlier ConnectionFactories.get(ConnectionFactoryOptions) or provider-construction hook associates the options with the returned factory, advice on create() has nothing from which to derive host, port, database, or user, and an implementation following this guidance will reproduce the missing metadata it is meant to fix. Specify the earlier options-to-factory context-store step and the subsequent propagation to the asynchronously emitted connection.

Useful? React with 👍 / 👎.


This is common whenever any instrumentation class in the module is compatible with versions below the declared min — `assertInverse` then contradicts that class's compatibility.

**Especially avoid defaulting `assertInverse = true` when hooking a concrete driver class** (as opposed to a JDK SPI — but note you usually should NOT be hooking a concrete driver at all; see instrumenter-module.md). Concrete driver classes tend to be structurally stable across a much wider version range than the `compileOnly`/`testImplementation` coordinate you happened to pin. Example: a PostgreSQL module declared `versions = "[42.0.0,)"` + `assertInverse = true`, but `org.postgresql.jdbc.PgStatement` is unchanged back through 9.2 (2013), so muzzle passed on 9.2/9.3/9.4 and the inverse-assertion failed for six old releases. The declared floor matched the pinned dependency, not any real API-shape boundary. Do not set `assertInverse` unless you can point to a specific API change at the declared minimum; otherwise omit it.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not treat inverse muzzle as proof the matcher target exists

When a concrete target class appears only as the string returned by instrumentedType() or in a name matcher, muzzle passing against an old artifact does not prove that class exists there: MuzzleGenerator.generateReferences() derives references from advice bytecode and explicit additional references, not matcher strings. The JDBC implementation also lists the older PostgreSQL jdbc2/jdbc3/jdbc4 statement classes separately from the newer org.postgresql.jdbc.PgPreparedStatement, contradicting the claim that the org.postgresql.jdbc.PgStatement shape is unchanged back through 9.2. Explain this matcher blind spot and require an explicit muzzle reference or runtime coverage instead; otherwise an agent may interpret the inverse result as compatibility with releases on which its target matcher never fires.

Useful? React with 👍 / 👎.

@datadog-datadog-us1-prod datadog-datadog-us1-prod Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: FAIL

The new universal rules overreach real repository structure: shared names such as jax-rs span unrelated families, while concrete SPI implementations still need vendor-only lifecycle and compatibility advice. Following either rule literally can misplace a module or omit advice required to produce and finish spans.

📊 Validated against 9 scenarios · Open Bits AI session

🤖 Datadog Autotest · Commit bb7c329 · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

- Implement the **narrowest** `Instrumenter` interface possible:
- Prefer `ForSingleType` > `ForKnownTypes` > `ForTypeHierarchy`
- **EXCEPTION — API specification / interface-only libraries**: when the target library is a specification JAR containing only interfaces (no concrete classes), `ForSingleType` does not work because there are no concrete types to instrument directly. You MUST use `ForTypeHierarchy` with `implementsInterface(named("the.interface.Fqn"))`. This is how vendor implementations of the specification (ActiveMQ, IBM MQ, EclipseLink, Hibernate, etc.) get instrumented through the common interface contract.
- **EXCEPTION applies even when you are handed a CONCRETE implementation, not the spec jar.** The trigger is "does this type implement a shared JDK/spec SPI that other vendors also implement?" — NOT "is the coordinate an interface-only jar?" If you are given a single concrete driver (e.g. `org.postgresql:postgresql`, whose `org.postgresql.jdbc.PgStatement` implements `java.sql.Statement`), you MUST still hook the SPI interface via `ForTypeHierarchy` + `implementsInterface(named("java.sql.Statement"))`, NOT the concrete class via `ForSingleType(named("org.postgresql.jdbc.PgStatement"))`. Hooking the concrete class (a) covers only that one vendor while the SPI hook covers all conforming drivers with one module, and (b) collides at runtime with the existing SPI module that already instruments the same interface — both fire on the same object and mutually suppress spans via the shared `CallDepthThreadLocalMap.incrementCallDepth(<SpiType>.class)` guard. Before instrumenting any concrete class, check whether it implements a type already listed below; if so, the existing SPI module already covers it — do not generate a parallel per-vendor module.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 SPI membership does not prove behavior is covered

Generated integrations can omit required vendor-only advice, causing missing or unfinished spans for affected libraries.

Assertion details
  • Input: A concrete SPI implementation with behavior outside the shared interface, such as Tomcat Request.recycle() or DB2-specific JDBC compatibility.
  • Expected: Reject a vendor module only when its target method and behavior are already covered by existing SPI advice; allow implementation-only lifecycle and compatibility hooks.
  • Actual: The instruction treats implementing a listed SPI as proof that existing SPI advice covers every relevant behavior. The repository contradicts this: PostgreSQL prepared statements are explicitly allowlisted, DB2 has vendor-specific JDBC modules, and Tomcat instruments concrete Request.recycle() alongside servlet SPI modules.
Suggested change
- **EXCEPTION applies even when you are handed a CONCRETE implementation, not the spec jar.** The trigger is "does this type implement a shared JDK/spec SPI that other vendors also implement?" — NOT "is the coordinate an interface-only jar?" If you are given a single concrete driver (e.g. `org.postgresql:postgresql`, whose `org.postgresql.jdbc.PgStatement` implements `java.sql.Statement`), you MUST still hook the SPI interface via `ForTypeHierarchy` + `implementsInterface(named("java.sql.Statement"))`, NOT the concrete class via `ForSingleType(named("org.postgresql.jdbc.PgStatement"))`. Hooking the concrete class (a) covers only that one vendor while the SPI hook covers all conforming drivers with one module, and (b) collides at runtime with the existing SPI module that already instruments the same interface — both fire on the same object and mutually suppress spans via the shared `CallDepthThreadLocalMap.incrementCallDepth(<SpiType>.class)` guard. Before instrumenting any concrete class, check whether it implements a type already listed below; if so, the existing SPI module already covers it — do not generate a parallel per-vendor module.
- **EXCEPTION applies even when you are handed a CONCRETE implementation, not the spec jar.** The trigger is whether the target behavior is declared by a shared JDK/spec SPI and already covered by existing SPI advice — NOT whether the coordinate is an interface-only jar. When given a concrete driver such as `org.postgresql:postgresql`, inspect the implemented SPI and the existing instrumentation first. If the target method is already advised through `ForTypeHierarchy` + `implementsInterface(...)`, reuse or extend that SPI module rather than adding overlapping vendor advice, which can duplicate or suppress spans through a shared call-depth guard. Do not infer coverage from `implements` alone: implementation-specific methods or lifecycle hooks that are not declared and advised on the SPI may still require a concrete/vendor module.

Was this helpful? React 👍 or 👎
🤖 Datadog Autotest · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

Comment on lines +51 to +53
**The integration name you are given may NOT match the existing family directory — and if it doesn't, the directory wins, not the name.** Before creating a module, grep the whole tree for your intended `super(...)` name: `grep -rn 'super("<name>"' dd-java-agent/instrumentation/`. If ANY existing module already declares that name — including version-sibling modules you are not touching — your module MUST join that family's directory as `<existing-family-dir>/<family>-<version>/`; it must NOT become a new top-level module under a different slug. Placement and name are ONE decision: a taken name dictates the directory.

**Concrete failure (Cassandra regen, R-DB-1):** the eval was given the integration slug `cassandra`, but dd-trace-java's family directory is `datastax-cassandra/` with siblings `datastax-cassandra-3.0/`, `-3.8/`, `-4.0/`, all declaring `super("cassandra")`. The agent created a new top-level `instrumentation/cassandra/` module that also declared `super("cassandra")`, producing two `@AutoService(InstrumenterModule.class)` registrations for the same name. Result: a **silent tracing outage** — ByteBuddy advice failed to apply, zero spans, all tests timed out, and there was no build error to catch it. This is especially dangerous under the blind protocol: if the same-version master module was deleted, "modify it in place" has no target — but the surviving siblings still hold the name, so grepping for the name (not looking for a same-version directory) is what tells you where the module belongs. When the name is taken and the correct family directory differs from the slug you were handed, place the module in the family directory and match the siblings' `super(...)` exactly; if there is genuinely no correct home without colliding, STOP and surface it rather than shipping a parallel registration.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Shared integration names do not identify one family

Following the rule can place generated code under an unrelated family or halt a valid integration.

Assertion details
  • Input: A new module using a shared primary name such as jax-rs or ci-visibility.
  • Expected: Use name matches as evidence, then place the module by library coordinates, packages, and compatible version family; allow shared configuration names across unrelated families.
  • Actual: The prescribed grep returns three legitimate top-level families for jax-rs and nine for ci-visibility, so a matching primary name does not dictate one directory. Many existing modules also share these names without registration failure.
Suggested change
**The integration name you are given may NOT match the existing family directory — and if it doesn't, the directory wins, not the name.** Before creating a module, grep the whole tree for your intended `super(...)` name: `grep -rn 'super("<name>"' dd-java-agent/instrumentation/`. If ANY existing module already declares that name — including version-sibling modules you are not touching — your module MUST join that family's directory as `<existing-family-dir>/<family>-<version>/`; it must NOT become a new top-level module under a different slug. Placement and name are ONE decision: a taken name dictates the directory.
**Concrete failure (Cassandra regen, R-DB-1):** the eval was given the integration slug `cassandra`, but dd-trace-java's family directory is `datastax-cassandra/` with siblings `datastax-cassandra-3.0/`, `-3.8/`, `-4.0/`, all declaring `super("cassandra")`. The agent created a new top-level `instrumentation/cassandra/` module that also declared `super("cassandra")`, producing two `@AutoService(InstrumenterModule.class)` registrations for the same name. Result: a **silent tracing outage** — ByteBuddy advice failed to apply, zero spans, all tests timed out, and there was no build error to catch it. This is especially dangerous under the blind protocol: if the same-version master module was deleted, "modify it in place" has no target — but the surviving siblings still hold the name, so grepping for the name (not looking for a same-version directory) is what tells you where the module belongs. When the name is taken and the correct family directory differs from the slug you were handed, place the module in the family directory and match the siblings' `super(...)` exactly; if there is genuinely no correct home without colliding, STOP and surface it rather than shipping a parallel registration.
**The integration name you are given may NOT match the existing family directory.** Before creating a module, grep the whole tree for the intended primary `super(...)` name: `grep -rn 'super("<name>"' dd-java-agent/instrumentation/`. Treat matches as evidence, not a unique placement key: inspect their library coordinates, target packages, and muzzle ranges. When matches for the same library form one version family, join that family as `<existing-family-dir>/<family>-<version>/`; when a shared config name spans unrelated families, choose the family from the target library rather than the name alone.
**Concrete failure (Cassandra regen, R-DB-1):** the eval was given the integration slug `cassandra`, but the surviving `datastax-cassandra-3.0/` and `-3.8/` siblings establish `datastax-cassandra/` as the family even if the same-version module was deleted. Recreating that version under a top-level `instrumentation/cassandra/` produces a parallel implementation instead of restoring the family member. Place it with the surviving siblings and preserve their `super(...)` value. If the name search points to multiple plausible families and the target coordinates do not disambiguate them, STOP and surface the ambiguity.

Was this helpful? React 👍 or 👎
🤖 Datadog Autotest · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tag: ai generated Largely based on code generated by an AI or LLM

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant