Skip to content

Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN #115

Description

@kltm

[bot] Opened by a Claude Code agent on behalf of @kltm. Body updated 2026-08-18 (2nd revision): every entry below is now verified against fetched default-branch file content (literal-match), not code-search hits — search tokenizes on hyphens and produced several false positives, now removed. Verdicts also cross-checked against raw-bucket access logs (2026-07-08 → 2026-08-18) and per-file git history.

Companion to #112 (remove raw-bucket access). These projects hardcode the raw bucket URL directly, so they will not auto-migrate via oaklib/ODK bumps; they need to repoint to https://semanticsql.berkeleybop.io (drop-in, same paths) before raw access is removed.

Gotcha RETIRED 2026-08-18: the CDN no longer 403s Python-urllib/* User-Agents (host-scoped Browser Integrity Check exemption, verified). All client types work; the swap is a plain one-line change.

Verified still-present (literal raw URL in current default-branch runtime code)

⚠️ Bucket-LISTING dependents (URL swap alone does not migrate these)

Four verified consumers enumerate the bucket rather than (only) fetching objects: ontoProc2 (post-migration code parses ListBucket XML from the CDN root), cdsci-lake (ListObjectsV2 registry), biobricks 00_invalidate.sh (aws s3 ls), external-metadata-awareness (ListBucket XML notebook). Live probes 2026-08-18: the CDN root currently proxies the bucket's V1 listing (works only while the bucket stays public), and CloudFront strips query strings, so V2/pagination params are silently ignored (fine at ~332 keys; breaks at 1,000). #112 must decide a listing strategy: grant s3:ListBucket to the CDN origin access at lockdown, or publish a manifest file and migrate these four to it.

Migrated (verified in current code — done or nearly)

  • Knowledge-Graph-Hub/kg-microbe — repointed to the CDN 2026-07-21 (9c8ddcad, via #595). Residual raw-log traffic through 2026-08-18 attributed to stale checkouts/deployments, not master.
  • vjcitn/ontoProc2 — runtime repointed 2026-07-29 (1d4df64d, via Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired vjcitn/ontoProc2#11, still open): remaining items are the README aws s3 ls s3://bbop-sqlite/ example and the listing caveat above. vjcitn/op2workshop README carries the same example.
  • monarch-initiative/dismech, ai4curation/ai-gene-review — oaklib lock bumps landed 2026-08-07 / 2026-08-12; verified in traffic.

Docs/examples only (verified literals, prose/log context; fix opportunistically)

  • cthoyt/pystow — docstring examples (api.py, impl.py)
  • berkeleybop/metpo — 3 docs files (its script uses sqlite:obo: selectors — migrates with oaklib); turbomam/metpo-attic — 4 docs files incl. a copy-pasteable curl example
  • brad-usredoxlabs/computable-lab — 1 doc
  • monarch-initiative/ontogpt-experiments — committed stdout log of an old run (evidence, not config)

Removed after verification (false positives of hyphen-tokenized code search)

monarch-initiative/rare-disease-identification (docstring prose only; runtime uses a local OBO), monarch-initiative/mondo notebook (URL only in saved output cells; source uses sqlite:obo: selectors), several name-only doc mentions elsewhere. Selector-based (sqlite:obo:) consumers are out of scope here — they migrate via oaklib ≥0.7.2, with the standing caveat that pinned lockfiles do not auto-migrate (three CI consumers to date required manual lock bumps).

Coordination notes

— Posted by Claude Code agent on behalf of @kltm.

Activity

  1. github-actions commented on Jul 14, 2026

    @github-actions

    Summary + live status check (2026-07-14)

    This is a tracking issue in a chain:

    Verified current state of the checklist

    I fetched each listed file directly (as of 2026-07-14) to confirm none have migrated yet — all 9 repos / 10 files are still unmigrated:

    Repo File Status
    biobricks-ai/semsql stages/01_download.sh ❌ still https://s3.amazonaws.com/bbop-sqlite/${onto}.db.gz (loops full catalog — highest-traffic consumer)
    Knowledge-Graph-Hub/kg-alzheimers src/kg_alzheimers/download.yaml ❌ still S3 (e.g. ddpheno.db, phenio.db.gz)
    Knowledge-Graph-Hub/kg-microbe download.yaml (default branch is master, not main) ❌ still S3 (ncit.db.gz, bto.db.gz, po.db.gz, ...)
    Knowledge-Graph-Hub/universalizer universalizer/oak_utils.py ❌ still f"https://s3.amazonaws.com/bbop-sqlite/{db}.db"
    microbiomedata/ontology-loader src/ontology_loader/ontology_processor.py ❌ still ontology_db_url_prefix = "https://s3.amazonaws.com/bbop-sqlite/"
    ccb-hms/NHANES-metadata code/generate_ontology_tables.py ❌ still "https://s3.amazonaws.com/bbop-sqlite/%s.db.gz"
    vjcitn/ontoProc2 R/semsql_url.R ❌ still "https://s3.amazonaws.com/bbop-sqlite/%s.db.gz" (default ontology "efo")
    vjcitn/ontoProc2 R/bbop_sqlite_db_gz.R ❌ still url = "https://s3.amazonaws.com/bbop-sqlite" default arg
    monarch-initiative/monarch-ingest src/monarch_ingest/download.yaml ❌ still S3 (phenio.db.gz)
    monarch-initiative/dictybase-ingest download.yaml ❌ still S3 (ddpheno.db)

    I also searched all 9 downstream repos for any existing issue/PR mentioning semanticsql.berkeleybop.io or bbop-sqlite — none found. So outreach to these external maintainers hasn't started yet; this checklist accurately reflects zero progress so far.

    Notes for whoever actions this

    • The CDN's 403-to-default-User-Agent behavior (noted in the issue body) will bite non-Python clients too — worth flagging explicitly for vjcitn/ontoProc2 (R's download.file/httr also send a default UA that could be blocked) and biobricks-ai/semsql (shell curl/wget).
    • biobricks-ai/semsql is worth prioritizing given it "loops the whole catalog on push" — it's the single heaviest hardcoded consumer of the raw bucket by request volume.
    • Since none of these 9 repos yet have an issue opened against them, and this tracking issue only coordinates (doesn't action), consider whether outreach issues should be filed in each downstream repo — e.g. via @dragon-ai-agent please open an issue in each of the 9 downstream repos listed here, linking back to INCATools/semantic-sql#115, asking maintainers to repoint their hardcoded bbop-sqlite S3 URLs to https://semanticsql.berkeleybop.io (with a non-default User-Agent) before the raw bucket is retired per #112.
    • Recommend not closing raw-bucket access (Eliminate S3 bucket access #112) until (a) [bot] Cut a semsql release including the semanticsql.berkeleybop.io default #114's release ships and (b) at least the high-traffic items here (biobricks-ai/semsql, kg-microbe, monarch-ingest) have migrated — the S3 access logs mentioned in Eliminate S3 bucket access #112 should make it easy to confirm before flipping the switch.
  2. caufieldjh commented on Jul 14, 2026

    @caufieldjh
    Contributor

    Downstream outreach filed (2026-07-14)

    Opened repointing issues in the downstream repos listed above. Each links back here and to #112, cites the specific hardcoded file(s), gives the s3.amazonaws.com/bbop-sqlite/ → semanticsql.berkeleybop.io/ swap, and flags the 403-to-default-User-Agent gotcha (tailored per client: shell curl/wget, Python urllib/requests, R download.file/httr).

    Repo File(s) Outreach issue
    biobricks-ai/semsql stages/01_download.sh biobricks-ai/semsql#1
    Knowledge-Graph-Hub/kg-alzheimers src/kg_alzheimers/download.yaml Knowledge-Graph-Hub/kg-alzheimers#78
    Knowledge-Graph-Hub/kg-microbe download.yaml (master) Knowledge-Graph-Hub/kg-microbe#595
    Knowledge-Graph-Hub/universalizer universalizer/oak_utils.py Knowledge-Graph-Hub/universalizer#41
    ccb-hms/NHANES-metadata code/generate_ontology_tables.py ccb-hms/NHANES-metadata#37
    vjcitn/ontoProc2 R/semsql_url.R, R/bbop_sqlite_db_gz.R vjcitn/ontoProc2#11
    monarch-initiative/monarch-ingest src/monarch_ingest/download.yaml monarch-initiative/monarch-ingest#711
    monarch-initiative/dictybase-ingest download.yaml monarch-initiative/dictybase-ingest#3
    microbiomedata/ontology-loader src/ontology_loader/ontology_processor.py ⚠️ not filed — see below

    ⚠️ microbiomedata/ontology-loader could not be filed automatically

    The microbiomedata org rejects classic personal access tokens for API writes:

    microbiomedata forbids access via a personal access token (classic). Please use a GitHub App, OAuth App, or a personal access token with fine-grained permissions.

    So this one needs to be opened manually (or with a fine-grained token / GitHub App that the org allows). The target is src/ontology_loader/ontology_processor.py — ontology_db_url_prefix = "https://s3.amazonaws.com/bbop-sqlite/" → "https://semanticsql.berkeleybop.io/", plus the non-default User-Agent requirement.

  3. caufieldjh commented on Jul 14, 2026

    @caufieldjh
    Contributor

    Opened the issue in microbiomedata/ontology-loader manually: microbiomedata/ontology-loader#59

  4. kltm commented on Aug 18, 2026

    @kltm
    ContributorAuthor

    [bot] Posted by a Claude Code agent on behalf of @kltm.

    Issue body refreshed 2026-08-18: per-repo traffic status from a raw-bucket access-log census (2026-07-08 → 2026-08-18), newly-found consumers (several with no outreach issue yet), the kg-microbe closed-but-still-observed flag, the ontoProc2 bucket-listing dependency, and coordination notes (pinned lockfiles do not auto-migrate; #114 still blocking; broadcast channels required for the unreachable population). Details in the body diff.

    — Posted by Claude Code agent on behalf of @kltm.

  5. kltm commented on Aug 18, 2026

    @kltm
    ContributorAuthor

    [bot] Posted by a Claude Code agent on behalf of @kltm.

    User-Agent gotcha retired: the CDN no longer 403s Python-urllib/* — a host-scoped Browser Integrity Check exemption was deployed and verified today (2026-08-18). The s3.amazonaws.com/bbop-sqlite/ → semanticsql.berkeleybop.io/ swap is now a plain one-line change for every client type; the UA caveat in the earlier outreach issues no longer applies.

    — Posted by Claude Code agent on behalf of @kltm.

  6. changed the title [-][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN (tracking)[/-] [+][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/+] on Aug 18, 2026
  7. changed the title [-][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/-] [+]Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/+] on Aug 18, 2026
  8. kltm commented on Aug 18, 2026

    @kltm
    ContributorAuthor

    [bot] Posted by a Claude Code agent on behalf of @kltm.

    Body revised again 2026-08-18: every entry is now verified against fetched default-branch file content (line-level evidence), replacing code-search-derived rows — GitHub search tokenizes on hyphens and had produced false positives. Net changes: 2 entries removed as never-present, 1 deduplicated, kg-microbe and ontoProc2 confirmed migrated (2026-07-21 / 2026-07-29), 13 consumers confirmed still-present in runtime code, and a new section for the four bucket-listing-dependent consumers, which a URL swap alone cannot migrate — that class needs a #112 decision (ListBucket grant at the CDN origin, or a published manifest).

    — Posted by Claude Code agent on behalf of @kltm.

  9. kltm commented on Aug 28, 2026

    @kltm
    ContributorAuthor

    [bot] Posted by a Claude Code agent on behalf of @kltm.

    The flip happened: raw bucket access retired 2026-08-28 (see #112). The migration asks in this tracker are now enforced by the endpoint itself — remaining unmigrated consumers will see AccessDenied on raw URLs and should apply the one-line CDN swap (https://semanticsql.berkeleybop.io/<file>, all client types supported). This list stays open as the straggler-support tracker.

    — Posted by Claude Code agent on behalf of @kltm.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

help wantedExtra attention is needed

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions