Repository navigation
Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN #115
Description
Activity
Summary + live status check (2026-07-14)
This is a tracking issue in a chain:
- Provide sqlite via a different URL #110 (closed) — established the vendor-neutral
semanticsql.berkeleybop.ioCDN URL. - Eliminate S3 bucket access #112 (open) — plan to eliminate raw S3 (
bbop-sqlite) bucket access once traffic has shifted to the CDN. - [bot] Cut a semsql release including the semanticsql.berkeleybop.io default #114 (open) — blocker: the CDN-default change landed on
masterbut hasn't been released to PyPI yet (semsqlis still0.4.0, from 2025-02-05, as of this check). Anyone relying onsemsql download/ the library default (rather than a hardcoded URL) won't get the CDN switch until that release ships. - Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN #115 (this issue) — the subset of consumers that bypass both OAK's selector and the
semsqlpackage entirely by hardcoding the raw S3 URL, so neither Eliminate S3 bucket access #112's traffic-shift nor [bot] Cut a semsql release including the semanticsql.berkeleybop.io default #114's release will fix them.
Verified current state of the checklist
I fetched each listed file directly (as of 2026-07-14) to confirm none have migrated yet — all 9 repos / 10 files are still unmigrated:
Repo File Status biobricks-ai/semsql stages/01_download.sh❌ still https://s3.amazonaws.com/bbop-sqlite/${onto}.db.gz(loops full catalog — highest-traffic consumer)Knowledge-Graph-Hub/kg-alzheimers src/kg_alzheimers/download.yaml❌ still S3 (e.g. ddpheno.db,phenio.db.gz)Knowledge-Graph-Hub/kg-microbe download.yaml(default branch ismaster, notmain)❌ still S3 ( ncit.db.gz,bto.db.gz,po.db.gz, ...)Knowledge-Graph-Hub/universalizer universalizer/oak_utils.py❌ still f"https://s3.amazonaws.com/bbop-sqlite/{db}.db"microbiomedata/ontology-loader src/ontology_loader/ontology_processor.py❌ still ontology_db_url_prefix = "https://s3.amazonaws.com/bbop-sqlite/"ccb-hms/NHANES-metadata code/generate_ontology_tables.py❌ still "https://s3.amazonaws.com/bbop-sqlite/%s.db.gz"vjcitn/ontoProc2 R/semsql_url.R❌ still "https://s3.amazonaws.com/bbop-sqlite/%s.db.gz"(default ontology"efo")vjcitn/ontoProc2 R/bbop_sqlite_db_gz.R❌ still url = "https://s3.amazonaws.com/bbop-sqlite"default argmonarch-initiative/monarch-ingest src/monarch_ingest/download.yaml❌ still S3 ( phenio.db.gz)monarch-initiative/dictybase-ingest download.yaml❌ still S3 ( ddpheno.db)I also searched all 9 downstream repos for any existing issue/PR mentioning
semanticsql.berkeleybop.ioorbbop-sqlite— none found. So outreach to these external maintainers hasn't started yet; this checklist accurately reflects zero progress so far.Notes for whoever actions this
- The CDN's 403-to-default-
User-Agentbehavior (noted in the issue body) will bite non-Python clients too — worth flagging explicitly forvjcitn/ontoProc2(R'sdownload.file/httralso send a default UA that could be blocked) andbiobricks-ai/semsql(shellcurl/wget). biobricks-ai/semsqlis worth prioritizing given it "loops the whole catalog on push" — it's the single heaviest hardcoded consumer of the raw bucket by request volume.- Since none of these 9 repos yet have an issue opened against them, and this tracking issue only coordinates (doesn't action), consider whether outreach issues should be filed in each downstream repo — e.g. via
@dragon-ai-agent please open an issue in each of the 9 downstream repos listed here, linking back to INCATools/semantic-sql#115, asking maintainers to repoint their hardcoded bbop-sqlite S3 URLs to https://semanticsql.berkeleybop.io (with a non-default User-Agent) before the raw bucket is retired per #112. - Recommend not closing raw-bucket access (Eliminate S3 bucket access #112) until (a) [bot] Cut a semsql release including the semanticsql.berkeleybop.io default #114's release ships and (b) at least the high-traffic items here (
biobricks-ai/semsql,kg-microbe,monarch-ingest) have migrated — the S3 access logs mentioned in Eliminate S3 bucket access #112 should make it easy to confirm before flipping the switch.
- Provide sqlite via a different URL #110 (closed) — established the vendor-neutral
Downstream outreach filed (2026-07-14)
Opened repointing issues in the downstream repos listed above. Each links back here and to #112, cites the specific hardcoded file(s), gives the
s3.amazonaws.com/bbop-sqlite/→semanticsql.berkeleybop.io/swap, and flags the 403-to-default-User-Agent gotcha (tailored per client: shellcurl/wget, Pythonurllib/requests, Rdownload.file/httr).Repo File(s) Outreach issue biobricks-ai/semsql stages/01_download.shbiobricks-ai/semsql#1 Knowledge-Graph-Hub/kg-alzheimers src/kg_alzheimers/download.yamlKnowledge-Graph-Hub/kg-alzheimers#78 Knowledge-Graph-Hub/kg-microbe download.yaml(master)Knowledge-Graph-Hub/kg-microbe#595 Knowledge-Graph-Hub/universalizer universalizer/oak_utils.pyKnowledge-Graph-Hub/universalizer#41 ccb-hms/NHANES-metadata code/generate_ontology_tables.pyccb-hms/NHANES-metadata#37 vjcitn/ontoProc2 R/semsql_url.R,R/bbop_sqlite_db_gz.Rvjcitn/ontoProc2#11 monarch-initiative/monarch-ingest src/monarch_ingest/download.yamlmonarch-initiative/monarch-ingest#711 monarch-initiative/dictybase-ingest download.yamlmonarch-initiative/dictybase-ingest#3 microbiomedata/ontology-loader src/ontology_loader/ontology_processor.py⚠️ not filed — see below⚠️ microbiomedata/ontology-loader could not be filed automaticallyThe
microbiomedataorg rejects classic personal access tokens for API writes:microbiomedataforbids access via a personal access token (classic). Please use a GitHub App, OAuth App, or a personal access token with fine-grained permissions.So this one needs to be opened manually (or with a fine-grained token / GitHub App that the org allows). The target is
src/ontology_loader/ontology_processor.py—ontology_db_url_prefix = "https://s3.amazonaws.com/bbop-sqlite/"→"https://semanticsql.berkeleybop.io/", plus the non-default User-Agent requirement.Opened the issue in
microbiomedata/ontology-loadermanually: microbiomedata/ontology-loader#59- added a commit that references this issue
on Aug 18, 2026 [bot] Posted by a Claude Code agent on behalf of @kltm.
Issue body refreshed 2026-08-18: per-repo traffic status from a raw-bucket access-log census (2026-07-08 → 2026-08-18), newly-found consumers (several with no outreach issue yet), the kg-microbe closed-but-still-observed flag, the ontoProc2 bucket-listing dependency, and coordination notes (pinned lockfiles do not auto-migrate; #114 still blocking; broadcast channels required for the unreachable population). Details in the body diff.
— Posted by Claude Code agent on behalf of @kltm.
[bot] Posted by a Claude Code agent on behalf of @kltm.
User-Agent gotcha retired: the CDN no longer 403s
Python-urllib/*— a host-scoped Browser Integrity Check exemption was deployed and verified today (2026-08-18). Thes3.amazonaws.com/bbop-sqlite/→semanticsql.berkeleybop.io/swap is now a plain one-line change for every client type; the UA caveat in the earlier outreach issues no longer applies.— Posted by Claude Code agent on behalf of @kltm.
- changed the title
[-][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN (tracking)[/-][+][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/+]on Aug 18, 2026 - changed the title
[-][bot] Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/-][+]Migrate direct raw-S3 bbop-sqlite consumers to the semanticsql CDN[/+]on Aug 18, 2026 [bot] Posted by a Claude Code agent on behalf of @kltm.
Body revised again 2026-08-18: every entry is now verified against fetched default-branch file content (line-level evidence), replacing code-search-derived rows — GitHub search tokenizes on hyphens and had produced false positives. Net changes: 2 entries removed as never-present, 1 deduplicated, kg-microbe and ontoProc2 confirmed migrated (2026-07-21 / 2026-07-29), 13 consumers confirmed still-present in runtime code, and a new section for the four bucket-listing-dependent consumers, which a URL swap alone cannot migrate — that class needs a #112 decision (ListBucket grant at the CDN origin, or a published manifest).
— Posted by Claude Code agent on behalf of @kltm.
- added this to Software essential and proactive maintenance and removed this from Software essential and proactive maintenance
on Aug 18, 2026 - added a commit that references this issue
on Aug 27, 2026 [bot] Posted by a Claude Code agent on behalf of @kltm.
The flip happened: raw bucket access retired 2026-08-28 (see #112). The migration asks in this tracker are now enforced by the endpoint itself — remaining unmigrated consumers will see AccessDenied on raw URLs and should apply the one-line CDN swap (
https://semanticsql.berkeleybop.io/<file>, all client types supported). This list stays open as the straggler-support tracker.— Posted by Claude Code agent on behalf of @kltm.
- added a commit that references this issue
on Sep 22, 2026
[bot] Opened by a Claude Code agent on behalf of @kltm. Body updated 2026-08-18 (2nd revision): every entry below is now verified against fetched default-branch file content (literal-match), not code-search hits — search tokenizes on hyphens and produced several false positives, now removed. Verdicts also cross-checked against raw-bucket access logs (2026-07-08 → 2026-08-18) and per-file git history.
Companion to #112 (remove raw-bucket access). These projects hardcode the raw bucket URL directly, so they will not auto-migrate via oaklib/ODK bumps; they need to repoint to
https://semanticsql.berkeleybop.io(drop-in, same paths) before raw access is removed.Gotcha RETIRED 2026-08-18: the CDN no longer 403s
Python-urllib/*User-Agents (host-scoped Browser Integrity Check exemption, verified). All client types work; the swap is a plain one-line change.Verified still-present (literal raw URL in current default-branch runtime code)
src/monarch_ingest/download.yamlL28 — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired monarch-initiative/monarch-ingest#711 (open, no engagement) — phenio pulls also confirmed in access logssrc/kg_alzheimers/download.yamlL186/L255 — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired Knowledge-Graph-Hub/kg-alzheimers#78 (open, no engagement) — traffic-confirmeddownload.yamlL10;src/versions.pyfilters oncontains=["bbop-sqlite"], so the URL fix must also update that selector or version-tracking silently breaks — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired monarch-initiative/dictybase-ingest#3 (open, no engagement) — traffic-confirmedstages/01_download.sh(URL generator);stages/00_invalidate.shdoes anonymousaws s3 ls(bucket-listing dependency, see below);.bb/source.jsonldmetadata — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired biobricks-ai/semsql#1 (open, no engagement)universalizer/oak_utils.pyL48 (note: fetches uncompressed.dbobjects) — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired Knowledge-Graph-Hub/universalizer#41 (open, no engagement)src/ontology_loader/ontology_processor.pyL100 — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired microbiomedata/ontology-loader#59 (open, no engagement). Covers microbiomedata/nmdc-runtime too (its notebook just calls this library).code/generate_ontology_tables.pyL30 — outreach Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired ccb-hms/NHANES-metadata#37 (open, no engagement)app/normalization/grounding/seed.pyL57 (SEMANTIC_SQL_BASE_URL) — outreach [bot] Repoint semantic-sql base URL to the semanticsql CDN before raw-bucket retirement waldronlab/bioanalyzer-backend#124src/metacurator/grounding/local_duckdb.pyL26 (single base-URL constant; cleanest possible migration) — outreach [bot] Repoint SEMSQL_BASE_URL to the semanticsql CDN before raw-bucket retirement seandavi/metacurator#29 — hancestro pulls confirmed in access logssrc/cdsci/lake/config.pyL141 +sources/ontology/ingest.py;scripts/check_chebi_currency.pyL53 — outreach via comment on check-chebi-currency infers 'a refresh would help' from byte size, not release CultureBotAI/MediaIngredientMech#206notebooks/environmental_context_value_sets/generate_voting_sheet.ipynb(source cell) +notebooks/multi-lexmatch/interleave_s3_catalog_yaml_registry_bioportal_obo.ipynbcalls the S3 ListBucket XML API; 2 docs files — outreach [bot] semsql pulls reference the retiring raw bbop-sqlite bucket (URL swap + one ListBucket dependency) microbiomedata/external-metadata-awareness#553resource/reactome/reactome*.mdproduct_url:frontmatter — machine-consumable registry metadata, not prose; directs downstream users to the raw URL — outreach [bot] reactome product_url points at the retiring raw bbop-sqlite bucket — repoint to the semanticsql CDN Knowledge-Graph-Hub/kg-registry#701Four verified consumers enumerate the bucket rather than (only) fetching objects: ontoProc2 (post-migration code parses ListBucket XML from the CDN root), cdsci-lake (ListObjectsV2 registry), biobricks
00_invalidate.sh(aws s3 ls), external-metadata-awareness (ListBucket XML notebook). Live probes 2026-08-18: the CDN root currently proxies the bucket's V1 listing (works only while the bucket stays public), and CloudFront strips query strings, so V2/pagination params are silently ignored (fine at ~332 keys; breaks at 1,000). #112 must decide a listing strategy: grants3:ListBucketto the CDN origin access at lockdown, or publish a manifest file and migrate these four to it.Migrated (verified in current code — done or nearly)
9c8ddcad, via #595). Residual raw-log traffic through 2026-08-18 attributed to stale checkouts/deployments, not master.1d4df64d, via Repoint hardcoded bbop-sqlite S3 URLs to the semanticsql CDN before raw-bucket access is retired vjcitn/ontoProc2#11, still open): remaining items are the READMEaws s3 ls s3://bbop-sqlite/example and the listing caveat above. vjcitn/op2workshop README carries the same example.Docs/examples only (verified literals, prose/log context; fix opportunistically)
api.py,impl.py)sqlite:obo:selectors — migrates with oaklib); turbomam/metpo-attic — 4 docs files incl. a copy-pasteablecurlexampleRemoved after verification (false positives of hyphen-tokenized code search)
monarch-initiative/rare-disease-identification (docstring prose only; runtime uses a local OBO), monarch-initiative/mondo notebook (URL only in saved output cells; source uses
sqlite:obo:selectors), several name-only doc mentions elsewhere. Selector-based (sqlite:obo:) consumers are out of scope here — they migrate via oaklib ≥0.7.2, with the standing caveat that pinned lockfiles do not auto-migrate (three CI consumers to date required manual lock bumps).Coordination notes
semsql download/ library-default users: semsql on PyPI is still 0.4.0 (2025-02-05).odkfull:v1.6.1) still bundles oaklib 0.6.23, so every ODK-based ontology repo's CI pulls from the raw bucket and cannot migrate by its own action until ODK ships (verified 2026-08-18: onlyodkfull:devcarries the fix). Release request filed: [bot] Released odkfull images bundle pre-CDN oaklib (0.6.23) — ODK-based repos cannot migrate off the retiring bbop-sqlite bucket until a release ships the #1354 bump ontology-development-kit#1368.— Posted by Claude Code agent on behalf of @kltm.