Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 37 additions & 0 deletions .github/workflows/adapter-contract.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
name: Adapter Contract

on:
pull_request:
paths:
- 'docs/adapters*.mdx'
- 'skills/create-adapter/**'
- 'skills/upload-parity-experiments/**'
- 'scripts/validate_adapter.py'
- 'scripts/tests/**'
- '.github/workflows/adapter-review.yml'
- '.github/workflows/adapter-contract.yml'
push:
branches: [main]
paths:
- 'docs/adapters*.mdx'
- 'skills/create-adapter/**'
- 'skills/upload-parity-experiments/**'
- 'scripts/validate_adapter.py'
- 'scripts/tests/**'
- '.github/workflows/adapter-review.yml'
- '.github/workflows/adapter-contract.yml'

permissions:
contents: read

jobs:
contract-examples:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
with:
persist-credentials: false
- uses: actions/setup-python@v6
with:
python-version: '3.12'
- run: python -m unittest discover -s scripts/tests -v
484 changes: 39 additions & 445 deletions .github/workflows/adapter-review.yml

Large diffs are not rendered by default.

54 changes: 29 additions & 25 deletions docs/adapters-human.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,8 @@ Do not use the tutorial below as your source of truth.
Join our [Discord](https://discord.com/invite/6xWPKhGDbA) (`#adapters-announcements`) and reach out to [Xiangning Lin](mailto:rosielin.xl@gmail.com). Check the [Adapter List](https://docs.google.com/spreadsheets/d/1mJbiASPm32DDNzEnV6eDGwpEf3FlMUe5dhkZmFjjSoo/edit?gid=0#gid=0) for available benchmarks. We cover API costs for parity experiments.
</Callout>

This walkthrough follows [Adapter contract v1](./adapters#contract-version-and-compatibility), including its pinned Harbor compatibility reference and stable requirement IDs. The agent guide takes precedence over generated scaffolds and external templates.

## Quick Start

```bash
Expand Down Expand Up @@ -122,14 +124,16 @@ Complete `adapter.py` and `main.py` so that running the adapter produces a valid
</Folder>
</Files>

Each `task.toml` must contain a valid, unique `name` field that identifies the task in the registry. Sanitize upstream identifiers (lowercase, replace special characters with hyphens) so the resulting names are stable and registry-safe. The author name and email in `task.toml` refer to the original benchmark authors, not the adapter contributor. See [§8 Tips](#8-register-the-dataset) for the full naming guidance.
Each `task.toml` must contain a valid, unique `name` field that identifies the task in the registry. Sanitize upstream identifiers (lowercase, replace special characters with hyphens) so the resulting names are stable and registry-safe. Credit the original benchmark authors in `[task].authors` (ADP-AUTHORS), for example `authors = [{ name = "Benchmark Author", email = "author@example.com" }]`; email may be omitted when unavailable. See [§8 Tips](#8-register-the-dataset) for the full naming guidance.

**Running the adapter:**
**Running the adapter (ADP-CLI):**
```bash
uv run python -m {adapter_name}.main --output-dir <path>
cd src/<adapter-name>
uv run <adapter-name> --output-dir <path>
```

**Tips:**
- **Windows (ADP-SCRIPTS):** Set `[environment].os = "windows"` and use `solve.bat` / `test.bat`. See the agent guide for verifier paths.
- Minor prompt tweaks (e.g., "write files in place without asking") are fine, as long as they apply to both the original benchmark and Harbor sides.
- Adapting only a subset of tasks is acceptable if documented in the README.
- If your benchmark requires GPU, add a `docker-compose.yaml` with nvidia device reservations in the task's `environment/` directory for Docker runs. For cloud/Modal runs, also set `gpus` in `task.toml`. See the [featurebench adapter](https://github.com/harbor-framework/adapters/tree/main/src/featurebench) for a comprehensive example with separate CPU/GPU/Modal configs.
Expand Down Expand Up @@ -211,7 +215,7 @@ For expensive benchmarks, you can run parity on a representative subset. Discuss

## 5. Run Parity Experiments

The purpose of parity experiments is to prove result equivalence between Harbor and the original benchmark. Run the **same agents, models, and settings** on both the original benchmark and your Harbor adapter, multiple times each. Report results as **mean ± sample SEM** (sample standard error of the mean) on both sides. They should be **comparable** to demonstrate equivalence.
The purpose of parity experiments is to prove result equivalence between Harbor and the original benchmark. Run the **same agents, models, and settings** on both the original benchmark and your Harbor adapter, at least twice each (three or more preferred). A single published original score does not satisfy this requirement. Report results as **mean ± sample SEM** (sample standard error of the mean) on both sides. They should be **comparable** to demonstrate equivalence.

```bash
# Harbor side
Expand All @@ -222,31 +226,31 @@ See the [AI adapter guide](https://harborframework.com/docs/datasets/adapters#re

## 6. Record Parity Results

Create `parity_experiment.json` in your adapter directory:
Create `parity_experiment.json` in your adapter directory (ADP-PARITY, ADP-STATS). This concrete example uses numeric counts and run scores:

```json
[
{
"adapter_name": "<adapter-name>",
"agent": "<agent>@<version>",
"model": "<model-version>",
"date": "<date>",
"adapted_benchmark_size": "<total-tasks-converted>",
"parity_benchmark_size": "<tasks-used-for-parity>",
"number_of_runs": "<runs-per-side>",
"notes": "<any special notes>",
"original_parity_repo": "<fork-url>",
"adapter_pr": ["<pr-url>"],
"dataset_pr": ["<pr-url>"],
"parity_pr": ["<hf-pr-url>"],
"adapter_name": "my-benchmark",
"agent": "codex@1.0",
"model": "gpt-5-2025-06-01",
"date": "2025-06-15",
"adapted_benchmark_size": 500,
"parity_benchmark_size": 500,
"number_of_runs": 3,
"notes": "None",
"original_parity_repo": "https://github.com/user/my-benchmark-fork",
"adapter_pr": ["https://github.com/harbor-framework/adapters/pull/123"],
"dataset_pr": ["https://github.com/laude-institute/harbor-datasets/pull/45"],
"parity_pr": ["https://huggingface.co/datasets/harborframework/parity-experiments/discussions/12"],
"metrics": [
{
"benchmark_name": "<name>",
"metric": "<metric>",
"original": "<mean ± sample SEM>",
"harbor": "<mean ± sample SEM>",
"original_runs": ["<run1>", "<run2>", "..."],
"harbor_runs": ["<run1>", "<run2>", "..."]
"benchmark_name": "my-benchmark",
"metric": "pass@1",
"original": "45.2 ± 0.6245",
"harbor": "44.8 ± 0.5292",
"original_runs": [44.0, 45.5, 46.1],
"harbor_runs": [43.8, 45.0, 45.6]
}
]
}
Expand All @@ -258,7 +262,7 @@ Also include a summary table in your README. Values are formatted as `mean ± sa
```markdown
| Agent | Model | Metric | Runs | Dataset Size | Original (mean ± SEM) | Harbor (mean ± SEM) |
|-------|-------|--------|------|--------------|-----------------------|---------------------|
| codex@0.1.2 | gpt-5 | pass@1 | 5 | 2000 (100%) | X ± Y | X ± Y |
| codex@1.0 | gpt-5 | pass@1 | 3 | 500 (100%) | 45.2 ± 0.6245 | 44.8 ± 0.5292 |
```

## 7. Upload Results
Expand Down Expand Up @@ -286,7 +290,7 @@ A dataset is a collection of tasks, and the two have a many-to-many relationship
```bash
git clone https://github.com/{you}/harbor-datasets.git
cd src/<adapter-name>
uv run python -m <adapter_name>.main --output-dir /path/to/harbor-datasets/datasets/<adapter-name>
uv run <adapter-name> --output-dir /path/to/harbor-datasets/datasets/<adapter-name>
```

**8.2.** Generate `dataset.toml` once your generated tasks are finalized.
Expand Down
Loading