Harbor Benchmarks

Use this guide to run Harbor Terminal-Bench Lite from a fresh Switchyard clone. It covers the two smoke paths most people need first:

  • Direct upstream: Harbor calls the provider directly. Switchyard is disabled.
  • Switchyard routing: Harbor calls Switchyard, and Switchyard routes across two model tiers.

Both paths use the same generated dataset, task proxy, pinned agent versions, and run artifact layout. Passing --server-config starts the Rust server; omitting it disables Switchyard and points Harbor directly at the upstream provider.

For a small automated MMLU-Redux example using NeMo Gym instead of Harbor, see Evaluate Switchyard routing with NeMo Gym.

Prerequisites

From the repo root:

uv sync

You also need Docker with Compose support, because baseline runs launch task containers and use the generated benchmark proxy topology. Runs with --server-config also start Switchyard inside Docker.

Harbor is installed as a dev dependency. Check that the CLI resolves from the uv environment:

uv run --no-sync harbor --help

Configure Your Provider

The checked-in smoke commands use OpenRouter's OpenAI-compatible endpoint by default:

export OPENROUTER_API_KEY="..."

To use another OpenAI-compatible provider, either export a generic upstream key:

export UPSTREAM_API_KEY="..."
export UPSTREAM_BASE_URL="https://provider.example/v1"

or pass the provider-specific key variable explicitly:

bash benchmark/run-baseline.sh \
  --upstream-base-url https://provider.example/v1 \
  --upstream-api-key-env PROVIDER_API_KEY \
  ...

Rust server TOML files refer to the credential through api_key_env. For another provider, copy a config and update its api_key_env, base_url, and model ids to match that provider.

One-Time Setup

run-baseline.sh has a blanket preflight check for the current patch file. It reverse-checks the exact diff against the installed Harbor tree, so stale or partial patch applications fail before launching Harbor. Apply the patch to the current uv environment:

REPO_ROOT="$(git rev-parse --show-toplevel)"
HARBOR_SITE="$(
  cd "$REPO_ROOT"
  uv run --no-sync python - <<'PY'
import sysconfig

print(sysconfig.get_paths()["purelib"])
PY
)"

cd "$HARBOR_SITE"
patch -p1 < "$REPO_ROOT/benchmark/patches/harbor-agent-patches.diff"
cd "$REPO_ROOT"

Reapply this after recreating the virtualenv, reinstalling Harbor, or running a forced dependency reinstall.

The generated dataset is local build output and is not committed. This command downloads and exports openthoughts-tblite@2.0, prebakes pinned agent versions into each task image, injects the benchmark proxy, and writes switchyard_dataset_manifest.json:

uv run --no-sync python benchmark/prepare_harbor_dataset.py --overwrite

Default output:

benchmark/datasets/openthoughts-tblite-closed-book

To reuse an already exported Harbor dataset instead of downloading again:

uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dir /path/to/exported/openthoughts-tblite \
  --overwrite

The pinned versions live in benchmark/agent-versions.env. To prepare a different Harbor dataset, see Benchmark Datasets.

Terminal-Bench 2.0 is supported through the same generated local proxy dataset path. The TB2 export keeps model/tool egress on the closed-book path while allowlisting the package and data sources required by the official Oracle solutions.

uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dataset terminal-bench/terminal-bench-2 \
  --output-dir benchmark/datasets/terminal-bench-2-closed-book \
  --overwrite

Terminal-Bench 2.1 (the verified iteration of 2.0) is supported the same way and shares the 2.0 Oracle allowlist:

uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dataset terminal-bench/terminal-bench-2-1 \
  --output-dir benchmark/datasets/terminal-bench-2-1-closed-book \
  --overwrite

SWE-Bench Pro is supported with the Harbor dataset cais/swebenchpro. The generated dataset uses the same pinned-agent and closed-book proxy path without opening dataset-specific agent egress.

uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dataset cais/swebenchpro \
  --output-dir benchmark/datasets/swebenchpro-closed-book \
  --overwrite

Run Without Switchyard

Omit --server-config to fully disable Switchyard. The runner still creates the benchmark Docker network for the generated proxy sidecar, but Harbor sends model calls straight to ${UPSTREAM_BASE_URL:-https://openrouter.ai/api/v1} using OPENROUTER_API_KEY by default:

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --model openai/gpt-5.5 \
  --agent codex \
  --reasoning-effort xhigh \
  --n-tasks 1 \
  --n-concurrent 1 \
  --max-retries 0

For another OpenAI-compatible upstream, pass --upstream-base-url and --upstream-api-key-env. Claude Code direct runs require an Anthropic-compatible upstream because Switchyard translation is disabled.

Run With Switchyard Routing

Pass --server-config to start switchyard-server and route Harbor traffic through it. This smoke test uses benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml, a Rust task-classifier configuration for coding-agent tasks:

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --server-config benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml \
  --model switchyard \
  --agent codex \
  --reasoning-effort xhigh \
  --n-tasks 1 \
  --n-concurrent 1 \
  --max-retries 0

Use the route id from the TOML as --model. In this config, the Gemini classifier selects the target tier for the task, then Switchyard routes to one of:

  • strong: anthropic/claude-opus-4.7
  • weak: moonshotai/kimi-k2.7-code

Classifier model: google/gemini-3.5-flash.

To smoke-test a single-model Switchyard path instead, use one of:

benchmark/server-configs/tb-lite-single-gpt-5-5.toml
benchmark/server-configs/tb-lite-single-opus-4-7.toml

By default, the runner starts in the background and prints the PID, log path, and kill command.

Book Modes

Both book modes use the same generated --harbor-path dataset, prebaked agent images, and proxy sidecar topology. Switchyard is Dockerized only when --server-config is provided.

Closed-book mode is the default:

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --server-config benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml \
  --model switchyard \
  --agent codex \
  --n-tasks 1

In closed-book mode, the proxy allows Switchyard/model traffic, blocks public cheat sources such as raw.githubusercontent.com, strips hosted web/search/code tools from model API payloads, and adds agent-specific web-disable settings where supported.

Open-book mode keeps the same proxy path but broadens egress:

bash benchmark/run-baseline.sh \
  --book-mode open \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --server-config benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml \
  --model switchyard \
  --agent codex \
  --n-tasks 1

Use open-book mode only when the evaluation intentionally allows internet access. The manifest records the mode, the local dataset digest, a snapshot of the server config, proxy metadata, upstream base URL for direct runs, and agent version pins in both modes.

Run A Full TB Lite Pass

After the smoke test succeeds, remove --n-tasks 1, raise concurrency to match your host and provider quota, and let the runner use the background wrapper.

Direct upstream:

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --model openai/gpt-5.5 \
  --agent codex \
  --reasoning-effort xhigh \
  --n-concurrent 8 \
  --max-retries 2

Switchyard LLM-classifier routing:

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --server-config benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml \
  --model switchyard \
  --agent codex \
  --reasoning-effort xhigh \
  --n-concurrent 8 \
  --max-retries 2

Tune --n-concurrent for your machine and provider quota. Use --task-id, --task-list-file, or --n-tasks for subsets.

Run with the pi coding agent

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/openthoughts-tblite-closed-book \
  --server-config benchmark/server-configs/tb-lite-llm-classifier-opus-kimi-gemini.toml \
  --agent pi \
  --model switchyard \
  --reasoning-effort high \
  --harbor-extra --ae --harbor-extra PI_CONTEXT_WINDOW=200000 \
  --harbor-extra --ae --harbor-extra PI_MAX_OUTPUT_TOKENS=32000 \
  --n-concurrent 8 \
  --max-retries 2

With --server-config, the script passes the model label switchyard/<route> to Harbor's pi agent. The patched agent then writes ~/.pi/agent/models.json inside the task container. That file defines a switchyard provider that points at OPENAI_BASE_URL and uses the openai-completions API. The agent environment variables PI_CONTEXT_WINDOW and PI_MAX_OUTPUT_TOKENS set contextWindow and maxTokens on that model entry. When they are unset, pi uses its defaults of 128000 and 16384. --reasoning-effort sets pi's --thinking level, so pass one of off, minimal, low, medium, high, or xhigh. Without --server-config, pass pi's own provider label as --model, for example openrouter/openai/gpt-5.5.

Inspect A Run

Run directories are created under benchmark/tb_runs/. The most useful artifacts are:

run_manifest.json
server.log
harbor.log
server_metrics_final.prom
routing_stats_final.json
jobs/<job-name>/result.json
jobs/<job-name>/<task-id>/agent/trajectory.json

The manifest records the command, git state, Harbor patch provenance, local dataset digest, copied server config, direct-upstream metadata when Switchyard is disabled, book-mode settings, agent version pins, log paths, and final Harbor status.

server_metrics_final.prom is the final /metrics snapshot. routing_stats_final.json is the final aggregate /v1/stats snapshot, including model and tier calls, errors, tokens, and latency. Neither artifact provides task or trial attribution. The runner writes them only after Harbor exits and while the Rust server is still reachable; otherwise the manifest records them as missing. routing_requests.jsonl and routing_stats_by_task.json are not produced by the Rust server.

Docker Image Notes

Baseline runs build switchyard-baseline:local from the repository-root Dockerfile. The default is to rebuild before each run so the container matches the current checkout.

To reuse an already built image:

SWITCHYARD_DOCKER_BUILD=0 bash benchmark/run-baseline.sh ...

Only reuse the image when you know it already contains the current Rust switchyard-server binary.

DeepSWE v1.1

DeepSWE uses Harbor's task format but its own runner, Pier (required since v1.1 for the separate-verifier/collect-hook pattern). It is not part of run-baseline.sh.

See DeepSWE v1.1 qualification settings for the exact versions, timeouts, scoring rules, and routing profiles used for qualification.

git clone https://github.com/datacurve-ai/deep-swe benchmark/datasets/deep-swe
uv tool install 'datacurve-pier>0.3.0'

Direct upstream:

export OPENAI_API_KEY="..."
pier run -p benchmark/datasets/deep-swe/tasks --agent mini-swe-agent --model openai/gpt-5.5

Switchyard routing: start the server, then point the agent's OpenAI client at it through --ae. Pier's Docker environment routes all agent traffic through a policy proxy whose Safe_ports ACL only allows ports 80/443, so bind Switchyard to 443:

switchyard-server --config benchmark/server-configs/deepswe-single.toml \
  --host 0.0.0.0 --port 443

pier run -p benchmark/datasets/deep-swe/tasks --agent mini-swe-agent \
  --model openai/deepswe-single \
  --ae OPENAI_BASE_URL=http://host.docker.internal:443/v1 \
  --ae OPENAI_API_KEY=unused

--agent codex reads the same OPENAI_BASE_URL/OPENAI_API_KEY pair through --ae:

pier run -p benchmark/datasets/deep-swe/tasks --agent codex \
  --model openai/deepswe-single \
  --ae OPENAI_BASE_URL=http://host.docker.internal:443/v1 \
  --ae OPENAI_API_KEY=unused

host.docker.internal requires Docker Desktop; on Linux, pass the host's Docker-bridge address instead. Binding port 443 needs elevated privileges on most Linux hosts (sudo, or setcap 'cap_net_bind_service=+ep' on the binary).

Advisor-gate routing (GPT-5.6 Luna executes; GPT-5.6 Sol reviews its "done" claims, up to three per task) measured 57.5% +/- 3.2 on the full 113 tasks (k=3, closed-book). The profile keeps the routing parameters exactly as run and reaches both models through OpenRouter, so it needs only OPENROUTER_API_KEY. The measured runs had Codex at reasoning effort max (model_reasoning_effort = "max" in the agent's Codex config); the profile forces max on both models server-side regardless:

export OPENROUTER_API_KEY="..."
switchyard-server --config benchmark/routing-profiles/deepswe-v11-advisor-gate-luna-sol.toml \
  --host 0.0.0.0 --port 443

pier run -p benchmark/datasets/deep-swe/tasks --agent codex \
  --model openai/switchyard \
  --ae OPENAI_BASE_URL=http://host.docker.internal:443/v1 \
  --ae OPENAI_API_KEY=unused

Plan/execute routing uses GPT-5.6 Sol for repository inspection and planning, then hands the full trajectory to GPT-5.6 Luna after the first mutation. It measured 60.8% +/- 5.3 at $180.51 +/- 9.43 per run on the full closed-book benchmark (k=3). The profile preserves the exact planning prompt and routing policy, with provider settings adapted from NVIDIA Inference Hub to OpenRouter:

export OPENROUTER_API_KEY="..."
switchyard-server \
  --config benchmark/routing-profiles/deepswe-v11-plan-execute-luna-sol.toml \
  --host 0.0.0.0 --port 443

pier run -p benchmark/datasets/deep-swe/tasks --agent codex \
  --model openai/switchyard \
  --ae OPENAI_BASE_URL=http://host.docker.internal:443/v1 \
  --ae OPENAI_API_KEY=unused

The qualified stage-router profile (GPT-5.6 Luna efficient tier, GPT-5.6 Sol capable tier) solved 76/113 tasks (67.3% strict) on the full closed-book benchmark. Its routing policy is published with OpenRouter provider settings so the file runs as-is with OPENROUTER_API_KEY. The header records the Switchyard commit, model ids, harness inputs, run id, and the one crashed task. The profile's public route id is gpt-5.6-luna, matching the qualification run:

export OPENROUTER_API_KEY="..."
switchyard-server \
  --config benchmark/routing-profiles/deepswe-v11-stage-router-luna-sol.toml \
  --host 0.0.0.0 --port 443

pier run -p benchmark/datasets/deep-swe/tasks --agent codex \
  --model openai/gpt-5.6-luna \
  --ae OPENAI_BASE_URL=http://host.docker.internal:443/v1 \
  --ae OPENAI_API_KEY=unused

Smoke subset:

pier run -p benchmark/datasets/deep-swe/tasks --agent mini-swe-agent \
  --model openai/gpt-5.5 --n-tasks 1 --sample-seed 0

Results land under jobs/<job-name>/, per Pier's own layout.

Troubleshooting

If the runner reports that the current Harbor patch is not applied cleanly, recreate or reinstall the uv environment and rerun the patch command from this README.

If port 4000 is busy, pass a different port:

bash benchmark/run-baseline.sh ... --port 4001

If the Docker reachability preflight fails, check Docker/Compose first. The preflight proves the task container can reach Switchyard through the benchmark Docker network. For local debugging only, it can be bypassed with SWITCHYARD_CLOSED_BOOK_PREFLIGHT=0.