harbor run -d looni-lab/smdd-benchharbor run -d looni-lab/smdd-benchSMDD-Bench: five small-molecule design task families with an agent harness, GPU oracle and setup guide.
harbor run -d looni-lab/smdd-benchSMDD-Bench: five small-molecule design task families with an agent harness, GPU oracle and setup guide.
harbor run -d looni-lab/smdd-benchPaper | SMDD-Bench website | Blog: Is Human Taste Overrated in Harness Engineering?
SMDD-Bench tests whether language-model agents can carry out long-horizon small-molecule drug design. Each task gives the agent a scientific objective, molecular data, and a fixed set of tools. The agent works inside an isolated task environment, produces a final artifact, and a separate verifier scores the result.
This Harbor release packages 502 tasks, the original agent harness, and the build sources needed to run the benchmark through Harbor.
SMDD-Bench contains five task families:
| Type | Task | What the agent produces |
|---|---|---|
| 1 | 2D Pharmacophore Identification | A pharmacophore that separates active and inactive molecules |
| 2 | Interaction Point Discovery | Conserved interaction points in a binding pocket |
| 3 | Scaffold Hopping | A molecule with a different scaffold that retains binding interactions |
| 4 | Lead Optimization | A molecule meeting property, binding, and structural constraints |
| 5 | Fragment Assembly | A drug-like binder assembled from the supplied fragments |
The exact output format is defined in each task's instruction.md.
A Harbor run has three moving pieces:
The Harbor host runs the Python agent harness, calls the language-model provider, and launches the task and verifier containers. This machine needs Docker, Python, and your model credentials. It can be a CPU host or the same GPU VM used by the oracle.
The task and verifier containers isolate what the agent sees from what the evaluator sees. Harbor stages the public task files in the task container, then evaluates the submitted artifact in a separate verifier container containing the grading files. Harbor creates and tears down these containers automatically.
The GPU oracle is a persistent Docker service that runs ADMET-AI and Boltz behind an authenticated HTTP endpoint. The agent uses it during a trajectory, and the Types 3-5 verifiers use it for Boltz scoring. Type 4 verification additionally computes ADMET metrics inside the CPU verifier environment.
In other words, Harbor owns the trial lifecycle and isolation; the oracle provides the expensive scientific models.
You need at least one Linux machine with an NVIDIA GPU for the oracle. Installing the Python package on a laptop is not enough: the deployed oracle requires CUDA, an NVIDIA driver, Docker, and NVIDIA Container Toolkit configured for docker run --gpus all.
We recommend at least 48 GB of VRAM per GPU to reduce Boltz out-of-memory failures. This applies to every GPU used by the oracle, including Ray heads and workers. Exact memory use depends on the protein, sampling settings, and concurrency, so 48 GB is a practical deployment target rather than a guarantee.
For higher-throughput runs, an 8 x NVIDIA A100 80 GB machine is one option. Each Boltz worker uses a single GPU; GPU memory is not pooled across cards, so an individual prediction still has to fit on one GPU. Start smaller if your workload allows it, then increase concurrency after checking GPU and host-memory use.
Also leave enough host RAM for model processes and concurrent verifier containers. Some task manifests request 32 GiB RAM per trial, so inspect task.toml before increasing concurrency.
The agent-facing interface is the same with or without Ray. Ray only changes how the oracle schedules Boltz work.
| Setup | Harbor host | GPU service |
|---|---|---|
| One VM, no Ray | On the GPU VM | One oracle container using its local GPUs |
| Separate hosts, no Ray | On a CPU host | One oracle container on a GPU VM |
| Multiple GPU VMs with Ray | On the head VM or a separate CPU host | One head container and one worker container per additional VM |
With Ray, ADMET stays on the head and Boltz workers run across the available GPUs. The head GPU can also host a Boltz worker. Without Ray, a standalone oracle starts one worker for each GPU visible to Docker.
Run the following on the Linux machine that will act as the Harbor host. If Harbor and the oracle share a GPU VM, this is that machine.
Create a Python 3.12 environment, install Harbor, authenticate, and download the release:
uv venv --python 3.12 "$HOME/.venvs/smdd-harbor"
source "$HOME/.venvs/smdd-harbor/bin/activate"
uv pip install harbor==0.20.0
harbor auth login
harbor download "looni-lab/smdd-bench@candidate" --export --output-dir ./smdd-download
Move into the downloaded dataset directory, install the bundled host-side package, and make sure the tasks are visible:
cd smdd-download/smdd-bench
unset SMDD_BENCH_ROOT
uv pip install ./smdd_harbor_release-0.1.2-py3-none-any.whl
python run.py --list-tasks
The wheel contains the host-side harness, Harbor adapter, and evaluation helpers. Scientific-model dependencies live in the Docker images, so you do not need to install Boltz or ADMET into this Python environment.
Keep this environment activated for the remaining commands.
The release has the following top-level structure:
downloaded-dataset/
|-- README.md # Complete setup and running guide
|-- SETUP.md # Checklist and links into this guide
|-- run.py # Task listing, reference tests and agent trials
|-- release.json # Version, image tags and source metadata
|-- smdd_harbor_release-*.whl # Installable Python harness and adapter
|-- smdd-runtime.zip # Docker and oracle source package
+-- smdd_004_A2A_Adenosine_0__v1/ # One of the task directories
|-- task.toml # Resources, images and verifier settings
|-- instruction.md # Agent-visible instructions
|-- environment/ # Public inputs and task build context
|-- tests/ # Verifier code and private grading inputs
+-- solution/ # Reference solution, where provided
The runner accepts canonical task IDs such as smdd_004_A2A_Adenosine_0 and resolves Harbor's __v1 directory suffix automatically.
The benchmark operator receives the grading files as part of the release, but Harbor stages them only inside the separate verifier runtime. If you modify or repackage a task, preserve that separation.
Before building images, extract the runtime sources on every machine that needs them:
python -m zipfile -e smdd-runtime.zip .
This creates smdd-runtime/:
| Path | Purpose |
|---|---|
build_images.sh |
Builds all images or only the Harbor/oracle group |
agent_env/ |
Agent science, GPU model, and HTTP-service Dockerfiles |
agent_harness/ |
Agent loop, tools, HTTP service, and Ray implementation |
smdd_harbor/ |
Adapter connecting the agent loop to Harbor |
evaluator/, metrics/ |
Evaluation code and scientific scoring functions |
smdd-harbor/docker/ |
Task runtime and five verifier-image build contexts |
scripts/smoke_oracle_ray.py |
Optional multi-node inference diagnostic |
The release ships the Docker build sources rather than prebuilt images. Build the image group needed by each machine.
| Image role | Where needed | How it is used |
|---|---|---|
| Agent science image (smdd-agent) | Harbor host | Build parent for task runtimes |
| Evaluator image (smdd-evals) | Harbor host | Build parent for the lead runtime and verifiers |
| Task science image | Harbor host | Harbor runs Types 1, 2, 3, and 5 in it |
| Task lead image | Harbor host | Harbor runs Type 4 in it |
| Five verifier images: 2d, 3d, scaffold, lead, fragment | Harbor host | Parents for the separate per-task verifier containers |
| Model image (smdd-oracle) | GPU image-build host | Build parent containing Boltz, ADMET, and Ray dependencies |
| HTTP oracle image (smdd-oracle-http) | Every GPU VM | Persistent oracle service started manually |
From the downloaded dataset directory, choose the command that matches the machine's role:
# Harbor and oracle on the same GPU VM:
bash smdd-runtime/build_images.sh all
# OR: a separate Harbor host needs the task/verifier image group:
bash smdd-runtime/build_images.sh harbor
# On the separate GPU VM, build the oracle image group:
bash smdd-runtime/build_images.sh oracle
If you are using Ray, every GPU VM must run the same oracle image. Build it once and transfer that exact image, or pull the same registry digest on every node. Worker VMs do not need the Harbor task or verifier images unless Harbor also runs there.
You do not manually start task or verifier containers. Harbor creates them from the task manifests when a trial runs.
There are two supported layouts. Both expose the same authenticated HTTP API on port 8090, so the Harbor-side commands later in this guide are identical.
Keep oracle traffic on a private network or VPN. The endpoint must be reachable from both the Harbor host and Harbor's verifier containers. In particular, 127.0.0.1 inside a verifier refers to that verifier container, not to the GPU host.
Use this when all oracle GPUs are on one machine.
Run on the GPU VM:
export ORACLE_IMAGE="smdd-oracle-http:latest"
export STATE="$HOME/smdd-oracle/state"
export CACHE="$HOME/smdd-oracle/cache"
install -d -m 700 "$STATE" "$CACHE"
test -s "$STATE/token" || openssl rand -hex 32 > "$STATE/token"
chmod 600 "$STATE/token"
docker run -d --name smdd-oracle-http --restart unless-stopped \
--gpus all --shm-size=16g -p 8090:8090 \
-e ORACLE_ROLE=standalone \
-v "$STATE:/srv/oracle_state" \
-v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" \
--max-samples 10 --max-parallel-samples 2 --microbatch-size 1
docker logs --tail 100 -f smdd-oracle-http
The service automatically discovers the GPUs exposed by Docker. To restrict it to GPU 0, replace --gpus all with --gpus '"device=0"'.
Initial model loading can take several minutes. Ctrl+C while following the logs only exits the log viewer; the detached oracle container keeps running. The 16g shared-memory value is a starting point and can be adjusted for the host and workload.
Use this when you want Boltz requests to be scheduled across multiple GPU machines.
The example below uses two VMs with one GPU each. Every node needs the same oracle image, mutually reachable private or VPN addresses, compatible GPUs, and bidirectional connectivity over the trusted network. Ray uses auxiliary ports in addition to 6379, 6380, 6381, and 20000-20063; opening only 6379 is not enough.
First copy the exact oracle image to each worker, substituting the worker's SSH address:
docker save "smdd-oracle-http:latest" | ssh ubuntu@WORKER_SSH_ADDRESS docker load
Set HEAD_IP to the head VM's own private or VPN address:
export HEAD_IP="10.77.0.1"
export ORACLE_IMAGE="smdd-oracle-http:latest"
export STATE="$HOME/smdd-ray/state"
export CACHE="$HOME/smdd-ray/cache"
install -d -m 700 "$STATE" "$CACHE"
test -s "$STATE/token" || openssl rand -hex 32 > "$STATE/token"
chmod 600 "$STATE/token"
docker run -d --name smdd-head --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=head -e RAY_NODE_IP="$HEAD_IP" \
-e RAY_EXPECTED_GPUS=2 -e ORACLE_HTTP_HOST=0.0.0.0 \
-v "$STATE:/srv/oracle_state" \
-v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" \
--max-samples 10 --max-parallel-samples 2 --microbatch-size 1
Start the worker after the head container has been created; you do not need to wait for the head HTTP endpoint to become ready.
export HEAD_IP="10.77.0.1"
export WORKER_IP="10.77.0.2"
export ORACLE_IMAGE="smdd-oracle-http:latest"
docker run -d --name smdd-worker --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=worker -e RAY_NODE_IP="$WORKER_IP" \
-e RAY_ADDRESS="$HEAD_IP:6379" \
"$ORACLE_IMAGE"
Repeat the worker command on additional VMs, changing only WORKER_IP for each machine while keeping the same HEAD_IP.
RAY_EXPECTED_GPUS is the total GPU count across the whole cluster, including the head. The head waits for that many Boltz actors before its HTTP endpoint becomes ready.
Workers do not serve HTTP and do not need the bearer token. Shared results and artifacts live on the head, so the VMs do not need a shared filesystem. Every node does need outbound access to the MSA service, normally api.colabfold.com.
Back on the head, check that Ray sees the expected nodes and GPUs:
docker exec smdd-head /opt/oracle/bin/ray status
docker logs --tail 100 smdd-head
For this two-GPU example, you should see two active nodes and two GPUs. If a worker is missing, inspect docker logs smdd-worker on that machine and verify that it can reach the head's private address. Do not substitute a public NAT address for a node's local bind address.
If you want the Ray dashboard, forward its loopback port from your computer:
ssh -N -L 8265:127.0.0.1:8265 ubuntu@HEAD_SSH_ADDRESS
Then open http://localhost:8265 while that SSH session remains active.
For either deployment, run the health check from the GPU host or Ray head using an address that the Harbor host can also reach:
export ORACLE_IP="10.77.0.1"
curl --fail --silent --show-error \
-H "Authorization: Bearer $(cat "$STATE/token")" \
"http://$ORACLE_IP:8090/v1/health" | python -m json.tool
Wait until the service and model report healthy before launching trials. A connection refusal during startup can simply mean that the models are still loading; check the oracle logs before assuming the service failed.
Keep --max-samples 10. This controls the maximum number of diffusion samples per request, not the number of GPUs. The benchmark verifiers request ten samples even when the oracle has only one GPU.
Run this on the Harbor host from the downloaded dataset directory.
If the oracle is on another machine, securely copy its token file to the Harbor host first and use that host-local absolute path here:
export SMDD_AGENT_ORACLE_HTTP_URL="http://10.77.0.1:8090"
export SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE="/absolute/path/to/token"
export SMDD_AGENT_ORACLE_CACHE_DIR="$HOME/.cache/smdd-harbor"
curl --fail --silent --show-error \
-H "Authorization: Bearer $(cat "$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE")" \
"$SMDD_AGENT_ORACLE_HTTP_URL/v1/health" | python -m json.tool
Replace the example IP and token path with the values for your deployment. With Ray, the URL always points to the head.
Your model-provider credentials also belong on this trusted Harbor host. For example, the Claude command later in this guide expects ANTHROPIC_API_KEY.
The runner configures the agent's oracle client and passes the required oracle credentials into Harbor's verifier environment without printing the token in generated commands. The task container itself remains network-isolated: the trusted host makes prediction requests on the agent's behalf, while verifier containers use their own network access to contact the oracle.
A successful host-side health check proves that the host can reach the oracle, but not necessarily that verifier containers can. The smoke tests below exercise the complete path.
Before spending model calls, verify that Harbor, the task images, the verifier images, and the oracle can complete a real evaluation.
Start with one supplied reference solution:
python run.py --task smdd_001_P39086_0 --smoke
Then run one reference task from each task family:
for task in smdd_001_P39086_0 smdd_002_O75874_0 smdd_003_3KWZ_2 smdd_004_A2A_Adenosine_0 smdd_005_O60674_0; do
python run.py --task "$task" --smoke || break
done
Types 1-2 can verify their reference solutions locally. Types 3-5 require the oracle.
Reference solutions are bundled only for selected tasks, so use these tests to validate the installation rather than as a way to run the full benchmark.
Once the smoke tests pass, export the model credential and launch a real task:
export ANTHROPIC_API_KEY="YOUR_ANTHROPIC_API_KEY"
python run.py --task smdd_004_A2A_Adenosine_0 --model anthropic/claude-sonnet-4-6
The same command works whether the oracle is standalone or backed by Ray. The harness coordinates the model and tool calls; the oracle handles scientific prediction scheduling.
Harbor writes results, submitted artifacts, and verifier logs under jobs/ next to run.py.
Useful runner options:
| Option | Purpose |
|---|---|
--list-tasks |
List canonical task IDs without starting containers |
--task ID |
Select one task |
--model NAME |
Choose the LiteLLM model for a real trial |
--smoke |
Run the reference solution instead of an LLM |
--dataset-dir PATH |
Use tasks from another downloaded directory |
--print-command |
Inspect the generated Harbor command without launching it |
--concurrency N |
Set Harbor trial concurrency; does not allocate GPUs |
The helper runs one selected task and one attempt. Increasing --concurrency does not create extra rollouts by itself; it only controls how many already-scheduled trials Harbor may run simultaneously.
Start with one concurrent trial. Once that works reliably, scale to larger batches while watching GPU memory, host memory, and oracle queue time.
When you are finished, stop only the persistent containers that you started manually:
# Without Ray, on the GPU VM:
docker stop smdd-oracle-http
# With Ray: first on the worker VM, then on the head VM:
docker stop smdd-worker
docker stop smdd-head
These commands leave the containers and their mounted state, cache, and token files in place.
For a standalone oracle, docker start smdd-oracle-http restarts the existing container. For Ray, restart the head and then the workers and wait for the cluster to become healthy again. If you need to change environment variables or startup arguments, recreate the containers while keeping the same persistent mounts.
Paper | SMDD-Bench website | Blog: Is Human Taste Overrated in Harness Engineering?
SMDD-Bench tests whether language-model agents can carry out long-horizon small-molecule drug design. Each task gives the agent a scientific objective, molecular data, and a fixed set of tools. The agent works inside an isolated task environment, produces a final artifact, and a separate verifier scores the result.
This Harbor release packages 502 tasks, the original agent harness, and the build sources needed to run the benchmark through Harbor.
SMDD-Bench contains five task families:
| Type | Task | What the agent produces |
|---|---|---|
| 1 | 2D Pharmacophore Identification | A pharmacophore that separates active and inactive molecules |
| 2 | Interaction Point Discovery | Conserved interaction points in a binding pocket |
| 3 | Scaffold Hopping | A molecule with a different scaffold that retains binding interactions |
| 4 | Lead Optimization | A molecule meeting property, binding, and structural constraints |
| 5 | Fragment Assembly | A drug-like binder assembled from the supplied fragments |
The exact output format is defined in each task's instruction.md.
A Harbor run has three moving pieces:
The Harbor host runs the Python agent harness, calls the language-model provider, and launches the task and verifier containers. This machine needs Docker, Python, and your model credentials. It can be a CPU host or the same GPU VM used by the oracle.
The task and verifier containers isolate what the agent sees from what the evaluator sees. Harbor stages the public task files in the task container, then evaluates the submitted artifact in a separate verifier container containing the grading files. Harbor creates and tears down these containers automatically.
The GPU oracle is a persistent Docker service that runs ADMET-AI and Boltz behind an authenticated HTTP endpoint. The agent uses it during a trajectory, and the Types 3-5 verifiers use it for Boltz scoring. Type 4 verification additionally computes ADMET metrics inside the CPU verifier environment.
In other words, Harbor owns the trial lifecycle and isolation; the oracle provides the expensive scientific models.
You need at least one Linux machine with an NVIDIA GPU for the oracle. Installing the Python package on a laptop is not enough: the deployed oracle requires CUDA, an NVIDIA driver, Docker, and NVIDIA Container Toolkit configured for docker run --gpus all.
We recommend at least 48 GB of VRAM per GPU to reduce Boltz out-of-memory failures. This applies to every GPU used by the oracle, including Ray heads and workers. Exact memory use depends on the protein, sampling settings, and concurrency, so 48 GB is a practical deployment target rather than a guarantee.
For higher-throughput runs, an 8 x NVIDIA A100 80 GB machine is one option. Each Boltz worker uses a single GPU; GPU memory is not pooled across cards, so an individual prediction still has to fit on one GPU. Start smaller if your workload allows it, then increase concurrency after checking GPU and host-memory use.
Also leave enough host RAM for model processes and concurrent verifier containers. Some task manifests request 32 GiB RAM per trial, so inspect task.toml before increasing concurrency.
The agent-facing interface is the same with or without Ray. Ray only changes how the oracle schedules Boltz work.
| Setup | Harbor host | GPU service |
|---|---|---|
| One VM, no Ray | On the GPU VM | One oracle container using its local GPUs |
| Separate hosts, no Ray | On a CPU host | One oracle container on a GPU VM |
| Multiple GPU VMs with Ray | On the head VM or a separate CPU host | One head container and one worker container per additional VM |
With Ray, ADMET stays on the head and Boltz workers run across the available GPUs. The head GPU can also host a Boltz worker. Without Ray, a standalone oracle starts one worker for each GPU visible to Docker.
Run the following on the Linux machine that will act as the Harbor host. If Harbor and the oracle share a GPU VM, this is that machine.
Create a Python 3.12 environment, install Harbor, authenticate, and download the release:
uv venv --python 3.12 "$HOME/.venvs/smdd-harbor"
source "$HOME/.venvs/smdd-harbor/bin/activate"
uv pip install harbor==0.20.0
harbor auth login
harbor download "looni-lab/smdd-bench@candidate" --export --output-dir ./smdd-download
Move into the downloaded dataset directory, install the bundled host-side package, and make sure the tasks are visible:
cd smdd-download/smdd-bench
unset SMDD_BENCH_ROOT
uv pip install ./smdd_harbor_release-0.1.2-py3-none-any.whl
python run.py --list-tasks
The wheel contains the host-side harness, Harbor adapter, and evaluation helpers. Scientific-model dependencies live in the Docker images, so you do not need to install Boltz or ADMET into this Python environment.
Keep this environment activated for the remaining commands.
The release has the following top-level structure:
downloaded-dataset/
|-- README.md # Complete setup and running guide
|-- SETUP.md # Checklist and links into this guide
|-- run.py # Task listing, reference tests and agent trials
|-- release.json # Version, image tags and source metadata
|-- smdd_harbor_release-*.whl # Installable Python harness and adapter
|-- smdd-runtime.zip # Docker and oracle source package
+-- smdd_004_A2A_Adenosine_0__v1/ # One of the task directories
|-- task.toml # Resources, images and verifier settings
|-- instruction.md # Agent-visible instructions
|-- environment/ # Public inputs and task build context
|-- tests/ # Verifier code and private grading inputs
+-- solution/ # Reference solution, where provided
The runner accepts canonical task IDs such as smdd_004_A2A_Adenosine_0 and resolves Harbor's __v1 directory suffix automatically.
The benchmark operator receives the grading files as part of the release, but Harbor stages them only inside the separate verifier runtime. If you modify or repackage a task, preserve that separation.
Before building images, extract the runtime sources on every machine that needs them:
python -m zipfile -e smdd-runtime.zip .
This creates smdd-runtime/:
| Path | Purpose |
|---|---|
build_images.sh |
Builds all images or only the Harbor/oracle group |
agent_env/ |
Agent science, GPU model, and HTTP-service Dockerfiles |
agent_harness/ |
Agent loop, tools, HTTP service, and Ray implementation |
smdd_harbor/ |
Adapter connecting the agent loop to Harbor |
evaluator/, metrics/ |
Evaluation code and scientific scoring functions |
smdd-harbor/docker/ |
Task runtime and five verifier-image build contexts |
scripts/smoke_oracle_ray.py |
Optional multi-node inference diagnostic |
The release ships the Docker build sources rather than prebuilt images. Build the image group needed by each machine.
| Image role | Where needed | How it is used |
|---|---|---|
| Agent science image (smdd-agent) | Harbor host | Build parent for task runtimes |
| Evaluator image (smdd-evals) | Harbor host | Build parent for the lead runtime and verifiers |
| Task science image | Harbor host | Harbor runs Types 1, 2, 3, and 5 in it |
| Task lead image | Harbor host | Harbor runs Type 4 in it |
| Five verifier images: 2d, 3d, scaffold, lead, fragment | Harbor host | Parents for the separate per-task verifier containers |
| Model image (smdd-oracle) | GPU image-build host | Build parent containing Boltz, ADMET, and Ray dependencies |
| HTTP oracle image (smdd-oracle-http) | Every GPU VM | Persistent oracle service started manually |
From the downloaded dataset directory, choose the command that matches the machine's role:
# Harbor and oracle on the same GPU VM:
bash smdd-runtime/build_images.sh all
# OR: a separate Harbor host needs the task/verifier image group:
bash smdd-runtime/build_images.sh harbor
# On the separate GPU VM, build the oracle image group:
bash smdd-runtime/build_images.sh oracle
If you are using Ray, every GPU VM must run the same oracle image. Build it once and transfer that exact image, or pull the same registry digest on every node. Worker VMs do not need the Harbor task or verifier images unless Harbor also runs there.
You do not manually start task or verifier containers. Harbor creates them from the task manifests when a trial runs.
There are two supported layouts. Both expose the same authenticated HTTP API on port 8090, so the Harbor-side commands later in this guide are identical.
Keep oracle traffic on a private network or VPN. The endpoint must be reachable from both the Harbor host and Harbor's verifier containers. In particular, 127.0.0.1 inside a verifier refers to that verifier container, not to the GPU host.
Use this when all oracle GPUs are on one machine.
Run on the GPU VM:
export ORACLE_IMAGE="smdd-oracle-http:latest"
export STATE="$HOME/smdd-oracle/state"
export CACHE="$HOME/smdd-oracle/cache"
install -d -m 700 "$STATE" "$CACHE"
test -s "$STATE/token" || openssl rand -hex 32 > "$STATE/token"
chmod 600 "$STATE/token"
docker run -d --name smdd-oracle-http --restart unless-stopped \
--gpus all --shm-size=16g -p 8090:8090 \
-e ORACLE_ROLE=standalone \
-v "$STATE:/srv/oracle_state" \
-v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" \
--max-samples 10 --max-parallel-samples 2 --microbatch-size 1
docker logs --tail 100 -f smdd-oracle-http
The service automatically discovers the GPUs exposed by Docker. To restrict it to GPU 0, replace --gpus all with --gpus '"device=0"'.
Initial model loading can take several minutes. Ctrl+C while following the logs only exits the log viewer; the detached oracle container keeps running. The 16g shared-memory value is a starting point and can be adjusted for the host and workload.
Use this when you want Boltz requests to be scheduled across multiple GPU machines.
The example below uses two VMs with one GPU each. Every node needs the same oracle image, mutually reachable private or VPN addresses, compatible GPUs, and bidirectional connectivity over the trusted network. Ray uses auxiliary ports in addition to 6379, 6380, 6381, and 20000-20063; opening only 6379 is not enough.
First copy the exact oracle image to each worker, substituting the worker's SSH address:
docker save "smdd-oracle-http:latest" | ssh ubuntu@WORKER_SSH_ADDRESS docker load
Set HEAD_IP to the head VM's own private or VPN address:
export HEAD_IP="10.77.0.1"
export ORACLE_IMAGE="smdd-oracle-http:latest"
export STATE="$HOME/smdd-ray/state"
export CACHE="$HOME/smdd-ray/cache"
install -d -m 700 "$STATE" "$CACHE"
test -s "$STATE/token" || openssl rand -hex 32 > "$STATE/token"
chmod 600 "$STATE/token"
docker run -d --name smdd-head --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=head -e RAY_NODE_IP="$HEAD_IP" \
-e RAY_EXPECTED_GPUS=2 -e ORACLE_HTTP_HOST=0.0.0.0 \
-v "$STATE:/srv/oracle_state" \
-v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" \
--max-samples 10 --max-parallel-samples 2 --microbatch-size 1
Start the worker after the head container has been created; you do not need to wait for the head HTTP endpoint to become ready.
export HEAD_IP="10.77.0.1"
export WORKER_IP="10.77.0.2"
export ORACLE_IMAGE="smdd-oracle-http:latest"
docker run -d --name smdd-worker --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=worker -e RAY_NODE_IP="$WORKER_IP" \
-e RAY_ADDRESS="$HEAD_IP:6379" \
"$ORACLE_IMAGE"
Repeat the worker command on additional VMs, changing only WORKER_IP for each machine while keeping the same HEAD_IP.
RAY_EXPECTED_GPUS is the total GPU count across the whole cluster, including the head. The head waits for that many Boltz actors before its HTTP endpoint becomes ready.
Workers do not serve HTTP and do not need the bearer token. Shared results and artifacts live on the head, so the VMs do not need a shared filesystem. Every node does need outbound access to the MSA service, normally api.colabfold.com.
Back on the head, check that Ray sees the expected nodes and GPUs:
docker exec smdd-head /opt/oracle/bin/ray status
docker logs --tail 100 smdd-head
For this two-GPU example, you should see two active nodes and two GPUs. If a worker is missing, inspect docker logs smdd-worker on that machine and verify that it can reach the head's private address. Do not substitute a public NAT address for a node's local bind address.
If you want the Ray dashboard, forward its loopback port from your computer:
ssh -N -L 8265:127.0.0.1:8265 ubuntu@HEAD_SSH_ADDRESS
Then open http://localhost:8265 while that SSH session remains active.
For either deployment, run the health check from the GPU host or Ray head using an address that the Harbor host can also reach:
export ORACLE_IP="10.77.0.1"
curl --fail --silent --show-error \
-H "Authorization: Bearer $(cat "$STATE/token")" \
"http://$ORACLE_IP:8090/v1/health" | python -m json.tool
Wait until the service and model report healthy before launching trials. A connection refusal during startup can simply mean that the models are still loading; check the oracle logs before assuming the service failed.
Keep --max-samples 10. This controls the maximum number of diffusion samples per request, not the number of GPUs. The benchmark verifiers request ten samples even when the oracle has only one GPU.
Run this on the Harbor host from the downloaded dataset directory.
If the oracle is on another machine, securely copy its token file to the Harbor host first and use that host-local absolute path here:
export SMDD_AGENT_ORACLE_HTTP_URL="http://10.77.0.1:8090"
export SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE="/absolute/path/to/token"
export SMDD_AGENT_ORACLE_CACHE_DIR="$HOME/.cache/smdd-harbor"
curl --fail --silent --show-error \
-H "Authorization: Bearer $(cat "$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE")" \
"$SMDD_AGENT_ORACLE_HTTP_URL/v1/health" | python -m json.tool
Replace the example IP and token path with the values for your deployment. With Ray, the URL always points to the head.
Your model-provider credentials also belong on this trusted Harbor host. For example, the Claude command later in this guide expects ANTHROPIC_API_KEY.
The runner configures the agent's oracle client and passes the required oracle credentials into Harbor's verifier environment without printing the token in generated commands. The task container itself remains network-isolated: the trusted host makes prediction requests on the agent's behalf, while verifier containers use their own network access to contact the oracle.
A successful host-side health check proves that the host can reach the oracle, but not necessarily that verifier containers can. The smoke tests below exercise the complete path.
Before spending model calls, verify that Harbor, the task images, the verifier images, and the oracle can complete a real evaluation.
Start with one supplied reference solution:
python run.py --task smdd_001_P39086_0 --smoke
Then run one reference task from each task family:
for task in smdd_001_P39086_0 smdd_002_O75874_0 smdd_003_3KWZ_2 smdd_004_A2A_Adenosine_0 smdd_005_O60674_0; do
python run.py --task "$task" --smoke || break
done
Types 1-2 can verify their reference solutions locally. Types 3-5 require the oracle.
Reference solutions are bundled only for selected tasks, so use these tests to validate the installation rather than as a way to run the full benchmark.
Once the smoke tests pass, export the model credential and launch a real task:
export ANTHROPIC_API_KEY="YOUR_ANTHROPIC_API_KEY"
python run.py --task smdd_004_A2A_Adenosine_0 --model anthropic/claude-sonnet-4-6
The same command works whether the oracle is standalone or backed by Ray. The harness coordinates the model and tool calls; the oracle handles scientific prediction scheduling.
Harbor writes results, submitted artifacts, and verifier logs under jobs/ next to run.py.
Useful runner options:
| Option | Purpose |
|---|---|
--list-tasks |
List canonical task IDs without starting containers |
--task ID |
Select one task |
--model NAME |
Choose the LiteLLM model for a real trial |
--smoke |
Run the reference solution instead of an LLM |
--dataset-dir PATH |
Use tasks from another downloaded directory |
--print-command |
Inspect the generated Harbor command without launching it |
--concurrency N |
Set Harbor trial concurrency; does not allocate GPUs |
The helper runs one selected task and one attempt. Increasing --concurrency does not create extra rollouts by itself; it only controls how many already-scheduled trials Harbor may run simultaneously.
Start with one concurrent trial. Once that works reliably, scale to larger batches while watching GPU memory, host memory, and oracle queue time.
When you are finished, stop only the persistent containers that you started manually:
# Without Ray, on the GPU VM:
docker stop smdd-oracle-http
# With Ray: first on the worker VM, then on the head VM:
docker stop smdd-worker
docker stop smdd-head
These commands leave the containers and their mounted state, cache, and token files in place.
For a standalone oracle, docker start smdd-oracle-http restarts the existing container. For Ray, restart the head and then the workers and wait for the cluster to become healthy again. If you need to change environment variables or startup arguments, recreate the containers while keeping the same persistent mounts.