Loading dataset details…
harbor run -d infra-bench/infra-bench-v120 AI-infra engineering tasks on AMD MI35X (5 categories: model deployment, profiling, kernel implementation, kernel tuning, debug triage)
harbor run -d infra-bench/infra-bench-v1Github repository: InfraBench
20 real-world AI-infrastructure engineering tasks on AMD MI35X (CDNA4 / gfx950) GPUs, across five categories: model deployment, profiling analysis, kernel implementation, kernel tuning, and debug triage. Each task runs an agent in a Docker sandbox and is scored by a deterministic per-task verifier.
Built on Harbor framework via the harbor CLI.
/dev/kfd + /dev/dri) and Dockerharbor CLI and an LLM-gateway-authenticated agent/data/models (mounted read-only into model-dependent tasks)export INFRABENCH_GPU_COUNT=8
harbor run -d infra-bench/infra-bench-v1 -a claude-code -m <model> -n 6
GPU Pool for Concurrency:
qr-rmsnorm-fusion, sglang-mmmu-ipc-crash);INFRA_PARALLEL ≤ INFRABENCH_GPU_COUNT - 2.flock).jobs/<timestamp>/. verifier/reward.txt is the score (1 = pass, 0 = fail);harbor view jobs to browse trajectories and harbor analyze jobs to analyze failures.