Organization inventory
Private datasets are available only to members of this organization.
| Dataset | AccessVisibility | Tasks |
|---|---|---|
actava-ai/chi-bench χ-Bench — long-horizon, policy-rich U.S. healthcare workflow agent benchmark across provider prior-auth, payer UM, and care management. 78 single-agent tasks on the hub (75 single-domain + 3 marathon); the 23 provider-payer E2E arena tasks need the two-agent dual-pa-e2e harness and run via the source repo, not the hub. Each task builds a self-contained image on demand (no hosted image/registry): Harbor clones the source repo and downloads the public fixtures dataset at build time. PREREQUISITES: Docker + Harbor CLI, and an APPROVED Hugging Face token for the gated Managed-Care Operations Handbook (every task needs it; fetched at container start). Request access: https://huggingface.co/datasets/actava/managed-care-operations-handbook RUN (HF_TOKEN from shell/--env-file; -y confirms, -i picks one task): HF_TOKEN=<approved-token> harbor run -d actava-ai/chi-bench@v1.0.1 -i actava-ai/pa_t016_t016_o001_p01_p2p_payer -a claude-code -m claude-opus-4-7 -y LINKS — Code/runner/CLI: https://github.com/actava-ai/chi-bench · Fixtures dataset: https://huggingface.co/datasets/actava/chi-bench · Handbook (gated): https://huggingface.co/datasets/actava/managed-care-operations-handbook · Leaderboard: https://actava.ai/benchmarks/leaderboards · Paper: https://arxiv.org/abs/2605.16679 · Harbor-hub guide: https://github.com/actava-ai/chi-bench/blob/main/docs/harbor-hub.md | Public | 78 |
1 dataset