harbor run -d blobfishai/domainbench-24harbor run -d blobfishai/domainbench-2424 deeply realistic stateful MCP tasks across six enterprise domains
harbor run -d blobfishai/domainbench-2424 deeply realistic stateful MCP tasks across six enterprise domains
harbor run -d blobfishai/domainbench-24DomainBench-24 v3.3.2 is a 24-task, six-domain benchmark for stateful tool-using agents. Every task is a high-level employee request with exactly 26 provider-native evidence reads plus 17–36 task-relevant live-domain reads, for 43–62 causal evidence reads before the decision. Each task also includes 28 agent-visible source files in nine native formats, three exact costed and dated operating options with authority status, 13 dependent derivations, 56–67 task-specific causal criteria, an isolated SQLite snapshot, a typed MCP tool surface, a semantically graded natural Slack stakeholder handoff, provider-native exact-record readbacks for every changed domain row, persisted result readback, and deterministic state/trace/workpaper verification. Hidden reference wording and JSON serialization are not graded. Exact total call order is not graded; causal read-before-write and post-write readback are.
The checked-in reference plans and exact replays pass 24/24 tasks. Thirteen task-applicable adversarial controls—including correct-state/wrong-reasoning, wrong-target, and keyword-stuffing controls—are rejected for every task: 312/312 correct rejections across 360 pristine executions. Reference plans are solvability controls, not model leaderboard entries.
harbor run -d blobfishai/domainbench-24 -a oracle
World source, task evidence, and public trajectories: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24/tree/main/worlds
Dataset mirror: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24
DomainBench-24 v3.3.2 is a 24-task, six-domain benchmark for stateful tool-using agents. Every task is a high-level employee request with exactly 26 provider-native evidence reads plus 17–36 task-relevant live-domain reads, for 43–62 causal evidence reads before the decision. Each task also includes 28 agent-visible source files in nine native formats, three exact costed and dated operating options with authority status, 13 dependent derivations, 56–67 task-specific causal criteria, an isolated SQLite snapshot, a typed MCP tool surface, a semantically graded natural Slack stakeholder handoff, provider-native exact-record readbacks for every changed domain row, persisted result readback, and deterministic state/trace/workpaper verification. Hidden reference wording and JSON serialization are not graded. Exact total call order is not graded; causal read-before-write and post-write readback are.
The checked-in reference plans and exact replays pass 24/24 tasks. Thirteen task-applicable adversarial controls—including correct-state/wrong-reasoning, wrong-target, and keyword-stuffing controls—are rejected for every task: 312/312 correct rejections across 360 pristine executions. Reference plans are solvability controls, not model leaderboard entries.
harbor run -d blobfishai/domainbench-24 -a oracle
World source, task evidence, and public trajectories: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24/tree/main/worlds
Dataset mirror: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24