Loading datasets…
Loading datasets…
Loading datasets…
Organization collection
Datasets published by this organization.
Organization inventory
Private datasets are available only to members of this organization.
| Dataset | AccessVisibility | Tasks |
|---|---|---|
blobfishai/factorybench-100 100 executable manufacturing ERP workflow tasks with deterministic FactoryScore grading | Public | 100 |
blobfishai/devopsbench-100 100 deterministic long-horizon DevOps/SRE agent tasks over an executable NovaCart world with 32-67 call reference trajectories and state-diff vcode verifiers. | Public | 100 |
blobfishai/salesbench-100 100 high-level sales operations requests with evidence-determined multi-system trajectories. | Public | 100 |
blobfishai/counselbench-100 100 high-level legal-agent tasks with 97-asset multi-provider evidence rooms, distinct deep trajectories, and native causal state verification. | Public | 100 |
blobfishai/ledgerbench-100 100 causal corporate-finance workflows over a closed ERP sandbox: provider-shaped MCP tools, distributed evidence, exact state transitions, and no LLM judge. | Public | 100 |
blobfishai/arc-crm-6 Six independently authored synthetic CRM workflows on one mocked CLI/web/REST/MCP world per episode. Partial Arc-inspired coverage, not source-row reproduction. | Public | 6 |
blobfishai/domainbench-24 24 deeply realistic stateful MCP tasks across six enterprise domains | Public | 24 |
blobfishai/webbench WebBench: browser-only correlated SaaS work over deterministic populated Blobfish worlds — 159 tasks graded executably against the applications backends (WebScore, no LLM judge) | Public | 159 |
blobfishai/hubbench HubBench: one oracle-proven, deterministically graded Blobfish benchmark family per Harbor Hub professional-domain cluster — 104 stateful multi-system employee-decision tasks over isolated SQLite worlds (HubScore, no LLM judge) | Public | 104 |
blobfishai/dealbench-100-suite 100 executable investment-banking workflows with deterministic DealScore grading | Public | 100 |
blobfishai/erpbench-100-suite 100 executable Oracle-Fusion-shaped ERP workflows with deterministic ERPScore grading | Public | 100 |
blobfishai/litigation Executable litigation environments for agent training and evaluation. 100 long-horizon tasks, a median of 108 verified tool-calling steps each, worked across eight systems of record (Clio Manage, iManage Work, CM/ECF, Google Calendar, Docusign, Relativity, CourtListener and nine litigation sub-systems). Graded by deterministic state-diff and trace assertions — no LLM judge on the reward path. | Public | 100 |
blobfishai/semikongbench-100 100 executable semiconductor exception workflows with deterministic SemiOpsScore grading | Public | 100 |
13 datasets
Organization collection
Datasets published by this organization.
Organization inventory
Private datasets are available only to members of this organization.
| Dataset | AccessVisibility | Tasks |
|---|---|---|
blobfishai/factorybench-100 100 executable manufacturing ERP workflow tasks with deterministic FactoryScore grading | Public | 100 |
blobfishai/devopsbench-100 100 deterministic long-horizon DevOps/SRE agent tasks over an executable NovaCart world with 32-67 call reference trajectories and state-diff vcode verifiers. | Public | 100 |
blobfishai/salesbench-100 100 high-level sales operations requests with evidence-determined multi-system trajectories. | Public | 100 |
blobfishai/counselbench-100 100 high-level legal-agent tasks with 97-asset multi-provider evidence rooms, distinct deep trajectories, and native causal state verification. | Public | 100 |
blobfishai/ledgerbench-100 100 causal corporate-finance workflows over a closed ERP sandbox: provider-shaped MCP tools, distributed evidence, exact state transitions, and no LLM judge. | Public | 100 |
blobfishai/arc-crm-6 Six independently authored synthetic CRM workflows on one mocked CLI/web/REST/MCP world per episode. Partial Arc-inspired coverage, not source-row reproduction. | Public | 6 |
blobfishai/domainbench-24 24 deeply realistic stateful MCP tasks across six enterprise domains | Public | 24 |
blobfishai/webbench WebBench: browser-only correlated SaaS work over deterministic populated Blobfish worlds — 159 tasks graded executably against the applications backends (WebScore, no LLM judge) | Public | 159 |
blobfishai/hubbench HubBench: one oracle-proven, deterministically graded Blobfish benchmark family per Harbor Hub professional-domain cluster — 104 stateful multi-system employee-decision tasks over isolated SQLite worlds (HubScore, no LLM judge) | Public | 104 |
blobfishai/dealbench-100-suite 100 executable investment-banking workflows with deterministic DealScore grading | Public | 100 |
blobfishai/erpbench-100-suite 100 executable Oracle-Fusion-shaped ERP workflows with deterministic ERPScore grading | Public | 100 |
blobfishai/litigation Executable litigation environments for agent training and evaluation. 100 long-horizon tasks, a median of 108 verified tool-calling steps each, worked across eight systems of record (Clio Manage, iManage Work, CM/ECF, Google Calendar, Docusign, Relativity, CourtListener and nine litigation sub-systems). Graded by deterministic state-diff and trace assertions — no LLM judge on the reward path. | Public | 100 |
blobfishai/semikongbench-100 100 executable semiconductor exception workflows with deterministic SemiOpsScore grading | Public | 100 |
13 datasets