blobfishai/factorybench-100
100 executable manufacturing ERP workflow tasks with deterministic FactoryScore grading
harbor run -d blobfishai/factorybench-100license: cc-by-4.0 task_categories:
- tool-use
- question-answering language:
- en tags:
- manufacturing
- erp
- agents
- mcp
- deterministic-evaluation pretty_name: FactoryBench-100 size_categories:
- n<1K configs:
- config_name: default
data_files:
- split: test path: data/tasks.jsonl
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for long-horizon enterprise ERP agents. Each independently authored task runs in an isolated SQLite snapshot and exposes documented Oracle Fusion Cloud 26a REST operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations over synthetic state.
Harbor runs the authoritative SQLite state and trace in a private root-owned sidecar. The agent container receives only the task instruction, typed tool CLI, and tool schemas; it does not receive the database, runtime, verifier, or gold state.
The data is entirely synthetic. No Oracle software, proprietary UI, customer record, or copied task is included.
Metric
The single metric is FactoryScore:
100 × deterministic workflow checks passed / checks available, averaged over
all evaluated tasks.
Checks cover read-before-write controls, exact ERP state transitions, exact answer fields, write-scope containment, and error-free execution. Strict pass is reported only as supporting evidence.
Qualification
- Reference oracle: 100.00 FactoryScore, 100/100 strict passes
- Incomplete-workflow control: 85.71
- Read-only control: 42.86
- No-control ablation: 14.29
- Deterministic replay sample: 10/10 matched
- Single-mutation omission checks: 300/300 detected
These rows are measured controls, not claims about frontier models.
Pinned model runs
| Model | Harness | Coverage | FactoryScore | Strict passes | Selection |
|---|---|---|---|---|---|
| gpt-5.6-luna | Harbor 0.21.0 / codex 0.150.1 / medium | 100/100 | 85.29 | 0/100 | Full v2 suite (100/100 tasks). Codex 0.150.1 was preinstalled in the agent image; all prompts, tools, private sidecars, and verifiers were byte-identical to the released tasks. |
Coverage is part of the result. A stratified subset is not presented as a
100-task score. Full manifests and task-level traces are mirrored under
model-runs/.
Fields
Each JSONL row includes the natural-language prompt, role, workflow family, 12 heterogeneous context files, required tools and reads, allowed write tables, human-readable rubric, and metric contract. Executable worlds, oracle traces, exact verifier specifications, and Harbor tasks live in the source repository.
Links
- Website: https://blobfish.ai/benchmarks/factorybench-100
- Source and executable world: https://github.com/blobfishai/factory-agent-simulation
- Harbor: https://hub.harborframework.com/datasets/blobfishai/factorybench-100/latest
Influences
The release takes design inspiration from Enterprise-Bench, ERP-Bench, Mercor APEX, and Archipelago. All task records and implementations in FactoryBench are independently authored.
| Task |
|---|
blobfishai/factorybench-037 |
blobfishai/factorybench-066 |
blobfishai/factorybench-098 |
blobfishai/factorybench-001 |
blobfishai/factorybench-096 |
blobfishai/factorybench-094 |
blobfishai/factorybench-049 |
blobfishai/factorybench-036 |
blobfishai/factorybench-010 |
blobfishai/factorybench-021 |
blobfishai/factorybench-055 |
blobfishai/factorybench-013 |
blobfishai/factorybench-018 |
blobfishai/factorybench-038 |
blobfishai/factorybench-019 |
blobfishai/factorybench-062 |
blobfishai/factorybench-003 |
blobfishai/factorybench-067 |
blobfishai/factorybench-064 |
blobfishai/factorybench-034 |
blobfishai/factorybench-054 |
blobfishai/factorybench-093 |
blobfishai/factorybench-028 |
blobfishai/factorybench-092 |
blobfishai/factorybench-084 |
blobfishai/factorybench-040 |
blobfishai/factorybench-023 |
blobfishai/factorybench-061 |
blobfishai/factorybench-077 |
blobfishai/factorybench-057 |
blobfishai/factorybench-053 |
blobfishai/factorybench-080 |
blobfishai/factorybench-004 |
blobfishai/factorybench-087 |
blobfishai/factorybench-012 |
blobfishai/factorybench-052 |
blobfishai/factorybench-051 |
blobfishai/factorybench-078 |
blobfishai/factorybench-043 |
blobfishai/factorybench-079 |
blobfishai/factorybench-015 |
blobfishai/factorybench-072 |
blobfishai/factorybench-083 |
blobfishai/factorybench-050 |
blobfishai/factorybench-075 |
blobfishai/factorybench-033 |
blobfishai/factorybench-002 |
blobfishai/factorybench-039 |
blobfishai/factorybench-035 |
blobfishai/factorybench-016 |
blobfishai/factorybench-068 |
blobfishai/factorybench-088 |
blobfishai/factorybench-089 |
blobfishai/factorybench-025 |
blobfishai/factorybench-086 |
blobfishai/factorybench-032 |
blobfishai/factorybench-056 |
blobfishai/factorybench-063 |
blobfishai/factorybench-081 |
blobfishai/factorybench-076 |
blobfishai/factorybench-007 |
blobfishai/factorybench-097 |
blobfishai/factorybench-011 |
blobfishai/factorybench-041 |
blobfishai/factorybench-031 |
blobfishai/factorybench-046 |
blobfishai/factorybench-095 |
blobfishai/factorybench-099 |
blobfishai/factorybench-026 |
blobfishai/factorybench-085 |
blobfishai/factorybench-074 |
blobfishai/factorybench-059 |
blobfishai/factorybench-047 |
blobfishai/factorybench-070 |
blobfishai/factorybench-024 |
blobfishai/factorybench-091 |
blobfishai/factorybench-022 |
blobfishai/factorybench-073 |
blobfishai/factorybench-048 |
blobfishai/factorybench-069 |
blobfishai/factorybench-060 |
blobfishai/factorybench-014 |
blobfishai/factorybench-065 |
blobfishai/factorybench-030 |
blobfishai/factorybench-082 |
blobfishai/factorybench-006 |
blobfishai/factorybench-071 |
blobfishai/factorybench-020 |
blobfishai/factorybench-029 |
blobfishai/factorybench-042 |
blobfishai/factorybench-005 |
blobfishai/factorybench-045 |
blobfishai/factorybench-017 |
blobfishai/factorybench-009 |
blobfishai/factorybench-058 |
blobfishai/factorybench-090 |
blobfishai/factorybench-task-100 |
blobfishai/factorybench-044 |
blobfishai/factorybench-008 |
blobfishai/factorybench-027 |
Displaying 100 of 100 tasks