blobfishai/arc-crm-6

Six independently authored synthetic CRM workflows on one mocked CLI/web/REST/MCP world per episode. Partial Arc-inspired coverage, not source-row reproduction.

harbor run -d blobfishai/arc-crm-6

Blobfish Arc CRM 6 · 0.1.2

Six independently authored synthetic CRM workflows with 31 domain/evidence tools on three mock servers, two engine controls, and 33 synthetic PDF/XLSX/EML assets. Each episode shares one stateful world across CLI, real HTML forms, REST and MCP.

Scope and results

All six frozen reference solutions pass local Docker execution with strict score 100 and exact final-state/trace parity. All 167 authoring negative controls reject. These are oracle solvability checks, not model leaderboard results. No model is ranked. Independent qualification admits the complete task set, not a sample.

The inspiration, Arc-Intelligence/arc-crm-benchmark, has 1,200 conversations and 27 observed tool names at its source pin. We inspected those rows but reproduce zero source rows. Six new workflows map 16 observed names behind 15 independent contracts; this is partial source-inspired coverage, not upstream API/ABI parity, and not adaptation of every Harbor or Hugging Face catalog record.

This dataset binds the locally qualified packages. Registry execution and remote-object receipts are separate evidence; this publication-time card does not claim those gates.

harbor run -d blobfishai/arc-crm-6@v0.1.2 -a oracle -e docker -n 1 -k 1 -r 0

Explore and reproduce

  • Clean implementation
  • Hugging Face dataset
  • Harbor dataset
  • release.json: qualified freeze identity and exact task digests.
  • source-lock.json: upstream pin, inspected fields, license and coverage limits.
  • qualification.json: authoring traces, replays and negative controls.
  • local-oracle-receipt.json: all-six local Docker artifact hashes; workstation paths omitted, original receipt digest retained.

The Hugging Face distribution includes frozen/ (the complete unmodified input freeze, including its original candidate manifest), data/tasks.jsonl (six flat preview rows), evidence/ (collected snapshots and oracle traces) and separate registry receipts. The frozen candidate labels are immutable historical inputs; the external publication receipts describe the later qualification stages.

Asset file_sha256 / file_bytes bind the native PDF/XLSX/EML download. content_sha256 and legacy sha256 instead bind the full UTF-8 evidence text. Native PDFs wrap and paginate without clipping; independent extraction tests cover every PDF, and frozen admission checks complete-text rendering. Version 0.1.1 passed runtime oracles but truncated nine PDF files; its immutable commit and receipts remain historical evidence, not this corrected release.

Isolation and grading

Harbor 0.21.0 and Docker are required. The non-root UID10001 agent has only public client/schema/evidence. Its image has no seed state, builder, oracle or verifier. An internal network connects main (1 CPU/1 GiB/pids256) to world (0.5 CPU/512 MiB/pids128). Zero GPUs or paid model calls are needed for oracles. The random private snapshot credential exists only in the world container.

Harbor stops main, collects the world snapshot, destroys the agent environment, then starts a fresh networkless verifier (1 CPU/1 GiB/pids256). Agent-writable log files are not trusted. State-diff, required reads/writes, protected rows and table counts are graded deterministically by HubScore; there is no LLM judge. Harbor reward is 1 only for strict pass. Recorded scores are diagnostic, not model ranks.

Build timeout: 600s; agent: 1200s; collection: 40s; verifier: 120s. The declared 2 GiB storage budget is metadata, not a Docker-enforced disk quota. Docker cleanup removes trial containers/networks/volumes, not the downloaded dataset or evidence. Do not mount the complete package into an agent. solution/ is oracle-only and tests/ belongs only in the separate grading container.

All organizations and evidence are synthetic. Code/fixtures are Apache-2.0; the upstream card's MIT label is source metadata, not a license assertion for its linked code. HubBench names in the shared engine are attribution, not membership in the separate HubBench benchmark. See LICENSE and ENGINE-NOTICE in frozen/.

Task
blobfishai/arc-crm-004
blobfishai/arc-crm-001
blobfishai/arc-crm-005
blobfishai/arc-crm-003
blobfishai/arc-crm-006
blobfishai/arc-crm-002

Displaying 6 of 6 tasks