Loading dataset details…
harbor run -d aryaniyaps/lamina-benchLaminaBench v1: SkillsBench-paired product implementation benchmark (10 tasks, control + treatment arms)
harbor run -d aryaniyaps/lamina-benchLaminaBench v1 — SkillsBench-paired product implementation benchmark.
taskNNN-control + taskNNN-treatment) for paired evaluationinstruction.md only (no Lamina skills)AGENTS.md/CLAUDE.mdharbor run -d "aryaniyaps/lamina-bench@v1" -a claude-code -m "<model>" \
--ak "prompt_template=benchmarks/harbor/prompt_template.j2"
Publishing uploads task definitions to the Harbor registry. You do not need bench:run, results/, or scored artifacts.
harbor auth login
npm run fixtures:vendor # once, for OSS task fixtures
npm run bench:harbor:publish
Answer y when Harbor asks to confirm making tasks public.
The publish script uploads tasks from benchmarks/harbor/tasks/, refreshes dataset.toml digests from the registry (harbor sync --upgrade), then publishes the dataset manifest.