harbor run -d shiv-eshwar/lb6-dev-pilot-issue18-rewardkitLamina Product Coding Pilot — 3-arm (direct / plan / lamina). Small product apps, LLM-judge on product code, median of 3 seeds. Development-only; not confirmatory LaminaBench-6.
harbor run -d shiv-eshwar/lb6-dev-pilot-issue18-rewardkitHarbor dataset: shiv-eshwar/lb6-dev-pilot-issue18-rewardkit
Hub: https://hub.harborframework.com/datasets/shiv-eshwar/lb6-dev-pilot-issue18-rewardkit
Compares three ways to build small product apps with the same agent and judge:
| Arm | What it means |
|---|---|
direct |
Code straight from the brief |
plan |
Short plan, then code |
lamina |
Lamina skills (/lamina-init + /lamina-design), then code |
Development-only · not confirmatory LaminaBench-6 · not a marketing claim.
12 public tasks = 4 product tasks × 3 arms:
| Task | Arms |
|---|---|
| loan-library | direct · plan · lamina |
| review-room | direct · plan · lamina |
| simple-list | direct · plan · lamina |
| toggle-preference | direct · plan · lamina |
/app/app.mjs + /app/ui.mjs (LLM-as-judge).lamina/ process notes; no /lamina-verify step in this pilot(Internals for operators: Harbor RewardKit judge openai/gpt-5.5; GitHub issue #18 shape.)
# Needs: CURSOR_API_KEY, OPENAI_API_KEY, REWARDKIT_JUDGE=openai/gpt-5.5, LITELLM_DROP_PARAMS=1
harbor run \
-d shiv-eshwar/lb6-dev-pilot-issue18-rewardkit@latest \
-a cursor-cli \
-m cursor/composer-2.5 \
--env-file .env \
-n 1
Pinned tag: @issue18-rewardkit-median-n3 (same content as latest at publish time).
For the full local 3-seed / median protocol:
benchmarks/lb6/pilot/publication/REPRODUCE.md
benchmarks/lb6/pilot/publication/harbor-job-urls.tsvOK: under this pilot harness, lamina’s median score is above plan and direct on all four tasks.
Not OK: confirmatory LaminaBench-6 proof; that process artifacts were graded; that every seed matched the same deltas.