userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/PJensen__8bfa5cf3 |
userbench/nathanbooth-konecta__d2090df8 |
userbench/thieso2__711016f6 |
userbench/marcus-sa__27ae199d |
userbench/cyyeh__3cc7d29d |
userbench/FSM1__1fe3ba80 |
userbench/henryph24__abd778f8 |
userbench/raman325__f4d22eb8 |
userbench/yyovil__8e748208 |
userbench/dc_000__b5514499 |
userbench/nathanbooth-konecta__33b00343 |
userbench/marcus-sa__0121e4af |
userbench/henryph24__7290ea77 |
userbench/johyunduk__263db43d |
userbench/winksaville__aad4a4f9 |
userbench/christso__0244b4d0 |
userbench/jobinlawrance__7cd7e40a |
userbench/scottdensmore__e7c0c3f1 |
userbench/khaong__d9031f8a |
userbench/kmiki0__672300e9 |
userbench/khaong__e1ad66fe |
userbench/armelhbobdad__dfebbec3 |
userbench/achildrenmile__2ff9e2ca |
userbench/Soph__7302b475 |
userbench/gabadi__7b40d2af |
userbench/KeKs0r__f7d88343 |
userbench/jakobtfaber__92492978 |
userbench/ET-NoahDolev__e4e19c4b |
userbench/penso__e8f7bfab |
userbench/dcambria__37a1120c |
userbench/ababushkin__c4f5c7c5 |
userbench/henryph24__32550e4b |
userbench/scottdensmore__878978e8 |
userbench/melagiri__e0dad522 |
userbench/cyyeh__a430a77c |
userbench/raman325__9dacf299 |
userbench/penso__ea51f5a6 |
userbench/penso__56c1b069 |
userbench/robouden__36ed725c |
userbench/yyovil__0acc2714 |
userbench/dc_010__ff7c22f6 |
userbench/christso__7616b4d3 |
userbench/lyston11__ec0ffdd5 |
userbench/armelhbobdad__7a6c95c5 |
userbench/penso__ccf3440a |
userbench/yutakobayashidev__eb53a405 |
userbench/admarble__af4442f0 |
userbench/dc_004__9a96d08d |
userbench/Stark-Industries0417__1c7b5273 |
userbench/FSM1__123d9f43 |
userbench/kmiki0__35023941 |
userbench/johyunduk__5a1334df |
userbench/ET-NoahDolev__d33c5ef7 |
userbench/achildrenmile__9b4c1a14 |
userbench/ET-NoahDolev__9f401a7b |
userbench/johyunduk__cb36cb1b |
userbench/dc_010__609f9b8a |
userbench/Poytr1__79b74c55 |
userbench/wildlily1021__8a0b4073 |
userbench/Poytr1__7704b402 |
userbench/johyunduk__ed18bb52 |
userbench/gabadi__5ff54caa |
userbench/Poytr1__a2acb2d7 |
userbench/khaong__3e62c23c |
userbench/junaid-appointy__beeefe8f |
userbench/jeevanpillay__71e57a23 |
userbench/nosman__490dad19 |
userbench/alishakawaguchi__c005ae79 |
userbench/cyyeh__87913fda |
userbench/KeKs0r__a3626c5d |
userbench/armelhbobdad__b387bc68 |
userbench/mvanhorn__47cf85ab |
userbench/kohaku500__9036119a |
userbench/4thwithme__e0b170b3 |
userbench/achildrenmile__43d956ea |
userbench/heddendorp__1866113f |
userbench/robouden__0104b37c |
userbench/oddessentials__0494e728 |
userbench/malkoG__c2d73c4f |
userbench/junaid-appointy__cee3b6be |
userbench/heddendorp__e9ee70fb |
userbench/penso__05181f3e |
userbench/KeKs0r__a3599b78 |
userbench/hutusi__332ec1ee |
userbench/wildlily1021__08bd9283 |
userbench/blackgirlbytes__c3c0d0c7 |
userbench/4thwithme__96d539a8 |
userbench/johyunduk__0daa688f |
userbench/khaong__8bd56ca2 |
userbench/winksaville__4f87e52f |
userbench/christso__c229da20 |
userbench/malkoG__654e600d |
userbench/ta93abe__e1d7b4a4 |
userbench/yutakobayashidev__c9c5ca60 |
userbench/marcus-sa__8f91d1b5 |
userbench/hutusi__ba6fef9a |
userbench/malkoG__d0e54495 |
userbench/jskswamy__a7f1470b |
userbench/MohammedMqat__eaaeae4f |
userbench/dc_010__71b80cd3 |
Displaying 100 of 620 tasks