userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/jeevanpillay__fa06bda8 |
userbench/penso__fe93fa2b |
userbench/malkoG__4ca6601d |
userbench/oddessentials__845a955a |
userbench/mvanhorn__89ae7ff7 |
userbench/kungfusaini__8314296d |
userbench/nathanbooth-konecta__35376685 |
userbench/dc_004__7c46d305 |
userbench/ababushkin__ec41cd5d |
userbench/jakobtfaber__8441d390 |
userbench/ta93abe__2f9b0c7c |
userbench/blackgirlbytes__64f0503c |
userbench/melagiri__476a562b |
userbench/nosman__a1b3187b |
userbench/PJensen__4ebea46d |
userbench/nosman__bd4601c6 |
userbench/135yshr__b4924a51 |
userbench/alishakawaguchi__3dc3da3f |
userbench/KeKs0r__7d6d7ccd |
userbench/jakobtfaber__2a61e65f |
userbench/dc_001__d5eb57d4 |
userbench/raman325__d86a71c6 |
userbench/scottdensmore__83347f87 |
userbench/jhoetter__8c270a1b |
userbench/ta93abe__2d529cb3 |
userbench/PJensen__d9c743d6 |
userbench/4thwithme__725a3b3c |
userbench/cyyeh__61ced54e |
userbench/hutusi__066149dd |
userbench/achildrenmile__97a0d30b |
userbench/ababushkin__fd476f56 |
userbench/ET-NoahDolev__744c71ca |
userbench/henryph24__e8f45c35 |
userbench/khaong__10e38ded |
userbench/Soph__33197b4e |
userbench/135yshr__d1edc0a3 |
userbench/jhoetter__60915a1f |
userbench/dc_001__aefe92aa |
userbench/jeevanpillay__454b9dde |
userbench/manderson240__9267ae98 |
userbench/henryph24__28f7ca31 |
userbench/MohammedMqat__8483a6a4 |
userbench/singampalliveerendra__dc989aa1 |
userbench/junaid-appointy__35091490 |
userbench/lyston11__182e418e |
userbench/jeevanpillay__342c66a8 |
userbench/nosman__416e9b74 |
userbench/TheurgicDuke771__1b8bda9a |
userbench/yyovil__6e71d5f5 |
userbench/malkoG__58805f63 |
userbench/dc_001__63608a3b |
userbench/admarble__63ccad05 |
userbench/heddendorp__6c4a125a |
userbench/jakobtfaber__bc4316a7 |
userbench/heddendorp__7d0856ac |
userbench/robouden__c6e9e844 |
userbench/oddessentials__379f489c |
userbench/mvanhorn__cf848111 |
userbench/admarble__347929ab |
userbench/malkoG__908ed694 |
userbench/khaong__3589873f |
userbench/manderson240__48f851af |
userbench/alishakawaguchi__8c8bffaf |
userbench/TheurgicDuke771__e0560a3c |
userbench/melagiri__25b8f546 |
userbench/TheurgicDuke771__b6a77404 |
userbench/wildlily1021__1261af91 |
userbench/dc_010__9ef010fa |
userbench/gabadi__cc9e1d70 |
userbench/PJensen__428a45fb |
userbench/ta93abe__22fa2a69 |
userbench/wildlily1021__a84e53a9 |
userbench/cyyeh__9a9888b5 |
userbench/manderson240__3c73d247 |
userbench/fcamblor__aa53f12b |
userbench/christso__dee1a922 |
userbench/4thwithme__c9912712 |
userbench/barbogast__5c442045 |
userbench/FSM1__53bf4e87 |
userbench/jhoetter__5acf62c0 |
userbench/manderson240__82601d5d |
userbench/MohammedMqat__f5ba11ff |
userbench/wildlily1021__85caed49 |
userbench/lyston11__733da7f7 |
userbench/christso__6f5610eb |
userbench/winksaville__94139dc3 |
userbench/achildrenmile__35c350fc |
userbench/barbogast__4509fc91 |
userbench/135yshr__e3ee00a3 |
userbench/scottdensmore__344cacf9 |
userbench/wildlily1021__ff1c6f83 |
userbench/Poytr1__9c1714f3 |
userbench/dc_000__d46bdcc2 |
userbench/christso__f277e5cc |
userbench/dc_000__4afde71e |
userbench/KeKs0r__638dbf35 |
userbench/junaid-appointy__cb0b492b |
userbench/kohaku500__24642283 |
userbench/barbogast__47bb82ec |
userbench/nosman__19985a6e |
Displaying 100 of 620 tasks