userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/kmiki0__2a36c043 |
userbench/melagiri__28dbbafd |
userbench/alishakawaguchi__bd8cd4f9 |
userbench/kmiki0__6260a388 |
userbench/manderson240__a3d3d466 |
userbench/cyyeh__0ef3c5c5 |
userbench/dc_004__90ee87f8 |
userbench/yutakobayashidev__95ae4c7c |
userbench/jhoetter__209999e7 |
userbench/ababushkin__b31eb66a |
userbench/jskswamy__9cd70f59 |
userbench/135yshr__cf5e60ca |
userbench/thieso2__2cd329c5 |
userbench/MohammedMqat__4e4e7193 |
userbench/marcus-sa__aafed4ab |
userbench/yyovil__357f2675 |
userbench/johyunduk__f296fc43 |
userbench/yyovil__262c7a5a |
userbench/4thwithme__d2c038e3 |
userbench/hutusi__bb886168 |
userbench/Soph__31443222 |
userbench/winksaville__0d00a824 |
userbench/thieso2__ac3f3046 |
userbench/dc_010__47ffc258 |
userbench/nathanbooth-konecta__74dc6f40 |
userbench/dcambria__8b203ca9 |
userbench/Soph__1a8c6a0c |
userbench/dc_001__0ced58ab |
userbench/alishakawaguchi__095c91a6 |
userbench/Poytr1__b720c576 |
userbench/mvanhorn__4e30e41e |
userbench/penso__89b3a2d0 |
userbench/heddendorp__b574678e |
userbench/singampalliveerendra__28776902 |
userbench/raman325__e67e0975 |
userbench/penso__395b6920 |
userbench/barbogast__2bfad951 |
userbench/achildrenmile__1a13a0ad |
userbench/PJensen__8bc81bf4 |
userbench/raman325__646fecc7 |
userbench/scottdensmore__96489d0d |
userbench/TheurgicDuke771__ad9ab945 |
userbench/jeevanpillay__7941e093 |
userbench/winksaville__e91d7fe5 |
userbench/Poytr1__d19883d2 |
userbench/TheurgicDuke771__01470379 |
userbench/jhoetter__919e665d |
userbench/armelhbobdad__9b514f47 |
userbench/blackgirlbytes__c3b99d32 |
userbench/melagiri__013c6017 |
userbench/dc_010__d717c75a |
userbench/kohaku500__c36c95c2 |
userbench/4thwithme__a2549fc8 |
userbench/kmiki0__f7732c44 |
userbench/yyovil__781caa58 |
userbench/malkoG__a7ddaeaf |
userbench/singampalliveerendra__3f75b898 |
userbench/heddendorp__1d900217 |
userbench/barbogast__4aa039d9 |
userbench/jeevanpillay__e77b5133 |
userbench/kohaku500__80f87c63 |
userbench/dc_010__4814d75a |
userbench/fcamblor__791476df |
userbench/barbogast__2627df09 |
userbench/admarble__2188f606 |
userbench/dcambria__b1663567 |
userbench/thieso2__ef3476ab |
userbench/dc_001__03bfd6bd |
userbench/dc_001__4e16b43d |
userbench/ta93abe__1d0c822b |
userbench/raman325__4750553d |
userbench/winksaville__dd69ca82 |
userbench/singampalliveerendra__b8fbec2d |
userbench/dc_000__52b607d4 |
userbench/135yshr__196168af |
userbench/junaid-appointy__af1dce79 |
userbench/dc_004__1b611cb2 |
userbench/jakobtfaber__bdaadc4b |
userbench/heddendorp__7c00d429 |
userbench/marcus-sa__e7c46cf0 |
userbench/jakobtfaber__8a4f165d |
userbench/nathanbooth-konecta__617a942a |
userbench/Poytr1__4ef838c8 |
userbench/Soph__fc2c3472 |
userbench/Stark-Industries0417__c6832f84 |
userbench/yutakobayashidev__3c5461f5 |
userbench/raman325__429c2fe6 |
userbench/FSM1__be1ec78f |
userbench/kohaku500__38862417 |
userbench/jeevanpillay__5904ac41 |
userbench/jobinlawrance__66e69569 |
userbench/KeKs0r__dbbc1fb2 |
userbench/christso__6858ee27 |
userbench/jdsingh122918__bb859790 |
userbench/jskswamy__89eb1820 |
userbench/singampalliveerendra__de5a3f57 |
userbench/yutakobayashidev__8e7c88b8 |
userbench/admarble__c7561573 |
userbench/fcamblor__7309daf9 |
userbench/jdsingh122918__92c17296 |
Displaying 100 of 620 tasks