userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/jakobtfaber__a0f851cc |
userbench/ta93abe__14a96679 |
userbench/robouden__cb40a5e2 |
userbench/kungfusaini__6409200a |
userbench/armelhbobdad__d8689cd9 |
userbench/khaong__06bbbed2 |
userbench/robouden__8b4b3c4a |
userbench/thieso2__3dbff2d6 |
userbench/oddessentials__eaf6331c |
userbench/johyunduk__4bab3ef7 |
userbench/Stark-Industries0417__9b04eb34 |
userbench/thieso2__67587edd |
userbench/henryph24__82af8b05 |
userbench/jeevanpillay__d951d14a |
userbench/armelhbobdad__248d4e03 |
userbench/kungfusaini__9a3333b8 |
userbench/jskswamy__2fd0350e |
userbench/scottdensmore__7b129602 |
userbench/dcambria__31d94d0a |
userbench/Soph__ed41ff61 |
userbench/TheurgicDuke771__fe9210ab |
userbench/malkoG__f0e27113 |
userbench/jdsingh122918__f6d90727 |
userbench/kmiki0__924a7b32 |
userbench/dcambria__5bdc81bf |
userbench/thieso2__5b029bc6 |
userbench/jhoetter__b49b7e7f |
userbench/achildrenmile__7dafef0e |
userbench/kungfusaini__f8d51bd4 |
userbench/ababushkin__0dafbf18 |
userbench/KeKs0r__a3482769 |
userbench/yyovil__149c4120 |
userbench/Poytr1__b2abde78 |
userbench/Poytr1__e42f2f97 |
userbench/gabadi__ae3008ee |
userbench/raman325__d4a2df8b |
userbench/manderson240__ec275534 |
userbench/gabadi__45c83dac |
userbench/jdsingh122918__96fc2d3b |
userbench/jhoetter__3517bfbd |
userbench/melagiri__3cf814f5 |
userbench/thieso2__8f357020 |
userbench/jakobtfaber__0fddf0ed |
userbench/singampalliveerendra__93f5702f |
userbench/dc_000__4b92ddba |
userbench/admarble__08f6ff97 |
userbench/jdsingh122918__f1bc9841 |
userbench/jhoetter__bcc3790b |
userbench/blackgirlbytes__32cf9092 |
userbench/winksaville__98ba6868 |
userbench/cyyeh__93e85db5 |
userbench/alishakawaguchi__585afa4a |
userbench/4thwithme__db958e1f |
userbench/kmiki0__603c6c1f |
userbench/malkoG__9b248afd |
userbench/yutakobayashidev__f67ef680 |
userbench/wildlily1021__c031a215 |
userbench/dc_000__aa12035f |
userbench/FSM1__8d0fb85e |
userbench/yyovil__a5d654fb |
userbench/Stark-Industries0417__69f46708 |
userbench/MohammedMqat__e972ce4f |
userbench/jskswamy__e7cb0835 |
userbench/jakobtfaber__1395ade5 |
userbench/robouden__f12f7493 |
userbench/fcamblor__5f26916b |
userbench/mvanhorn__716a17ba |
userbench/dc_000__d5528693 |
userbench/scottdensmore__b162aa5c |
userbench/jobinlawrance__a748127e |
userbench/singampalliveerendra__13c48199 |
userbench/fcamblor__10653b23 |
userbench/admarble__1ffea695 |
userbench/fcamblor__41e0acf2 |
userbench/yutakobayashidev__23d09ae6 |
userbench/melagiri__7a34ab4a |
userbench/junaid-appointy__9ae3d2e7 |
userbench/jobinlawrance__69d520ea |
userbench/nathanbooth-konecta__273ef024 |
userbench/jobinlawrance__a1dc6ac0 |
userbench/jdsingh122918__235948db |
userbench/khaong__7e2efac8 |
userbench/oddessentials__b4bcc213 |
userbench/lyston11__c145a945 |
userbench/wildlily1021__7748b2c8 |
userbench/henryph24__d2760f4a |
userbench/Stark-Industries0417__8c531ba1 |
userbench/singampalliveerendra__8dc175b0 |
userbench/penso__d1d5889f |
userbench/achildrenmile__210b9808 |
userbench/hutusi__490bc1e1 |
userbench/blackgirlbytes__077f6188 |
userbench/jobinlawrance__723f7ef8 |
userbench/Stark-Industries0417__fde8b674 |
userbench/barbogast__07017b77 |
userbench/raman325__9730850f |
userbench/dc_004__633347c9 |
userbench/TheurgicDuke771__53dd0e94 |
userbench/gabadi__82b5fa6a |
userbench/nosman__7c59d97b |
Displaying 100 of 620 tasks