userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/Soph__d4aeb5f7 |
userbench/fcamblor__b7014237 |
userbench/thieso2__f58a3bca |
userbench/Soph__a13f4134 |
userbench/kohaku500__cc54a542 |
userbench/PJensen__2366c990 |
userbench/Poytr1__792397bd |
userbench/kmiki0__364954f2 |
userbench/winksaville__7d7bb7df |
userbench/jdsingh122918__9d260baa |
userbench/winksaville__0b1794e9 |
userbench/barbogast__9c63e30b |
userbench/dc_004__b1a1b133 |
userbench/yutakobayashidev__32f6e7cb |
userbench/marcus-sa__60346b6a |
userbench/scottdensmore__17cdc6a5 |
userbench/junaid-appointy__311a12f4 |
userbench/winksaville__0c62d5b9 |
userbench/blackgirlbytes__df26765a |
userbench/dc_001__7381cb8f |
userbench/cyyeh__5f330796 |
userbench/armelhbobdad__1ca1c892 |
userbench/dc_004__ce4ab540 |
userbench/MohammedMqat__eb4f7f08 |
userbench/gabadi__aba2eda8 |
userbench/4thwithme__56e34cff |
userbench/135yshr__824ae900 |
userbench/marcus-sa__56b8c4e7 |
userbench/MohammedMqat__9ce28792 |
userbench/135yshr__488a7841 |
userbench/blackgirlbytes__42416c99 |
userbench/Stark-Industries0417__3c378e34 |
userbench/FSM1__9ca02217 |
userbench/Soph__d2bccb86 |
userbench/ET-NoahDolev__d2ea0f07 |
userbench/oddessentials__49939de2 |
userbench/4thwithme__1e9b9622 |
userbench/135yshr__676c5289 |
userbench/robouden__e896e954 |
userbench/melagiri__4203583e |
userbench/henryph24__c2387b32 |
userbench/marcus-sa__b4fa4432 |
userbench/jskswamy__e6f3d857 |
userbench/marcus-sa__81996502 |
userbench/FSM1__a876f41a |
userbench/jskswamy__636b5fea |
userbench/cyyeh__af946993 |
userbench/christso__e5946a37 |
userbench/barbogast__601f787e |
userbench/dc_001__d88c2c32 |
userbench/KeKs0r__c7dd8b9b |
userbench/dcambria__91b175cb |
userbench/TheurgicDuke771__ec7192a4 |
userbench/ababushkin__15c2bcc4 |
userbench/khaong__18ba6389 |
userbench/yyovil__3b746732 |
userbench/kohaku500__d39c7601 |
userbench/scottdensmore__42381bd0 |
userbench/alishakawaguchi__54a1160b |
userbench/junaid-appointy__d8cc1ea6 |
userbench/dc_004__a833e6e4 |
userbench/lyston11__5cd5012e |
userbench/PJensen__2744e784 |
userbench/malkoG__9b476884 |
userbench/dc_000__25e7cf92 |
userbench/armelhbobdad__d91ed9bd |
userbench/kohaku500__d7998363 |
userbench/lyston11__3d28b804 |
userbench/dc_000__49970ee1 |
userbench/henryph24__c072ea86 |
userbench/robouden__b294b008 |
userbench/jobinlawrance__015f6990 |
userbench/dc_010__fa91db13 |
userbench/nathanbooth-konecta__8e4c4321 |
userbench/jskswamy__45376c4b |
userbench/Stark-Industries0417__3c662166 |
userbench/yyovil__61ad7991 |
userbench/ta93abe__a68fb6da |
userbench/oddessentials__74f6a0b5 |
userbench/heddendorp__9c1e8c6a |
userbench/MohammedMqat__27bcfe81 |
userbench/mvanhorn__8a8dc66c |
userbench/henryph24__f0fc860a |
userbench/dc_010__5a9e3805 |
userbench/PJensen__409f7f83 |
userbench/nosman__5b7c798e |
userbench/jobinlawrance__2edaea1b |
userbench/kungfusaini__051dc260 |
userbench/nathanbooth-konecta__62642c4b |
userbench/ET-NoahDolev__6ab6543d |
userbench/barbogast__015a4d1e |
userbench/melagiri__130f32d7 |
userbench/kungfusaini__9d0a7d9e |
userbench/ET-NoahDolev__a018815b |
userbench/raman325__a9ce30a5 |
userbench/mvanhorn__7921dff4 |
userbench/johyunduk__8bd19a58 |
userbench/Stark-Industries0417__e6fddb86 |
userbench/manderson240__2bac58a0 |
userbench/khaong__0b023426 |
Displaying 100 of 620 tasks