userbench/UserBench
UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.
Published 7/21/2026 by Kevin Xiang Li
harbor run -d userbench/UserBenchUserBench
UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.
At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.
This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.
Scale (this revision)
| Developers | 62 (everyone with ≥10 eval points; 6 with <10 dropped) |
| Tasks | 620 (exactly 10 per developer) |
| Condition | noprofile |
Task selection rule
Same rule as build_agentic.py --per-dev N:
- Start from the held-out eval points in the clean UserBench v2 cohort.
- Keep developers with ≥10 points; drop the rest.
- For each remaining developer, sort points by
environment/history.mdbyte size ascending (equivalent to sorting bylen(context)in the builder) and take the first 10. - Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.
Task naming
Harbor requires org/name package refs. Tasks are published as:
userbench/<username>__<hash>
Examples:
- GitHub:
userbench/winksaville__aad4a4f9 - DataClaw (anonymized):
userbench/dc_004__05fe03e7
Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.
Keywords
userbench · user-simulation · coding-agents · move-prediction · noprofile
Taxonomy / judging
Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).
The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.
Train / held-out split
Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).
How to reference
| What | Ref |
|---|---|
| Dataset | userbench/UserBench@v2 (also latest) |
| Task | userbench/<username>__<hash>@v2 |
| Hub | https://hub.harborframework.com/datasets/userbench/UserBench |
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>
Links
- Site / dataset EDA: https://userbench.vercel.app
- Agentic results (10-dev / 226-point slice of the prior package): https://userbench.vercel.app/results
| Task |
|---|
userbench/kohaku500__31ef4143 |
userbench/KeKs0r__90c80d6d |
userbench/PJensen__d78917b2 |
userbench/mvanhorn__a4b76ce8 |
userbench/marcus-sa__288f068e |
userbench/4thwithme__404e13d3 |
userbench/jeevanpillay__2d2d580e |
userbench/admarble__b47ec152 |
userbench/singampalliveerendra__bfe1ab76 |
userbench/christso__d8ca1042 |
userbench/blackgirlbytes__eb3010a3 |
userbench/PJensen__1a77191b |
userbench/dcambria__311fbdc2 |
userbench/mvanhorn__c15f2981 |
userbench/thieso2__491f5eda |
userbench/kohaku500__829d75fd |
userbench/robouden__158fa32c |
userbench/manderson240__5c697026 |
userbench/robouden__cfbea6a0 |
userbench/jskswamy__81abacf0 |
userbench/dc_004__74515f8d |
userbench/achildrenmile__71a8a165 |
userbench/jobinlawrance__8d3fd0f5 |
userbench/admarble__bacf1d41 |
userbench/melagiri__af1394ab |
userbench/singampalliveerendra__91bc17ba |
userbench/ta93abe__3b70eafa |
userbench/junaid-appointy__cac807fd |
userbench/heddendorp__8a98faf8 |
userbench/kungfusaini__9f8077cf |
userbench/ababushkin__7f121b4c |
userbench/nosman__458811d2 |
userbench/jhoetter__2da05fce |
userbench/jhoetter__eb3780f4 |
userbench/alishakawaguchi__e546a69e |
userbench/jdsingh122918__44679c06 |
userbench/ET-NoahDolev__cf912278 |
userbench/fcamblor__962f5e71 |
userbench/135yshr__446e7a75 |
userbench/FSM1__d3bad63b |
userbench/TheurgicDuke771__dd18f7dc |
userbench/oddessentials__90f1206d |
userbench/dcambria__3462644a |
userbench/dcambria__cff3e93f |
userbench/135yshr__e4b9747c |
userbench/hutusi__67ea1d83 |
userbench/hutusi__d53dbb64 |
userbench/kungfusaini__0dcc9805 |
userbench/admarble__dbb54fbd |
userbench/manderson240__c62b91cf |
userbench/dc_010__b9fc11b2 |
userbench/armelhbobdad__8d99ddd6 |
userbench/johyunduk__b15e8506 |
userbench/KeKs0r__ac28eed7 |
userbench/jobinlawrance__ba8c72c7 |
userbench/scottdensmore__6d44807a |
userbench/dc_001__536cd200 |
userbench/TheurgicDuke771__99a7df5d |
userbench/nathanbooth-konecta__30525504 |
userbench/yutakobayashidev__52dbe4cd |
userbench/armelhbobdad__2856056c |
userbench/jakobtfaber__fbdba2d7 |
userbench/wildlily1021__2e8bcca3 |
userbench/hutusi__aa2cfed0 |
userbench/fcamblor__d55d0502 |
userbench/cyyeh__8615d833 |
userbench/ababushkin__55f888fa |
userbench/gabadi__de9f94c6 |
userbench/heddendorp__8c0f454e |
userbench/kmiki0__a9bb3ed2 |
userbench/lyston11__5fac2a4d |
userbench/Soph__ec880138 |
userbench/ta93abe__25e7c1b9 |
userbench/ta93abe__8cd6cb7b |
userbench/MohammedMqat__09e14a28 |
userbench/johyunduk__75a332c5 |
userbench/jskswamy__264ee35c |
userbench/kmiki0__da54e58e |
userbench/junaid-appointy__915b9090 |
userbench/oddessentials__0558d1c7 |
userbench/ababushkin__3c505e52 |
userbench/fcamblor__4c9bb47a |
userbench/mvanhorn__8c97d4de |
userbench/achildrenmile__5d92c930 |
userbench/MohammedMqat__81504953 |
userbench/dc_000__e5a61347 |
userbench/wildlily1021__e4c4eb11 |
userbench/manderson240__4d6f4d55 |
userbench/kungfusaini__e97ae043 |
userbench/nathanbooth-konecta__8fa7119c |
userbench/dc_004__667e2977 |
userbench/lyston11__c3e8ff4c |
userbench/jdsingh122918__d44c4e2b |
userbench/alishakawaguchi__fcb29da8 |
userbench/jeevanpillay__335616a6 |
userbench/jdsingh122918__170092ed |
userbench/FSM1__271c699b |
userbench/nosman__d7d30e52 |
userbench/blackgirlbytes__116d4a16 |
userbench/alishakawaguchi__04c32a20 |
Displaying 100 of 620 tasks