userbench/UserBench

UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench

UserBench

UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.

At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.

This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.

Scale (this revision)

Developers 62 (everyone with ≥10 eval points; 6 with <10 dropped)
Tasks 620 (exactly 10 per developer)
Condition noprofile

Task selection rule

Same rule as build_agentic.py --per-dev N:

  1. Start from the held-out eval points in the clean UserBench v2 cohort.
  2. Keep developers with ≥10 points; drop the rest.
  3. For each remaining developer, sort points by environment/history.md byte size ascending (equivalent to sorting by len(context) in the builder) and take the first 10.
  4. Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.

Task naming

Harbor requires org/name package refs. Tasks are published as:

userbench/<username>__<hash>

Examples:

  • GitHub: userbench/winksaville__aad4a4f9
  • DataClaw (anonymized): userbench/dc_004__05fe03e7

Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.

Keywords

userbench · user-simulation · coding-agents · move-prediction · noprofile

Taxonomy / judging

Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).

The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.

Train / held-out split

Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).

How to reference

What Ref
Dataset userbench/UserBench@v2 (also latest)
Task userbench/<username>__<hash>@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>

Links

Task
userbench/Soph__d4aeb5f7
userbench/fcamblor__b7014237
userbench/thieso2__f58a3bca
userbench/Soph__a13f4134
userbench/kohaku500__cc54a542
userbench/PJensen__2366c990
userbench/Poytr1__792397bd
userbench/kmiki0__364954f2
userbench/winksaville__7d7bb7df
userbench/jdsingh122918__9d260baa
userbench/winksaville__0b1794e9
userbench/barbogast__9c63e30b
userbench/dc_004__b1a1b133
userbench/yutakobayashidev__32f6e7cb
userbench/marcus-sa__60346b6a
userbench/scottdensmore__17cdc6a5
userbench/junaid-appointy__311a12f4
userbench/winksaville__0c62d5b9
userbench/blackgirlbytes__df26765a
userbench/dc_001__7381cb8f
userbench/cyyeh__5f330796
userbench/armelhbobdad__1ca1c892
userbench/dc_004__ce4ab540
userbench/MohammedMqat__eb4f7f08
userbench/gabadi__aba2eda8
userbench/4thwithme__56e34cff
userbench/135yshr__824ae900
userbench/marcus-sa__56b8c4e7
userbench/MohammedMqat__9ce28792
userbench/135yshr__488a7841
userbench/blackgirlbytes__42416c99
userbench/Stark-Industries0417__3c378e34
userbench/FSM1__9ca02217
userbench/Soph__d2bccb86
userbench/ET-NoahDolev__d2ea0f07
userbench/oddessentials__49939de2
userbench/4thwithme__1e9b9622
userbench/135yshr__676c5289
userbench/robouden__e896e954
userbench/melagiri__4203583e
userbench/henryph24__c2387b32
userbench/marcus-sa__b4fa4432
userbench/jskswamy__e6f3d857
userbench/marcus-sa__81996502
userbench/FSM1__a876f41a
userbench/jskswamy__636b5fea
userbench/cyyeh__af946993
userbench/christso__e5946a37
userbench/barbogast__601f787e
userbench/dc_001__d88c2c32
userbench/KeKs0r__c7dd8b9b
userbench/dcambria__91b175cb
userbench/TheurgicDuke771__ec7192a4
userbench/ababushkin__15c2bcc4
userbench/khaong__18ba6389
userbench/yyovil__3b746732
userbench/kohaku500__d39c7601
userbench/scottdensmore__42381bd0
userbench/alishakawaguchi__54a1160b
userbench/junaid-appointy__d8cc1ea6
userbench/dc_004__a833e6e4
userbench/lyston11__5cd5012e
userbench/PJensen__2744e784
userbench/malkoG__9b476884
userbench/dc_000__25e7cf92
userbench/armelhbobdad__d91ed9bd
userbench/kohaku500__d7998363
userbench/lyston11__3d28b804
userbench/dc_000__49970ee1
userbench/henryph24__c072ea86
userbench/robouden__b294b008
userbench/jobinlawrance__015f6990
userbench/dc_010__fa91db13
userbench/nathanbooth-konecta__8e4c4321
userbench/jskswamy__45376c4b
userbench/Stark-Industries0417__3c662166
userbench/yyovil__61ad7991
userbench/ta93abe__a68fb6da
userbench/oddessentials__74f6a0b5
userbench/heddendorp__9c1e8c6a
userbench/MohammedMqat__27bcfe81
userbench/mvanhorn__8a8dc66c
userbench/henryph24__f0fc860a
userbench/dc_010__5a9e3805
userbench/PJensen__409f7f83
userbench/nosman__5b7c798e
userbench/jobinlawrance__2edaea1b
userbench/kungfusaini__051dc260
userbench/nathanbooth-konecta__62642c4b
userbench/ET-NoahDolev__6ab6543d
userbench/barbogast__015a4d1e
userbench/melagiri__130f32d7
userbench/kungfusaini__9d0a7d9e
userbench/ET-NoahDolev__a018815b
userbench/raman325__a9ce30a5
userbench/mvanhorn__7921dff4
userbench/johyunduk__8bd19a58
userbench/Stark-Industries0417__e6fddb86
userbench/manderson240__2bac58a0
userbench/khaong__0b023426

Displaying 100 of 620 tasks