userbench/UserBench

UserBench eval: 620 next-message tasks across 62 developers. Composer 2.5 classifies gold and predicted messages into multi-label acts; reward is set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench

UserBench

UserBench is a Harbor eval for how well an agent can stand in for a real software engineer mid coding-agent session.

At each held-out turn, the agent reads the conversation so far (/sim/history.md) and writes the single next developer message to /sim/answer.txt. A judge labels that message into a 4-way move taxonomy; reward is 1.0 iff the predicted move matches the gold move (move-match), else 0.0.

This package is the public noprofile cut: tasks ship history only. Developer style profiles are optional Harbor skills at job time (--skill / agents[].skills), not part of the task instruction.

Scale (this revision)

Developers 62 (everyone with ≥10 eval points; 6 with <10 dropped)
Tasks 620 (exactly 10 per developer)
Condition noprofile

Task selection rule

Same rule as build_agentic.py --per-dev N:

  1. Start from the held-out eval points in the clean UserBench v2 cohort.
  2. Keep developers with ≥10 points; drop the rest.
  3. For each remaining developer, sort points by environment/history.md byte size ascending (equivalent to sorting by len(context) in the builder) and take the first 10.
  4. Sessions may be shared across selected points for a developer. Ties break by legacy task directory name ascending.

Task naming

Harbor requires org/name package refs. Tasks are published as:

userbench/<username>__<hash>

Examples:

  • GitHub: userbench/winksaville__aad4a4f9
  • DataClaw (anonymized): userbench/dc_004__05fe03e7

Versioning is via Harbor tags only (e.g. @v2), not embedded in the task name.

Keywords

userbench · user-simulation · coding-agents · move-prediction · noprofile

Taxonomy / judging

Moves are one of: approve · critical · directive · inquiry (fault-first decision rule).

The task verifier classifies the agent’s predicted message (and gold real when gold_move is null) with a configurable judge (SIMBENCH_JUDGE, default Gemini). Published agentic numbers on userbench.vercel.app/results used Composer 2.5 as the job-time judge on a 226-point / 10-developer subset of the older package — those runs are not full-620 results.

Train / held-out split

Each developer has a deep train history and a strictly later held-out set. Eval tasks are prediction points carved from held-out sessions (noprofile: history for the held conversation only).

How to reference

What Ref
Dataset userbench/UserBench@v2 (also latest)
Task userbench/<username>__<hash>@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench
harbor run -d userbench/UserBench@v2 -a <agent> -m <model>
harbor run -p userbench/winksaville__aad4a4f9@v2 -a <agent> -m <model>

Links

Task
userbench/PJensen__8bfa5cf3
userbench/nathanbooth-konecta__d2090df8
userbench/thieso2__711016f6
userbench/marcus-sa__27ae199d
userbench/cyyeh__3cc7d29d
userbench/FSM1__1fe3ba80
userbench/henryph24__abd778f8
userbench/raman325__f4d22eb8
userbench/yyovil__8e748208
userbench/dc_000__b5514499
userbench/nathanbooth-konecta__33b00343
userbench/marcus-sa__0121e4af
userbench/henryph24__7290ea77
userbench/johyunduk__263db43d
userbench/winksaville__aad4a4f9
userbench/christso__0244b4d0
userbench/jobinlawrance__7cd7e40a
userbench/scottdensmore__e7c0c3f1
userbench/khaong__d9031f8a
userbench/kmiki0__672300e9
userbench/khaong__e1ad66fe
userbench/armelhbobdad__dfebbec3
userbench/achildrenmile__2ff9e2ca
userbench/Soph__7302b475
userbench/gabadi__7b40d2af
userbench/KeKs0r__f7d88343
userbench/jakobtfaber__92492978
userbench/ET-NoahDolev__e4e19c4b
userbench/penso__e8f7bfab
userbench/dcambria__37a1120c
userbench/ababushkin__c4f5c7c5
userbench/henryph24__32550e4b
userbench/scottdensmore__878978e8
userbench/melagiri__e0dad522
userbench/cyyeh__a430a77c
userbench/raman325__9dacf299
userbench/penso__ea51f5a6
userbench/penso__56c1b069
userbench/robouden__36ed725c
userbench/yyovil__0acc2714
userbench/dc_010__ff7c22f6
userbench/christso__7616b4d3
userbench/lyston11__ec0ffdd5
userbench/armelhbobdad__7a6c95c5
userbench/penso__ccf3440a
userbench/yutakobayashidev__eb53a405
userbench/admarble__af4442f0
userbench/dc_004__9a96d08d
userbench/Stark-Industries0417__1c7b5273
userbench/FSM1__123d9f43
userbench/kmiki0__35023941
userbench/johyunduk__5a1334df
userbench/ET-NoahDolev__d33c5ef7
userbench/achildrenmile__9b4c1a14
userbench/ET-NoahDolev__9f401a7b
userbench/johyunduk__cb36cb1b
userbench/dc_010__609f9b8a
userbench/Poytr1__79b74c55
userbench/wildlily1021__8a0b4073
userbench/Poytr1__7704b402
userbench/johyunduk__ed18bb52
userbench/gabadi__5ff54caa
userbench/Poytr1__a2acb2d7
userbench/khaong__3e62c23c
userbench/junaid-appointy__beeefe8f
userbench/jeevanpillay__71e57a23
userbench/nosman__490dad19
userbench/alishakawaguchi__c005ae79
userbench/cyyeh__87913fda
userbench/KeKs0r__a3626c5d
userbench/armelhbobdad__b387bc68
userbench/mvanhorn__47cf85ab
userbench/kohaku500__9036119a
userbench/4thwithme__e0b170b3
userbench/achildrenmile__43d956ea
userbench/heddendorp__1866113f
userbench/robouden__0104b37c
userbench/oddessentials__0494e728
userbench/malkoG__c2d73c4f
userbench/junaid-appointy__cee3b6be
userbench/heddendorp__e9ee70fb
userbench/penso__05181f3e
userbench/KeKs0r__a3599b78
userbench/hutusi__332ec1ee
userbench/wildlily1021__08bd9283
userbench/blackgirlbytes__c3c0d0c7
userbench/4thwithme__96d539a8
userbench/johyunduk__0daa688f
userbench/khaong__8bd56ca2
userbench/winksaville__4f87e52f
userbench/christso__c229da20
userbench/malkoG__654e600d
userbench/ta93abe__e1d7b4a4
userbench/yutakobayashidev__c9c5ca60
userbench/marcus-sa__8f91d1b5
userbench/hutusi__ba6fef9a
userbench/malkoG__d0e54495
userbench/jskswamy__a7f1470b
userbench/MohammedMqat__eaaeae4f
userbench/dc_010__71b80cd3

Displaying 100 of 620 tasks