userbench/UserBench-train400

UserBench train400 twin: the same 620 held tasks plus 400 leak-safe prior turns per developer. Composer 2.5 multi-label acts score by set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench-train400

UserBench-train400

Twin of the public UserBench eval with leak-safe prior-session context.

Same 620 held tasks / 62 developers as userbench/UserBench@v2 (history, gold, verifier unchanged). Each task additionally ships earlier sessions for that developer under /sim/train/, totaling 400 human user-turns sampled with sqrt two-stage.

What the agent sees

Path Role
/sim/history.md Held conversation so far (identical to baseline)
/sim/train/_index.json Index of packed prior sessions (ids, timestamps, repos, turn takes, paths)
/sim/train/<sid>.md Contiguous prefixes of those sessions (same > DEVELOPER / > AGENT format)
/sim/answer.txt Where the agent writes the next developer message

The instruction is count-agnostic: it points at /sim/train/ without hardcoding “400”. The turn budget is a publish-time sampler parameter.

Leakage invariant

For each developer let T* = min(held_sessions[].ts) (manifest session timestamp).

Every packed train session satisfies:

  1. session.ts < T*
  2. max(turn.ts) < T*

This is stricter than the cohort’s session-start train/held split alone. Sessions that wall-clock-overlap the earliest held session are never included.

QC on this package: 0 / 620 tasks violate max(turn.ts) < T*.

Sampling

  1. Start from the developer’s full leak-safe train pool (all sessions/turns passing the T* rule).
  2. Allocate exactly 400 human-turns with sqrt_two_stage (time-stratified session pick, then contiguous-prefix waterfill).
  3. Emit prefixes through the k-th human turn (including intervening AGENT/TOOL/SYSTEM), with the same ~200-word-per-turn truncation as eval history.

The full pools (not just the 400-turn subsample) live in the GitHub repo under train_pools/ so future Hub variants (e.g. train1000) can resample from the same T*-filtered corpus without changing the leakage rule.

Scale

Developers 62
Tasks 620 (10 per developer; same selection as baseline)
Condition train400
Train budget 400 human-turns / developer (shared across that developer’s 10 tasks)

Task naming

Same short IDs as baseline (Harbor org/name):

userbench/<username>__<hash>

Versioning is via Harbor tags on this package (e.g. @v2). Digests differ from the noprofile baseline; consume via the dataset ref below so baseline UserBench@v2 stays pinned to zero-train digests.

Keywords

userbench · user-simulation · coding-agents · move-prediction · train400

How to reference

What Ref
This twin userbench/UserBench-train400@v2 (also latest)
Zero-train baseline userbench/UserBench@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench-train400
Full train pools (GitHub) https://github.com/AlienKevin/user-simulator/tree/main/train_pools
harbor run -d userbench/UserBench-train400@v2 -a <agent> -m <model>

Links

Task
userbench/dc_004__90ee87f8
userbench/Poytr1__e42f2f97
userbench/penso__ea51f5a6
userbench/jobinlawrance__8d3fd0f5
userbench/malkoG__d0e54495
userbench/malkoG__58805f63
userbench/cyyeh__9a9888b5
userbench/135yshr__446e7a75
userbench/heddendorp__9c1e8c6a
userbench/dc_010__4814d75a
userbench/Poytr1__b720c576
userbench/dcambria__10d6ba09
userbench/jskswamy__636b5fea
userbench/jeevanpillay__342c66a8
userbench/christso__0244b4d0
userbench/manderson240__4d6f4d55
userbench/henryph24__28f7ca31
userbench/johyunduk__263db43d
userbench/lyston11__c145a945
userbench/yutakobayashidev__95ae4c7c
userbench/henryph24__e8f45c35
userbench/dc_010__ff7c22f6
userbench/khaong__06bbbed2
userbench/ET-NoahDolev__e4e19c4b
userbench/jhoetter__5acf62c0
userbench/ababushkin__ec41cd5d
userbench/nathanbooth-konecta__617a942a
userbench/nathanbooth-konecta__62642c4b
userbench/PJensen__4ebea46d
userbench/Poytr1__a2acb2d7
userbench/marcus-sa__288f068e
userbench/4thwithme__725a3b3c
userbench/henryph24__d2760f4a
userbench/thieso2__5b029bc6
userbench/alishakawaguchi__fcb29da8
userbench/lyston11__581de3ef
userbench/penso__99f8e6bd
userbench/oddessentials__90f1206d
userbench/yyovil__3b746732
userbench/armelhbobdad__d91ed9bd
userbench/henryph24__c2387b32
userbench/marcus-sa__b4fa4432
userbench/khaong__10e38ded
userbench/marcus-sa__e7c46cf0
userbench/armelhbobdad__8d99ddd6
userbench/mvanhorn__c15f2981
userbench/ta93abe__1d0c822b
userbench/kohaku500__24642283
userbench/kmiki0__364954f2
userbench/jskswamy__9cd70f59
userbench/penso__56c1b069
userbench/nosman__490dad19
userbench/winksaville__94139dc3
userbench/cyyeh__87913fda
userbench/jakobtfaber__fbdba2d7
userbench/KeKs0r__f7d88343
userbench/malkoG__c2d73c4f
userbench/ET-NoahDolev__9f401a7b
userbench/TheurgicDuke771__fe9210ab
userbench/nosman__a1b3187b
userbench/135yshr__b4924a51
userbench/dc_010__609f9b8a
userbench/alishakawaguchi__e546a69e
userbench/mvanhorn__47cf85ab
userbench/FSM1__d3bad63b
userbench/PJensen__d78917b2
userbench/singampalliveerendra__bfe1ab76
userbench/barbogast__2bfad951
userbench/FSM1__be1ec78f
userbench/kmiki0__da54e58e
userbench/jskswamy__e6f3d857
userbench/ababushkin__7f121b4c
userbench/lyston11__3d28b804
userbench/blackgirlbytes__c3b99d32
userbench/hutusi__332ec1ee
userbench/khaong__3589873f
userbench/penso__89b3a2d0
userbench/yutakobayashidev__eb53a405
userbench/dc_000__aa12035f
userbench/jhoetter__60915a1f
userbench/thieso2__67587edd
userbench/Stark-Industries0417__e6fddb86
userbench/lyston11__5fac2a4d
userbench/dc_010__5a9e3805
userbench/dc_000__4afde71e
userbench/johyunduk__8bd19a58
userbench/raman325__9dacf299
userbench/gabadi__7b40d2af
userbench/henryph24__7290ea77
userbench/Stark-Industries0417__8c531ba1
userbench/wildlily1021__08bd9283
userbench/armelhbobdad__248d4e03
userbench/yutakobayashidev__0341b973
userbench/johyunduk__0daa688f
userbench/johyunduk__ed18bb52
userbench/yyovil__357f2675
userbench/TheurgicDuke771__53dd0e94
userbench/dc_001__4e16b43d
userbench/135yshr__cf5e60ca
userbench/hutusi__d53dbb64

Displaying 100 of 620 tasks