userbench/UserBench-train400

UserBench train400 twin: the same 620 held tasks plus 400 leak-safe prior turns per developer. Composer 2.5 multi-label acts score by set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench-train400

UserBench-train400

Twin of the public UserBench eval with leak-safe prior-session context.

Same 620 held tasks / 62 developers as userbench/UserBench@v2 (history, gold, verifier unchanged). Each task additionally ships earlier sessions for that developer under /sim/train/, totaling 400 human user-turns sampled with sqrt two-stage.

What the agent sees

Path Role
/sim/history.md Held conversation so far (identical to baseline)
/sim/train/_index.json Index of packed prior sessions (ids, timestamps, repos, turn takes, paths)
/sim/train/<sid>.md Contiguous prefixes of those sessions (same > DEVELOPER / > AGENT format)
/sim/answer.txt Where the agent writes the next developer message

The instruction is count-agnostic: it points at /sim/train/ without hardcoding “400”. The turn budget is a publish-time sampler parameter.

Leakage invariant

For each developer let T* = min(held_sessions[].ts) (manifest session timestamp).

Every packed train session satisfies:

  1. session.ts < T*
  2. max(turn.ts) < T*

This is stricter than the cohort’s session-start train/held split alone. Sessions that wall-clock-overlap the earliest held session are never included.

QC on this package: 0 / 620 tasks violate max(turn.ts) < T*.

Sampling

  1. Start from the developer’s full leak-safe train pool (all sessions/turns passing the T* rule).
  2. Allocate exactly 400 human-turns with sqrt_two_stage (time-stratified session pick, then contiguous-prefix waterfill).
  3. Emit prefixes through the k-th human turn (including intervening AGENT/TOOL/SYSTEM), with the same ~200-word-per-turn truncation as eval history.

The full pools (not just the 400-turn subsample) live in the GitHub repo under train_pools/ so future Hub variants (e.g. train1000) can resample from the same T*-filtered corpus without changing the leakage rule.

Scale

Developers 62
Tasks 620 (10 per developer; same selection as baseline)
Condition train400
Train budget 400 human-turns / developer (shared across that developer’s 10 tasks)

Task naming

Same short IDs as baseline (Harbor org/name):

userbench/<username>__<hash>

Versioning is via Harbor tags on this package (e.g. @v2). Digests differ from the noprofile baseline; consume via the dataset ref below so baseline UserBench@v2 stays pinned to zero-train digests.

Keywords

userbench · user-simulation · coding-agents · move-prediction · train400

How to reference

What Ref
This twin userbench/UserBench-train400@v2 (also latest)
Zero-train baseline userbench/UserBench@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench-train400
Full train pools (GitHub) https://github.com/AlienKevin/user-simulator/tree/main/train_pools
harbor run -d userbench/UserBench-train400@v2 -a <agent> -m <model>

Links

Task
userbench/lyston11__733da7f7
userbench/winksaville__dd69ca82
userbench/dc_001__7381cb8f
userbench/ta93abe__8cd6cb7b
userbench/dcambria__cff3e93f
userbench/heddendorp__7c00d429
userbench/alishakawaguchi__54a1160b
userbench/FSM1__271c699b
userbench/PJensen__1a77191b
userbench/4thwithme__d2c038e3
userbench/malkoG__908ed694
userbench/Soph__ec880138
userbench/jeevanpillay__454b9dde
userbench/Poytr1__b2abde78
userbench/yyovil__781caa58
userbench/PJensen__428a45fb
userbench/Stark-Industries0417__69f46708
userbench/oddessentials__74f6a0b5
userbench/admarble__bacf1d41
userbench/scottdensmore__b162aa5c
userbench/jeevanpillay__2d2d580e
userbench/manderson240__82601d5d
userbench/thieso2__2cd329c5
userbench/jobinlawrance__2edaea1b
userbench/TheurgicDuke771__ad9ab945
userbench/Soph__ed41ff61
userbench/ET-NoahDolev__d33c5ef7
userbench/wildlily1021__e4c4eb11
userbench/raman325__f4d22eb8
userbench/MohammedMqat__9ce28792
userbench/heddendorp__e9ee70fb
userbench/marcus-sa__27ae199d
userbench/ababushkin__b31eb66a
userbench/yutakobayashidev__3c5461f5
userbench/armelhbobdad__d8689cd9
userbench/ababushkin__55f888fa
userbench/kmiki0__924a7b32
userbench/fcamblor__41e0acf2
userbench/kohaku500__31ef4143
userbench/melagiri__013c6017
userbench/cyyeh__0ef3c5c5
userbench/jakobtfaber__bc4316a7
userbench/manderson240__3c73d247
userbench/wildlily1021__a84e53a9
userbench/ta93abe__a68fb6da
userbench/christso__d8ca1042
userbench/fcamblor__b7014237
userbench/jdsingh122918__170092ed
userbench/jhoetter__8c270a1b
userbench/winksaville__0c62d5b9
userbench/christso__e5946a37
userbench/Stark-Industries0417__fde8b674
userbench/jhoetter__209999e7
userbench/FSM1__a876f41a
userbench/ET-NoahDolev__2951d3d1
userbench/barbogast__2627df09
userbench/penso__d1d5889f
userbench/barbogast__9c63e30b
userbench/Soph__7302b475
userbench/dc_001__d5eb57d4
userbench/raman325__d4a2df8b
userbench/mvanhorn__cf848111
userbench/achildrenmile__35c350fc
userbench/kungfusaini__051dc260
userbench/penso__05181f3e
userbench/ta93abe__14a96679
userbench/jskswamy__e7cb0835
userbench/heddendorp__6c4a125a
userbench/singampalliveerendra__b8fbec2d
userbench/heddendorp__8c0f454e
userbench/TheurgicDuke771__ec7192a4
userbench/MohammedMqat__27bcfe81
userbench/achildrenmile__2ff9e2ca
userbench/hutusi__bb886168
userbench/dcambria__5bdc81bf
userbench/kohaku500__9036119a
userbench/dc_010__d717c75a
userbench/mvanhorn__7921dff4
userbench/robouden__8b4b3c4a
userbench/4thwithme__db958e1f
userbench/raman325__4750553d
userbench/dc_004__7c46d305
userbench/Soph__d4aeb5f7
userbench/fcamblor__962f5e71
userbench/PJensen__8bfa5cf3
userbench/ta93abe__2d529cb3
userbench/hutusi__490bc1e1
userbench/jobinlawrance__7cd7e40a
userbench/TheurgicDuke771__dd18f7dc
userbench/Poytr1__9c1714f3
userbench/admarble__c7561573
userbench/hutusi__ba6fef9a
userbench/marcus-sa__0121e4af
userbench/singampalliveerendra__93f5702f
userbench/dc_010__9ef010fa
userbench/nathanbooth-konecta__35376685
userbench/jeevanpillay__335616a6
userbench/PJensen__2366c990
userbench/nathanbooth-konecta__273ef024
userbench/admarble__63ccad05

Displaying 100 of 620 tasks