userbench/UserBench-train400

UserBench train400 twin: the same 620 held tasks plus 400 leak-safe prior turns per developer. Composer 2.5 multi-label acts score by set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench-train400

UserBench-train400

Twin of the public UserBench eval with leak-safe prior-session context.

Same 620 held tasks / 62 developers as userbench/UserBench@v2 (history, gold, verifier unchanged). Each task additionally ships earlier sessions for that developer under /sim/train/, totaling 400 human user-turns sampled with sqrt two-stage.

What the agent sees

Path Role
/sim/history.md Held conversation so far (identical to baseline)
/sim/train/_index.json Index of packed prior sessions (ids, timestamps, repos, turn takes, paths)
/sim/train/<sid>.md Contiguous prefixes of those sessions (same > DEVELOPER / > AGENT format)
/sim/answer.txt Where the agent writes the next developer message

The instruction is count-agnostic: it points at /sim/train/ without hardcoding “400”. The turn budget is a publish-time sampler parameter.

Leakage invariant

For each developer let T* = min(held_sessions[].ts) (manifest session timestamp).

Every packed train session satisfies:

  1. session.ts < T*
  2. max(turn.ts) < T*

This is stricter than the cohort’s session-start train/held split alone. Sessions that wall-clock-overlap the earliest held session are never included.

QC on this package: 0 / 620 tasks violate max(turn.ts) < T*.

Sampling

  1. Start from the developer’s full leak-safe train pool (all sessions/turns passing the T* rule).
  2. Allocate exactly 400 human-turns with sqrt_two_stage (time-stratified session pick, then contiguous-prefix waterfill).
  3. Emit prefixes through the k-th human turn (including intervening AGENT/TOOL/SYSTEM), with the same ~200-word-per-turn truncation as eval history.

The full pools (not just the 400-turn subsample) live in the GitHub repo under train_pools/ so future Hub variants (e.g. train1000) can resample from the same T*-filtered corpus without changing the leakage rule.

Scale

Developers 62
Tasks 620 (10 per developer; same selection as baseline)
Condition train400
Train budget 400 human-turns / developer (shared across that developer’s 10 tasks)

Task naming

Same short IDs as baseline (Harbor org/name):

userbench/<username>__<hash>

Versioning is via Harbor tags on this package (e.g. @v2). Digests differ from the noprofile baseline; consume via the dataset ref below so baseline UserBench@v2 stays pinned to zero-train digests.

Keywords

userbench · user-simulation · coding-agents · move-prediction · train400

How to reference

What Ref
This twin userbench/UserBench-train400@v2 (also latest)
Zero-train baseline userbench/UserBench@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench-train400
Full train pools (GitHub) https://github.com/AlienKevin/user-simulator/tree/main/train_pools
harbor run -d userbench/UserBench-train400@v2 -a <agent> -m <model>

Links

Task
userbench/ababushkin__15c2bcc4
userbench/khaong__18ba6389
userbench/dc_000__52b607d4
userbench/cyyeh__3cc7d29d
userbench/ta93abe__e1d7b4a4
userbench/mvanhorn__4e30e41e
userbench/dc_010__b9fc11b2
userbench/jskswamy__264ee35c
userbench/oddessentials__0b10a78b
userbench/dc_000__49970ee1
userbench/alishakawaguchi__c005ae79
userbench/KeKs0r__ac28eed7
userbench/jobinlawrance__69d520ea
userbench/blackgirlbytes__b2b09b06
userbench/ET-NoahDolev__63f9b284
userbench/4thwithme__404e13d3
userbench/jeevanpillay__fa06bda8
userbench/cyyeh__93e85db5
userbench/jskswamy__45376c4b
userbench/mvanhorn__8a8dc66c
userbench/FSM1__8d0fb85e
userbench/thieso2__ef3476ab
userbench/dcambria__b1663567
userbench/junaid-appointy__cac807fd
userbench/nosman__416e9b74
userbench/malkoG__4ca6601d
userbench/khaong__0b023426
userbench/raman325__429c2fe6
userbench/winksaville__0b1794e9
userbench/KeKs0r__a3626c5d
userbench/ababushkin__fd476f56
userbench/armelhbobdad__dfebbec3
userbench/winksaville__4f87e52f
userbench/ET-NoahDolev__d2ea0f07
userbench/gabadi__aba2eda8
userbench/ta93abe__2f9b0c7c
userbench/135yshr__676c5289
userbench/yutakobayashidev__23d09ae6
userbench/nosman__458811d2
userbench/jskswamy__2fd0350e
userbench/kohaku500__d39c7601
userbench/nathanbooth-konecta__30525504
userbench/lyston11__c12cd364
userbench/melagiri__28dbbafd
userbench/achildrenmile__71a8a165
userbench/scottdensmore__878978e8
userbench/MohammedMqat__eb4f7f08
userbench/heddendorp__b574678e
userbench/thieso2__f58a3bca
userbench/melagiri__476a562b
userbench/scottdensmore__344cacf9
userbench/junaid-appointy__af1dce79
userbench/jobinlawrance__ba8c72c7
userbench/KeKs0r__c7dd8b9b
userbench/winksaville__0d00a824
userbench/raman325__646fecc7
userbench/khaong__3e62c23c
userbench/oddessentials__49939de2
userbench/kohaku500__829d75fd
userbench/robouden__b294b008
userbench/alishakawaguchi__bd8cd4f9
userbench/Poytr1__79b74c55
userbench/jeevanpillay__d951d14a
userbench/nosman__d7d30e52
userbench/manderson240__5c697026
userbench/yyovil__149c4120
userbench/manderson240__c62b91cf
userbench/johyunduk__5a1334df
userbench/Soph__a13f4134
userbench/PJensen__8bc81bf4
userbench/kohaku500__38862417
userbench/christso__c229da20
userbench/MohammedMqat__f5ba11ff
userbench/blackgirlbytes__077f6188
userbench/jhoetter__b49b7e7f
userbench/melagiri__4203583e
userbench/scottdensmore__e7c0c3f1
userbench/scottdensmore__96489d0d
userbench/marcus-sa__8f91d1b5
userbench/melagiri__3cf814f5
userbench/gabadi__9e498963
userbench/yyovil__6e71d5f5
userbench/robouden__0104b37c
userbench/gabadi__cc9e1d70
userbench/admarble__08f6ff97
userbench/jakobtfaber__8a4f165d
userbench/jakobtfaber__2a61e65f
userbench/jdsingh122918__92c17296
userbench/Stark-Industries0417__f054d028
userbench/kohaku500__c36c95c2
userbench/winksaville__7d7bb7df
userbench/barbogast__601f787e
userbench/MohammedMqat__8483a6a4
userbench/scottdensmore__83347f87
userbench/gabadi__de9f94c6
userbench/wildlily1021__8a0b4073
userbench/Soph__d2bccb86
userbench/Poytr1__7704b402
userbench/jobinlawrance__015f6990
userbench/yutakobayashidev__32f6e7cb

Displaying 100 of 620 tasks