userbench/UserBench-train400

UserBench train400 twin: the same 620 held tasks plus 400 leak-safe prior turns per developer. Composer 2.5 multi-label acts score by set Jaccard/IoU.

Published 7/21/2026 by

harbor run -d userbench/UserBench-train400

UserBench-train400

Twin of the public UserBench eval with leak-safe prior-session context.

Same 620 held tasks / 62 developers as userbench/UserBench@v2 (history, gold, verifier unchanged). Each task additionally ships earlier sessions for that developer under /sim/train/, totaling 400 human user-turns sampled with sqrt two-stage.

What the agent sees

Path Role
/sim/history.md Held conversation so far (identical to baseline)
/sim/train/_index.json Index of packed prior sessions (ids, timestamps, repos, turn takes, paths)
/sim/train/<sid>.md Contiguous prefixes of those sessions (same > DEVELOPER / > AGENT format)
/sim/answer.txt Where the agent writes the next developer message

The instruction is count-agnostic: it points at /sim/train/ without hardcoding “400”. The turn budget is a publish-time sampler parameter.

Leakage invariant

For each developer let T* = min(held_sessions[].ts) (manifest session timestamp).

Every packed train session satisfies:

  1. session.ts < T*
  2. max(turn.ts) < T*

This is stricter than the cohort’s session-start train/held split alone. Sessions that wall-clock-overlap the earliest held session are never included.

QC on this package: 0 / 620 tasks violate max(turn.ts) < T*.

Sampling

  1. Start from the developer’s full leak-safe train pool (all sessions/turns passing the T* rule).
  2. Allocate exactly 400 human-turns with sqrt_two_stage (time-stratified session pick, then contiguous-prefix waterfill).
  3. Emit prefixes through the k-th human turn (including intervening AGENT/TOOL/SYSTEM), with the same ~200-word-per-turn truncation as eval history.

The full pools (not just the 400-turn subsample) live in the GitHub repo under train_pools/ so future Hub variants (e.g. train1000) can resample from the same T*-filtered corpus without changing the leakage rule.

Scale

Developers 62
Tasks 620 (10 per developer; same selection as baseline)
Condition train400
Train budget 400 human-turns / developer (shared across that developer’s 10 tasks)

Task naming

Same short IDs as baseline (Harbor org/name):

userbench/<username>__<hash>

Versioning is via Harbor tags on this package (e.g. @v2). Digests differ from the noprofile baseline; consume via the dataset ref below so baseline UserBench@v2 stays pinned to zero-train digests.

Keywords

userbench · user-simulation · coding-agents · move-prediction · train400

How to reference

What Ref
This twin userbench/UserBench-train400@v2 (also latest)
Zero-train baseline userbench/UserBench@v2
Hub https://hub.harborframework.com/datasets/userbench/UserBench-train400
Full train pools (GitHub) https://github.com/AlienKevin/user-simulator/tree/main/train_pools
harbor run -d userbench/UserBench-train400@v2 -a <agent> -m <model>

Links

Task
userbench/dc_010__fa91db13
userbench/4thwithme__56e34cff
userbench/thieso2__3dbff2d6
userbench/jdsingh122918__bb859790
userbench/kungfusaini__76670d24
userbench/FSM1__123d9f43
userbench/wildlily1021__85caed49
userbench/TheurgicDuke771__01470379
userbench/Soph__1a8c6a0c
userbench/4thwithme__1e9b9622
userbench/jdsingh122918__d44c4e2b
userbench/hutusi__67ea1d83
userbench/barbogast__4aa039d9
userbench/cyyeh__61ced54e
userbench/135yshr__488a7841
userbench/nathanbooth-konecta__8e4c4321
userbench/malkoG__f0e27113
userbench/jhoetter__eb3780f4
userbench/dc_004__74515f8d
userbench/wildlily1021__c031a215
userbench/gabadi__ae3008ee
userbench/cyyeh__a430a77c
userbench/jdsingh122918__9d260baa
userbench/kohaku500__cc54a542
userbench/scottdensmore__17cdc6a5
userbench/dc_004__667e2977
userbench/TheurgicDuke771__99a7df5d
userbench/nathanbooth-konecta__33b00343
userbench/robouden__f12f7493
userbench/henryph24__f0fc860a
userbench/fcamblor__10653b23
userbench/blackgirlbytes__df26765a
userbench/dc_000__d5528693
userbench/cyyeh__5f330796
userbench/admarble__b47ec152
userbench/oddessentials__b4bcc213
userbench/4thwithme__a2549fc8
userbench/admarble__dbb54fbd
userbench/kmiki0__672300e9
userbench/alishakawaguchi__3dc3da3f
userbench/MohammedMqat__81504953
userbench/junaid-appointy__beeefe8f
userbench/raman325__9730850f
userbench/oddessentials__0558d1c7
userbench/achildrenmile__97a0d30b
userbench/nathanbooth-konecta__d2090df8
userbench/jeevanpillay__e77b5133
userbench/kungfusaini__9f8077cf
userbench/robouden__e896e954
userbench/PJensen__409f7f83
userbench/melagiri__af1394ab
userbench/achildrenmile__7dafef0e
userbench/marcus-sa__81996502
userbench/ta93abe__25e7c1b9
userbench/lyston11__182e418e
userbench/wildlily1021__ff1c6f83
userbench/dcambria__3462644a
userbench/kmiki0__a9bb3ed2
userbench/FSM1__e064c641
userbench/ababushkin__0dafbf18
userbench/4thwithme__e0b170b3
userbench/dcambria__8b203ca9
userbench/FSM1__9ca02217
userbench/dc_004__b1a1b133
userbench/melagiri__e0dad522
userbench/hutusi__2751eb8c
userbench/robouden__c6e9e844
userbench/kungfusaini__6409200a
userbench/alishakawaguchi__095c91a6
userbench/KeKs0r__7d6d7ccd
userbench/christso__7616b4d3
userbench/marcus-sa__60346b6a
userbench/johyunduk__4bab3ef7
userbench/dc_001__4b92ecaf
userbench/khaong__8bd56ca2
userbench/blackgirlbytes__64f0503c
userbench/dc_004__ce4ab540
userbench/fcamblor__4c9bb47a
userbench/blackgirlbytes__116d4a16
userbench/135yshr__e4b9747c
userbench/scottdensmore__42381bd0
userbench/KeKs0r__638dbf35
userbench/jhoetter__bcc3790b
userbench/singampalliveerendra__13c48199
userbench/marcus-sa__56b8c4e7
userbench/TheurgicDuke771__b6a77404
userbench/jhoetter__2da05fce
userbench/kungfusaini__f8d51bd4
userbench/robouden__36ed725c
userbench/barbogast__5c442045
userbench/jakobtfaber__1395ade5
userbench/singampalliveerendra__de5a3f57
userbench/KeKs0r__dbbc1fb2
userbench/yyovil__a5d654fb
userbench/nosman__7c59d97b
userbench/junaid-appointy__311a12f4
userbench/mvanhorn__89ae7ff7
userbench/malkoG__a7ddaeaf
userbench/henryph24__c072ea86
userbench/dc_004__1b611cb2

Displaying 100 of 620 tasks