harbor run -d dissei/financial-judgment-fullSeven runnable point-in-time financial judgment tasks with evidence, a separate verifier and historical evaluation records.
harbor run -d dissei/financial-judgment-fullHarbor · GitHub · Website · Hugging Face
Dissei explores financial reasoning and analysis behind institutional investment and credit decisions: interpreting evidence, weighing trade-offs and reaching well-supported conclusions.
Financial Judgment contains seven runnable tasks drawn from one completed real-world deal. Each asks for a written judgment at a decision date in December 2021, June 2022 or September 2022, using supporting evidence and a task-specific rubric. The release includes 79 rubric criteria, reference answers and the tools needed to retrieve evidence and evaluate submissions.
Version 2.0.0 is available publicly on GitHub and Harbor. The Hugging Face mirror is public with manual gating; log in and receive approved access to download its files or use the gated viewer.
Use Git, the GitHub CLI (gh), Docker with Compose support and uv. Start in a parent directory where financial-judgment-v2.0.0 does not already exist. The commands create a new checkout and download the three Docker image archives.
git clone --branch v2.0.0 --depth 1 \
https://github.com/Dissei-org/financial-judgment.git financial-judgment-v2.0.0 &&
cd financial-judgment-v2.0.0 &&
gh release download v2.0.0 --repo Dissei-org/financial-judgment \
--pattern 'financial-judgment-*.tar.gz' --dir images
Continue after the clone and all downloads succeed. Local execution requires both the repository's task files and the image archives containing runtime and evidence. The repository also provides images/SHA256SUMS, the run guide and historical attachments. Docker loads the compressed archives directly; keep them intact.
For direct downloads, save each archive below under the cloned repository's images/ directory. The matching SHA256SUMS is included at images/SHA256SUMS; image-inventory.json records image/platform digests. Keep files from different releases in separate directories.
| Snapshot | Docker image archive | SHA-256 |
|---|---|---|
| 2021-12 | financial-judgment-2021-12.tar.gz | 5f945820221ca7b3972b3044b2c5ac337eb6e39b0aff6e56bd1e7dfabb932edf |
| 2022-06 | financial-judgment-2022-06.tar.gz | f61750f2e4a1e84a3f5c80293c8758e7304d564808e92d16b7e62ba8b6180675 |
| 2022-09 | financial-judgment-2022-09.tar.gz | b82a59b92f2bb1285cefff6923d01e9b5b3c7d3d1283173f991e0f4f66f39d6a |
Files are also available from the GitHub release and Hugging Face mirror. GitHub and Harbor downloads are public and independent of Hugging Face access.
Use the complete release layout with tasks/fab01-*/, the root attachments and the three images/ archives. Each task has eight files: instruction.md, task.toml, environment/Dockerfile, tests/Dockerfile, tests/verify.py, tests/test.sh, tests/task.yaml and solution/solve.sh. The host-side task YAML contains the rubric, and the solution contains the reference answer.
Harbor 0.23.0 requires environment/. Its one-line Dockerfile references the supplied image, which contains the runtime and evidence. tests/Dockerfile builds the separate verifier from the same loaded image. Download and load the archives before running tasks so that Docker can resolve the local image tags used by the task files.
From the repository root, install Harbor and verify the archive hashes:
uv tool install harbor==0.23.0
shasum -a 256 -c images/SHA256SUMS
Continue only if all three checksums pass, then load:
docker load --input images/financial-judgment-2021-12.tar.gz
docker load --input images/financial-judgment-2022-06.tar.gz
docker load --input images/financial-judgment-2022-09.tar.gz
Supply your own authorized agent/provider access and judge credentials. From the mirror root:
harbor run --path ./tasks \
-a terminus-2 -m "${DISSEI_AGENT_MODEL:?Set your authorized provider/model ID}" \
-n 1 -k 1 --ak max_turns=40 --ak reasoning_effort=high \
--ve DISSEI_JUDGE_BASE_URL="${DISSEI_JUDGE_BASE_URL:?Set your judge base URL}" \
--ve DISSEI_JUDGE_MODEL="${DISSEI_JUDGE_MODEL:?Set your judge model}" \
--ve DISSEI_JUDGE_API_KEY="${DISSEI_JUDGE_API_KEY:-}"
Use --path ./tasks/fab01-2209-st06 to start with one task. Here -n 1 limits concurrency to one trial and -k 1 selects one attempt per task. The manifest pins the published task packages by content digest. The command above runs local task directories in the documented release layout; a direct Harbor export can have a different layout. Hosted New Job execution and remote --dataset runs have not been validated. USAGE.md covers images, credentials, outputs and failures.
The separate verifier requires an explicit judge endpoint and model; --ve forwards your settings into it. Configure your own provider access. The package includes no provider account, bundled commercial agent CLI or hosted scoring service. Calls may incur charges and send evidence and answers to the provider you select. Local jobs are not automatically uploaded.
Each task has a 20-call dataroom retrieval budget and a 30-minute agent timeout. The historical harness used 40 agent turns. The verifier checks submission provenance and replays credited retrievals before applying the rubric. A valid non-submission receives zero; an execution error may leave no reward. Missing required judge configuration produces exit 3. A scoring exception reaching the entry point or top-level graded_unavailable / contradiction_unavailable produces exit 4. The tested whole-endpoint outage and wholly unusable decisive verdicts followed the exit-4 path. Other judge failures can still leave a reward, as described below.
Unavailable critical verdicts are excluded from the gate denominator. If no critical verdicts are available, gate_score defaults to 1.0. Unavailable pitfall verdicts add no penalty. These per-row failures do not trigger the top-level unavailable check and can still produce a reward. Missing or malformed graded criterion diagnostics can coexist with a usable holistic score, and multiple-vote aggregation can retain a usable verdict despite some unusable votes.
These are the evaluator's existing scoring semantics. Inspect per-row unavailable flags in score_breakdown.json and trial errors alongside rewards; a reward file alone does not establish that every criterion was judged. See USAGE.md for details.
The public task packages are version 2.0.0. Local instance IDs match the historical records.
| Harbor task package | Local instance | Anchor | Family | Criteria |
|---|---|---|---|---|
dissei/financial-judgment-dg04 |
fab01-2112-dg04 |
2021-12 | Diagnostic | 12 |
dissei/financial-judgment-pr02 |
fab01-2112-pr02 |
2021-12 | Predictive | 11 |
dissei/financial-judgment-ex01 |
fab01-2206-ex01 |
2022-06 | Explanatory | 13 |
dissei/financial-judgment-qn02 |
fab01-2206-qn02 |
2022-06 | Quantitative | 10 |
dissei/financial-judgment-cf01 |
fab01-2209-cf01 |
2022-09 | Counterfactual | 13 |
dissei/financial-judgment-cp02 |
fab01-2209-cp02 |
2022-09 | Comparative | 10 |
dissei/financial-judgment-st06 |
fab01-2209-st06 |
2022-09 | Strategic | 10 |
Each task is a single question at one decision date in the same transaction. Three image snapshots support linux/amd64 and linux/arm64. Evidence, exact financial facts and source notices are preserved.
The release includes the evaluator, rubrics, judge prompts and reference answers. Recipients can also extract runtime source and evidence from the images. The separate verifier limits ordinary task-agent access to task-specific grading files; recipients controlling the host can inspect them. Task-authoring machinery and unrelated private production workflows are excluded.
Case names are suppressed (name_suppressed). Dates, amounts and market figures are retained and may permit re-identification.
The results below were recorded on 2026-09-24–25 (UTC), including a GLM-5.3 recovery completed after midnight UTC. Version 2.0.0 preserves these historical outcomes. Reward / 100 is the equal-weight mean of seven continuous task rewards, multiplied by 100. The score reflects the documented retrieval, reasoning, submission and judging setup. Accuracy requires a separate measurement.
| Model | Historical reward / 100 | Submitted / 7 |
|---|---|---|
| Claude Fable 5.1 | 60.73 | 7 |
| GPT-6 Astra | 52.45 | 7 |
| Claude Opus 5.5 | 52.17 | 7 |
| GLM-5.3 | 44.91 | 7 |
| Kimi K3 | 35.79 | 7 |
| Gemini 3.8 Flash | 9.01 | 1 |
The 42 selected outcomes used the same seven historical task versions, Harbor 0.23.0, Terminus-2 JSON, high agent reasoning, 40 turns, 16,384 output tokens per request and a gemini-3.1-pro-preview judge with low reasoning and one vote. Historical pins document the versions evaluated. Current execution uses the version-2.0.0 packages. The Quickstart covers local execution with your chosen credentials and omits some historical provider and token settings recorded in evaluation-protocol.json. Exact replication has not been established.
Gemini 3.8 Flash is protocol-limited: one submission, six turn-limit outcomes and 50 rejected JSON responses; five infrastructure-failed tasks were recovered without replacing the original valid outcomes. GLM-5.3's completed recovery is retained alongside its original timeout and interrupted recovery; six original successful outcomes were preserved. No successful trial was rerun to seek a higher score. This single-case pilot does not establish repeated-run variance or significance of small ranking differences.
| File | Contents |
|---|---|
run-records.tar.gz |
Sanitized derivative of 16 jobs and 86 attempts, including submitted answers, trajectories, verifier outputs, failures/recoveries and separately labeled execution QA/no-op/stub runs. |
run-index.json |
Stable pseudonymous job/trial relationships, archive paths and selected outcomes. |
evaluation-protocol.json |
Historical task provenance, exact model identities, parameter values and selection rules; private endpoints are redacted. |
redaction-summary.json |
Declared identity, infrastructure and historical source-excerpt transformations and limitations. |
image-inventory.json |
Measured archive checksums and clean image index/platform descriptors for the three supplied images; the minimal-task layout reuses the archives unchanged. |
The comparison uses 42 selected outcomes from the 86 recorded attempts. Seven additional selected outcomes belong to a separately labeled historical Gemini 3.1 Pro baseline: seven submissions, mean reward 0.3124. All numerical outcomes are preserved. Logs, including duplicated trajectory and terminal representations, contain the redactions declared in redaction-summary.json. Native Harbor jobs and interactive trajectories remain private.
Release preparation used deterministic local judge smoke checks to test dataroom and verifier mechanics. Those checks provide no new measurement of financial grading accuracy or model quality.
License category: Other / restricted. The owner confirmed source redistribution rights. Downstream training, adaptation and redistribution remain subject to LICENSE.md. Preserve copyright and source notices.
This release uses new Harbor package identities. Earlier package histories, images and source records remain private. The publication plan records the plan preceding release.
Contact tech@dissei.credit for licensing, questions and private security reports.