harbor run -d dissei/financial-judgmentDissei Financial Judgment explores financial reasoning and analysis for institutional investment and credit decisions.
harbor run -d dissei/financial-judgmentDissei explores financial reasoning and analysis behind institutional investment and credit decisions: interpreting evidence, weighing trade-offs and reaching well-supported conclusions.
This sample presents questions drawn from a completed private-equity transaction.
Open a task to read its original question and brief. These public previews come from the seven-task evaluated sample; the case evidence and scoring materials require approved access.
Preview downloads contain the question, brief and metadata only. They are not runnable evaluations. The scores below come from the complete controlled-access tasks.
Reward / 100 is the equal-weight mean of the seven saved continuous task rewards, multiplied by 100. It is not accuracy. Values use unrounded rewards before display. Scores come from the complete evaluation tasks, not the read-only public previews. No task data, rubric or recorded reward changed for this publication.
| Model | Reward / 100 | Submitted | Run notes |
|---|---|---|---|
| Claude Fable 5.1 | 60.73 | 7/7 | 7 submitted; no infrastructure recovery |
| GPT-6 Astra | 52.45 | 7/7 | 7 submitted; no infrastructure recovery |
| Claude Opus 5.5 | 52.17 | 7/7 | 7 submitted; no infrastructure recovery |
| GLM-5.3 | 44.91 | 7/7 | 7 submitted; 1 task recovered; earlier recovery interrupted by the VM limit |
| Kimi K3 | 35.79 | 7/7 | 7 submitted; no infrastructure recovery |
| Gemini 3.8 Flash | 9.01 | 1/7 | Protocol-limited: 1 submitted, 6 turn limits; 5 tasks recovered after infrastructure failures |
Gemini 3.8 Flash's result is protocol-limited: six selected attempts reached the turn limit, and the JSON command parser rejected 50 responses. A non-submission receives zero without an LLM grade. The number therefore measures this harness interaction as well as the model. GLM-5.3's recovery completed after a previous recovery was interrupted by the VM's runtime limit; the other six successful original trials were retained. No successful trial was rerun to seek a higher score.
gemini-3.1-pro-preview judge, low reasoning and one vote, applies the authored rubric.This is a small, single-case pilot with no repeated-run variance estimate or error bars. Judge choice, evidence coverage and harness behavior can affect results. The scores are not an external certification, proof of absence of benchmark defects, or evidence that small score differences are significant.
Independent reproduction requires approved access to the evaluation materials. Reviewers can request the exact task/version manifest, grading specification, selected outputs and failure/recovery records under agreed terms. Public summary metrics are reported by Dissei; Harbor hosts them and does not independently recompute or certify them.
The seven question briefs, task metadata, results and methodology are public. Case evidence, execution tools, rubrics, reference answers, evaluator implementation and private run records remain available through approved evaluation or research access.
Each public preview contains only its question, brief and metadata. This dataset starts at revision 1 and does not inherit the original benchmark's historical revisions or grant access to its protected files.
Evaluation, independent review, retention, redistribution and training rights require agreement. Publishing these question previews does not grant rights to the underlying benchmark assets.
Request evaluation or review access: tech@dissei.credit · dissei.ai/contact.