Keeping the answer out of the question.
A historical test can look impressive for the wrong reason. Here’s the evaluation rule we’re designing around.
The problem with hindsight
A historical prediction task asks a model to reason as if an outcome has not happened. But a legal corpus can contain the final opinion, the vote, or a later summary of the same case. If that material reaches the model, the test no longer measures the ability we meant to test.
Separate evidence from answers
Our documented evaluation design keeps case inputs separate from held-out labels. The index builder defaults to the training split and excludes benchmark inputs, benchmark labels, holdout cases, and records marked as containing outcomes. Related documents need a shared leakage group so a near-duplicate cannot silently cross the split.
A date matters as much as a split
The methodology also requires a cutoff: material published after the simulated decision date must be excluded. These are design requirements, not proof that every leakage route has already been closed. Temporal filtering and end-to-end evaluation still need to be implemented and tested.
What success will have to mean
Vote agreement alone will not be enough. We plan to examine whether citations support the reasoning, whether uncertainty is calibrated, and where the system fails. No benchmark results or accuracy percentages have been produced yet.