A court of language models, tried on fifty real disputes
Background. Language models are increasingly proposed as adjudicators, advocates and decision aids in disputes. Their accuracy on legal questions has been measured; how they behave as a court — as interacting attorneys, judge and jury — has not.
Methods. Fifty real disputes from courts, ombudsmen, regulators, sports bodies, domain-name panels, tribunals and de-identified interpersonal records were frozen into outcome-blind packets and tried by a simulated court in which two attorneys, a judge and a jury were each played by a language model. Each case received one canonical trial under a fixed configuration. Twenty paired trials compared same-model with mixed-model juries, and fifty controlled interventions changed one factor at a time: exact rerun, attorney substitution, judge substitution, jury enlargement from three to nine, and mixed-model jury. Verdicts were compared with the historical disposition where a clean binary comparison existed (n = 40). Every trial was reviewed for legitimacy and fairness and measured from its transcript for language overlap between juror rationales, deliberation dynamics and burden-of-proof language (504 ballots; 2,467 recorded decisions).
Results. The canonical verdict agreed with history in 35 of 40 cases (87.5%; 95% CI 73.9%–94.5%). Exact reruns reproduced the winner in 12/12 pairs, and 33 of 36 cases re-tried under any condition were stable; the three unstable cases were multi-question disputes. Four of the five disagreements were defendant verdicts in claimant-aligned cases (agreement 81% when history favoured the claimant vs 95% the defendant; Fisher p = 0.35), and 90% of defendant ballots invoked the burden of proof against 64% of plaintiff ballots. Same-model juries were unanimous in 98.9% of trials; their rationales overlapped at 0.37 against 0.23 for mixed juries, lower in 19 of 20 paired cases (Wilcoxon p = 3.8e-06), while enlarging a same-model jury changed no verdict and no overlap. Substituting an attorney changed one verdict in ten and a judge two in ten, in every case by reframing which question was decisive rather than through evidence.
Conclusions. Simulated adjudication is repeatable and record-responsive, but its errors are structured: burden-driven, framing-sensitive and correlated across same-model jurors. These properties bear directly on how such systems should be designed and evaluated.
Accuracy is not the whole question
Evaluations of language models on legal tasks have concentrated on single-model accuracy against known outcomes. A court, however, is a process: advocates select and frame evidence, a judge fixes the decision question and the standard of proof, and a jury deliberates and votes. When every role is played by a model, the behaviour of the process — not only the accuracy of the verdict — determines whether the result is legitimate, fair and stable.
This study asks four questions of a simulated court run on real disputes. Does its disposition agree with history, and when it does not, is the error random or directional? Is the process fair and consistent across reruns and configurations? Does substituting the attorney, judge or jury, enlarging the jury, or mixing jury models change the verdict, the reasoning, or only the style? And what does the behaviour imply for the design of such systems? The simulation platform is the instrument; the behaviour of model adjudication is the subject.
Frozen records, blinded outcomes, one factor changed at a time
Corpus
Fifty disputes were selected purposively for behavioural diversity across twenty domains (Appendix D). Source families comprised published court judgments, ombudsman determinations, regulatory adjudications, sports appeals, domain-name decisions, tribunal records and de-identified interpersonal dispute records. Each case was frozen into an outcome-blind packet containing a neutral brief, forum-specific rules, both parties’ positions and concessions, requested remedies, disputed facts and a normalized exhibit record delivered as documentary text. Forty cases had a clean binary historical disposition; nine were mixed or partial and one was settled without a merits decision.
Simulation protocol
Each trial proceeded through opening statements, direct and cross examination, closing statements, the judge’s instructions, deliberation, closed ballots and verdict. Every canonical trial used one configuration: GPT-5.6 Terra for both attorneys, Claude Opus 5 as judge, and Gemini 3.8 Flash for every juror. Seven further models entered only through interventions: Claude Fable 5.1, GPT-5.6 Sol, Grok 4.6, Kimi K3, DeepSeek V4 Pro, Qwen 3.8 Max and GLM-5.3. Historical outcomes were withheld from every actor. A trial was accepted only if it produced a verdict and passed actor-order, model-route and evidence-delivery audits; no accepted trial truncated the evidence record.
Experimental design
| Component | Trials | Design |
|---|---|---|
| Canonical sample | 50 | One audited trial per case under the fixed configuration; the reference for all comparisons. |
| Jury-composition gate | 20 | Ten cases, each tried once with a same-model jury and once with a mixed-model jury of equal size (3–11 seats). |
| Controlled interventions | 50 | Ten cases per panel, paired to the canonical trial: exact rerun; attorney substitution (five plaintiff, five defence); judge substitution; jury enlarged from three to nine of the same model; mixed-model jury of three. |
| Retained repeat baselines | 2 | Two additional exact reruns preserved from the primary run. |
Measures
Historical agreement was assessed on the forty binary cases after transcript-level review of party-role mapping, which corrected one automated label. Each trial received a structured single-reviewer assessment of outcome legitimacy, procedural fairness, factual integrity, evidence use and advocacy balance (high, medium, low) and a set of non-exclusive failure flags. From each trial’s event log and decision traces we computed the language overlap between juror ballot rationales — the cosine similarity of term-weighted word and word-pair vectors, averaged over all juror pairs (0 = no shared wording, 1 = identical) — and the overlap between each rationale and the judge’s instructions, the winning closing and the losing closing; lexicon counts of certainty, hedging, concession and burden-of-proof markers per 100 words; deliberation order, explicitly stated side-leaning where present, and references to other jurors; and, for every decision, the strategy candidates listed and the one chosen.
Statistical analysis
Proportions are reported with Wilson 95% confidence intervals. Paired within-case comparisons use the Wilcoxon signed-rank test; the asymmetry of disagreements uses Fisher’s exact test; the composition gate uses McNemar’s exact test. Ten cases per intervention panel detect only large sensitivities. No model was randomised across roles, so per-model comparisons are descriptive and confined to the appendix.
What the court did
The verdict matched the historical outcome in 35 of 40 clean cases
The canonical verdict agreed with the historical disposition in 35 of 40 binary cases (Figure 1). Three of the five disagreements were sports disputes in which a question of jurisdiction, rule interpretation or burden was compressed into an adversarial winner. In the darts appeal the packet inverted the party roles, so the automated winner label masked a merits result opposite to history; the reviewed figure is used throughout.
When the court disagreed with history, it favoured the defendant
History favoured the claimant in 21 of the forty cases and the defendant in 19; the court returned 22 defendant verdicts. 4 of the 5 disagreements were defendant verdicts in claimant-aligned cases (Figure 2; Fisher exact p = 0.35). The mechanism is visible in the ballots: 90% of defendant ballots invoked the burden of proof against 64% of plaintiff ballots (mean 1.18 vs 0.69 references), deliberation statements that named a side leaned to the defendant 145 times and to the plaintiff 60 times, and plaintiff ballots carried twice the certainty markers of defendant ballots (0.25 vs 0.12 per 100 words). Where the record was thin, the standard contested or the remedy partial, “not proved” was the path of least resistance. The ten mixed or settled histories, forced through a binary form, produced 6 defendant and 4 plaintiff winners.
Matching history was easier than earning a high review
Agreement with history was more common than a high review rating (Figure 3). Outcome legitimacy was rated high in 6 cases and procedural fairness in 3; the modal assessment was a defensible result reached through thinner reasoning, a cruder remedy or a less forum-specific procedure than the real adjudication. Speaking volume did not decide cases: plaintiff counsel spoke more in 47 of 50 trials and won 22, and the side citing more unique exhibits won 6 of the 12 trials in which it did so. Winners conceded less (1.6 vs 2.4 concession markers per side) and asserted more (5.7 vs 4.4 certainty markers).
The same record produced the same answer almost everywhere
All 12 exact reruns reproduced the canonical winner and were unanimous. Thirty-six cases were tried more than once under some combination of rerun, attorney substitution, judge substitution, jury enlargement and jury mixing; 33 returned the same winner under every condition. The three exceptions share a structure: a decision with several parts forced into one winner. In the former spouse’s repayment claim, a Kimi K3 defence attorney and, separately, a Kimi K3 judge moved the decisive question from acknowledgment of the debt to quantum and limitation, and the verdict reversed. In the chess team-selection jurisdiction appeal, a mixed jury and a GPT-5.6 Sol judge each placed federation autonomy first and reversed the canonical answer toward history. In the commercial grower’s wrong-seed claim, a mixed jury accepted breach and foreseeability but refused an unproven quantum and reversed away from history. No reversal involved evidence the canonical trial had not seen.
Same-model juries were unanimous because their members thought alike
Same-model juries were unanimous in 91 of 92 trials (1 dissent in 404 ballots, in the role-inverted darts appeal); mixed-model juries in 17 of 20 (4 dissents in 100 ballots). In the composition gate both jury types agreed with history in 8 of 10 cases; two verdicts changed, one toward and one away from history (McNemar p = 1.0).
Rationale overlap was 0.37 on same-model juries and 0.23 on mixed juries; in the 20 paired cases the mixed panel was lower in 19 (Figure 4; Wilcoxon p = 3.8e-06). Deliberation statements showed the same shift (0.38 → 0.27). Enlarging a Gemini jury from three to nine seats produced 9–0 in all ten pairs and left overlap essentially unchanged (0.38 → 0.35). Same-model jurors in the same trial overlapped more (0.36) than the same model on the same case in a different trial (0.31), indicating that part of the convergence arises from shared in-trial context rather than a shared prior alone. Ballots overlapped with the winning closing (0.20) more than with the losing closing (0.15) or the instructions (0.16).
Deliberation functioned as sequential monologue. 5.4% of same-model and 10.0% of mixed deliberation statements referenced another juror. Where a juror stated a lean (205 of 504 deliberations), its ballot followed 93% of the time, and the 14 exceptions all but once joined the eventual majority. The first speaker’s stated lean matched the verdict in 93% of trials. No juror used the available option to pass.
Changing an actor changed the argument more often than the verdict
| Intervention | Verdicts changed | 95% CI | Unanimous | Rationale overlap | Agreed with history |
|---|---|---|---|---|---|
| Exact rerun | 0 / 10 | 0%–28% | 10 / 10 | 0.38 | 7 / 8 |
| Attorney substitution | 1 / 10 | 2%–40% | 10 / 10 | 0.38 | 6 / 8 |
| Judge substitution | 2 / 10 | 6%–51% | 10 / 10 | 0.34 | 7 / 8 |
| Jury enlarged 3 → 9, same model | 0 / 10 | 0%–28% | 10 / 10 | 0.35 | 7 / 8 |
| Mixed-model jury | 0 / 10 | 0%–28% | 9 / 10 | 0.25 | 7 / 8 |
| Mixed-model jury (composition gate) | 2 / 10 | 6%–51% | 8 / 10 | 0.21 | 8 / 10 |
Attorney substitution altered volume and posture far more than evidence (Figure 5). Every substitute cited the canonical exhibit set within one item; Claude Fable 5.1 and Claude Opus 5 argued at roughly twice the length of GPT-5.6 Terra and conceded several times more often, Qwen 3.8 Max conceded nothing, and Kimi K3 produced the single reversal by reframing limitation and quantum. Judge substitution changed the length and hedging of the instructions (Figure 6) but not the degree to which jurors echoed them (0.12–0.25 across judges). Both judge-related reversals were framing events: Kimi K3 converted a potentially partial award into an all-or-nothing $30,000 proof threshold, and GPT-5.6 Sol placed federation autonomy before the jurisdiction question. One substitute judge produced no instructions at all and the jury voted regardless (Appendix C).
Repeatable, directional, and correlated
Three properties of the simulated court emerge consistently. First, it is repeatable: identical inputs produced identical winners in every rerun, and the great majority of cases were invariant to who argued, who instructed and who decided. Where the winner did change, the case had more than one dispositive question and the change was one of framing. A single-winner verdict form therefore does not merely lose information about partial outcomes; it makes the result hostage to whichever actor fixes the order of questions.
Second, its errors have a direction. Under a burden-of-proof lens a thin, contested or partial record resolves as “not proved”, and the ballots show that lens applied far more often when voting for the defendant. The asymmetry is not statistically established at forty cases, but its mechanism is directly observable, and it is the hypothesis most worth a preregistered test with issue-level outcomes.
Third, unanimity on same-model juries reflects correlated cognition rather than persuasion. Jurors did not reference one another, arrived at deliberation already decided, and wrote rationales whose overlap fell sharply when the panel mixed models but not at all when it added seats. Diversity of model, not number of jurors, is what changed the reasoning; it did not detectably change accuracy.
Simulated adjudication is therefore best understood as a repeatable, evidence-responsive instrument whose failure modes are structural and designable: issue-level verdicts rather than a single winner, heterogeneous juries rather than larger ones, and instructions that must exist before a vote. The secondary behaviours recorded in the appendix — an absolute preference for the first strategy considered, certainty rising where the court was wrong, and a jury voting on without instructions — point the same way.
The court will keep being studied, and next it will be tried under pressure
Every trial in this study was run in good faith: both sides argued from a frozen, honest record, and the court was never deliberately misled. That is the right first test of a simulated court and the wrong last one. The behaviours documented here — burden-driven defaults, framing sensitivity, correlated juror cognition, actors that never revisit their first strategy — are exactly the properties an adversary would exploit. The programme continues along the following vectors, in roughly this order.
- Adversarial evidence. Trials in which one side submits fabricated exhibits, altered dates, or documents that contradict the shared record, to measure whether attorneys detect the conflict, whether judges instruct on it, and how jurors resolve two exhibits that cannot both be true.
- Manipulation of the court itself. Evidence and submissions carrying instructions addressed to the models rather than to the tribunal — the courtroom form of prompt injection — and rhetorical devices aimed at the known biases: appeals to certainty, invented precedent, burden inversion, and volume.
- Conflicting and incomplete records. Cases where the two sides hold genuinely inconsistent accounts with no tiebreaker, where key exhibits are missing, or where evidence arrives late, to characterise how the court behaves at the edge of proof rather than in its centre.
- Issue-level verdicts. Replacing the single winner with structured findings — liability, causation, quantum, limitation, remedy — and re-running the multi-question cases that were unstable here, to test whether the instability disappears when the verdict form can express the answer.
- The defendant-ward burden effect. A preregistered test on a balanced corpus with issue-level outcomes, varying the standard of proof and the wording of the instructions, to establish whether the asymmetry observed here is a property of the court or of this sample.
- Jury design. Heterogeneous panels by default; structured deliberation with mandated cross-reference, sequential versus simultaneous ballots, an appointed dissenter, and randomised speaking order, measured against the same overlap and dissent metrics used here.
- Judicial variation. Instructions that differ in standard, in the order of questions, and in the presence or absence of a charge on burden, to isolate how much of a verdict the judge writes before the jury sits.
- Appeal and review. A second-instance court reviewing first-instance transcripts for error, to test whether a model tribunal can detect the failure modes catalogued here when they occur in another model tribunal.
- Human comparison. The same frozen records put to practising advocates and lay panels under matched procedure, so that agreement, direction of error and convergence can be compared between human and model courts rather than only against history.
- Multimodal and native evidence. Original images, audio and video in place of documentary reconstructions, to test whether perception changes what the court finds proved.
- Scale. A larger, stratified corpus with repeated randomised role assignment, so that per-model and per-domain effects can be estimated rather than described.
Each of these will be run on frozen records under the same audit discipline, with failed and surprising trials preserved rather than rerun, and published against the exact environment version they describe.
Bounded claims
- The corpus is purposively diverse, not a probability sample; forty cases support binary comparison; history is a benchmark, not ground truth.
- Ten cases per intervention panel detect only large sensitivities; the defendant-ward asymmetry is directional.
- Cases recur across panels, so trials are not independent observations.
- Reviews are structured single-reviewer judgments; language, lean and overlap measures are lexical approximations.
- No model was randomised across roles and ballot counts are unbalanced; model profiles describe style, not accuracy.
- Evidence was delivered as documentary text; multimodal perception was not tested.
- A binary verdict cannot express partial remedies, regulatory decisions, consent outcomes or appeals, and every observed instability occurred in such cases.
Properties of the process, not of a model
A simulated court of language models agreed with the historical outcome in 35 of 40 real disputes, reproduced its own verdicts exactly, and remained stable under most substitutions of attorney, judge and jury. Its disagreements ran toward the defendant through burden framing, its reversals occurred only where a multi-part decision was forced into one winner, and its juries converged because their members shared a model, not because they deliberated. These are properties of the process rather than of any single model, and they define what such a system can be relied on for and what it cannot.
Secondary findings and coverage
Models differed in style, length and confidence, not detectably in accuracy
Gemini 3.8 Flash cast 419 of 504 ballots; the other nine juror models cast 5–13 each, and no model was randomised across roles. Per-model agreement with history ranged from 75% to 90% on these denominators and cannot be distinguished. Behavioural differences were nonetheless consistent (Table 3, Figure 7). Gemini wrote the shortest rationales and almost never named an exhibit in its ballot; Claude Opus 5 and Grok 4.6 anchored ballots to named exhibits; the OpenAI and Qwen jurors used two to three times more certainty language per word than the Anthropic, Google or DeepSeek jurors. As attorneys, the Anthropic models argued longest and conceded most, Qwen shortest and never. Instruction length ranged from about 500 words (DeepSeek, GPT-5.6, Gemini) to over 1,100 (Grok, Claude Fable 5.1).
| Model | Ballots | Words per rationale | Exhibits named | Certainty per 100 words | Voted plaintiff | Dissents |
|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 419 | 74 | 0.10 | 0.17 | 42% | 1 |
| Claude Fable 5.1 | 13 | 134 | 0.85 | 0.12 | 38% | 1 |
| GPT-5.6 Sol | 13 | 68 | 0.00 | 0.48 | 46% | 0 |
| Qwen 3.8 Max | 11 | 67 | 0.00 | 0.39 | 45% | 0 |
| Grok 4.6 | 11 | 104 | 1.91 | 0.10 | 36% | 1 |
| DeepSeek V4 Pro | 9 | 95 | 0.22 | 0.08 | 33% | 0 |
| Kimi K3 | 9 | 162 | 0.89 | 0.14 | 44% | 1 |
| Claude Opus 5 | 7 | 231 | 2.86 | 0.17 | 43% | 0 |
| GLM-5.3 | 7 | 202 | 0.57 | 0.13 | 43% | 0 |
| GPT-5.6 Terra | 5 | 60 | 0.00 | 0.24 | 40% | 1 |
Patterns in how the actors decided
Strategy selection. Each actor listed strategy candidates with an associated risk before acting. Across 2,467 decisions the first candidate was chosen 2,459 times (99.7%); the remaining 8 were unparseable outputs. The stated risk never changed a choice.
Certainty and error. Ballots in disagreeing trials carried 2.1 times the certainty markers per 100 words of ballots in agreeing trials (0.35 vs 0.16) and 2.5 times the hedges; advocates’ certainty per 100 words rose from 0.09 at opening to 0.21 at closing while hedging remained flat.
Reviewer flags. Correlated juror framing (39/50) and a verdict form unable to express partial liability, separate issues, constrained remedies, consent or appellate relief (34/50) were flagged at similar rates in agreeing and disagreeing cases; burden and causation errors were the flags that tracked disagreement (Figure 8).
When an actor produced nothing, the court carried on
On 8 occasions an actor’s output carried no usable answer, 4 of them because the model reasoned until its output limit without reaching a decision. In the PPI-commission limitation appeal the Qwen 3.8 Max judge delivered no instruction text; the three Gemini jurors deliberated and voted 3–0 for the historically aligned side from the closings alone, none remarked on the absence, and their rationales invoked a burden of proof that had not been charged. The reviewer rated that trial low on legitimacy and fairness. A Gemini juror and a DeepSeek juror each submitted an empty ballot after reasoning to the limit; both were rejected and both jurors voted validly on retry. One DeepSeek juror proposed a second ballot after a valid one, which the court refused.
Case coverage by domain
Domain-level reviewed comparison counts are unavailable, not zero. The table reports case coverage and simulated winners only; the study-wide historical comparison remains 35 of 40 clean binary cases.
| Domain | Cases | Sim. winners P / D |
|---|---|---|
| Consumer Financial | 6 | 2 / 4 |
| Neighbor Property | 6 | 4 / 2 |
| Court | 5 | 1 / 4 |
| Interpersonal Family | 4 | 2 / 2 |
| Betting Gambling | 3 | 1 / 2 |
| Domain Ip | 3 | 2 / 1 |
| Interpersonal Friendship | 3 | 1 / 2 |
| Interpersonal Romantic | 3 | 1 / 2 |
| Marketplace Ecommerce | 3 | 1 / 2 |
| Sports | 3 | 1 / 2 |
| Advertising Claims | 2 | 1 / 1 |
| Commercial Supply | 1 | 1 / 0 |
| Consumer Services | 1 | 1 / 0 |
| Education Services | 1 | 0 / 1 |
| Insurance | 1 | 0 / 1 |
| Interpersonal Household | 1 | 0 / 1 |
| Interpersonal Recreational | 1 | 1 / 0 |
| Recreational Services | 1 | 1 / 0 |
| Transport Tort | 1 | 0 / 1 |
| Workplace Organizational | 1 | 1 / 0 |
| Total | 50 | 22 / 28 |
The research archive preserves every frozen packet and hash, event transcript, decision trace, actor and evidence audit, juror ballot, case review and cross-case table, together with the run-, statement-, ballot-, trace- and instruction-level tables and scripts used here. Failed or incomplete attempts are stored separately and were never substituted.