Legal & Litigation

Can a simulated court resolve real-world disputes?

A behavioural study of language-model adjudication

Abstract

A court of language models, tried on fifty real disputes

Background. Language models are increasingly proposed as adjudicators, advocates and decision aids in disputes. Their accuracy on legal questions has been measured; how they behave as a court — as interacting attorneys, judge and jury — has not.

Methods. Fifty real disputes from courts, ombudsmen, regulators, sports bodies, domain-name panels, tribunals and de-identified interpersonal records were frozen into outcome-blind packets and tried by a simulated court in which two attorneys, a judge and a jury were each played by a language model. Each case received one canonical trial under a fixed configuration. Twenty paired trials compared same-model with mixed-model juries, and fifty controlled interventions changed one factor at a time: exact rerun, attorney substitution, judge substitution, jury enlargement from three to nine, and mixed-model jury. Verdicts were compared with the historical disposition where a clean binary comparison existed (n = 40). Every trial was reviewed for legitimacy and fairness and measured from its transcript for language overlap between juror rationales, deliberation dynamics and burden-of-proof language (504 ballots; 2,467 recorded decisions).

Results. The canonical verdict agreed with history in 35 of 40 cases (87.5%; 95% CI 73.9%–94.5%). Exact reruns reproduced the winner in 12/12 pairs, and 33 of 36 cases re-tried under any condition were stable; the three unstable cases were multi-question disputes. Four of the five disagreements were defendant verdicts in claimant-aligned cases (agreement 81% when history favoured the claimant vs 95% the defendant; Fisher p = 0.35), and 90% of defendant ballots invoked the burden of proof against 64% of plaintiff ballots. Same-model juries were unanimous in 98.9% of trials; their rationales overlapped at 0.37 against 0.23 for mixed juries, lower in 19 of 20 paired cases (Wilcoxon p = 3.8e-06), while enlarging a same-model jury changed no verdict and no overlap. Substituting an attorney changed one verdict in ten and a judge two in ten, in every case by reframing which question was decisive rather than through evidence.

Conclusions. Simulated adjudication is repeatable and record-responsive, but its errors are structured: burden-driven, framing-sensitive and correlated across same-model jurors. These properties bear directly on how such systems should be designed and evaluated.

Section 01 · Introduction

Accuracy is not the whole question

Evaluations of language models on legal tasks have concentrated on single-model accuracy against known outcomes. A court, however, is a process: advocates select and frame evidence, a judge fixes the decision question and the standard of proof, and a jury deliberates and votes. When every role is played by a model, the behaviour of the process — not only the accuracy of the verdict — determines whether the result is legitimate, fair and stable.

This study asks four questions of a simulated court run on real disputes. Does its disposition agree with history, and when it does not, is the error random or directional? Is the process fair and consistent across reruns and configurations? Does substituting the attorney, judge or jury, enlarging the jury, or mixing jury models change the verdict, the reasoning, or only the style? And what does the behaviour imply for the design of such systems? The simulation platform is the instrument; the behaviour of model adjudication is the subject.

Section 02 · Methods

Frozen records, blinded outcomes, one factor changed at a time

Corpus

Fifty disputes were selected purposively for behavioural diversity across twenty domains (Appendix D). Source families comprised published court judgments, ombudsman determinations, regulatory adjudications, sports appeals, domain-name decisions, tribunal records and de-identified interpersonal dispute records. Each case was frozen into an outcome-blind packet containing a neutral brief, forum-specific rules, both parties’ positions and concessions, requested remedies, disputed facts and a normalized exhibit record delivered as documentary text. Forty cases had a clean binary historical disposition; nine were mixed or partial and one was settled without a merits decision.

Simulation protocol

Each trial proceeded through opening statements, direct and cross examination, closing statements, the judge’s instructions, deliberation, closed ballots and verdict. Every canonical trial used one configuration: GPT-5.6 Terra for both attorneys, Claude Opus 5 as judge, and Gemini 3.8 Flash for every juror. Seven further models entered only through interventions: Claude Fable 5.1, GPT-5.6 Sol, Grok 4.6, Kimi K3, DeepSeek V4 Pro, Qwen 3.8 Max and GLM-5.3. Historical outcomes were withheld from every actor. A trial was accepted only if it produced a verdict and passed actor-order, model-route and evidence-delivery audits; no accepted trial truncated the evidence record.

Experimental design

Table 1. Study components and trial counts.
ComponentTrialsDesign
Canonical sample50One audited trial per case under the fixed configuration; the reference for all comparisons.
Jury-composition gate20Ten cases, each tried once with a same-model jury and once with a mixed-model jury of equal size (3–11 seats).
Controlled interventions50Ten cases per panel, paired to the canonical trial: exact rerun; attorney substitution (five plaintiff, five defence); judge substitution; jury enlarged from three to nine of the same model; mixed-model jury of three.
Retained repeat baselines2Two additional exact reruns preserved from the primary run.

Measures

Historical agreement was assessed on the forty binary cases after transcript-level review of party-role mapping, which corrected one automated label. Each trial received a structured single-reviewer assessment of outcome legitimacy, procedural fairness, factual integrity, evidence use and advocacy balance (high, medium, low) and a set of non-exclusive failure flags. From each trial’s event log and decision traces we computed the language overlap between juror ballot rationales — the cosine similarity of term-weighted word and word-pair vectors, averaged over all juror pairs (0 = no shared wording, 1 = identical) — and the overlap between each rationale and the judge’s instructions, the winning closing and the losing closing; lexicon counts of certainty, hedging, concession and burden-of-proof markers per 100 words; deliberation order, explicitly stated side-leaning where present, and references to other jurors; and, for every decision, the strategy candidates listed and the one chosen.

Statistical analysis

Proportions are reported with Wilson 95% confidence intervals. Paired within-case comparisons use the Wilcoxon signed-rank test; the asymmetry of disagreements uses Fisher’s exact test; the composition gate uses McNemar’s exact test. Ten cases per intervention panel detect only large sensitivities. No model was randomised across roles, so per-model comparisons are descriptive and confined to the appendix.

Section 03 · Results

What the court did

Result 3.1 · Agreement with history

The verdict matched the historical outcome in 35 of 40 clean cases

The canonical verdict agreed with the historical disposition in 35 of 40 binary cases (Figure 1). Three of the five disagreements were sports disputes in which a question of jurisdiction, rule interpretation or burden was compressed into an adversarial winner. In the darts appeal the packet inverted the party roles, so the automated winner label masked a merits result opposite to history; the reviewed figure is used throughout.

87.5%
0%25%50%75%100%
Figure 1. Reviewed agreement between the canonical verdict and the historical disposition. Point estimate 35/40; band, Wilson 95% interval 73.9%–94.5%. Ten mixed or settled cases are excluded.
Result 3.2 · Direction of disagreement

When the court disagreed with history, it favoured the defendant

History favoured the claimant in 21 of the forty cases and the defendant in 19; the court returned 22 defendant verdicts. 4 of the 5 disagreements were defendant verdicts in claimant-aligned cases (Figure 2; Fisher exact p = 0.35). The mechanism is visible in the ballots: 90% of defendant ballots invoked the burden of proof against 64% of plaintiff ballots (mean 1.18 vs 0.69 references), deliberation statements that named a side leaned to the defendant 145 times and to the plaintiff 60 times, and plaintiff ballots carried twice the certainty markers of defendant ballots (0.25 vs 0.12 per 100 words). Where the record was thin, the standard contested or the remedy partial, “not proved” was the path of least resistance. The ten mixed or settled histories, forced through a binary form, produced 6 defendant and 4 plaintiff winners.

History favoured the claimant (n = 21)81%
History favoured the defendant (n = 19)95%
Defendant ballots citing the burden of proof90%
Plaintiff ballots citing the burden of proof64%
Figure 2. Top: agreement with history by the side history favoured. Bottom: share of juror ballots that explicitly invoked the burden of proof, by the side voted for (all 504 ballots).
Result 3.3 · Legitimacy and fairness

Matching history was easier than earning a high review

Agreement with history was more common than a high review rating (Figure 3). Outcome legitimacy was rated high in 6 cases and procedural fairness in 3; the modal assessment was a defensible result reached through thinner reasoning, a cruder remedy or a less forum-specific procedure than the real adjudication. Speaking volume did not decide cases: plaintiff counsel spoke more in 47 of 50 trials and won 22, and the side citing more unique exhibits won 6 of the 12 trials in which it did so. Winners conceded less (1.6 vs 2.4 concession markers per side) and asserted more (5.7 vs 4.4 certainty markers).

Outcome legitimacyHigh 6 · Medium 43 · Low 1
Procedural fairnessHigh 3 · Medium 44 · Low 3
Factual integrityHigh 3 · Medium 47
Evidence useHigh 9 · Medium 41
Advocacy balanceHigh 36 · Medium 13 · Low 1
Figure 3. Structured review ratings of the fifty canonical trials (teal = high, blue = medium, rose = low).
Result 3.4 · Repeatability

The same record produced the same answer almost everywhere

All 12 exact reruns reproduced the canonical winner and were unanimous. Thirty-six cases were tried more than once under some combination of rerun, attorney substitution, judge substitution, jury enlargement and jury mixing; 33 returned the same winner under every condition. The three exceptions share a structure: a decision with several parts forced into one winner. In the former spouse’s repayment claim, a Kimi K3 defence attorney and, separately, a Kimi K3 judge moved the decisive question from acknowledgment of the debt to quantum and limitation, and the verdict reversed. In the chess team-selection jurisdiction appeal, a mixed jury and a GPT-5.6 Sol judge each placed federation autonomy first and reversed the canonical answer toward history. In the commercial grower’s wrong-seed claim, a mixed jury accepted breach and foreseeability but refused an unproven quantum and reversed away from history. No reversal involved evidence the canonical trial had not seen.

Result 3.5 · Jury composition

Same-model juries were unanimous because their members thought alike

Same-model juries were unanimous in 91 of 92 trials (1 dissent in 404 ballots, in the role-inverted darts appeal); mixed-model juries in 17 of 20 (4 dissents in 100 ballots). In the composition gate both jury types agreed with history in 8 of 10 cases; two verdicts changed, one toward and one away from history (McNemar p = 1.0).

Rationale overlap was 0.37 on same-model juries and 0.23 on mixed juries; in the 20 paired cases the mixed panel was lower in 19 (Figure 4; Wilcoxon p = 3.8e-06). Deliberation statements showed the same shift (0.38 → 0.27). Enlarging a Gemini jury from three to nine seats produced 9–0 in all ten pairs and left overlap essentially unchanged (0.38 → 0.35). Same-model jurors in the same trial overlapped more (0.36) than the same model on the same case in a different trial (0.31), indicating that part of the convergence arises from shared in-trial context rather than a shared prior alone. Ballots overlapped with the winning closing (0.20) more than with the losing closing (0.15) or the instructions (0.16).

Sitewide discount claims excluded already-re0.36 → 0.25
Popeyes influencer and paid Instagram food a0.38 → 0.20
Pub conversation and alleged share-price bon0.33 → 0.18
Red Bull protest of Car 63's yellow-flag res0.49 → 0.22
Sudanese Online Olympiad team-selection juri0.40 → 0.26
Damp and mould response by London & Quadrant0.42 → 0.23
Puppy buyer disputes purebred description us0.26 → 0.22
Commercial grower receives the wrong chilli 0.29 → 0.20
Conjoined clinical-negligence appeals on psy0.31 → 0.14
Canal owner challenges statutory bar to poll0.33 → 0.22
Small-business APP fraud reimbursement compl0.39 → 0.27
Football playoff expulsion and post-match be0.48 → 0.28
Former partner seeks repayment for family-pr0.27 → 0.22
Neighboring farmers dispute the balance of a0.45 → 0.28
Father and adult son dispute whether a handm0.31 → 0.16
Phone loan is due only when the borrower rec0.36 → 0.32
Training provider seeks course fees after jo0.34 → 0.34
Lender appeals limitation ruling on undisclo0.30 → 0.08
Viewing-platform visitors looking and photog0.45 → 0.23
Former manager disputes transfer of kaimun.c0.37 → 0.28
same-model jury mixed-model jury
Figure 4. Mean pairwise language overlap between juror rationales in the same case, same-model jury versus mixed-model jury (n = 20 paired cases; ten from the composition gate, ten from the intervention panel).

Deliberation functioned as sequential monologue. 5.4% of same-model and 10.0% of mixed deliberation statements referenced another juror. Where a juror stated a lean (205 of 504 deliberations), its ballot followed 93% of the time, and the 14 exceptions all but once joined the eventual majority. The first speaker’s stated lean matched the verdict in 93% of trials. No juror used the available option to pass.

Result 3.6 · Actor substitution

Changing an actor changed the argument more often than the verdict

Table 2. Outcomes of the paired intervention panels (ten cases each). Agreement with history is reported over the historically comparable cases in each panel.
InterventionVerdicts changed95% CIUnanimousRationale overlapAgreed with history
Exact rerun0 / 100%–28%10 / 100.387 / 8
Attorney substitution1 / 102%–40%10 / 100.386 / 8
Judge substitution2 / 106%–51%10 / 100.347 / 8
Jury enlarged 3 → 9, same model0 / 100%–28%10 / 100.357 / 8
Mixed-model jury0 / 100%–28%9 / 100.257 / 8
Mixed-model jury (composition gate)2 / 106%–51%8 / 100.218 / 10

Attorney substitution altered volume and posture far more than evidence (Figure 5). Every substitute cited the canonical exhibit set within one item; Claude Fable 5.1 and Claude Opus 5 argued at roughly twice the length of GPT-5.6 Terra and conceded several times more often, Qwen 3.8 Max conceded nothing, and Kimi K3 produced the single reversal by reframing limitation and quantum. Judge substitution changed the length and hedging of the instructions (Figure 6) but not the degree to which jurors echoed them (0.12–0.25 across judges). Both judge-related reversals were framing events: Kimi K3 converted a potentially partial award into an all-or-nothing $30,000 proof threshold, and GPT-5.6 Sol placed federation autonomy before the jurisdiction question. One substitute judge produced no instructions at all and the jury voted regardless (Appendix C).

Claude Fable 5.1 · Red Bull protest of Car 63's yello2,367 → 5,011
Claude Opus 5 · Used-car wet-belt durability and e3,362 → 5,841
GPT-5.6 Sol · Sheltered-housing roof leaks prope3,810 → 3,340
Gemini 3.8 Flash · Grouped football substitutions cha2,656 → 2,987
Grok 4.6 · Puppy buyer disputes purebred desc3,688 → 3,756
Kimi K3 · Former spouse seeks repayment of rverdict changed3,265 → 4,740
DeepSeek V4 Pro · Former partners dispute ownership 3,231 → 4,056
Qwen 3.8 Max · Mare owner refuses breeding-servic2,666 → 2,687
GLM-5.3 · Conjoined clinical-negligence appe3,424 → 3,739
Claude Fable 5.1 · Former manager disputes transfer o2,953 → 5,881
GPT-5.6 Terra (canonical) substitute
Figure 5. Words spoken by the substituted side in the ten attorney-substitution trials, canonical attorney versus substitute.
Claude Fable 5.1 · Popeyes influencer and paid Instag715 → 688
GPT-5.6 Sol · Sudanese Online Olympiad team-seleverdict changed598 → 507
GPT-5.6 Terra · Small-business APP fraud reimburse744 → 511
Gemini 3.8 Flash · Sheltered-housing roof leaks prope702 → 538
Grok 4.6 · Puppy buyer disputes purebred desc731 → 1,147
Kimi K3 · Former spouse seeks repayment of rverdict changed749 → 808
DeepSeek V4 Pro · Phone loan is due only when the bo734 → 506
Qwen 3.8 Max · Lender appeals limitation ruling ono instructions708 → 0
GLM-5.3 · Family-farm inheritance promise an607 → 557
Claude Fable 5.1 · Former manager disputes transfer o1,111 → 1,156
Claude Opus 5 (canonical) substitute
Figure 6. Length of the jury instructions in the ten judge-substitution trials, canonical judge versus substitute.
Section 04 · Discussion

Repeatable, directional, and correlated

Three properties of the simulated court emerge consistently. First, it is repeatable: identical inputs produced identical winners in every rerun, and the great majority of cases were invariant to who argued, who instructed and who decided. Where the winner did change, the case had more than one dispositive question and the change was one of framing. A single-winner verdict form therefore does not merely lose information about partial outcomes; it makes the result hostage to whichever actor fixes the order of questions.

Second, its errors have a direction. Under a burden-of-proof lens a thin, contested or partial record resolves as “not proved”, and the ballots show that lens applied far more often when voting for the defendant. The asymmetry is not statistically established at forty cases, but its mechanism is directly observable, and it is the hypothesis most worth a preregistered test with issue-level outcomes.

Third, unanimity on same-model juries reflects correlated cognition rather than persuasion. Jurors did not reference one another, arrived at deliberation already decided, and wrote rationales whose overlap fell sharply when the panel mixed models but not at all when it added seats. Diversity of model, not number of jurors, is what changed the reasoning; it did not detectably change accuracy.

Simulated adjudication is therefore best understood as a repeatable, evidence-responsive instrument whose failure modes are structural and designable: issue-level verdicts rather than a single winner, heterogeneous juries rather than larger ones, and instructions that must exist before a vote. The secondary behaviours recorded in the appendix — an absolute preference for the first strategy considered, certainty rising where the court was wrong, and a jury voting on without instructions — point the same way.

Section 05 · Next steps

The court will keep being studied, and next it will be tried under pressure

Every trial in this study was run in good faith: both sides argued from a frozen, honest record, and the court was never deliberately misled. That is the right first test of a simulated court and the wrong last one. The behaviours documented here — burden-driven defaults, framing sensitivity, correlated juror cognition, actors that never revisit their first strategy — are exactly the properties an adversary would exploit. The programme continues along the following vectors, in roughly this order.

  1. Adversarial evidence. Trials in which one side submits fabricated exhibits, altered dates, or documents that contradict the shared record, to measure whether attorneys detect the conflict, whether judges instruct on it, and how jurors resolve two exhibits that cannot both be true.
  2. Manipulation of the court itself. Evidence and submissions carrying instructions addressed to the models rather than to the tribunal — the courtroom form of prompt injection — and rhetorical devices aimed at the known biases: appeals to certainty, invented precedent, burden inversion, and volume.
  3. Conflicting and incomplete records. Cases where the two sides hold genuinely inconsistent accounts with no tiebreaker, where key exhibits are missing, or where evidence arrives late, to characterise how the court behaves at the edge of proof rather than in its centre.
  4. Issue-level verdicts. Replacing the single winner with structured findings — liability, causation, quantum, limitation, remedy — and re-running the multi-question cases that were unstable here, to test whether the instability disappears when the verdict form can express the answer.
  5. The defendant-ward burden effect. A preregistered test on a balanced corpus with issue-level outcomes, varying the standard of proof and the wording of the instructions, to establish whether the asymmetry observed here is a property of the court or of this sample.
  6. Jury design. Heterogeneous panels by default; structured deliberation with mandated cross-reference, sequential versus simultaneous ballots, an appointed dissenter, and randomised speaking order, measured against the same overlap and dissent metrics used here.
  7. Judicial variation. Instructions that differ in standard, in the order of questions, and in the presence or absence of a charge on burden, to isolate how much of a verdict the judge writes before the jury sits.
  8. Appeal and review. A second-instance court reviewing first-instance transcripts for error, to test whether a model tribunal can detect the failure modes catalogued here when they occur in another model tribunal.
  9. Human comparison. The same frozen records put to practising advocates and lay panels under matched procedure, so that agreement, direction of error and convergence can be compared between human and model courts rather than only against history.
  10. Multimodal and native evidence. Original images, audio and video in place of documentary reconstructions, to test whether perception changes what the court finds proved.
  11. Scale. A larger, stratified corpus with repeated randomised role assignment, so that per-model and per-domain effects can be estimated rather than described.

Each of these will be run on frozen records under the same audit discipline, with failed and surprising trials preserved rather than rerun, and published against the exact environment version they describe.

Section 06 · Limitations

Bounded claims

  • The corpus is purposively diverse, not a probability sample; forty cases support binary comparison; history is a benchmark, not ground truth.
  • Ten cases per intervention panel detect only large sensitivities; the defendant-ward asymmetry is directional.
  • Cases recur across panels, so trials are not independent observations.
  • Reviews are structured single-reviewer judgments; language, lean and overlap measures are lexical approximations.
  • No model was randomised across roles and ballot counts are unbalanced; model profiles describe style, not accuracy.
  • Evidence was delivered as documentary text; multimodal perception was not tested.
  • A binary verdict cannot express partial remedies, regulatory decisions, consent outcomes or appeals, and every observed instability occurred in such cases.
Section 07 · Conclusion

Properties of the process, not of a model

A simulated court of language models agreed with the historical outcome in 35 of 40 real disputes, reproduced its own verdicts exactly, and remained stable under most substitutions of attorney, judge and jury. Its disagreements ran toward the defendant through burden framing, its reversals occurred only where a multi-part decision was forced into one winner, and its juries converged because their members shared a model, not because they deliberated. These are properties of the process rather than of any single model, and they define what such a system can be relied on for and what it cannot.

Appendix

Secondary findings and coverage

Appendix A · Model behaviour by role

Models differed in style, length and confidence, not detectably in accuracy

Gemini 3.8 Flash cast 419 of 504 ballots; the other nine juror models cast 5–13 each, and no model was randomised across roles. Per-model agreement with history ranged from 75% to 90% on these denominators and cannot be distinguished. Behavioural differences were nonetheless consistent (Table 3, Figure 7). Gemini wrote the shortest rationales and almost never named an exhibit in its ballot; Claude Opus 5 and Grok 4.6 anchored ballots to named exhibits; the OpenAI and Qwen jurors used two to three times more certainty language per word than the Anthropic, Google or DeepSeek jurors. As attorneys, the Anthropic models argued longest and conceded most, Qwen shortest and never. Instruction length ranged from about 500 words (DeepSeek, GPT-5.6, Gemini) to over 1,100 (Grok, Claude Fable 5.1).

Table 3. Juror behaviour by model across all 504 ballots.
ModelBallotsWords per rationaleExhibits namedCertainty per 100 wordsVoted plaintiffDissents
Gemini 3.8 Flash419740.100.1742%1
Claude Fable 5.1131340.850.1238%1
GPT-5.6 Sol13680.000.4846%0
Qwen 3.8 Max11670.000.3945%0
Grok 4.6111041.910.1036%1
DeepSeek V4 Pro9950.220.0833%0
Kimi K391620.890.1444%1
Claude Opus 572312.860.1743%0
GLM-5.372020.570.1343%0
GPT-5.6 Terra5600.000.2440%1
Attorneys: words per statement
GPT-5.6 Terra512
Claude Fable 5.1908
Claude Opus 5974
DeepSeek V4 Pro676
Gemini 3.8 Flash498
Kimi K3790
GPT-5.6 Sol557
Qwen 3.8 Max448
Grok 4.6626
GLM-5.3623
Attorneys: concessions per statement
GPT-5.6 Terra0.33
Claude Fable 5.12.75
Claude Opus 53.33
DeepSeek V4 Pro0.50
Gemini 3.8 Flash0.50
Kimi K31.83
GPT-5.6 Sol0.50
Qwen 3.8 Max0.00
Grok 4.61.00
GLM-5.31.33
Judges: words of jury instructions
Claude Opus 5723
Claude Fable 5.1922
DeepSeek V4 Pro506
Gemini 3.8 Flash538
Kimi K3808
GPT-5.6 Sol507
GPT-5.6 Terra511
Qwen 3.8 Max0
Grok 4.61147
GLM-5.3557
Jurors: certainty markers per 100 words
Gemini 3.8 Flash0.17
Claude Fable 5.10.12
GPT-5.6 Sol0.48
Qwen 3.8 Max0.39
Grok 4.60.10
DeepSeek V4 Pro0.08
Kimi K30.14
Claude Opus 50.17
GLM-5.30.13
GPT-5.6 Terra0.24
Figure 7. Behavioural profiles by model and role. Attorney and judge panels include the canonical model (1,284 statements; 102 instruction sets) and each substitute (six statements or one instruction set each).
Appendix B · Decision traces

Patterns in how the actors decided

Strategy selection. Each actor listed strategy candidates with an associated risk before acting. Across 2,467 decisions the first candidate was chosen 2,459 times (99.7%); the remaining 8 were unparseable outputs. The stated risk never changed a choice.

Certainty and error. Ballots in disagreeing trials carried 2.1 times the certainty markers per 100 words of ballots in agreeing trials (0.35 vs 0.16) and 2.5 times the hedges; advocates’ certainty per 100 words rose from 0.09 at opening to 0.21 at closing while hedging remained flat.

Reviewer flags. Correlated juror framing (39/50) and a verdict form unable to express partial liability, separate issues, constrained remedies, consent or appellate relief (34/50) were flagged at similar rates in agreeing and disagreeing cases; burden and causation errors were the flags that tracked disagreement (Figure 8).

Shared model convergence89% → 80%
Outcome schema compression66% → 40%
Overconfidence43% → 40%
Burden or causation error20% → 60%
Evidence resolution or coverage14% → 20%
Procedural forum mismatch17% → 20%
Reasoning or instruction error11% → 0%
agreed with history disagreed
Figure 8. Share of the forty binary cases carrying each non-exclusive reviewer flag, by agreement with history.
Appendix C · Missing output

When an actor produced nothing, the court carried on

On 8 occasions an actor’s output carried no usable answer, 4 of them because the model reasoned until its output limit without reaching a decision. In the PPI-commission limitation appeal the Qwen 3.8 Max judge delivered no instruction text; the three Gemini jurors deliberated and voted 3–0 for the historically aligned side from the closings alone, none remarked on the absence, and their rationales invoked a burden of proof that had not been charged. The reviewer rated that trial low on legitimacy and fairness. A Gemini juror and a DeepSeek juror each submitted an empty ballot after reasoning to the limit; both were rejected and both jurors voted validly on retry. One DeepSeek juror proposed a second ballot after a valid one, which the court refused.

Appendix D · Coverage

Case coverage by domain

Domain-level reviewed comparison counts are unavailable, not zero. The table reports case coverage and simulated winners only; the study-wide historical comparison remains 35 of 40 clean binary cases.

Table 4. Case coverage and simulated winners by domain. Small cells are descriptive only.
DomainCasesSim. winners P / D
Consumer Financial62 / 4
Neighbor Property64 / 2
Court51 / 4
Interpersonal Family42 / 2
Betting Gambling31 / 2
Domain Ip32 / 1
Interpersonal Friendship31 / 2
Interpersonal Romantic31 / 2
Marketplace Ecommerce31 / 2
Sports31 / 2
Advertising Claims21 / 1
Commercial Supply11 / 0
Consumer Services11 / 0
Education Services10 / 1
Insurance10 / 1
Interpersonal Household10 / 1
Interpersonal Recreational11 / 0
Recreational Services11 / 0
Transport Tort10 / 1
Workplace Organizational11 / 0
Total5022 / 28

The research archive preserves every frozen packet and hash, event transcript, decision trace, actor and evidence audit, juror ballot, case review and cross-case table, together with the run-, statement-, ballot-, trace- and instruction-level tables and scripts used here. Failed or incomplete attempts are stored separately and were never substituted.