CrossLaw.
A Cross-Jurisdiction Framework for Evaluating Large Language Models in Common-Law Systems
Evaluating what large language models get right, what they get wrong, and how confidently they say it.
Results describe the models and interfaces tested in 2026 and are not a recommendation to use or avoid any product. Reasoning-quality results await lawyer verification.
01 / The research
One study.
Two connected parts.
CrossLaw tests commercial language models on appellate outcome prediction. Part I compares six models across Nigeria, Australia, the United Kingdom and the United States. Part II examines how supplying Nigerian legal materials in different forms is associated with prediction performance.
Across both parts, the thesis reports 2,680 scored predictions: 960 in the primary benchmark and 1,720 further predictions in the Nigerian extension. The same Nigerian baseline predictions are counted once.
Source: Thesis §§1.2–1.3, Chapter 11 · source-reportedCross-jurisdiction benchmark
Four legal systems. Two prompting conditions.
80 final appellate cases, six commercial models, with and without citation-based context.
Thesis Chapters 3–4Nigerian extension · new predictions
One jurisdiction. Four stages of inquiry.
General materials, preloading, summary length and summarisation tools tested on the Nigerian cases.
Thesis Chapters 5–1002 / Part I · Primary benchmark
The context mattered.
The caveats do too.
346/480 → 390/480 · +9.2 percentage points
Source: Thesis Table 4.4 · source-reportedIncorrect verdicts stated with more than 50% confidence (EV < −0.5): 23.1% of primary predictions.
Source: Thesis §4.5 · source-reportedCitation context was associated with higher accuracy for five of six models; Grok was the exception. These results describe the tested models and interfaces, not legal competence.
03 / Part II · Nigerian extension
Not just more information.
How information is supplied.
Supply
General Nigerian legal materials
Simultaneous loadingPreload
Materials before the case
Plus identical rerunsCompress
Relational summary length
Five tested lengthsCompare
Summarisation tools
Fixed prediction modelSource: Thesis Chapters 5–9 · Experimental sequence, not a single pre-registered design.
04 / At a glance
Eleven major findings.
Each with a source.
The thesis’s own summary, kept in its original wording. Select a finding to read it in full.
01No single model led in every settingPart I
No single model led in every setting. Grok had the highest aggregate accuracy under No Clue (67/80, 83.8%) and Gemini under With Clue (75/80, 93.8%). Accuracy also varied across jurisdictions within individual models. DeepSeek, for example, was correct on 20/20 United Kingdom cases under No Clue but on 9/20 to 10/20 cases in the other jurisdictions.
Thesis §1.5, Tables 4.1–4.3 · source-reportedA score in one country or prompting condition cannot stand in for evidence on Nigerian cases; compare the relevant model and condition before drawing conclusions.
02Citation context was associated with higher accuracyPart I
Citation context was associated with higher accuracy for five of six models. Pooled accuracy rose from 346/480 (72.1%) to 390/480 (81.3%). Grok was the exception (67/80 to 64/80). The paired change was statistically significant for Gemini (p = 0.0044) and ChatGPT (p = 0.0347).
Thesis §1.5, Table 4.4 · source-reportedCase-specific citation context changed results, but it did not help every model. A supplied citation is context to check, not independent confirmation.
03Confident errors were commonPart I
Confident errors were common. Of the 960 primary predictions, 222 (23.1%) were Silent Failures, that is incorrect verdicts stated with more than 50% confidence. The count fell from 134 under No Clue to 88 under With Clue. The jurisdictional pattern did not follow a simple resource ordering, since Australia had the highest Silent Failure rate (30.8%).
Thesis §1.5, Tables 4.7–4.10 · source-reportedAn authoritative-sounding answer may be wrong. Verify the outcome and cited authority against primary legal sources rather than relying on stated confidence.
04Accuracy and confidence reliability were separablePart I
Accuracy and confidence reliability were separable. DeepSeek’s 20/20 United Kingdom result coexisted with the highest Expected Calibration Error and Brier score of the six models, and fell to 16/20 on a rerun.
Thesis §1.5, §§4.4, 4.9 · source-reportedAccuracy alone hides unreliable confidence and rerun instability. Examine calibration and repeated-case results alongside correct-verdict counts.
05General legal materials helped modestlyPart II
General legal materials helped modestly, clues helped most. Adding the fixed 11-source Nigerian corpus raised facts-only (No Clue) accuracy from 67.5% to 72.5% under simultaneous loading and to 75.0% and 74.2% when preloaded, but no material-based condition reached the original With Clue baseline (84.2%).
Thesis §1.5, Chapters 6–7 · source-reportedAdding general Nigerian legal material had a smaller observed association than case-specific context in these tests; more material alone is not a dependable remedy.
06Preloading was modestly betterPart II
Preloading was modestly better than simultaneous loading (+1.7 to +4.2 points), a consistent direction but within run-to-run variation.
Thesis §1.5, Chapter 7 · source-reportedPreloading deserves controlled testing, not a blanket recommendation: its modest observed advantage was within run-to-run variation.
07Model-constructed summaries did not match fuller materialsPart II
Model-constructed relational summaries did not match the fuller materials under preloading. Across five preloaded lengths from 50–100 to 500–750 words per source, pooled accuracy was flat (64.2%–67.5%), below the preloaded fuller materials (74.6%) and below the 70.0% always-Dismissed reference.
Thesis §1.5, Chapters 7–8 · source-reportedCheck which legal rules a summary retains before treating a shorter, model-written version as equivalent to fuller materials.
08Each model had its own length curvePart II
Each model had its own length curve. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT (90.0%) and at 150–300 words or longer for Claude, while Grok and Perplexity were below their facts-only baselines at every length. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length.
Thesis §1.5, Chapter 8 · source-reportedTest summary lengths separately for each model; a pooled length result concealed different model-level patterns.
09The summarisation tool matteredPart II
The summarisation tool mattered when the predictor was fixed. With ChatGPT as the predictor, stored summaries from ChatGPT and Claude gave 90.0% at 100–150 words and 85.0% at 150–300 words, against 80.0% for Gemini Notebook and 75.0% for Microsoft Copilot and Gemini. Every tool gave equal or higher accuracy at the shorter length. Claude’s summaries at 100–150 words gave the highest two-run mean (92.5%, the mean of 90.0% and 95.0%). This is descriptively above ChatGPT’s accuracy with the preloaded full materials in the second stage (85.0%), although that comparison crosses stages and is not controlled.
Thesis §1.5, Chapter 9 · source-reportedThe tool creating a summary can affect downstream answers even with the predictor held fixed. This small, descriptive comparison needs replication before tool-selection advice.
10Run-to-run consistency is itself a findingPart II
Run-to-run consistency is itself a finding. Identical reruns of the six-model GLMP experiment agreed on only about three verdicts in four, with swings of up to 55 points for a single model (Grok) and 40 points for another (DeepSeek), whereas the fixed-predictor reruns of the fourth stage agreed on 18–19 of 20 verdicts.
Thesis §1.5, Chapters 7 and 9 · source-reportedRepeat identical trials and report variation; a single run can give a misleading impression of how a system performs.
11Confidence was not a reliable signalPart II
Confidence was not a reliable signal of correctness, and accuracy in all conditions was markedly lower on Allowed than on Dismissed cases. Two Nigerian cases, NG 009 and NG 020, were predicted incorrectly by all six models under No Clue in the primary benchmark and remained incorrect in every run-a condition of the fourth stage.
Thesis §1.5, Chapters 9–10 · source-reportedCheck performance on both Allowed and Dismissed outcomes and examine difficult cases; a high aggregate score can conceal asymmetric errors.
Note: Finding 11 is the thesis’s broad summary; the fixed-predictor tool results in Chapter 9 include stronger Allowed-case performance in particular conditions. The pattern is not uniform across individual conditions.
05 / The real-world question
Who needs this evidence,
and why?
These are possible uses of the findings, not measured changes in legal practice. The study does not establish client outcomes or endorse a product.
Primary · Nigerian legal practitioners
AI-assisted research is growing, while independent evidence on Nigerian appellate cases and suitable verification practices is limited.
The Nigerian results show how the tested models, citation context and different forms of local legal material behaved on the same cases. This can inform questions to ask and what to verify; it is not a current vendor recommendation.
Findings 1–3 and 5–11 · Thesis §§1.1, 1.4, 1.5, 10.12; proposal Impact StatementSecondary · Australian practitioners, law firms and courts
A plausible citation or confident legal answer can still be wrong, including on Australian legal questions.
The Australian comparisons make confident error and jurisdiction-sensitive performance visible, supporting independent checking of authorities and outcomes rather than reliance on a confidence score.
Findings 1–4 and 11 · Thesis §§1.1, 1.4, 1.5; proposal Impact StatementWider · researchers, benchmark builders, developers and policymakers
Evaluations concentrated in better-resourced legal systems may miss local variation, calibration problems and unstable reruns.
The shared four-jurisdiction design and Nigerian extension offer testable evidence for broader evaluation, including model-by-jurisdiction comparisons, confidence reliability and repeatability. Governance tools are still planned.
Findings 1–4 and 7–11 · Thesis §§1.1, 1.4, 1.5, 10.10–10.12; proposal Impact StatementPractitioner’s Manual, Prompt Optimisation Playbook, Model Selection Matrix, Duty of Inquiry Checklist and Traceable Governance Framework remain planned. Lawyer verification of reasoning is in progress. Read the next steps
Where the research stands
The evidence has a boundary.
The work continues.
Phases 1–3 · benchmark and four-stage Nigerian extension
Phase 4 · independent lawyer verification of legal reasoning
Phase 5 · practitioner resources; Phase 6 · thesis and publication
The next five minutes
Take the research
into the room.
Start presentation Explore the full research
