Research presentation · 202601 / 04

CrossLaw.

A Cross-Jurisdiction Framework for Evaluating Large Language Models in Common-Law Systems

Evaluating what large language models get right, what they get wrong, and how confidently they say it.

Ebere Josephine Uba
Master of Artificial Intelligence (Research)
Supervised by A/Prof Dongmo Zhang and Dr Jolin Qu
School of Computer, Data and Mathematical Sciences · Western Sydney University
Evidence boundary

Results describe the models and interfaces tested in 2026 and are not a recommendation to use or avoid any product. Reasoning-quality results await lawyer verification.

01 / The research

One study.
Two connected parts.

CrossLaw tests commercial language models on appellate outcome prediction. Part I compares six models across Nigeria, Australia, the United Kingdom and the United States. Part II examines how supplying Nigerian legal materials in different forms is associated with prediction performance.

Across both parts, the thesis reports 2,680 scored predictions: 960 in the primary benchmark and 1,720 further predictions in the Nigerian extension. The same Nigerian baseline predictions are counted once.

Source: Thesis §§1.2–1.3, Chapter 11 · source-reported
I / 960

Cross-jurisdiction benchmark

Four legal systems. Two prompting conditions.

80 final appellate cases, six commercial models, with and without citation-based context.

Thesis Chapters 3–4
II / 1,720

Nigerian extension · new predictions

One jurisdiction. Four stages of inquiry.

General materials, preloading, summary length and summarisation tools tested on the Nigerian cases.

Thesis Chapters 5–10

02 / Part I · Primary benchmark

The context mattered.
The caveats do too.

Pooled accuracy · No Clue → With Clue
72.1%81.3%

346/480 → 390/480 · +9.2 percentage points

Source: Thesis Table 4.4 · source-reported
Silent Failure · study-specific definition222/960

Incorrect verdicts stated with more than 50% confidence (EV < −0.5): 23.1% of primary predictions.

Source: Thesis §4.5 · source-reported

Citation context was associated with higher accuracy for five of six models; Grok was the exception. These results describe the tested models and interfaces, not legal competence.

03 / Part II · Nigerian extension

Not just more information.
How information is supplied.

01

Supply

General Nigerian legal materials

Simultaneous loading
02

Preload

Materials before the case

Plus identical reruns
03

Compress

Relational summary length

Five tested lengths
04

Compare

Summarisation tools

Fixed prediction model

Source: Thesis Chapters 5–9 · Experimental sequence, not a single pre-registered design.

04 / At a glance

Eleven major findings.
Each with a source.

The thesis’s own summary, kept in its original wording. Select a finding to read it in full.

01No single model led in every settingPart I

No single model led in every setting. Grok had the highest aggregate accuracy under No Clue (67/80, 83.8%) and Gemini under With Clue (75/80, 93.8%). Accuracy also varied across jurisdictions within individual models. DeepSeek, for example, was correct on 20/20 United Kingdom cases under No Clue but on 9/20 to 10/20 cases in the other jurisdictions.

Thesis §1.5, Tables 4.1–4.3 · source-reported
What this could mean · editorial interpretationNigerian legal practitioners

A score in one country or prompting condition cannot stand in for evidence on Nigerian cases; compare the relevant model and condition before drawing conclusions.

02Citation context was associated with higher accuracyPart I

Citation context was associated with higher accuracy for five of six models. Pooled accuracy rose from 346/480 (72.1%) to 390/480 (81.3%). Grok was the exception (67/80 to 64/80). The paired change was statistically significant for Gemini (p = 0.0044) and ChatGPT (p = 0.0347).

Thesis §1.5, Table 4.4 · source-reported
What this could mean · editorial interpretationPractitioners across the four jurisdictions

Case-specific citation context changed results, but it did not help every model. A supplied citation is context to check, not independent confirmation.

03Confident errors were commonPart I

Confident errors were common. Of the 960 primary predictions, 222 (23.1%) were Silent Failures, that is incorrect verdicts stated with more than 50% confidence. The count fell from 134 under No Clue to 88 under With Clue. The jurisdictional pattern did not follow a simple resource ordering, since Australia had the highest Silent Failure rate (30.8%).

Thesis §1.5, Tables 4.7–4.10 · source-reported
What this could mean · editorial interpretationLawyers and courts

An authoritative-sounding answer may be wrong. Verify the outcome and cited authority against primary legal sources rather than relying on stated confidence.

04Accuracy and confidence reliability were separablePart I

Accuracy and confidence reliability were separable. DeepSeek’s 20/20 United Kingdom result coexisted with the highest Expected Calibration Error and Brier score of the six models, and fell to 16/20 on a rerun.

Thesis §1.5, §§4.4, 4.9 · source-reported
What this could mean · editorial interpretationLaw firms evaluating tools

Accuracy alone hides unreliable confidence and rerun instability. Examine calibration and repeated-case results alongside correct-verdict counts.

05General legal materials helped modestlyPart II

General legal materials helped modestly, clues helped most. Adding the fixed 11-source Nigerian corpus raised facts-only (No Clue) accuracy from 67.5% to 72.5% under simultaneous loading and to 75.0% and 74.2% when preloaded, but no material-based condition reached the original With Clue baseline (84.2%).

Thesis §1.5, Chapters 6–7 · source-reported
What this could mean · editorial interpretationNigerian practitioners

Adding general Nigerian legal material had a smaller observed association than case-specific context in these tests; more material alone is not a dependable remedy.

06Preloading was modestly betterPart II

Preloading was modestly better than simultaneous loading (+1.7 to +4.2 points), a consistent direction but within run-to-run variation.

Thesis §1.5, Chapter 7 · source-reported
What this could mean · editorial interpretationTeams designing legal research workflows

Preloading deserves controlled testing, not a blanket recommendation: its modest observed advantage was within run-to-run variation.

07Model-constructed summaries did not match fuller materialsPart II

Model-constructed relational summaries did not match the fuller materials under preloading. Across five preloaded lengths from 50–100 to 500–750 words per source, pooled accuracy was flat (64.2%–67.5%), below the preloaded fuller materials (74.6%) and below the 70.0% always-Dismissed reference.

Thesis §1.5, Chapters 7–8 · source-reported
What this could mean · editorial interpretationTeams preparing Nigerian legal summaries

Check which legal rules a summary retains before treating a shorter, model-written version as equivalent to fuller materials.

08Each model had its own length curvePart II

Each model had its own length curve. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT (90.0%) and at 150–300 words or longer for Claude, while Grok and Perplexity were below their facts-only baselines at every length. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length.

Thesis §1.5, Chapter 8 · source-reported
What this could mean · editorial interpretationBenchmark designers and developers

Test summary lengths separately for each model; a pooled length result concealed different model-level patterns.

09The summarisation tool matteredPart II

The summarisation tool mattered when the predictor was fixed. With ChatGPT as the predictor, stored summaries from ChatGPT and Claude gave 90.0% at 100–150 words and 85.0% at 150–300 words, against 80.0% for Gemini Notebook and 75.0% for Microsoft Copilot and Gemini. Every tool gave equal or higher accuracy at the shorter length. Claude’s summaries at 100–150 words gave the highest two-run mean (92.5%, the mean of 90.0% and 95.0%). This is descriptively above ChatGPT’s accuracy with the preloaded full materials in the second stage (85.0%), although that comparison crosses stages and is not controlled.

Thesis §1.5, Chapter 9 · source-reported
What this could mean · editorial interpretationTeams comparing summarisation tools

The tool creating a summary can affect downstream answers even with the predictor held fixed. This small, descriptive comparison needs replication before tool-selection advice.

10Run-to-run consistency is itself a findingPart II

Run-to-run consistency is itself a finding. Identical reruns of the six-model GLMP experiment agreed on only about three verdicts in four, with swings of up to 55 points for a single model (Grok) and 40 points for another (DeepSeek), whereas the fixed-predictor reruns of the fourth stage agreed on 18–19 of 20 verdicts.

Thesis §1.5, Chapters 7 and 9 · source-reported
What this could mean · editorial interpretationResearchers and procurement teams

Repeat identical trials and report variation; a single run can give a misleading impression of how a system performs.

11Confidence was not a reliable signalPart II

Confidence was not a reliable signal of correctness, and accuracy in all conditions was markedly lower on Allowed than on Dismissed cases. Two Nigerian cases, NG 009 and NG 020, were predicted incorrectly by all six models under No Clue in the primary benchmark and remained incorrect in every run-a condition of the fourth stage.

Thesis §1.5, Chapters 9–10 · source-reported
What this could mean · editorial interpretationLawyers, courts and evaluators

Check performance on both Allowed and Dismissed outcomes and examine difficult cases; a high aggregate score can conceal asymmetric errors.

Note: Finding 11 is the thesis’s broad summary; the fixed-predictor tool results in Chapter 9 include stronger Allowed-case performance in particular conditions. The pattern is not uniform across individual conditions.

05 / The real-world question

Who needs this evidence,
and why?

These are possible uses of the findings, not measured changes in legal practice. The study does not establish client outcomes or endorse a product.

Primary · Nigerian legal practitioners

The gap

AI-assisted research is growing, while independent evidence on Nigerian appellate cases and suitable verification practices is limited.

What the research offers

The Nigerian results show how the tested models, citation context and different forms of local legal material behaved on the same cases. This can inform questions to ask and what to verify; it is not a current vendor recommendation.

Findings 1–3 and 5–11 · Thesis §§1.1, 1.4, 1.5, 10.12; proposal Impact Statement

Secondary · Australian practitioners, law firms and courts

The gap

A plausible citation or confident legal answer can still be wrong, including on Australian legal questions.

What the research offers

The Australian comparisons make confident error and jurisdiction-sensitive performance visible, supporting independent checking of authorities and outcomes rather than reliance on a confidence score.

Findings 1–4 and 11 · Thesis §§1.1, 1.4, 1.5; proposal Impact Statement

Wider · researchers, benchmark builders, developers and policymakers

The gap

Evaluations concentrated in better-resourced legal systems may miss local variation, calibration problems and unstable reruns.

What the research offers

The shared four-jurisdiction design and Nigerian extension offer testable evidence for broader evaluation, including model-by-jurisdiction comparisons, confidence reliability and repeatability. Governance tools are still planned.

Findings 1–4 and 7–11 · Thesis §§1.1, 1.4, 1.5, 10.10–10.12; proposal Impact Statement

Practitioner’s Manual, Prompt Optimisation Playbook, Model Selection Matrix, Duty of Inquiry Checklist and Traceable Governance Framework remain planned. Lawyer verification of reasoning is in progress. Read the next steps

Where the research stands

The evidence has a boundary.
The work continues.

Complete

Phases 1–3 · benchmark and four-stage Nigerian extension

In progress

Phase 4 · independent lawyer verification of legal reasoning

Planned

Phase 5 · practitioner resources; Phase 6 · thesis and publication

Source: Thesis §§3.12, 10.11, Appendix 12.6 · planned work where indicated

The next five minutes

Take the research
into the room.

Start presentation

Explore the full research

Follow the evidence, chapter by chapter.