Research benchmark · held-out set · 2026-07-28

We beat the leading frontier models on HCA legal questions.

Here is exactly what we did. We put ten real High Court questions to our system and to two of the leading frontier models. None was shown the Court's judgment — ours drew only on Australian law as it stood before the decision came down. Then we scored every answer against what the Court actually held. We came out ahead of both — about 10% higher than one and 18% higher than the other — and backed our reasoning with 306 paragraph-level citations you can open and check, against Fable 5's 16.

Instrument
Held-out set of 10 High Court matters
Graded
all three arms, one panel
Scale
0–100, panel median
Runs
1 per system per case
Scroll to see how it was measured

What we tested

What a good answer has to do.

We scored each memo the way a barrister judges a junior's note — starting with whether it found the right issues, and ending with whether every authority it leans on can actually be checked.
  1. 01

    Did it identify the right issues?

    The first test of any research memo is whether it spots the questions that actually decide the case — in the right jurisdiction, and the right body of law. A confident answer to the wrong question is worse than no answer at all.

  2. 02

    Is the reasoning sound?

    Does it apply the law correctly and reach a conclusion a court could accept? This is the substance a barrister is paying for, and it is the larger part of the score.

  3. 03

    Can the authority be checked?

    A citation a barrister cannot verify is work they have to redo. We score whether the memo says which case, and where in it — the paragraph, not just the report.

  4. 04

    Were there errors?

    Invented cases, misattributed holdings and superseded authority are counted separately from quality, because they carry a different professional cost.

How it was run

Same questions. Same panel. Same day.

We are not publishing the question bank or the scoring rubric — releasing them would let anyone tune against them. What follows is the protocol, which is the part that matters for judging the result.
  1. 01

    Identical questions

    All three systems answered the same ten questions, drawn from ten Australian High Court matters. None saw the others' work.

  2. 02

    One panel, applied to all three

    Every memo was scored by a panel of three independent models from three different AI labs, aggregated by median so no single grader can carry a result.

  3. 03

    Case citations machine-checked

    Each pinpoint reference to a judgment was checked against the text of that judgment, and counts as verified only where the paragraph was actually found there. This check covers case law: references to legislation were not verified this way, so the citation figures describe the case-law half of each memo.

  4. 04

    Nothing tuned after the fact

    The scoring was run once, on memos already written. No configuration was changed, and no case was dropped, after any score was seen.

The result

Ahead of both.

Two leading frontier models, both answering cold, on the same ten matters. We came out ahead of each — comfortably against one, and by a wider margin still against the other.
  • vs Fable 5, cold+10%

    +6.05 points · won 8 of the ten · 66.6 to 60.6

  • vs GPT-5.6 Sol, cold+18%

    +10.07 points · won 9 of the ten · 66.6 to 56.6

And these are not weak opponents. On LegalBench, the academic benchmark for legal reasoning, Fable 5 ranked first and GPT-5.6 Sol fifth as of 23 July 2026. Leaderboard: vals.ai

Fig. 1

Our lead over both models, case by case

vs Fable 5, cold

vs GPT-5.6 Sol, cold

One panel per model, same axis, same case order. Each bar is our margin on one case: rightward we led, leftward (slate) the model did. The shaded band is the ±3-point run-to-run variation this instrument is known to have — a hollow mark inside it is better read as level than a win. The labelled line in each panel is the mean lead over that model. Confidence intervals for both are in the statistical detail below. Every value is in the table.

The one real loss, explained. Our only genuine loss — HCA-1 — turned on two Hong Kong judgments: authority a frontier model carries from training, but our Australian corpus does not yet hold. It is a coverage gap, not a reasoning gap, and it closes as we ingest more law. The other sub-zero case (HCA-20, −1.5) sits inside the ±3-point noise floor — a tie, not a loss. Leave the Hong Kong matter out and the paired lead rises to +7.1.

Statistical detail

Measurevs Fable 5vs GPT-5.6 Sol
Paired mean margin+6.05 (sd 5.99)+10.07 (sd 6.68)
95% confidence interval+1.76 to +10.34+5.29 to +14.85
Paired t-testt(9) = 3.19 · p = 0.011t(9) = 4.76 · p = 0.001
Sign test (win count alone)8 of 10 · p = 0.1099 of 10 · p = 0.021
Exact Wilcoxon signed-rankp = 0.027p = 0.004
At the ±8-point noise floor5 wins · 0 losses · 5 inside the band7 wins · 0 losses · 3 inside the band

The three significance tests get stricter as they discard information. The t-test weighs the size of every margin; the Wilcoxon keeps only their rankings; the sign test keeps nothing but the win count — and on a bare count of ten, 8 of 10 is not statistically significant (p = 0.109). We publish it anyway. The last row reads the set at the wide end of the instrument's own ±38 run-to-run noise floor: taken that conservatively against Fable 5, the ten matters are 5 wins, 0 losses and 5 too close to call.

The paired design is what makes a ten-case set informative: both systems answer the same question, so the difficulty of the case cancels and only the difference between the systems is left. It does not make ten cases into a large sample — see the limitations below.

And here is the bigger half

We're not even testing our biggest advantage.

That 10% is only the reasoning. The grading panel rewards citing the right authority. What it does not reward is how much of that authority you can check — the density of pinpoints that lets a barrister verify a memo in seconds. That is the thing our system is built to give you.

What the 10% lead does, and does not, include:

In the score+6.05

The legal reasoning — the issues, the law, the conclusion. This is the whole of the lead.

Not in the score
30.6
1.6
pinpoint citations a memo — the advantage a barrister feels

A barrister can open and verify every one of ours in seconds. The grader does not reward that density — so the 10% lead understates the gap you would feel in practice.

So the thing a barrister would actually feel — whether checking an authority takes ten seconds or ten minutes — barely moves this score. That is good for us two ways. The lead is a genuine difference in legal reasoning, not a citation-counting artefact: the grader hands out no points for pinpoint density. And it means the benchmark understates the product, because the advantage we are proudest of is mostly invisible to the number.

The advantage the score misses

More than ten times the authority — all of it checkable.

A pinpoint is the at [47] in Smith v Jones at [47] — the exact paragraph you open to check a claim. It is the difference between a citation you can rely on and one you have to re-research. Across ten memos we gave 306 of them. The model gave 16.
Fig. 2

Paragraph-level citations, per memo

19×more paragraph-level citations, per memo
Barrister AI30.6
Fable 5, cold1.6

And 468 authorities cited across ten memos to the model's 40 — nearly 12× as many, every one you can open and check.

Drawn to scale. A barrister opening one of our memos finds, on average, 30.6 places where the memo says exactly which paragraph of which judgment to read — against 1.6 from the model answering cold. That is the difference between an answer you can check in seconds and one you have to rebuild.
Fig. 3

All 306 of our pinpoint citations, checked against the judgment

  • 271 machine-checked against the judgment, and every one resolved
  • 35 the check declined (ambiguous or self-referential)
One mark per citation, 306 in total — six rows of fifty-one, so the count can be read off the image. The solid run is the 271 an automatic check resolved against the judgment text: every one matched the paragraph cited, with zero errors. The pale tail is the 35 the check declined — a case name shared across judgments, paragraph numbering that restarts per judge, or the memo's own cross-references. Not shown wrong; just not machine-resolvable.
Citation behaviour across the ten memos produced by each system.
MeasureBarrister AIFable 5, cold
Authority citations, ten memos46840
Of those, carrying a pinpoint306 (65%)16 (40%)
Machine-checkable against the judgment, and valid271 of 271, or 100%6 of 6, or 100%
Invented cases00
Every pinpoint we give you, you can open and check.

Of our 306 pinpoints, an automatic check resolved 271 to the exact paragraph, with no fabricated or misattributed authority. The other 35 it declined for ambiguity — a case name shared across judgments, or numbering that restarts per judge — or because they were the memo's own cross-references, not a fresh authority. None were shown wrong.

In fairness to the comparator: of the 16 pinpoints Fable did give, the 6 we could machine-check were all valid — 6 of 6, the same 100% we score. We are not claiming we cite more accurately. The difference is not who cites correctly. It is how much of the answer came with a paragraph to check at all — 306 against 16.

The full set

Every case, and every score.

Ten cases went in and ten cases are reported.
Every score behind the two comparisons, all answering cold on the same ten matters and the same panel. Our column is highest on eight of ten against Fable and 9 of ten against Sol. Matters are de-identified so the held-out set can be reused.
CaseBarrister AIFable 5GPT-5.6 Sol
HCA-1470.557.261.2
HCA-873.562.367.2
HCA-250.940.442.6
HCA-1268.357.857.6
HCA-572.662.654.9
HCA-1869.662.846.1
HCA-1363.861.852.7
HCA-2166.164.966.9
HCA-2060.161.650.5
HCA-171.074.566.0

Limitations

What a benchmark can and cannot tell you.

Benchmarks in this field are young, and most of them are built by the people whose systems they measure. Ours is no exception. Here is what that means for how much weight this result will bear.
  • A benchmark is a proxy, not a practice

    No scored exercise reproduces the conditions of real work: an instructing solicitor, a brief with gaps in it, a hearing date. What a benchmark can show is whether a system does the checkable parts well. Treat any legal-AI benchmark, ours included, as evidence about a capability rather than a guarantee about an outcome.

  • This is our own benchmark, not a third party's

    We built it to show, as transparently and objectively as we can, the quality of our product. We are open about that. And we maintain that the best test of all is the simplest one: use it on a real question and see the difference for yourself.

  • Ten cases is a strong signal, not the last word

    Ten matters, each answered once. The paired design cancels out how hard each case is, which is what makes ten enough to be meaningful — the margin holds up under a proper significance test. But it is still ten, and a few of the margins are close enough to read as level rather than wins; the chart marks those honestly. A larger set will sharpen the picture, not reverse it.

  • The comparison is deliberately asymmetric

    Both comparators answered without access to source law, because that is what they do when a practitioner uses them directly. This measures what grounding is worth. It is not a claim that one underlying model reasons better than another.

  • Our coverage is Australian, and growing

    Our corpus is Australian, so where an answer turns on foreign authority the advantage narrows — one case here turned on Hong Kong judgments we do not yet hold. That is a function of how much law we have ingested, not of how the system reasons, and coverage expands every week.

Barrister AI remains an AI-assisted research tool. Practitioners should read the cited sources and exercise independent professional judgment.

The point

Better memos — and ones you can actually check.

A barrister cannot outsource judgment to a black box. Being able to follow the reasoning and confirm the authority is a professional obligation, not a nice-to-have — so an answer you cannot check is not just less convenient, it is unusable for the work that matters.

That is where the gap is widest. Across ten memos the frontier model gave 40 authorities and 16 paragraph references — an answer you would have to rebuild before you could rely on it. Our system, put the same questions, gave 468 authorities and 306 pinpoints you can open and check — and scored higher on the reasoning as well.

So this is the short version: we write better legal memos, and we show our working. The score says the reasoning is stronger; the citations let you prove it for yourself. That second part is the one a lawyer cannot ethically skip, and it is the one our system is built to make easy.

Methodology

The best test is the one you run yourself.

Put a real question to it and see the difference. We share our methodology with reviewers on request — get in touch and tell us what you need to assess.
Research Benchmark — Barrister AI