Research benchmark · held-out set · 2026-07-28
We beat the leading frontier models on HCA legal questions.
Here is exactly what we did. We put ten real High Court questions to our system and to two of the leading frontier models. None was shown the Court's judgment — ours drew only on Australian law as it stood before the decision came down. Then we scored every answer against what the Court actually held. We came out ahead of both — about 10% higher than one and 18% higher than the other — and backed our reasoning with 306 paragraph-level citations you can open and check, against Fable 5's 16.
- Instrument
- Held-out set of 10 High Court matters
- Graded
- all three arms, one panel
- Scale
- 0–100, panel median
- Runs
- 1 per system per case
What we tested
What a good answer has to do.
We scored each memo the way a barrister judges a junior's note — starting with whether it found the right issues, and ending with whether every authority it leans on can actually be checked.- 01
Did it identify the right issues?
The first test of any research memo is whether it spots the questions that actually decide the case — in the right jurisdiction, and the right body of law. A confident answer to the wrong question is worse than no answer at all.
- 02
Is the reasoning sound?
Does it apply the law correctly and reach a conclusion a court could accept? This is the substance a barrister is paying for, and it is the larger part of the score.
- 03
Can the authority be checked?
A citation a barrister cannot verify is work they have to redo. We score whether the memo says which case, and where in it — the paragraph, not just the report.
- 04
Were there errors?
Invented cases, misattributed holdings and superseded authority are counted separately from quality, because they carry a different professional cost.
How it was run
Same questions. Same panel. Same day.
We are not publishing the question bank or the scoring rubric — releasing them would let anyone tune against them. What follows is the protocol, which is the part that matters for judging the result.- 01
Identical questions
All three systems answered the same ten questions, drawn from ten Australian High Court matters. None saw the others' work.
- 02
One panel, applied to all three
Every memo was scored by a panel of three independent models from three different AI labs, aggregated by median so no single grader can carry a result.
- 03
Case citations machine-checked
Each pinpoint reference to a judgment was checked against the text of that judgment, and counts as verified only where the paragraph was actually found there. This check covers case law: references to legislation were not verified this way, so the citation figures describe the case-law half of each memo.
- 04
Nothing tuned after the fact
The scoring was run once, on memos already written. No configuration was changed, and no case was dropped, after any score was seen.
The result
Ahead of both.
Two leading frontier models, both answering cold, on the same ten matters. We came out ahead of each — comfortably against one, and by a wider margin still against the other.- vs Fable 5, cold+10%
+6.05 points · won 8 of the ten · 66.6 to 60.6
- vs GPT-5.6 Sol, cold+18%
+10.07 points · won 9 of the ten · 66.6 to 56.6
And these are not weak opponents. On LegalBench, the academic benchmark for legal reasoning, Fable 5 ranked first and GPT-5.6 Sol fifth as of 23 July 2026. Leaderboard: vals.ai
Our lead over both models, case by case
vs Fable 5, cold
vs GPT-5.6 Sol, cold
The one real loss, explained. Our only genuine loss — HCA-1 — turned on two Hong Kong judgments: authority a frontier model carries from training, but our Australian corpus does not yet hold. It is a coverage gap, not a reasoning gap, and it closes as we ingest more law. The other sub-zero case (HCA-20, −1.5) sits inside the ±3-point noise floor — a tie, not a loss. Leave the Hong Kong matter out and the paired lead rises to +7.1.
Statistical detail
| Measure | vs Fable 5 | vs GPT-5.6 Sol |
|---|---|---|
| Paired mean margin | +6.05 (sd 5.99) | +10.07 (sd 6.68) |
| 95% confidence interval | +1.76 to +10.34 | +5.29 to +14.85 |
| Paired t-test | t(9) = 3.19 · p = 0.011 | t(9) = 4.76 · p = 0.001 |
| Sign test (win count alone) | 8 of 10 · p = 0.109 | 9 of 10 · p = 0.021 |
| Exact Wilcoxon signed-rank | p = 0.027 | p = 0.004 |
| At the ±8-point noise floor | 5 wins · 0 losses · 5 inside the band | 7 wins · 0 losses · 3 inside the band |
The three significance tests get stricter as they discard information. The t-test weighs the size of every margin; the Wilcoxon keeps only their rankings; the sign test keeps nothing but the win count — and on a bare count of ten, 8 of 10 is not statistically significant (p = 0.109). We publish it anyway. The last row reads the set at the wide end of the instrument's own ±3–8 run-to-run noise floor: taken that conservatively against Fable 5, the ten matters are 5 wins, 0 losses and 5 too close to call.
The paired design is what makes a ten-case set informative: both systems answer the same question, so the difficulty of the case cancels and only the difference between the systems is left. It does not make ten cases into a large sample — see the limitations below.
And here is the bigger half
We're not even testing our biggest advantage.
That 10% is only the reasoning. The grading panel rewards citing the right authority. What it does not reward is how much of that authority you can check — the density of pinpoints that lets a barrister verify a memo in seconds. That is the thing our system is built to give you.What the 10% lead does, and does not, include:
The legal reasoning — the issues, the law, the conclusion. This is the whole of the lead.
A barrister can open and verify every one of ours in seconds. The grader does not reward that density — so the 10% lead understates the gap you would feel in practice.
So the thing a barrister would actually feel — whether checking an authority takes ten seconds or ten minutes — barely moves this score. That is good for us two ways. The lead is a genuine difference in legal reasoning, not a citation-counting artefact: the grader hands out no points for pinpoint density. And it means the benchmark understates the product, because the advantage we are proudest of is mostly invisible to the number.
The advantage the score misses
More than ten times the authority — all of it checkable.
A pinpoint is the at [47] in Smith v Jones at [47] — the exact paragraph you open to check a claim. It is the difference between a citation you can rely on and one you have to re-research. Across ten memos we gave 306 of them. The model gave 16.Paragraph-level citations, per memo
And 468 authorities cited across ten memos to the model's 40 — nearly 12× as many, every one you can open and check.
All 306 of our pinpoint citations, checked against the judgment
- 271 machine-checked against the judgment, and every one resolved
- 35 the check declined (ambiguous or self-referential)
| Measure | Barrister AI | Fable 5, cold |
|---|---|---|
| Authority citations, ten memos | 468 | 40 |
| Of those, carrying a pinpoint | 306 (65%) | 16 (40%) |
| Machine-checkable against the judgment, and valid | 271 of 271, or 100% | 6 of 6, or 100% |
| Invented cases | 0 | 0 |
Of our 306 pinpoints, an automatic check resolved 271 to the exact paragraph, with no fabricated or misattributed authority. The other 35 it declined for ambiguity — a case name shared across judgments, or numbering that restarts per judge — or because they were the memo's own cross-references, not a fresh authority. None were shown wrong.
In fairness to the comparator: of the 16 pinpoints Fable did give, the 6 we could machine-check were all valid — 6 of 6, the same 100% we score. We are not claiming we cite more accurately. The difference is not who cites correctly. It is how much of the answer came with a paragraph to check at all — 306 against 16.
The full set
Every case, and every score.
Ten cases went in and ten cases are reported.| Case | Barrister AI | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| HCA-14 | 70.5 | 57.2 | 61.2 |
| HCA-8 | 73.5 | 62.3 | 67.2 |
| HCA-2 | 50.9 | 40.4 | 42.6 |
| HCA-12 | 68.3 | 57.8 | 57.6 |
| HCA-5 | 72.6 | 62.6 | 54.9 |
| HCA-18 | 69.6 | 62.8 | 46.1 |
| HCA-13 | 63.8 | 61.8 | 52.7 |
| HCA-21 | 66.1 | 64.9 | 66.9 |
| HCA-20 | 60.1 | 61.6 | 50.5 |
| HCA-1 | 71.0 | 74.5 | 66.0 |
Limitations
What a benchmark can and cannot tell you.
Benchmarks in this field are young, and most of them are built by the people whose systems they measure. Ours is no exception. Here is what that means for how much weight this result will bear.- A benchmark is a proxy, not a practice
No scored exercise reproduces the conditions of real work: an instructing solicitor, a brief with gaps in it, a hearing date. What a benchmark can show is whether a system does the checkable parts well. Treat any legal-AI benchmark, ours included, as evidence about a capability rather than a guarantee about an outcome.
- This is our own benchmark, not a third party's
We built it to show, as transparently and objectively as we can, the quality of our product. We are open about that. And we maintain that the best test of all is the simplest one: use it on a real question and see the difference for yourself.
- Ten cases is a strong signal, not the last word
Ten matters, each answered once. The paired design cancels out how hard each case is, which is what makes ten enough to be meaningful — the margin holds up under a proper significance test. But it is still ten, and a few of the margins are close enough to read as level rather than wins; the chart marks those honestly. A larger set will sharpen the picture, not reverse it.
- The comparison is deliberately asymmetric
Both comparators answered without access to source law, because that is what they do when a practitioner uses them directly. This measures what grounding is worth. It is not a claim that one underlying model reasons better than another.
- Our coverage is Australian, and growing
Our corpus is Australian, so where an answer turns on foreign authority the advantage narrows — one case here turned on Hong Kong judgments we do not yet hold. That is a function of how much law we have ingested, not of how the system reasons, and coverage expands every week.
Barrister AI remains an AI-assisted research tool. Practitioners should read the cited sources and exercise independent professional judgment.
The point
Better memos — and ones you can actually check.
A barrister cannot outsource judgment to a black box. Being able to follow the reasoning and confirm the authority is a professional obligation, not a nice-to-have — so an answer you cannot check is not just less convenient, it is unusable for the work that matters.That is where the gap is widest. Across ten memos the frontier model gave 40 authorities and 16 paragraph references — an answer you would have to rebuild before you could rely on it. Our system, put the same questions, gave 468 authorities and 306 pinpoints you can open and check — and scored higher on the reasoning as well.
So this is the short version: we write better legal memos, and we show our working. The score says the reasoning is stronger; the citations let you prove it for yourself. That second part is the one a lawyer cannot ethically skip, and it is the one our system is built to make easy.
Methodology