Official launch partners
Chapter 2 · Can I trust it?
00 Introduction 3 lessons
  1. 0.1 Who this guide is for and how to read it
  2. 0.2 The eight questions every lawyer asks
  3. 0.3 Survival glossary
01 How can AI help my practice and day-to-day work? 5 lessons
  1. 1.1 What is AI really?
  2. 1.2 Chat vs. agent
  3. 1.3 Where skills, MCP and plugins fit in
  4. 1.4 What lawyers are actually using AI for right now
  5. 1.5 Does it really save time?
02 Can I trust it? 5 lessons
  1. 2.1 Hallucinations
  2. 2.2 What is the answer based on?
  3. 2.3 When the model tells you what you want to hear
  4. 2.4 What a benchmark score really tells you
  5. 2.5 Test it on your own matters
03 Am I allowed to use it? 7 lessons
  1. 3.1 The three questions behind the question
  2. 3.2 What happens to the documents I upload?
  3. 3.3 Who can access, store or reuse my data?
  4. 3.4 Anonymisation and pseudonymisation
  5. 3.5 Professional rules by jurisdiction
  6. 3.6 Sovereignty and compliance
  7. 3.7 Other risks
04 How do I choose the best tools? 6 lessons
  1. 4.1 Same brain, different bodies
  2. 4.2 Subscription vs. API access
  3. 4.3 Legal AI tool or general-purpose assistant?
  4. 4.4 Open-source vs. closed, from the buyer's seat
  5. 4.5 Panorama of tools
  6. 4.6 Questions to ask before choosing a legal AI tool
05 How do we make it work across the firm or legal team? 7 lessons
  1. 5.1 Why one enthusiast is not an adoption strategy
  2. 5.2 Choosing a first pilot
  3. 5.3 Training the team
  4. 5.4 Measuring time saved and quality
  5. 5.5 Who maintains the tools and the shared know-how?
  6. 5.6 Change management
  7. 5.7 Working with IT, security and procurement
06 How do I get better results? 3 lessons
  1. 6.1 Give it better context
  2. 6.2 Turning your methods into reusable instructions
  3. 6.3 Watch out for AI slop
07 What does this mean for my career and my firm? 4 lessons
  1. 7.1 Which skills should lawyers develop?
  2. 7.2 What does a legal engineer do?
  3. 7.3 How AI may change fees, staffing and client expectations
  4. 7.4 What an AI-native firm might look like
08 Where and how do I start now? 4 lessons
  1. 8.1 Starter kits by profile
  2. 8.2 Classic mistakes to avoid
  3. 8.3 Section 1 recap
  4. 8.4 Staying current without drowning

Lesson 2.2 Can I trust it? 7 min read

What the Answer Is Based On (Memory, Documents & Search)

DEFINITIONSSERVICESTERMFEESCLAUSE 12LIABILITYCONFIDENTIALSCHEDULE 1SCHEDULE 2

A client wants to end a supplier contract early. You ask an AI assistant whether the client can. With nothing else to go on, it answers from the general law: Yes, on reasonable notice. That is a fair rule for a contract that says nothing on the point. This one doesn't.

You attach supply-agreement.pdf and ask again. The answer flips: No, not before March 2027. Clause 12.1 allows thirty days' notice, but Schedule 2 bars it in the first year, and leaving early after that costs three months' fees.

Same model, same question. The only difference is what it had in front of it. What the model reads matters more than what it remembers, and you decide what it reads.

Three sources, one answer

An assistant can draw on three sources. The first is its memory: what the model learned in training. Ask it something with nothing attached, and that is all it has. It is enough to explain a rule, structure an argument or draft a clause. But it does not look anything up, and what it knows stops at a date (lesson 1.1).

The second is the documents you give it. Attach the client's contract or paste the other side's letter, and it answers from facts no model has ever seen. It can also point to the passage it relied on.

The third is a search. Turn it on, and it looks for web pages or judgments, reads some of them and answers from what it found. Legal research tools do the same in their own database of case law and legislation.

The three add up. When the model reads your contract, it uses what it learned to understand what a termination clause does. Your documents bring the facts of the matter. A search finds the law and the authorities to cite.

Switch a source off to see what the answer loses.

Can the client end the supply agreement early?

A contractual right to terminate is applied as written, so the wording of the agreement decides. Clause 12.1 allows thirty days' written notice, but Schedule 2 bars that notice before 1 March 2027. After that, leaving during the Initial Term costs three months' fees. A fee for exercising a right to leave is the price of that right, not a sum due on breach, so the rule against penalties is unlikely to apply (Cavendish v Makdessi [2015] UKSC 67).

supply-agreement.pdfSearched · 2 sources
  • Memory brings how the law works, and never switches off.
  • Your documents brings the facts of this contract.
  • Search brings the authority to cite.

The best answer draws on all three sources.

Illustrative answer about a fictional contract. Cavendish v Makdessi is a real UK Supreme Court decision on penalty clauses.

Adding sources makes a large difference. On legal research questions written by lawyers, the best model got 8.7% of answers from memory alone, and 42.9% once it could search case law and the web arXiv Legal Research Bench, Sep 2026 (opens in a new tab). Every model did better with tools than any model without them. For anything you will rely on, don't leave the model with only its memory.

Answering alone With tools to search case law and the web

Claude Opus 4.8

GPT-5.5

Claude Sonnet 4.6

GLM-5.2

Gemini 3.5 Flash

MiniMax-M3

GLM-5.1

Qwen3.7-Max

DeepSeek-V4-Pro

Gemini 3.1 Pro Preview

Kimi K2.6

Grok 4.3

GPT-5.4-mini

→ Answers fully correct (higher is better)

Every model answers far better when it can search for and read the law.

Vals AI, Legal Research Bench, Sep 2026: 413 US legal research questions written by lawyers.

From invention to omission

A source does not make the model error-free. It changes the error. Without one, the model may invent (lesson 2.1). With a source, it may leave out what matters.

In a 2026 study, 12 models summarised the full files of 100 US civil rights cases. Compared with the summaries lawyers had written, their main failure was leaving points out, not inventing them: even the best model covered fewer than 6 in 10 of the points the lawyers included arXiv Gavel, EMNLP 2026 (opens in a new tab).

Covered in the generated case summary Left out
  • Gemini 2.5 Flash
  • GPT-5
  • Gemini 2.5 Pro
  • Claude Sonnet 4
  • Claude Opus 4.1
  • GPT-4.1

No model came close to the lawyers' summaries.

Dou, Mamut and Xu, Gavel, EMNLP 2026: 100 US civil rights case files; the six proprietary 2025 models tested are shown, and the longest files were cut to fit some models. Lawve analysis.

Omission is harder to catch than invention. You catch an invented case by looking it up. A missing condition leaves nothing to look up: an answer built on clause 12.1 gives no hint that Schedule 2 exists. You only find it by reading the source. When the answer has a source, check what is missing, not only what is there.

Attaching is not reading

Attaching a file feels like handing it to a colleague who will read it. With a short file, that is roughly what happens. With a long one, or a whole bundle, it rarely is. Most assistants now work as agents: they have search tools, and they use them on your documents as they would on the web. They search the passages that seem relevant, read those, and answer from them.

There are good reasons for this. A model only works from what fits in its , and the largest hold about 1,500 pages (lesson 1.1). A data room of a few hundred contracts holds far more.

Even when everything fits, reading it all has a price. The model processes every word you send on powerful computers that someone pays to run, so more text means a longer wait and a higher bill.

Results also get worse as the input grows. In a 2026 test, recent frontier models missed a single risky step 2 to 30 times more often when it was buried after about 1,200 pages of routine log than when it stood on its own arXiv Classifier Context Rot, May 2026 (opens in a new tab). So agents are built and instructed to search first.

When the search misses, the passage never comes back, and the model answers from an incomplete file without knowing it. Whether the right passage comes back depends on the search tool, and the one most likely to stay behind is one that matters but looks unlike your question. Here is what that does to the client's contract.

Can the Customer end this agreement in November 2026?

Yes. Under clause 12.1, either party can end the agreement by giving at least thirty days' written notice. The agreement started on 1 March 2026.

Wrong, and every sentence in it is true.

What the search tool passed to the model

  1. Clause 12.1 Either party may terminate this Agreement by giving the other at least thirty days' written notice.

  2. Clause 2.1 This Agreement starts on 1 March 2026 (the Start Date) and continues for 24 months (the Initial Term).

  3. Clause 15.1 Each party shall keep the other's Confidential Information confidential during this Agreement and for three years after it ends. (not passed to the model)

  4. Clause 12.2 Clause 12.1 is subject to Schedule 2. (not passed to the model)

  5. Schedule 2, para 2 A party that terminates under clause 12.1 during the Initial Term shall pay the other an Early Termination Fee equal to three months' fees. (not passed to the model)

  6. Schedule 2, para 1 No notice under clause 12.1 may be given before the first anniversary of the Start Date. (not passed to the model)

The model answers from what the search hands it.

Illustrative example with a fictional contract.

Reading everything is not a simple fix either. Researchers who tried both on the case files found a trade-off: searching was cheaper but found less, while reading every file in full found more but cost more and added errors of its own arXiv Gavel, EMNLP 2026 (opens in a new tab). The answer is only as complete as what the search brought back.

What this means for your work

You cannot see inside the search, but you can steer it. Start by choosing what the model reads: the whole contract for a question about the contract, and a search in a legal database for the law, preferring the text of the law or the judgment to a commentary. For a large bundle, give it the documents that matter, not everything you have (lesson 1.1).

Then make your question do the work. Name the parts that matter, such as schedules, definitions and amendments, so the search brings them back even if they look nothing like your question. Ask for the passages it relied on, so you can see what came back. And before relying on the answer, ask what else bears on the question. Name what matters, and ask to see what came back. Try each habit on the client's question.

Add to the question:

Can the client end this agreement in November 2026?

Yes. Under clause 12.1, either party can end the agreement on thirty days' written notice.

Wrong, and nothing in it tells you so.

A sharper question brings back the passages that matter.

Illustrative example with the same fictional contract.

Next, lesson 2.3, "When the model tells you what you want to hear", looks at why a model can sound sure of an answer, then drop it as soon as you push back.