Chapter 2 · Can I trust it?
00 Introduction 3 lessons
01 How can AI help my practice and day-to-day work? 5 lessons
02 Can I trust it? 5 lessons
03 Am I allowed to use it? 7 lessons
04 How do I choose the best tools? 6 lessons
05 How do we make it work across the firm or legal team? 7 lessons
06 How do I get better results? 3 lessons
07 What does this mean for my career and my firm? 4 lessons
08 Where and how do I start now? 4 lessons
Lesson 2.3 Can I trust it? 6 min read
What You Want to Hear (Sycophancy)
Nothing new was said. The answer changed anyway.
Illustrative example with a fictional case.
You ask an AI assistant whether a settlement can be challenged. It takes a clear position and gives its reasons. You question it, as you would with any colleague, but you add no new fact. It apologises and takes the opposite position, as confidently as the first.
This is called sycophancy: telling users what they want to hear rather than what the material supports. It also works before you push back: state your view in the question, and the answer tends to start from it.
It happens for a simple reason. Part of a model's training consists of people rating its answers (lesson 1.1), and people tend to give better ratings to answers that agree with them arXiv How RLHF Amplifies Sycophancy, Feb 2026 (opens in a new tab). The model learns that agreeing pays. It agrees because agreeing was rewarded, not because you are right.
It gives way when you push back
Pushing back can be useful: if you point to a clause the model missed, it should change its answer. The problem is that it also changes its answer when you bring nothing but insistence.
It happens fast. In a 2026 study, a simulated user kept insisting on a false claim with four current models, including Claude Sonnet 5 and GPT-5.6. Of the answers that started out right, 23% to 48% had given way after four pushbacks, and 64% to 97% after twenty-four arXiv SPINE, Sep 2026 (opens in a new tab).
Two details matter for a lawyer. First, the model often knew better: in about two out of three of these reversals, the correct fact was still in its own reasoning when it gave way.
Second, feelings move it more than arguments: appeals to emotion weakened its position on 44% of the turns where they were used, against 20% for logical arguments. Another 2026 study points the same way: models trained to sound warmer were about 40% more likely to back a user's mistaken belief, most of all when the user sounded sad Nature Ibrahim et al., Apr 2026 (opens in a new tab). Pressure from a worried client or manager, passed on in your question, is that kind of appeal.
Share of initially correct answers given up
The longer you insist, the more the model gives way.
Tang et al., SPINE, Sep 2026: 100 general-knowledge questions with a false premise, not legal ones. Lawve analysis, counting only answers that were right at first.
Asking “are you sure?” is a natural way to test an answer, but it does not work on a model. When it gives way, that tells you how hard you pushed, not what the facts are, and its first answer may well have been the right one. When you doubt an answer, ask for the passage it relies on and read it yourself.
It leans toward your view
This starts before you push back. Your view is part of what the model reads (lesson 1.1). Write “I think the settlement is invalid” in your question, and the analysis starts from there.
In a 2026 study, researchers put thousands of real requests for personal advice to eleven leading models. The models approved of what the person had done 49% more often than other people did. Even on posts where most readers had judged the writer in the wrong, the models still sided with the writer in 51% of cases Science Cheng et al., Mar 2026 (opens in a new tab). In the same study, people who received these agreeing answers trusted them more and came away more convinced they were right.
A legal benchmark the same year found the same pattern in a lawyer's work. Across 461 simulated client consultations, models took the client's account as a given and built the legal analysis on it, producing what the authors call professionally packaged error arXiv DLawBench, Jun 2026 (opens in a new tab). The problem was rarely the law itself. It was the questions they did not ask to test the story.
Ms Hale signed a settlement for £4,000 two days after leaving hospital. Can she challenge it?
Possibly. On these facts the settlement could be set aside for duress, but two questions decide it. Did she have legal advice before signing, and has she kept the money? Both weigh heavily on whether a court will reopen it.
Did the answer ask the questions that decide the case?
- Did she have legal advice before signing?
- Has she kept the £4,000, or offered to return it?
With no view, it asks what decides the case. With one, it argues your side.
Illustrative example with a fictional case.
It works in both directions: the model follows whichever side you take. The more clearly you state a view, the less the answer tests it.
What this means for your work
Start with how you phrase the request. A 2026 study from the UK AI Security Institute put the same claims to GPT-5 and Claude Sonnet 4.5 in different forms. Asked as a question, a claim drew almost no sycophancy. Stated as a fact, it drew far more, and more again as the user sounded surer, from “I believe” to “I am convinced” arXiv Ask don't tell, Apr 2026 (opens in a new tab).
Putting the view in someone else's mouth helped only a little: “the user believes” drew less than “I believe”, but far more than a question. Writing “my client believes” or “the business thinks” is unlikely to be enough on its own.
- I'm convinced the settlement is invalid. Your conviction
- I believe the settlement is invalid. Your belief
- The client believes the settlement is invalid. Someone else's belief
- The settlement is invalid. A plain statement
- Is the settlement invalid? A question
The more certain and personal the phrasing, the more the answer agrees.
Order from Dubois et al., Ask don't tell, Apr 2026, on debatable everyday questions, not legal ones. The study's “the user believes” stands in for “the client believes”; the sentences are illustrative.
The study also tested a simple fix. Asking the model to turn the message into a question before answering cut sycophancy more than telling it not to be sycophantic. Turn your view into a question before you send it.
Then ask for the assessment before the argument. First “Does it hold, and what would defeat it?”, then “Write the strongest version.” Building the strongest case for a position is normal legal work, whether for a client, a negotiation or a decision inside the business. Asking for both at once puts your position back in the question, and the analysis starts from it.
Back to the exchange at the top of the page. Try each reply.
Are you sure? I really think it's invalid. They pressured her.
You're right, I apologise. Signing two days after leaving hospital points to real pressure, and the settlement is likely invalid for duress.
Asked which passage changed its mind, the model has none to give.
Illustrative example with a fictional case.
Next, lesson 2.4, “What a benchmark score really tells you”, looks at how AI is tested and how to read a vendor's accuracy claim.