How well AI models do the work lawyers actually do.
Lawve-bench-1 scores AI models on in-house legal tasks drawn from real practice: the drafting, review, research and multi-step matter work legal teams
handle every day. Each model gets a Lawve Index, a single score out of 100. The higher it is, the better the model handled the work.
Best right now:Claude Opus 5.5max leads with a Lawve Index of 63.
Effort, score and cost
Most AI models let you choose how hard they think before answering, from low to max effort. More effort usually gives better answers, but each task then costs
more and takes longer. In the chart below, each line is one model and each dot is one effort setting. Higher means a better score; further right means cheaper. A line that sits high and far to the right gives strong results without a large bill.
Take Claude Opus 5.5low: it scores 52 for $0.27 a task. At max effort the same model
reaches 63, but each task costs $4.11.
Overview
Lawve Index by model and effort level
Getting the most for your money
The next chart shows every model and effort setting we tested. The dashed line is the : the highest score you can get at each price. A dot below the line is outscored by a cheaper or equally priced option, so it is rarely the best pick. The cost axis grows tenfold at each
step, which keeps very cheap and very expensive models readable on the same chart.
Cost
Lawve Index vs. cost per task
Anthropic
OpenAI
xAI
Google, Meta, Zhipu
Pareto frontier
Getting answers quickly
Speed matters when someone is waiting on the answer. This chart works the same way, with the average time a model takes per task instead of its cost. Dots on the dashed line are the fastest
way to reach each score. Further left is faster. GPT-6.1 Sollow, for example, reaches 50 in about 30s a task.
Speed
Lawve Index vs. time per task
Anthropic
OpenAI
xAI
Google, Meta, Zhipu
Pareto frontier
What the results mean
Which AI is best for lawyers?
Claude Opus 5.5max, if budget allows. It leads Lawve-bench-1 with a Lawve Index of 63 at max effort. The right choice also depends on budget and turnaround: GPT-6.1 Solxhigh reaches 57 to 58 for under $0.50 a task, and GPT-6 Astramax posts the strongest Non-Hallucination scores in the top ten.
Is there an AI benchmark for legal work?
Yes, this one. Lawve-bench-1 measures how AI models perform on real legal tasks across six areas: legal knowledge, agentic knowledge work, reasoning, long-context, non-hallucination and agentic tool use. The tasks
are built in-house and kept private, so models cannot be trained on them.
Which model has the highest Lawve Index?
Claude Opus 5.5max, with 63. It is followed by Claude Opus 5.5xhigh at 61. GPT-6 Astramax and Claude Opus 5.5high share third place at 59. The Claude runs listed here use fallback.
Why aren't the tasks public?
To keep the scores honest. Public test sets end up in training data and reward models tuned to the test rather than to the work. Keeping ours private means a score reflects the work, not memorization.
How should I read the Lawve Index?
As a starting point, not a verdict. It is a single score out of 100 that summarizes the six areas; higher is better. A gap of a point or two rarely matters in practice. Weigh it against cost and
time per task, and always review AI output before relying on it.