← All guides

Tooling · 12 min read · Updated July 2026

How to evaluate and choose legal AI tools

Most legal AI evaluations test the wrong thing. Here is what to actually test, how to run a trial that produces a decision, and the questions vendors would rather you skipped.

Written for: Managing partners, legal ops, general counsel, anyone running a tool evaluation.

Key takeaways

  • Choose the action before you choose the tool. A tool evaluated against a named, high-frequency action produces a decision; one evaluated against 'legal work' produces a demo.
  • Test on your own real matters, including the awkward ones, never on the vendor's curated examples.
  • Data handling questions should be answered in writing before a trial starts, not after a preference has formed.
  • Exit cost is a buying criterion. If your process and data cannot leave, you have bought a dependency rather than a capability.
  • The best tool for an unsystematised practice is no tool. Sequence matters more than selection.

Before you evaluate anything

The single most consequential decision in a legal AI purchase happens before any vendor is contacted: which named action are you trying to automate? A firm that has done its workflow documentation can answer precisely - 'first-draft production on our commercial lease reviews, roughly forty a month, currently four hours each'. A firm that has not will describe a desire rather than a requirement, and will end up evaluating products against each other rather than against a job.

This matters because legal AI products are not general-purpose in practice even when they claim to be. A tool that is excellent at contract review may be mediocre at litigation drafting and useless at research-stock retrieval. Without a named action, evaluation degenerates into comparing feature lists and demo polish, which correlate weakly with whether the thing will work on your matters.

The prerequisite

If you cannot state the action, its monthly volume, and its current duration, stop the evaluation and go do the workflow documentation first. It takes a fortnight and it will change what you buy.

Running a trial that produces a decision

Most trials fail to produce a decision because they were never designed to. A licence is issued to a handful of people with an instruction to 'have a play', usage is sporadic, opinions at the end are impressionistic, and the decision defaults to whoever is most enthusiastic. Design the trial to answer a question instead.

  1. 01

    Fix the scope to one action and one matter type

    Not the whole practice. One named action, on one matter type, for a defined period. Breadth in a trial guarantees shallow evidence on everything.

  2. 02

    Assemble a test set of ten real matters

    Seven typical, three awkward - the ones with poor document quality, unusual structure, or a fact pattern that caused trouble. Vendors demonstrate on clean inputs; your practice does not run on clean inputs.

  3. 03

    Define the pass mark before you start

    In writing. For example: 'a usable first draft requiring under thirty minutes of revision on at least seven of ten matters, with zero unanchored assertions surviving preflight'. Setting this after seeing results is how enthusiasm becomes a purchase order.

  4. 04

    Have two people run the same matters independently

    Tool performance varies enormously with how it is used. If one person gets good results and another does not, you have learned something important about the training and process burden, not just the product.

  5. 05

    Time everything, including the checking

    The relevant figure is total time from start to shippable output, including review and correction. A tool that produces a draft in ninety seconds and requires two hours of verification has not saved time.

  6. 06

    Write the decision down against the pass mark

    One page: what was tested, against what standard, what happened, and the decision. This document is worth keeping - in a year you will want to know why you chose what you chose, and re-evaluations start from it.

The questions to get answered in writing

Ask these before the trial, in writing, and treat evasion as an answer. A vendor who cannot give a clear response on data handling to a law firm has told you something about their readiness for legal customers.

  • Is our data used to train your models, or any third party's models? A qualified yes is a no for most legal work.
  • Where is data processed and stored, in which jurisdictions, and can that be constrained contractually?
  • What is the retention period, and can we require deletion on a defined schedule and on exit?
  • Which sub-processors touch our data, and how are we notified when that list changes?
  • What happens to our configuration, prompts, playbooks, and outputs if we leave? In what format, and at what cost?
  • What is the incident-notification commitment, and has there been a security incident in the last twenty-four months?
  • Is there a contractual position on confidentiality and privilege that your legal team has reviewed for legal-sector customers specifically?

Build, buy, or keep what you have

The default assumption in most evaluations is that the answer is a purchase. It frequently is not. Three options are genuinely on the table and they have different profiles.

Choosing between the three real options
OptionBest whenWatch for
Keep and configure what you ownThe capability exists in a system already approved and deployed but was never configured properly.Underestimating configuration effort; assuming a product failure where there was a setup failure.
Buy market softwareThe action is common across firms, the volume justifies a licence, and a mature product exists.Exit cost, data-handling terms, and the tool dictating your process rather than serving it.
Build customA genuine gap exists that nothing on the market covers, and the action is high-volume and specific to your practice.Maintenance burden. Custom software you cannot maintain is a liability that arrives eighteen months later.

Treat exit cost as a buying criterion

The question to ask about any tool is not only how well it works but how expensive it would be to stop using. A system holding your matter data, your action library, your configured playbooks, and your accumulated duration history, with no clean export, has converted your operating model into its own asset.

This is the practical meaning of the principle that a practice should own its primitives. Your matter definition, your action library, your role chart, your evidence rule, and your research stock are the firm's assets. Tools are how they run today. If a tool cannot be swapped without rebuilding the operating model, the dependency has been inverted, and the price of that will be paid at renewal.

Frequently asked

  • How do I evaluate legal AI software?

    Name the action first - what it is, its monthly volume, and its current duration. Then trial one tool against one named action on ten of your own real matters, seven typical and three awkward, with a written pass mark set before you start and total time-to-shippable measured including review. Evaluations run against 'legal work' in general produce demos rather than decisions.

  • What questions should I ask a legal AI vendor?

    Whether your data trains their or anyone's models; where it is processed and stored and whether that can be contractually constrained; retention and deletion terms; the sub-processor list and change notification; what you can export on exit and in what format; incident-notification commitments and recent incident history; and whether there is a reviewed contractual position on confidentiality and privilege for legal customers. Get answers in writing before the trial.

  • Should a law firm build its own AI tools?

    Only where a genuine gap exists that nothing on the market covers and the action is high-volume and specific to your practice. The dominant risk is maintenance - custom software without a maintenance plan becomes a liability roughly eighteen months in. For most firms most of the time, configuring what you already own or buying a mature product is the better decision.

  • How long should a legal AI trial run?

    Long enough to process your ten-matter test set with two independent users, which is usually two to four weeks. Longer trials rarely produce better evidence; they produce familiarity, which people mistake for evidence. What determines trial quality is the test-set design and the pre-agreed pass mark, not duration.

  • What is the most common mistake when buying legal AI?

    Buying before the process is documented. A tool pointed at an unnamed, unsystematised action will underperform regardless of its quality, and the firm will conclude the technology does not work. The second most common mistake is ignoring exit cost, which converts a capability purchase into a dependency you discover at renewal.