Evaluating tools

How to evaluate legal AI: questions to ask before a pilot

Most evaluations of legal AI go wrong in the same way: the office watches a demo on the vendor’s file, likes it, and buys. The demo is chosen. Your files are not. Here is a shorter route.

Nine questions for the first meeting

1. What did you measure, on what set, and can I see it? Not a customer quote — a number, with the test set, the date it was run, and what counted as a failure. Researchers who tried to evaluate the leading legal research tools found the vendors provide no systematic access, publish few details about their models, and report no evaluation results at all [1]. If a vendor cannot produce a measurement, you are the measurement.

2. What happens when it cannot answer? The most important behaviour in the product. Does it return nothing and say why, or does it return its best guess in the same confident typeface as everything else? Ask for a live example of the system declining.

3. Is it deterministic? Run the same file twice, an hour apart, and compare the output byte for byte. If the two runs differ, you cannot reproduce a result for a judge, and you cannot tell a real change from sampling noise.

4. Can every sentence be traced to a page? Ask to click a line in the output and land on the source — the document, the page, the paragraph. Ask whether the trace is generated with the sentence or computed against the document afterwards. Only the second one is evidence.

5. Where does our data go, and who else can read it? Get this in writing, not in a slide: retention period, whether inputs or outputs train anyone’s model, which subprocessors touch it, and what happens on termination. Formal Opinion 512 treats this as a confidentiality question under Rule 1.6, and warns that boilerplate consent language added to engagement letters is not sufficient [3].

6. How long does verification take, and who does it? Ask the vendor to time it. The large firm Paul Weiss spent nearly a year and a half testing one product and did not develop hard metrics, because checking the system was so involved that it made any efficiency gains difficult to measure [1]. If your lawyers must re-verify every proposition, the hours you were buying are gone.

7. What can it do without a person in the loop, and what is it prevented from doing? California’s guidance is explicit: lawyers must not deploy agentic systems in a way that lets the system make substantive legal determinations, communicate legal advice, prepare and file pleadings, or otherwise act in a representative capacity without meaningful lawyer supervision [4]. A vendor should be able to say where its own boundary sits.

8. What do the judges in our courts require? As of May 2024 more than twenty-five federal judges had issued standing orders requiring attorneys to disclose or monitor AI use in their courtrooms [1], and the Chief Justice had already devoted part of a year-end report to briefs citing non-existent cases [5]. Check your division before the pilot, not after the filing.

9. What do we keep if we stop? The memos, the citations, the receipts, the extracted record. Get the export format in the agreement.

Two tests to run before you sign

The closed-file test. Take one case you have already finished, where you know the answer and the outcome. Hand over the discovery with no question attached and no coaching. Then compare the output against what actually happened. You are not measuring whether it sounds like a lawyer. You are measuring three things: what it found that your team found, what it missed, and what it asserted that is not in the file.

The adversarial test. Try to make it fail in front of you. Ask about a case that does not exist. Ask a question with a false premise — one system, when asked, agreed that Justice Ginsburg dissented in Obergefell and invented a copyright rationale for it [1]. Feed it a document that contradicts itself. A product that holds up under this in a sales meeting will hold up on a Friday afternoon.

How to read the answers

Two failure patterns are worth naming in advance, because both look like success in a demo.

The first is the tool that is accurate and unhelpful — it declines so often that the office stops opening it. The second is the tool that is fluent and unverifiable, which is worse, because the cost lands on whoever signs. The purpose-built legal research products measured by Stanford still hallucinated between 17 and 33 percent of the time — the study’s word for an answer that is either incorrect or cites a source that does not support it [2] — and they are the serious end of the market.

Formal Opinion 512 sets the floor either way: a lawyer who uses a generative tool without an appropriate degree of independent verification or review of its output may be violating the duty of competence [3]. The only product that changes your workload is one whose verification you do not have to repeat.

That is the standard worth holding a vendor to, and it is the one Apodicta is built against: every rendered line already checked against its source, and everything that failed listed by name instead of dropped.

Sources

  1. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking QueriesStanford Institute for Human-Centered AI · 23 May 2024, updated 30 May 2024
  2. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh, Surani, Dahl, Suzgun, Manning & Ho (Stanford RegLab / HAI), arXiv:2405.20362 · 30 May 2024
  3. Formal Opinion 512: Generative Artificial Intelligence ToolsABA Standing Committee on Ethics and Professional Responsibility · 29 July 2024
  4. Practical Guidance for the Use of Generative Artificial Intelligence in the Practice of LawState Bar of California, Standing Committee on Professional Responsibility and Conduct · 2026
  5. 2023 Year-End Report on the Federal JudiciaryChief Justice John G. Roberts, Jr., Supreme Court of the United States · 31 December 2023

More: all articles · questions defenders ask · the numbers, with their receipts

See the gate on your own record.

Show me on my files