Key takeaways
- AI governance is a delegation problem. Firms cannot safely delegate to machines because they never had an explicit delegation structure for humans either.
- Every automated action needs a named doer, a named human reviewer of record, and a defined standard for what review means.
- A prohibition-only policy drives usage underground, which is strictly worse than governed usage. Assume adoption and govern it.
- Client consent and confidentiality positions should be written down once, centrally, rather than decided ad hoc by whoever is on the matter.
- The single question that unblocks most risk conversations: who is responsible for catching an error in this output, by name, before it ships?
Why AI governance is really a delegation problem
The reason most firms cannot safely hand work to an AI system is not the technology. It is that they never had an explicit delegation structure for human work either. Ask a typical practice who reviews a first-year associate's draft before it goes to a client, and the honest answer is 'whoever is supervising, usually, if there is time'. Responsibility is discovered retroactively, generally when something has gone wrong.
That arrangement survives with humans because humans compensate. A junior lawyer who is unsure asks. A supervisor who notices something odd intervenes. The informal system has a great deal of slack in it, and the slack is invisible until you remove the humans who were providing it.
Hand a named action to an agent inside that same informal structure and the slack is gone. The agent does not ask, does not hesitate, and does not flag its own uncertainty in a way anyone will notice. Every governance failure people attribute to AI is really this: an unstated delegation structure being asked to carry a load it was never designed for.
The unblocking question
For every automated output: who, by name, is responsible for catching an error in this before it ships? If there is no answer, you do not have a governance gap - you have an undeployed system that should stay undeployed.
The framework: four things per automated action
Governance does not need to be a document. It needs to be four attributes, recorded against each action you automate, and kept current. If you can produce this table for every deployed agent, you have a governance framework, whatever else you do or do not write down.
| Attribute | What it specifies | Failure mode if left blank |
|---|---|---|
| Doer | Which agent or system performs the action, and on which matter types. | Scope creep - the tool is used on work it was never assessed for. |
| Reviewer of record | The named human who reviews and takes responsibility before anything ships. | Diffused responsibility; errors surface at the client rather than internally. |
| Review standard | What review means concretely - anchors checked, facts verified, or full substantive rewrite. | Review degrades to a glance under time pressure, invisibly. |
| Escalation trigger | The conditions under which the action stops and goes to a human immediately. | Edge cases processed as routine, which is where the serious errors live. |
Defining what review actually means
'A lawyer reviews the output' is not a standard. It is a sentence that everyone agrees with and that means something different to each person who reads it. Under deadline pressure it silently collapses into a scan for obvious errors, which is exactly the review that misses confident, fluent, wrong output.
Define review at three levels and assign one to each automated action explicitly. Most practices find that the majority of automated actions sit at level two, and that naming the level is what prevents the quiet drift down to level one.
- 01
Level 1 - Verification review
Every anchor is opened and confirmed; every factual assertion is checked against its source. The reviewer is not re-forming the judgment, but is confirming the output is accurate. Appropriate for high-volume routine drafting where the structure is known good.
- 02
Level 2 - Substantive review
Verification plus an independent assessment of whether the output is the right answer, not merely an accurate one. The reviewer forms their own view and would have been comfortable producing this themselves. Appropriate for most client-facing first drafts.
- 03
Level 3 - Rewrite from source
The output is treated as a research aid only. The lawyer produces the work product themselves. Appropriate for novel questions, contested law, high-value advice, and anything where the reasoning must demonstrably be the lawyer's own.
Writing a policy people actually follow
The typical legal AI policy is a page of prohibitions written by someone with no operational stake, circulated once, and thereafter ignored. Its practical effect is to drive usage underground - lawyers use consumer AI tools on personal accounts and do not mention it, which is strictly worse than governed usage on approved systems.
Assume adoption. Your people are already using these tools or will be within months. The policy's job is to make the safe path the easy path, not to pretend the unsafe path can be legislated away.
- State which systems are approved, for which categories of work, in plain language with examples. Ambiguity here is what pushes people to unapproved tools.
- State what must never be pasted into an unapproved system, specifically and concretely - client-identifying material, privileged communications, unredacted personal data.
- State the review standard per action type, using the three levels above, not a general exhortation to review carefully.
- State the anchoring rule and confirm it applies identically to human and machine output.
- Name an owner for the policy and a review date. An unowned policy is a document, not a control.
- Include a no-blame reporting route for mistakes. You want to hear about the near-miss; a punitive posture guarantees you will not.
Clients, confidentiality, and consent
Decide the firm's position centrally and write it down, rather than leaving each matter team to work it out under pressure. The questions are finite and the answers should be consistent across the practice.
The core positions to settle: whether client material may be processed by third-party AI systems and under what contractual terms; whether outputs are used to train external models, which for most legal work should be a hard no; what the retention and deletion position is; whether and how clients are informed; and how privilege is preserved across the processing chain.
On disclosure, the trend across the profession is toward transparency, and the commercial reality is that it is increasingly a positive signal rather than a risk. A firm that can explain that AI runs a systematic first pass under a strict anchoring rule, with a named lawyer reviewing everything, is describing a more rigorous process than most competitors run - and clients, particularly sophisticated in-house buyers, increasingly recognise that.
Frequently asked
What should a law firm AI policy contain?
Approved systems by work category, an explicit list of what must never be entered into unapproved tools, a defined review standard per action type rather than a general instruction to review carefully, the evidence-anchoring rule applied to humans and machines identically, a named policy owner with a review date, and a no-blame route for reporting mistakes. Prohibition-only policies drive usage onto personal accounts, which is worse than governed usage.
Who is responsible when AI makes a mistake in legal work?
The lawyer who reviewed and signed the output, exactly as with a draft from a junior colleague. Delegation does not transfer professional responsibility. This is why a named reviewer of record for every automated action is the foundational governance control - not a formality, but the mechanism that makes responsibility locatable before something goes wrong rather than after.
Do we need to tell clients we use AI?
Decide it centrally and apply it consistently rather than matter by matter. Check your jurisdiction's professional rules and any client engagement terms, then take a position. Practically, disclosure is increasingly an advantage: describing a systematic first pass under a strict anchoring rule with named human review is describing a more rigorous process than most firms run.
How do we stop staff using unapproved AI tools with client data?
Give them an approved tool that is genuinely good enough for the work, make the approved path the path of least resistance, and be specific about what must never be entered anywhere unapproved. Shadow usage is almost always a symptom of an approved option that is missing, slow to access, or worse than the free alternative - not of poor discipline.
What is the difference between an AI policy and AI governance?
A policy is a document. Governance is the operating structure: for each automated action, a named doer, a named human reviewer of record, a defined review standard, and an explicit escalation trigger, all kept current. A firm with the four-attribute record and no policy document is meaningfully governed. A firm with a policy and no delegation structure is not.