The workflow you already run#
Associates spend hours summarising contracts. So a model does the first read. It takes the signed agreement and returns the key terms, obligations and dates. It saves real time, so the habit spreads across the practice — NDAs, MSAs, side letters, all through the same convenient step.
The risk you can't yet see#
But a contract is not neutral text. It carries counterparty names, deal terms under NDA, and language that may be privileged. Send it to a model and the firm has shared client-confidential material with an outside system. Then a client asks a fair question: who saw our agreement, and can you prove it stayed inside the firm? Today nobody wrote it down. Confidentiality you cannot show is confidentiality a client will not trust.
Why nobody misconfigured anything#
Nothing here was a policy breach at the point it happened. An associate summarised a contract, which is the job. The tool was approved. The document was one they were entitled to read.
What changed is where the reading happened. A firm's confidentiality obligations attach to the document, not to the person, and they do not stop at the boundary of an approved tool. Once the habit spreads across the practice — NDAs, MSAs, side letters — the exposure is no longer a single document. It is a pattern nobody chose and nobody can describe.
That is the difficult part for a firm. The question a client asks is not "did anyone act improperly." It is "what happened to our agreement." Those need different answers, and only the second one needs a record.
What a wrong one costs here#
Contract summarisation has an unusual risk shape. The output is rarely relied on directly — an associate checks it. So the accuracy exposure is lower than it first appears.
The confidentiality exposure is the opposite. It does not need the summary to be wrong. It happens correctly, every time, in the ordinary course of the work being done well.
That inverts the usual advice. Here the thing to instrument is not the answer. It is the input: what left, when, under whose instruction, and whether it can be shown.
The smallest version that produces a number#
One practice group, one agreement type, six weeks. The output is not an accuracy score. It is an inventory: how many documents went to the model, what categories of content they contained, and what would have been withheld under a rule the firm writes itself.
That inventory is the deliverable, because it is the thing a client can be shown. A firm that can answer "here is what was sent, here is what was held back, here is who instructed it" is in a very different conversation from one offering reassurance.
It is also the cheapest way to find out whether the rule you would write is one you can actually apply. Most firms discover their first draft of it is too broad to be useful, and that is much better to learn in week three than in a client meeting.
What an Inspector caught#
Run in Watch mode, inside the firm's own environment, an Inspector observed the same summarization call:
On record#
Counterparty names and NDA terms it would have redacted before the model read them. Fourteen clauses segmented cleanly. One passage of likely-privileged language flagged for a human rather than summarized silently. And a provenance trail the firm could show a client — signed, on the firm's own stack, nothing run. Watch first: prove the confidentiality holds before an Inspector touches live matters.
The associates carried on summarising. The tool stayed approved. Nothing was blocked and no work slowed down.
What the firm gained is the ability to answer a client precisely. Not "we take confidentiality seriously", which every firm says, but "here is what was sent, here is what was withheld, and here is the rule that decided it." That is a different kind of sentence, and it is the only one that survives being asked twice.
What to check in your own stack#
Four questions. None of them needs a tool to answer, and all four are worth asking before you buy one.
- What is in the payload, not what is in the prompt? The prompt is written by a person and reviewed. The payload is assembled by a query and never read by anyone.
- Who would notice a wrong one, and when? Name the person and the moment. If you cannot, the review path does not exist yet.
- Where does the model run, and what leaves the tenant? "In our cloud" and "in our tenant" are different sentences.
- What record survives the conversation? If the answer is application logs, the answer is no.
What this does not tell you#
An Inspector in Watch mode reports what it would have caught. It does not tell you whether the workflow was worth automating, and it does not find a rule you have not written. Your policy is the ceiling.
It also will not fix data it cannot reach. If the access is not there, Watch mode says so in week one — an uncomfortable finding rather than a deliverable, and better in week one than month nine.
Related