This post is scheduled to publish on 15 September 2026, after Legal AI has been live on real solicitor matters for roughly six months. The content below is a template. The numbers and specific incidents will be filled with actual production data drawn from that deployment window.
What follows is the structural narrative. The specifics harden as production experience accumulates.
What Legal AI saw in production
Across six months of deployment, Legal AI processed contract analyses, risk-scored clauses, and produced draft position memos for real commercial matters at pilot firms. The deployment surfaced three patterns that were not obvious from architectural design.
Pattern one: the high-confidence hallucination class. The drafting agent occasionally produced clause re-writes that were plausible in English and legally wrong. The verification agent caught most, but not all. Post-hoc: we tuned the verification agent’s rule-set for specific clause types — indemnity caps, limitation splits — where the failure mode is the most expensive.
Pattern two: the client-context-dependent clause. A clause that’s safe for Client A is a red flag for Client B, because Client B has a prior dispute pattern, an unusual corporate structure, or a specific regulatory posture. The per-matter source-of-truth memory surfaced these. Without it, the agent would have mis-scored risk on 10–15% of matters.
Pattern three: the "we don’t know what we don’t know" class. Cases where the matter type didn’t match the classification tree, the contract was in a format outside the training distribution, or the jurisdictional context included rules the agents weren’t aware of. The QA gate rejected these, which was the right behaviour — but surfaced that the product needed a "flag for human triage" output alongside "score with confidence."
What the partners actually said
Placeholder: this section will hold a short set of partner quotes from pilot firms, gathered at the 3 and 6 month mark. Expect something in the shape of:
"The tool catches things our associates miss. It does not catch things partners would catch. That split is the right one — we pay partners to notice what isn’t there, and we need juniors fast-tracked through the mechanical work."
Why wrappers still don’t compete
Over the six-month deployment period, we also watched pilot firms evaluate ChatGPT Legal, Harvey, and a couple of the larger UK-specific wrapper products. The evaluation pattern was consistent: the wrappers produced plausible output, the partners didn’t trust the output, and the conversation returned to "how do we know it’s right?"
Without a verification gate, without per-matter memory, without inspectable reasoning, that question has no answer. Legal AI’s deployment taught us that the question is the product — every lawyer we met needed to know why, and the wrapper products could not tell them.
What we changed between month 1 and month 6
- Split the verification agent into two specialised gates — one for risk-scoring integrity, one for draft-memo factual accuracy.
- Added a "confidence below threshold" output class that surfaces to the partner as "flag for human triage" rather than a scored result.
- Tuned per-practice-area rubrics — commercial dispute work has different failure modes than corporate-transaction work.
- Extended the matter-source-of-truth memory to include historic dispute patterns, not just prior contracts.
What this means for firms evaluating Legal AI
If you’re a firm considering Legal AI or any other legal-vertical AI product, the questions we’d ask your vendor:
- What happens when the agent doesn’t know? Does it produce confident-sounding nonsense, or does it flag for human triage?
- How does it preserve per-matter privilege? Is the memory isolated matter-by-matter, or is it a shared corpus?
- Can the reasoning be inspected if HMRC’s equivalent of a professional-conduct enquiry opens? What does the audit trail look like?
- How do failure modes come back to you? Do you see them in production, or does the vendor fix them invisibly?
We think these questions are the ones a solicitor buying enterprise legal AI should be asking. We can answer them for Legal AI. We’re not certain the wrappers can.
Details: Legal AI product page. Pilot enquiries: enterprise@blackflake.com.
— Bartek Kubas · Founder-architect · Blackflake 15 September 2026 · Łódź, Poland