Home / Case studies / AI Contract Review in Three Languages for Legal Services
Case study · Legal services · Document AI & generative AI
AI Contract Review in Three Languages for Legal Services
A Spanish legal services organization needed AI contract review because it checked corporate contracts by hand in Spanish, Catalan & English for hours. Brainy Neurals built a review system that uses document AI & a language model to check every clause against known laws. Reviewers at the firm open each contract with every flag already placed on its exact line & page. Review that once took hours now takes minutes, on the client’s own report.
3
Languages in one review
Minutes
Per contract, once hours
Line & page
Where every flag lands
Published October 2026
At a glance
Four short answers sum up the contract review build for a Spanish legal services organization.
What problem did this solve?
A Spanish legal services organization reviewed corporate contracts by hand, in Spanish, Catalan & English. Long or scanned agreements took hours, & findings sat scattered across separate notes.
What did Brainy Neurals build?
Brainy Neurals built a document AI review system for a Spanish legal services organization. The system reads text & scanned PDFs alike, then flags each clause that breaks a known law on its exact line.
What changed after it went live?
Review moved from hours per contract to minutes, on the client’s own report. Reviewers now open each contract with every suspect line already flagged, in all three languages.
Who else could use this?
Any team that checks large volumes of documents against fixed standards could use the same pattern. Insurers, banks, hospitals & procurement desks all read agreements this way.
Engagement facts
- Industry · Legal services
- Client type · Spanish legal services organization
- Engagement · Contract review system build
- Focus · Contract review & compliance
- Timeline · Not disclosed
- Capabilities · Document AI & generative AI
- Delivery model · Project-based delivery
Why did contract review take hours?
Contract review took hours at a Spanish legal services organization because every
agreement was read by hand before anyone could call it clean. No AI contract review
or intelligent document processing sat behind the work, in any of its three languages.
Where the manual process broke
- Long agreements were read line by line, so one contract could swallow a lawyer’s afternoon.
- Scanned PDFs had to be read by eye, because software can’t search a picture of text.
- Legal references were checked from memory or a lookup, & wrong citations still slipped past.
- Each language had its own habits, so a rule written for Spanish missed the Catalan version.
- Findings landed in separate notes, & nobody held one view of contract quality.
contract review dataset budgeted 5 to 10 expert minutes per page, at law-firm rates.[1]
Why do common review approaches stall?
Common review approaches stall because each one breaks on a different part of the job. Teams usually try three of them before a grounded check, which keeps going but needs steady upkeep.
| Approach | What holds | Where it fails | Still suits |
|---|---|---|---|
| Manual review | Judgment on every line | Hours per contract | One-off, high-stakes agreements |
| Keyword & rule checks | Fast & consistent | Blind to meaning & scans | Fixed templates in one language |
| A general chatbot | Reads any language | Invents references, loses lines | A first rough summary |
| Grounded analyzer, our route | Checks clauses against known laws | Needs law set upkeep | High volumes, fixed standards |
The chatbot failure has been measured in published research. A 2024 study of legal questions found public language models hallucinating 58 to 88 percent of the time.[2]
What we built
We built the analyzer as one pipeline with four passes & one standing rule, so a lawyer makes every final call. Text contracts & scanned contracts meet in the same positioned text stream, which keeps the line & page of every word.
What counts as legal truth
The first decision set what counts as legal truth. We refused to let the model judge legal references from its own memory, because a model answering from memory invents authority.
So the law check reads from a maintained reference set of laws, & it flags whatever that set contradicts. The model reads every clause, & the reference set decides what the law actually says.
The model reads every clause, & the reference set decides what the law actually says.
How small the model’s role stays
The second decision kept generative AI development to a small, deliberate role. The language model translates flagged passages & drafts suggested fixes, but it never rules on whether a clause is valid.
The plainest pass earns its keep in production. Confirming clean sections tells a reviewer where reading is safe to skip, & skipped reading is where the hours went.
Two document paths feed one positioned stream, & four passes check it against sets the client controls. The final decision stays outside the system boundary.
The stack we used
The stack we used earned each place by keeping the line & page of whatever it touched. The law check runs on retrieval augmented generation, which means a reference set does the deciding instead of the model’s memory. Model & vendor names stay off this page on purpose.
Contract parsing
A PDF parsing layer keeps line & page positions, so every flag stays anchored. We ruled out a plain text dump.
Scanned input
An OCR engine, the software that turns page images into text, brings scans into the pipeline with positions. We ruled out manual retyping for scanned pages.
Mistake detection
A pattern set covering all three languages turns the firm’s standards into rules software can run. We ruled out hard-coded rules tied to one template.
Law verification
A maintained law reference set means no check ever runs on model memory. We ruled out trusting the model’s own recall.
Interpretation
One large language model reads all three languages & drafts suggested wording. We ruled out separate rule engines per language.
Review surface
A web view puts each flag beside the contract, so lawyers act where issues sit. We ruled out exporting flags to spreadsheets.
Orchestration
A workflow backend moves every contract through the passes with no hand-offs. We ruled out manual switching between tools.
How does one contract get reviewed?
One contract gets reviewed in six steps, from intake to a flagged view, in the order the system runs them.
- The system takes the contract in & detects whether the pages hold live text or scanned images.
- A parsing layer extracts the text with line & page positions, & the OCR engine does the same for scans.
- The mistake check compares every section against a maintained set of known error patterns, in the contract’s own language.
- The law check maps each clause to the reference set & flags any legal reference the set contradicts.
- The language model interprets flagged passages across Spanish, Catalan & English, then drafts a suggested correction for each.
- The reviewer view sets the contract beside its flags, anchored by line & page, with a summary on top.
Six steps belong to the system, & a lawyer’s judgment enters at the seventh.
What broke & how we fixed it
Three things broke during the build, & each one needed a different fix.
Scanned contracts broke first, coming back as text with no trustworthy lines. OCR read the
words & dropped the layout, so early flags pointed where reviewers couldn’t follow. A 2020
study of OCR quality found the same pattern, with language tasks getting worse as recognition errors rise.[3]
The law check broke next, & more quietly. Asked to confirm a reference from memory, the
model agreed with citations that only looked real. Plausible & correct are different
properties, & only one of them can be checked.
The pattern set broke last, on Catalan wording. Rules written from Spanish examples
missed the Catalan wording of the same mistake. Coverage looked complete right up until a
reviewer read a Catalan agreement.
Weeks like these are when teams hire AI developers to add hands without losing pace.
How each fix held up
Lines on scans
We rebuilt the OCR path to keep word positions & rebuild lines from them. A flag on a scanned page now lands exactly where a text-page flag does.
Invented references
We took the model out of the judging seat. It proposes a clause-to-law match, the reference set confirms or rejects it, & anything uncovered goes to a person unanswered.
Catalan coverage
We rebuilt the pattern set so every mistake carries its Spanish, Catalan & English forms together. A pattern missing a language now fails review before it ships.
Boilerplate, as it happens, started life as a printing word. Syndicated text once reached newspapers on steel plates, ready to run unchanged.
An AI proof of concept exists to surface work like this before anything gets promised.
What changed after go-live?
After go-live, the AI contract review pipeline from Brainy Neurals runs in production at the client, reading corporate contracts in three languages. Contracts that took hours to review now finish in minutes, on the client’s own report. We haven’t published an error-catch rate or a measured review time, because we won’t print a number we haven’t measured.
Day to day, a reviewer opens a contract & starts at the flags. The three languages stopped being three separate processes, so multilingual contract review now runs as one flow. The scanned archive reads like the rest of the pile.
| Dimension | Before | Now |
|---|---|---|
| Reading a contract | Line by line, in full | Flags first, then judgment |
| Time per review | Hours, sometimes days | Minutes, on the client’s report |
| Scanned PDFs | Read by eye | Parsed like any other contract |
| Legal references | Checked from memory & lookups | Checked against a reference set |
| Findings | Scattered across notes | One view, anchored to lines |
The mistake & law sets remain the client’s own standards, written down in a form software can run. Nothing in the pipeline drafts or signs a contract, so the lawyer’s authority stayed exactly where it was.
Walk one of your contracts through this pipeline
Send us one contract type with its languages & the rules it must meet. We’ll show where each check would land & what it would likely flag.
What would we do differently?
We’d do four things differently, & each lesson pairs the mistake we made with the practice we now repeat.
Treat position as data
- We flattened contracts to plain text early, & adding position tracking midway cost more than starting with it.
- Next time, keep line & page positions from the very first parse.
Ground every check before you scale
- A language model asked to confirm a legal reference will confirm a convincing one.
- Ground every check in a source you control, & let the model do the reading.
Budget the third language early
- Catalan arrived after the Spanish patterns had settled, so every rule needed reopening.
- Write all the languages into the pattern set together from the start.
Test degraded scans first
- We tuned on clean contracts & met the real archive later than we should have.
- Send heavily degraded documents through the pipeline before anything clean.
A language model asked to confirm a legal reference will confirm a convincing one.
Where else does AI contract review fit?
AI contract review is software that reads legal documents & flags clauses that break known rules or laws.
Insurance
Policy wordings & claims files checked against filed terms.
Patterns retrain on policy language.
Banking
Loan & mortgage files read against regulatory clauses.
Regulator texts form the reference set.
Capital markets
Fund documents & trading agreements held to disclosure rules.
Filings feed the reference set.
Healthcare
Payer contracts & consent forms against required clauses.
Patient identifiers get stricter handling.
Pharma
Quality agreements & batch records against good manufacturing practice clauses.
Pattern set learns GMP language.
Manufacturing
Supplier & quality agreements read against export terms.
Patterns learn engineering vocabulary.
Construction
Subcontracts & tender packs checked against contract conditions.
Standard contract forms anchor the checks.
Logistics
Customs paperwork & bills of lading against trade rules.
Customs codes replace the law set.
Real estate
Leases read against local tenancy law.
One reference set per jurisdiction.
Sports
Player & sponsorship contracts arriving in several languages.
Federation rulebooks become the reference set.
Every card runs the same line-level review pattern on a different document.
A port needs the target documents, a rebuilt pattern set, a fresh law set & confirming reviewers.
Questions buyers usually ask
How accurate is AI contract review?
Accuracy depends on the reference material the system checks against. This analyzer compares clauses to a maintained set of laws & error patterns instead of judging from model memory. A lawyer still rules on every flag it raises.
Can AI review contracts in languages other than English?
Yes, as long as the pipeline is built for each language you work in. The analyzer here reads Spanish, Catalan & English in one pass, with every mistake pattern carried in all three. Translation sits beside each flag, so reviewers can work across languages.
Will AI replace lawyers for contract review?
No, the analyzer removes the reading & leaves the judgment with the lawyer. It reads every line & places every flag, then a lawyer takes the decision from there. Responsibility for the contract never moved at all.
How long does it take to build an AI contract review system?
A proof of concept on your own contracts usually takes a few weeks. The pattern set & the law set need their own passes, & so does each language. Production follows once your reviewers trust the flags.
How much does AI review of contracts cost?
Cost depends on how many documents you process in how many languages, & on how much of your review standard is written down. Brainy Neurals scopes it from a short conversation & a contract sample, then quotes a fixed price. An AI readiness assessment tells you first whether your review rules can be encoded.
Is it safe to put confidential contracts through an AI tool?
Confidential contracts can be safe in an AI tool when the deployment matches your obligations. The analyzer can run against hosted models or entirely inside your own environment, & your obligations decide which. Settle that during scoping, before any document moves.
Tell us what your contracts need checked
Share your document volumes & languages, plus the PDF types you receive. One reply comes from the person who would architect the build.
Services behind this case study
Five Brainy Neurals services sit behind this case study, each one covering a layer of the build.
Document AI services
Contract parsing & OCR pipelines that keep every line & page position intact.
Generative AI applications
Language models held to narrow duties, interpreting & suggesting under human review.
RAG development
Retrieval layers that ground model answers in reference sets you control.
AI agent development
Workflow automation that routes documents through checks & queues results for people.
Hire AI developers
Engineers who join your team for the weeks when languages & layouts pile up.
An AI proof of concept tests this on your own contracts within weeks. AI consulting helps you choose rule sets before any build starts.
An AI readiness assessment checks whether your standards can be encoded, & our AI industry solutions show where this pattern already runs.
Similar case studies
Similar case studies from Brainy Neurals show the same review-gate idea at work in other settings.
AI Diet Assistant for Gastroenterology
Clinical dietary guidance generated under review gates, live in a healthcare setting.
All Brainy Neurals case studies
Every published Brainy Neurals build, gathered in one place for easy reading.
Cite this case study
Barot, Nandni. AI Contract Review in Three Languages for Legal Services. Brainy Neurals, September 2026. https://brainyneurals.com/case-studies/ai-legal-contract-review/
Sources cited on this page
[1] Hendrycks D, Burns C, Chen A, Ball S. CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. NeurIPS Datasets and Benchmarks Track. 2021. DOI 10.48550/arXiv.2103.06268
[2] Dahl M, Magesh V, Suzgun M, Ho DE. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis. 2024. Volume 16, pages 64 to 93. DOI 10.1093/jla/laae003
[3] van Strien D, Beelen K, Coll Ardanuy M, Hosseini K, McGillivray B, Colavizza G. Assessing the Impact of OCR Quality on Downstream NLP Tasks. Proceedings of ICAART 2020. Pages 484 to 496. DOI 10.5220/0009169004840496








