How it started
A lunch conversation about TV shows where lawyers get legal documents dumped on them.
I was finishing a contract with a legal team of more than a hundred people when someone described, over lunch, how they prepared for a case: hundreds of scanned PDF documents, read line by line, with notes taken by hand. Every time a record was updated, the reading started over. It sounded horrible, and solvable.
I knew there were AI tools for document analysis. I also knew from experience that I didn't have nearly enough information to start looking yet. Looking for solutions too early is the classic mistake. So I asked whether anyone would show me the documents after lunch.
Define the Problem—No Solutions!
Interviews first, not a solution search
Before thinking of specific solutions, I spent time with the people who do this work, in working sessions and hallway conversations. I asked them to walk me through how they worked with these documents.
“I'm going to spend eight hours reading these child interview documents, trying to identify whether any instances of fear were noted by the providers.”
“I can't have it guessing or hallucinating. Lawyers have already been accused of using AI in court, and opposing counsel has asked them to verify their statements.”
“I cannot send information to ChatGPT! These are sensitive documents that contain hospital records, abuse details and mental health assessments. No way.”
“We've been talking about AI. Some guys tried to explain it. It went over our heads. He just used a bunch of jargon.”
Two things were non-negotiable: confidential documents would not go to any outside service, and nothing the tool produced could be used unless a person could verify it against the page it came from. The last quote explained the third: earlier attempts to explain AI had gone over people's heads, so whatever I built had to be understandable to the people relying on it.
Define requirements—Still No Solutions!
Summarize the problem, known constraints, and definition of success.
A small legal team must process a growing volume of lengthy, complex and frequently changing medical and legal records. Critical information is found by reading every line and taking notes. Updates force the same reading again, which kills productivity and morale, and the information is highly confidential, with strict access control even between lawyers in the same office. The business cost: staff time devoted to intake and review instead of clients.
Requirements
- Deterministic
- The same inputs produce the same results, every time.
- Provenance
- Every artifact records how it was made: source, model, version, configuration and processing logic.
- Traceable
- Every extracted item points back to its source document and page.
- Testable
- Automated coverage of extraction and processing, so models and logic can be upgraded safely.
- Regenerable
- When a model, prompt or rule changes, provenance says exactly which documents to reprocess.
- Secure
- The number one concern, and the one that could stop the project before it started.
Constraints
- Free and open source, end to end
- Runs entirely on the machine. Air-gap capable, no external services
- Develops and runs on one MacBook M1 with 32 GB of memory
- Simple enough for a two-person team, one developer and one manager, who also support the core legal application
- About a month to a working solution
What success looked like
Staff spending less time on intake and review and more time with clients. In the interviews, four outputs were named as immediately useful:
- 1A list of every provider named in the documents
- 2A time-ordered list of events, each with its source document and page
- 3A table of problems and medications
- 4A keyword index for the things the team must never miss, with the exact page for each hit
Analyze and Prototype Solutions
The constraints guided the design.
Security and funding were the two questions that could stall everything, so I took them off the table immediately: the solution would be free, open source and entirely local. That one decision shaped every choice that followed.
Upload the PDFs straight into a local RAG chat
AnythingLLM and PrivateGPT on top of Ollama, with Mistral 7B and Qwen 2.5. I varied chunk size, overlap, temperature, chat mode and prompts exhaustively. Results were varied, incomplete and often wrong. The documents were too dense, OCR errors poisoned the answers, and relationships across documents were unreliable. Temperature 0 and query mode were needed for any determinism at all.
Lesson. I needed accurate text from these image PDFs before going any further.
Convert PDF to text with Unstructured.io
An open-source ETL for messy documents. Adding document-specific configuration and prompts improved things, and for the first time I could trace a bad answer back to its cause. But the per-document-type prompt engineering grew complicated fast, tables inside tables came out unusable, and OCR still made mistakes on the highest-quality setting. A "printed on" date being read as the appointment date was the kind of error that ended it.
Lesson. Hard stop. Every document carries the same kinds of information: who, what, when, where, why. If I could get a 100% accurate transcription, standardizing it was a job for code, not prompts.
A vision model that only transcribes, combined with Python that standardizes and makes targeted AI calls.
Unstructured's own documentation pointed the way: for scanned documents with complex layouts, use a vision language model. I gave qwen2.5-VL a deliberately narrow prompt: "Reproduce exactly what is printed. Don't correct, don't infer, write [illegible] rather than guess, never reorder rows." At temperature 0 it produced faithful Markdown page after page. Once every document was accurate text, turning it into a common "canonical" shape was ordinary, testable Python.
Lesson. Use Python to pre-process as much as possible. Use AI for things code can't do. Merge the results in a deterministic way.
Solution—Python/AI Pipeline
Small, deterministic steps. AI only where it earns its place.
The core is four steps (Figure 1). Only one of them uses a vision language model, and that model is asked to do exactly one thing: read what is on the page. Everything after it is configuration and code, which means it can be tested, diffed and reasoned about by the team that owns it.
The client's requirements were met by taking canonical.md from the core steps and processing it to generate the documents that meet them (Figure 2).
Keyword index
The legal team defined a strict keyword list per category, such as alcohol, housing, fear or non-compliance. Hits from that list are deterministic. The model may add phrases that match a category, but every model hit is labeled as such and carries the verbatim text that triggered it, so it can be checked against the page.
Document diff
When a record is updated, a plain deterministic diff finds what changed and where. The model writes a short summary of each change on top of that diff, and the raw diff stays in the file for validation. No one rereads 160 pages to find the two that changed.
Unified case file
Every document in a case merged into one file: a topic index, the people involved, a complete timeline in date order, and the problems and medications table. Each entry names its source file and page, and the timeline entries are the actual transcribed text, not a paraphrase.
Solution—Validation and Control
Determinism and guardrails are the product.
A golden document, run nightly
The team assembled a 60-page PDF from the worst of the worst. Its verified transcription is a permanent test case that runs every night, because a prompt tweak or a settings change can silently degrade transcription.
Model settings treated as engineering
Temperature 0 for determinism. A mild repeat penalty of 1.1, because medical records legitimately repeat themselves and a harsher penalty changes the words. A three-second cooldown between pages so the laptop's GPU doesn't mangle a transcription under contention.
The right model, not the biggest
The 7B vision model was promising; the 32B version was needed for pages where the model has to reason over a large area, and it is borderline on the hardware. Local models also need to be asked for what they were trained to produce: this one writes Markdown, so asking for JSON was a mistake.
Provenance in every file
Each output starts with a record of its source file and hash, document type, model name and parameters, generation time and page count. When anything changes, the pipeline knows what to regenerate.
An AI skill so the team can extend it
Adding a new document type is a five-stage, checklisted procedure captured as a skill for AI-assisted development, so the two-person team can extend the pipeline without me. It insists on scrubbing personal information after the vision step.
Results
Were the goals met?
Yes. All four outputs the team asked for in the interviews are produced: a provider list, a sourced timeline, problems and medications, and the keyword index. Each is easily traced to its source PDF, received during discovery.
Lesson: Had I jumped to solutions, I might have spent a lot of time investigating cloud-based approaches before really understanding the users' problems and concerns.
This solution keeps the control in the hands of the team. It can also be modified to meet future requirements, such as uploading to a database or using a secure cloud-based approach.
- Deterministic
- Yes, for all practical purposes. The vision model can add stray spaces or pipes; the canonical step normalizes them.
- Provenance
- Yes. Every document records what produced it.
- Traceable
- Yes. Source file and page are 100% accurate.
- Testable
- Yes. Tests on all the Python, the golden PDF for the vision step, more model tests as needed.
- Regenerable
- Yes, and with provenance it can be automated.
- Free, open source, local only
- Yes, on a single MacBook M1.
- Maintainable by two people
- Yes, with the skill and checklist.
What I took away
There is no substitute for direct interaction with clients. Without open, candid conversations, I never would have arrived at this solution.
The determinism requirement and the air-gap constraint (discovered early in the project) pushed the design toward small local models chained together, with ordinary code between them, instead of one large prompt sent to a paid cloud model. That turned out to be a better design at this time, not just a permitted one.
But the bigger takeaway was that they really liked this approach. The generated documents fit their existing policies for document control and security. They uploaded them into their current case management system, meeting all the requirements for document retention and tracking with no extra work.
The team now has a stable, local starting point for deterministic AI that no vendor can take away or reprice. If they later gain trust in AI services, or if new systems appear that they are comfortable with, this solution can be extended.
Presented atTwin Cities Solution Architecture Meetup, September 2026
The same approach, twenty years apart
I ran a wafer fab project the same way in 2005.
Different industry, different stack, same shape: start with the people, prove the measurement, take the constraints from the room, make the smallest change that fits, prove it, and hand it over to stay. That system is still in production twenty years later.