Skip to content
Five ChipLLC

Case study

A local, air-gapped AI pipeline for legal document discovery

Hundreds of pages of scanned records, read line by line. A month later: indexed, searchable and diffed case files that never leave one MacBook, with every fact traceable to its page.

Client
A county government legal team
Role
Principal consultant, Improving
Timeline
About a month to a working pipeline
Stack
Python, Ollama, Qwen vision and text models, Markdown. All local.

How it started

A lunch conversation about TV shows where lawyers get legal documents dumped on them.

I was finishing a contract with a legal team of more than a hundred people when someone described, over lunch, how they prepared for a case: hundreds of scanned PDF documents, read line by line, with notes taken by hand. Every time a record was updated, the reading started over. It sounded horrible, and solvable.

I knew there were AI tools for document analysis. I also knew from experience that I didn't have nearly enough information to start looking yet. Looking for solutions too early is the classic mistake. So I asked whether anyone would show me the documents after lunch.

Define the Problem—No Solutions!

Interviews first, not a solution search

Before thinking of specific solutions, I spent time with the people who do this work, in working sessions and hallway conversations. I asked them to walk me through how they worked with these documents.

  • “I'm going to spend eight hours reading these child interview documents, trying to identify whether any instances of fear were noted by the providers.”
  • “I can't have it guessing or hallucinating. Lawyers have already been accused of using AI in court, and opposing counsel has asked them to verify their statements.”
  • “I cannot send information to ChatGPT! These are sensitive documents that contain hospital records, abuse details and mental health assessments. No way.”
  • “We've been talking about AI. Some guys tried to explain it. It went over our heads. He just used a bunch of jargon.”

Two things were non-negotiable: confidential documents would not go to any outside service, and nothing the tool produced could be used unless a person could verify it against the page it came from. The last quote explained the third: earlier attempts to explain AI had gone over people's heads, so whatever I built had to be understandable to the people relying on it.

Define requirements—Still No Solutions!

Summarize the problem, known constraints, and definition of success.

A small legal team must process a growing volume of lengthy, complex and frequently changing medical and legal records. Critical information is found by reading every line and taking notes. Updates force the same reading again, which kills productivity and morale, and the information is highly confidential, with strict access control even between lawyers in the same office. The business cost: staff time devoted to intake and review instead of clients.

Requirements

Deterministic
The same inputs produce the same results, every time.
Provenance
Every artifact records how it was made: source, model, version, configuration and processing logic.
Traceable
Every extracted item points back to its source document and page.
Testable
Automated coverage of extraction and processing, so models and logic can be upgraded safely.
Regenerable
When a model, prompt or rule changes, provenance says exactly which documents to reprocess.
Secure
The number one concern, and the one that could stop the project before it started.

Constraints

  • Free and open source, end to end
  • Runs entirely on the machine. Air-gap capable, no external services
  • Develops and runs on one MacBook M1 with 32 GB of memory
  • Simple enough for a two-person team, one developer and one manager, who also support the core legal application
  • About a month to a working solution

What success looked like

Staff spending less time on intake and review and more time with clients. In the interviews, four outputs were named as immediately useful:

  1. 1A list of every provider named in the documents
  2. 2A time-ordered list of events, each with its source document and page
  3. 3A table of problems and medications
  4. 4A keyword index for the things the team must never miss, with the exact page for each hit

Analyze and Prototype Solutions

The constraints guided the design.

Security and funding were the two questions that could stall everything, so I took them off the table immediately: the solution would be free, open source and entirely local. That one decision shaped every choice that followed.

  1. Upload the PDFs straight into a local RAG chat

    AnythingLLM and PrivateGPT on top of Ollama, with Mistral 7B and Qwen 2.5. I varied chunk size, overlap, temperature, chat mode and prompts exhaustively. Results were varied, incomplete and often wrong. The documents were too dense, OCR errors poisoned the answers, and relationships across documents were unreliable. Temperature 0 and query mode were needed for any determinism at all.

    Lesson. I needed accurate text from these image PDFs before going any further.

  2. Convert PDF to text with Unstructured.io

    An open-source ETL for messy documents. Adding document-specific configuration and prompts improved things, and for the first time I could trace a bad answer back to its cause. But the per-document-type prompt engineering grew complicated fast, tables inside tables came out unusable, and OCR still made mistakes on the highest-quality setting. A "printed on" date being read as the appointment date was the kind of error that ended it.

    Lesson. Hard stop. Every document carries the same kinds of information: who, what, when, where, why. If I could get a 100% accurate transcription, standardizing it was a job for code, not prompts.

  3. A vision model that only transcribes, combined with Python that standardizes and makes targeted AI calls.

    Unstructured's own documentation pointed the way: for scanned documents with complex layouts, use a vision language model. I gave qwen2.5-VL a deliberately narrow prompt: "Reproduce exactly what is printed. Don't correct, don't infer, write [illegible] rather than guess, never reorder rows." At temperature 0 it produced faithful Markdown page after page. Once every document was accurate text, turning it into a common "canonical" shape was ordinary, testable Python.

    Lesson. Use Python to pre-process as much as possible. Use AI for things code can't do. Merge the results in a deterministic way.

Solution—Python/AI Pipeline

Small, deterministic steps. AI only where it earns its place.

The core is four steps (Figure 1). Only one of them uses a vision language model, and that model is asked to do exactly one thing: read what is on the page. Everything after it is configuration and code, which means it can be tested, diffed and reasoned about by the team that owns it.

The client's requirements were met by taking canonical.md from the core steps and processing it to generate the documents that meet them (Figure 2).

Core pipeline: scanned PDFs are split into page images, transcribed by a local vision model at temperature 0, assembled into a vision Markdown file, then converted by configuration rules with no AI into a canonical Markdown file. One MacBook. Nothing leaves the machine.Scanned PDFsScanned medical and case records1 · PDF to PNGOne page at a time2 · PNG to textA vision model transcribes each pageOllama, localqwen2.5-VL 32BTemperature 0Transcribe, never guess3 · Assemble pagesWrites the vision filevision.mdExact page text4 · Derive canonical MarkdownConfiguration rules. No AI.canonical.mdStandard shape per type
Figure 1. Core pipeline. Every document, whatever it looked like on paper, comes out as the same canonical Markdown shape for its type.
Output pipeline: each canonical Markdown file feeds a keyword index, a diff against the previous version and a unified case file, with targeted calls to a small local model, producing keywords, diff and unified Markdown documents. Still the same MacBook. Add steps as the team asks for them.canonical.mdSame shape for every document5 · Keyword indexTeam's list, plus flagged model hits6 · Diff and summaryAgainst the previous version7 · Unified case fileEvery document on one timelineOllama, localqwen2.5 7BTargeted calls onlyEach hit cites its pageCode does the restkeywords.mdHits, file and pagediff.mdWhat changed, whereunified.mdPeople and timeline
Figure 2. Outputs. Steps are added as the team asks for them, each a small Python step with a targeted model call where useful.

Keyword index

The legal team defined a strict keyword list per category, such as alcohol, housing, fear or non-compliance. Hits from that list are deterministic. The model may add phrases that match a category, but every model hit is labeled as such and carries the verbatim text that triggered it, so it can be checked against the page.

Document diff

When a record is updated, a plain deterministic diff finds what changed and where. The model writes a short summary of each change on top of that diff, and the raw diff stays in the file for validation. No one rereads 160 pages to find the two that changed.

Unified case file

Every document in a case merged into one file: a topic index, the people involved, a complete timeline in date order, and the problems and medications table. Each entry names its source file and page, and the timeline entries are the actual transcribed text, not a paraphrase.

Solution—Validation and Control

Determinism and guardrails are the product.

  • A golden document, run nightly

    The team assembled a 60-page PDF from the worst of the worst. Its verified transcription is a permanent test case that runs every night, because a prompt tweak or a settings change can silently degrade transcription.

  • Model settings treated as engineering

    Temperature 0 for determinism. A mild repeat penalty of 1.1, because medical records legitimately repeat themselves and a harsher penalty changes the words. A three-second cooldown between pages so the laptop's GPU doesn't mangle a transcription under contention.

  • The right model, not the biggest

    The 7B vision model was promising; the 32B version was needed for pages where the model has to reason over a large area, and it is borderline on the hardware. Local models also need to be asked for what they were trained to produce: this one writes Markdown, so asking for JSON was a mistake.

  • Provenance in every file

    Each output starts with a record of its source file and hash, document type, model name and parameters, generation time and page count. When anything changes, the pipeline knows what to regenerate.

  • An AI skill so the team can extend it

    Adding a new document type is a five-stage, checklisted procedure captured as a skill for AI-assisted development, so the two-person team can extend the pipeline without me. It insists on scrubbing personal information after the vision step.

Results

Were the goals met?

Yes. All four outputs the team asked for in the interviews are produced: a provider list, a sourced timeline, problems and medications, and the keyword index. Each is easily traced to its source PDF, received during discovery.

Lesson: Had I jumped to solutions, I might have spent a lot of time investigating cloud-based approaches before really understanding the users' problems and concerns.

This solution keeps the control in the hands of the team. It can also be modified to meet future requirements, such as uploading to a database or using a secure cloud-based approach.

Deterministic
Yes, for all practical purposes. The vision model can add stray spaces or pipes; the canonical step normalizes them.
Provenance
Yes. Every document records what produced it.
Traceable
Yes. Source file and page are 100% accurate.
Testable
Yes. Tests on all the Python, the golden PDF for the vision step, more model tests as needed.
Regenerable
Yes, and with provenance it can be automated.
Free, open source, local only
Yes, on a single MacBook M1.
Maintainable by two people
Yes, with the skill and checklist.

What I took away

There is no substitute for direct interaction with clients. Without open, candid conversations, I never would have arrived at this solution.

The determinism requirement and the air-gap constraint (discovered early in the project) pushed the design toward small local models chained together, with ordinary code between them, instead of one large prompt sent to a paid cloud model. That turned out to be a better design at this time, not just a permitted one.

But the bigger takeaway was that they really liked this approach. The generated documents fit their existing policies for document control and security. They uploaded them into their current case management system, meeting all the requirements for document retention and tracking with no extra work.

The team now has a stable, local starting point for deterministic AI that no vendor can take away or reprice. If they later gain trust in AI services, or if new systems appear that they are comfortable with, this solution can be extended.

Presented atTwin Cities Solution Architecture Meetup, September 2026

The same approach, twenty years apart

I ran a wafer fab project the same way in 2005.

Different industry, different stack, same shape: start with the people, prove the measurement, take the constraints from the room, make the smallest change that fits, prove it, and hand it over to stay. That system is still in production twenty years later.

Get in touch

What's on your mind?

A few sentences is enough to start a useful conversation. Just say what your thinking and we'll take it from there.