receipt-ocr

VLM-powered receipt extraction · local demo → 550 GB scale

arslankazmi/receipt-ocr

Architecture

Lane 1 — Local demo (200 images) Lane 2 — At scale (550 GB+) Phone / Image LocalBackend (disk) Donut + LoRA / Claude fallback Schema Validator JSON Phone / Image S3Backend + WebDataset shards Multi-GPU Donut / Async Claude 50x parallel Schema Validator JSON

Directory Structure

receipt-ocr/
├── app/
│   ├── __init__.py
│   ├── main.py              # FastAPI app (sync + async endpoints)
│   └── worker.py            # Celery worker for async extraction
├── data/
│   ├── storage.py           # StorageBackend interface (local / S3)
│   └── dataset.py           # CordDataset + WebDataset loaders
├── train/
│   ├── finetune_donut.py    # LoRA finetuning script
│   └── checkpoints/
│       └── donut-receipt-lora/
├── eval/
│   ├── run_eval.py          # SyncEvalRunner / AsyncEvalRunner
│   └── metrics.py           # field-level precision / recall / F1
├── scripts/
│   └── create_shards.py     # WebDataset tar shard creation + S3 sync
├── demo/
│   └── index.html           # this file — interactive project docs
├── tests/
│   └── test_api.py          # FastAPI TestClient tests (mocked model)
├── config.yaml
├── schema.json              # GroceryReceipt JSON Schema
├── pyproject.toml
├── Dockerfile
└── docker-compose.yml

Local Apps

Three FastAPI single-file apps drive the loop: the annotator captures ground truth from phone photos, the comparator A/B's every wired OCR / VLM pipeline on the same image, and the batch runner processes a folder of receipts end-to-end through a chosen pipeline and persists results to JSONL + SQLite for downstream tools.

Annotator UI — receipt image on the left, structured form on the right click to enlarge

Annotator

annotator/app.py · localhost:7860

capture GT

Drop a receipt photo in, pick a model to auto-suggest, hand-correct the JSON, hit save. Result lands in annotations.jsonl. Tracks which fields you corrected vs accepted — useful signal later.

  • Auto-suggest via Claude (haiku/sonnet), OCR→LLM pipelines, or Ollama qwen3:8b
  • Inline JSON editor with schema validation and rapidfuzz-based field-diff highlights
  • Tracks _fields_corrected per record for active-learning later
uv run uvicorn annotator.app:app --port 7860
Comparator UI — 19 receipt thumbnails, image preview, and model picker click to enlarge

Comparator

comparator/app.py · localhost:7861

A/B models

Pick any subset of the 15 wired models, drop in an image, stream results as each runner finishes (SSE). Each card has four toggle-views — Receipt JSON, Markdown (clean & raw), HTML, and the layout-detection JSON when the model emits one — rendered in sandboxed iframes.

  • 15 model_ids covering raw OCR, OCR→LLM, VLM image→JSON, layout pipelines, finetuned Donut/LayoutLMv3
  • Markdown-output models (PaddleOCR-VL, Nanonets, PP-StructureV3) auto-augmented with Haiku to also yield a Receipt JSON view
  • Per-model latency, error capture, and a slide-out log drawer tailing /tmp/comparator.log
  • Backed by the same _run_one_model() the benchmark harness uses
uv run uvicorn comparator.app:app --port 7861
Batch Runner UI — drop zone + queue, live result viewer with Receipt JSON, run history sidebar click to enlarge

Batch Runner

batch_runner/app.py · localhost:7862 · also a CLI: python -m pipeline …

production pipeline

Drop a folder of phone receipts, pick a pipeline (PaddleOCR-VL v1.6 → Haiku, Claude Haiku image→JSON, EasyOCR → Ollama, …), watch each one finish as Receipt JSON, then query everything from SQLite. Designed as a paperless-ngx external tool and as the backing data source for a separate expense tracker — both consume the same data/batch_runs/*/results.jsonl + data/receipts.db contract.

  • PaddleOCR-VL v1.6 wired as the headline pipeline (markdown + Haiku-augmented JSON)
  • 7 pipelines today; composable stage chains queued as a follow-up
  • Per-image SSE streaming, queue with status badges, click any row to see Receipt JSON / Markdown / raw payload
  • Every batch lands on disk: manifest.json, results.jsonl, errors.jsonl, summary.csv, plus symlinks to the original images
  • SQLite is a derived index — delete receipts.db and python -m pipeline reindex rebuilds it
uv run python batch_runner/app.py --port 7862 uv run python -m pipeline run --pipeline paddleocr-vl --input ./receipts/ --output data/batch_runs/ # paperless-ngx external-tool signature: single image → JSON on stdout uv run python -m pipeline run --pipeline paddleocr-vl --input ./one.jpg --output -

These close the loop: capture GT in the annotator → compare candidate models in the comparator → run the chosen pipeline in batch through the runner / CLI → SQLite + JSONL persistence ready for downstream tools.

Benchmark — 19 PKR Receipts

run 2026-06-04 · MPS (Apple Silicon)

Held-out evaluation of 11 OCR / VLM pipelines on a personal dataset of 19 grocery-and-pharmacy receipts captured in Lahore, PK (PKR currency). The annotator captures ground truth — store, date, line items with qty & total, currency, and overall total — and the benchmark scores each model against it via fuzzy field matching (rapidfuzz ≥80) and bipartite item matching (name ≥70 + total ±2%).

Best Overall F1
0.614
Claude Haiku · image → JSON
Best Local Stack
0.494
PaddleOCR → Ollama qwen3:8b
Fastest Local
5s
EasyOCR raw · p50 latency
Model n Store Total Items F1 Overall F1 p50
Claude Haiku (image) 19 0.72 0.61 0.35 0.614 34s
PaddleOCR → Ollama (qwen3:8b) 19 0.53 0.63 0.31 0.494 62s
EasyOCR → Ollama (qwen3:8b) 19 0.53 0.26 0.17 0.391 12s
PaddleOCR → Claude Haiku 19 0.37 0.37 0.21 0.326 75s
EasyOCR → Claude Haiku 19 0.26 0.11 0.00 0.189 49s
EasyOCR (raw text) 19 0.00 0.00 0.00 0.000 5s
PaddleOCR (raw text) 19 0.00 0.00 0.00 0.000 17s
PP-StructureV3 19 0.00 0.00 0.00 0.000 120s

Raw-OCR rows score 0 by design — they emit text not structured JSON, so the field-level scorer has nothing to match. Their value comes when piped into an LLM (rows 2-5 above). claude-sonnet, nanonets-ocr-s, and qwen2vl-2b are omitted: incomplete runs (2/19) or dependency / latency blockers.

Donut finetune experiment

LoRA finetune of naver-clova-ix/donut-base-finetuned-cord-v2 on 14 PKR train receipts (4 held out). Two attempts; neither converged on JSON output — honest result.

Run LoRA targets Trainable Epochs Val loss Output
v1 [query, value] (encoder only) 319K (0.16%) 5 8.08→7.93 "modulmodulmodul…"
v2 [q_proj,k_proj,v_proj,out_proj,query,value] 1.69M (0.83%) 15 6.13→2.42 "Sales Invoice Serveild Pharmacy…"

Key insight: Donut's decoder uses BART-style module names (q_proj etc.), not query — v1's LoRA was applied to the encoder only, so the decoder couldn't learn JSON generation. v2 wires the decoder LoRA correctly and the model starts reading the receipts, but 14 examples is below the floor needed for schema output. Estimate ≥100 receipts before a Donut LoRA pays off.

Reproduce: uv run python eval/benchmark.py --quick. Raw data: eval/results/benchmark-2026-06-04/.

Try it Locally

This is a local demo — the FastAPI server must be running on your machine. The form below calls localhost:8000.

Run locally in 3 steps

git clone https://github.com/arslankazmi/receipt-ocr
cd receipt-ocr
pip install -e .
uvicorn app.main:app   # starts on localhost:8000
# then open demo/index.html in your browser

Upload a receipt image. Requires the local server running at localhost:8000.

Local vs. Scale

Component Local demo At scale (550 GB+)
Storage LocalBackend (disk) S3Backend (boto3, LRU cache)
Data loading CordDataset (JSONL) WebDataset (tar shards, streaming)
Training python train/finetune_donut.py accelerate launch --multi_gpu
Eval SyncEvalRunner (sequential) AsyncEvalRunner (50x concurrent)
API Single uvicorn process Docker + Celery workers + Redis
Sharding N/A scripts/create_shards.py → aws s3 sync