Architecture
Directory Structure
receipt-ocr/ ├── app/ │ ├── __init__.py │ ├── main.py # FastAPI app (sync + async endpoints) │ └── worker.py # Celery worker for async extraction ├── data/ │ ├── storage.py # StorageBackend interface (local / S3) │ └── dataset.py # CordDataset + WebDataset loaders ├── train/ │ ├── finetune_donut.py # LoRA finetuning script │ └── checkpoints/ │ └── donut-receipt-lora/ ├── eval/ │ ├── run_eval.py # SyncEvalRunner / AsyncEvalRunner │ └── metrics.py # field-level precision / recall / F1 ├── scripts/ │ └── create_shards.py # WebDataset tar shard creation + S3 sync ├── demo/ │ └── index.html # this file — interactive project docs ├── tests/ │ └── test_api.py # FastAPI TestClient tests (mocked model) ├── config.yaml ├── schema.json # GroceryReceipt JSON Schema ├── pyproject.toml ├── Dockerfile └── docker-compose.yml
Local Apps
Three FastAPI single-file apps drive the loop: the annotator captures ground truth from phone photos, the comparator A/B's every wired OCR / VLM pipeline on the same image, and the batch runner processes a folder of receipts end-to-end through a chosen pipeline and persists results to JSONL + SQLite for downstream tools.
click to enlarge
Annotator
annotator/app.py ·
localhost:7860
Drop a receipt photo in, pick a model to auto-suggest, hand-correct the JSON,
hit save. Result lands in annotations.jsonl.
Tracks which fields you corrected vs accepted — useful signal later.
- Auto-suggest via Claude (haiku/sonnet), OCR→LLM pipelines, or Ollama qwen3:8b
- Inline JSON editor with schema validation and rapidfuzz-based field-diff highlights
- Tracks
_fields_correctedper record for active-learning later
uv run uvicorn annotator.app:app --port 7860
click to enlarge
Comparator
comparator/app.py ·
localhost:7861
Pick any subset of the 15 wired models, drop in an image, stream results as each runner finishes (SSE). Each card has four toggle-views — Receipt JSON, Markdown (clean & raw), HTML, and the layout-detection JSON when the model emits one — rendered in sandboxed iframes.
- 15 model_ids covering raw OCR, OCR→LLM, VLM image→JSON, layout pipelines, finetuned Donut/LayoutLMv3
- Markdown-output models (PaddleOCR-VL, Nanonets, PP-StructureV3) auto-augmented with Haiku to also yield a Receipt JSON view
- Per-model latency, error capture, and a slide-out log drawer tailing
/tmp/comparator.log - Backed by the same
_run_one_model()the benchmark harness uses
uv run uvicorn comparator.app:app --port 7861
click to enlarge
Batch Runner
batch_runner/app.py ·
localhost:7862 ·
also a CLI: python -m pipeline …
Drop a folder of phone receipts, pick a pipeline (PaddleOCR-VL v1.6 → Haiku, Claude
Haiku image→JSON, EasyOCR → Ollama, …), watch each one finish as Receipt JSON, then
query everything from SQLite. Designed as a paperless-ngx external tool and as the
backing data source for a separate expense tracker — both consume the same
data/batch_runs/*/results.jsonl + data/receipts.db contract.
- PaddleOCR-VL v1.6 wired as the headline pipeline (markdown + Haiku-augmented JSON)
- 7 pipelines today; composable stage chains queued as a follow-up
- Per-image SSE streaming, queue with status badges, click any row to see Receipt JSON / Markdown / raw payload
- Every batch lands on disk:
manifest.json,results.jsonl,errors.jsonl,summary.csv, plus symlinks to the original images - SQLite is a derived index — delete
receipts.dbandpython -m pipeline reindexrebuilds it
uv run python batch_runner/app.py --port 7862
uv run python -m pipeline run --pipeline paddleocr-vl --input ./receipts/ --output data/batch_runs/
# paperless-ngx external-tool signature: single image → JSON on stdout
uv run python -m pipeline run --pipeline paddleocr-vl --input ./one.jpg --output -
These close the loop: capture GT in the annotator → compare candidate models in the comparator → run the chosen pipeline in batch through the runner / CLI → SQLite + JSONL persistence ready for downstream tools.
Benchmark — 19 PKR Receipts
run 2026-06-04 · MPS (Apple Silicon)Held-out evaluation of 11 OCR / VLM pipelines on a personal dataset of 19 grocery-and-pharmacy receipts captured in Lahore, PK (PKR currency). The annotator captures ground truth — store, date, line items with qty & total, currency, and overall total — and the benchmark scores each model against it via fuzzy field matching (rapidfuzz ≥80) and bipartite item matching (name ≥70 + total ±2%).
| Model | n | Store | Total | Items F1 | Overall F1 | p50 |
|---|---|---|---|---|---|---|
| Claude Haiku (image) | 19 | 0.72 | 0.61 | 0.35 | 0.614 | 34s |
| PaddleOCR → Ollama (qwen3:8b) | 19 | 0.53 | 0.63 | 0.31 | 0.494 | 62s |
| EasyOCR → Ollama (qwen3:8b) | 19 | 0.53 | 0.26 | 0.17 | 0.391 | 12s |
| PaddleOCR → Claude Haiku | 19 | 0.37 | 0.37 | 0.21 | 0.326 | 75s |
| EasyOCR → Claude Haiku | 19 | 0.26 | 0.11 | 0.00 | 0.189 | 49s |
| EasyOCR (raw text) | 19 | 0.00 | 0.00 | 0.00 | 0.000 | 5s |
| PaddleOCR (raw text) | 19 | 0.00 | 0.00 | 0.00 | 0.000 | 17s |
| PP-StructureV3 | 19 | 0.00 | 0.00 | 0.00 | 0.000 | 120s |
Raw-OCR rows score 0 by design — they emit text not structured JSON, so the field-level
scorer has nothing to match. Their value comes when piped into an LLM (rows 2-5 above).
claude-sonnet, nanonets-ocr-s,
and qwen2vl-2b are omitted: incomplete runs (2/19) or
dependency / latency blockers.
Donut finetune experiment
LoRA finetune of naver-clova-ix/donut-base-finetuned-cord-v2
on 14 PKR train receipts (4 held out). Two attempts; neither converged on JSON output —
honest result.
| Run | LoRA targets | Trainable | Epochs | Val loss | Output |
|---|---|---|---|---|---|
| v1 | [query, value] (encoder only) |
319K (0.16%) | 5 | 8.08→7.93 | "modulmodulmodul…" |
| v2 | [q_proj,k_proj,v_proj,out_proj,query,value] |
1.69M (0.83%) | 15 | 6.13→2.42 | "Sales Invoice Serveild Pharmacy…" |
Key insight:
Donut's decoder uses BART-style module names (q_proj etc.),
not query — v1's LoRA was applied to the encoder only,
so the decoder couldn't learn JSON generation. v2 wires the decoder LoRA correctly and the
model starts reading the receipts, but 14 examples is below the floor needed for schema
output. Estimate ≥100 receipts before a Donut LoRA pays off.
Reproduce: uv run python eval/benchmark.py --quick.
Raw data: eval/results/benchmark-2026-06-04/.
Try it Locally
This is a local demo — the FastAPI server must be running on your machine. The form below calls localhost:8000.
Run locally in 3 steps
git clone https://github.com/arslankazmi/receipt-ocr cd receipt-ocr pip install -e . uvicorn app.main:app # starts on localhost:8000 # then open demo/index.html in your browser
Upload a receipt image. Requires the local server running at localhost:8000.
Preview
Extracted JSON
Required fields
Local vs. Scale
| Component | Local demo | At scale (550 GB+) |
|---|---|---|
| Storage | LocalBackend (disk) | S3Backend (boto3, LRU cache) |
| Data loading | CordDataset (JSONL) | WebDataset (tar shards, streaming) |
| Training | python train/finetune_donut.py |
accelerate launch --multi_gpu |
| Eval | SyncEvalRunner (sequential) | AsyncEvalRunner (50x concurrent) |
| API | Single uvicorn process | Docker + Celery workers + Redis |
| Sharding | N/A | scripts/create_shards.py → aws s3 sync |