Receipt field extraction with calibrated confidence, selective prediction, and
an append-only human-review workflow. On the 200-document held-out test split,
the frozen score can review 75.7% of normalized field predictions and retain
24.3% at 95.3% observed accuracy. See RESULTS.md for the complete evaluation
and its limitations.
The repository contains three layers:
- Stage 1: structured receipt extraction with source spans.
- Stage 2: reproducible confidence calibration and held-out evaluation.
- Stage 3: a FastAPI result service and Next.js review dashboard over a frozen serving export. Stage 3 does not call the model, recompute signals, refit the calibrator, or change the 0.90 threshold.
The committed data/stage3_serving.json contains the frozen calibration/test
field results needed by the application. Running the website does not require a
Gemini key, downloading CORD-v2, or rerunning Stages 1 and 2.
Requirements:
- Python 3.10 or newer
- Node.js 20.9 or newer
From the repository root, install the API dependencies:
python -m pip install -r requirements.txtStart the API in terminal 1:
python -m uvicorn docextract.api:app --host 127.0.0.1 --port 8000The health/provenance response is at http://127.0.0.1:8000/health; interactive
API documentation is at http://127.0.0.1:8000/docs.
Start the dashboard in terminal 2:
cd dashboard
Copy-Item .env.example .env.local -Force
npm install
npm run dev -- -p 3001Open http://127.0.0.1:3001. Development mode uses a WebSocket only for Next.js
hot reloading; it is not a DocExtract or model request.
For a production-mode local run without the development WebSocket:
cd dashboard
Copy-Item .env.example .env.local -Force
npm install
npm run build
npm run start -- -p 3001The dashboard renders queue filtering, detail pages, and correction submission
server-side. Corrections append to the ignored local file
data/corrections.sqlite; they never overwrite the committed model output or
change its calibrated score or abstention decision.
GET /health: frozen-artifact provenance and test field/document counts.GET /results: paginated fields, filterable by matcher, abstention, correction state, confidence range, and split.GET /results/{document_id}: source text, original values, ground truth, five signals, calibrated scores, decisions, and correction history.GET /review-queue: abstained fields ordered by lowest confidence first.POST /results/{document_id}/fields/{field_id}/correction: append a human correction without rescoring the field.
Use Python 3.10 or newer. From the repository root:
python -m pip install -r requirements.txt
python scripts/build_dataset.py
python scripts/run_stage1.py --offline
python scripts/render_results.py
python scripts/verify_stage1.pyDataset building downloads public annotations. Stage 1 offline mode uses committed model responses and records missing coverage without making API calls. To resume any missing extractions, configure GEMINI_API_KEY in .env using .env.example, then run python scripts/run_stage1.py (makes model API calls).
Stage 1 extracts receipt totals and line items, stores their source spans in SQLite, and measures exact and normalized accuracy, ECE, and AUROC. The fixed splits contain 150 development, 250 calibration, and 200 test documents. Only the complete test split is a held-out result. Confidence calibration and discrimination answer separate questions and are reported separately.
python scripts/verify_stage1.py requires complete coverage, creates a fresh temporary SQLite database, rebuilds every extraction from cache, and compares the resulting report with results/stage1.json, excluding timestamps, elapsed time, and call statistics. API access is explicitly blocked during verification.
Stage 2 combines self-reported confidence with self-consistency, source spans, value presence, and type validity in logistic calibration models. It includes an isotonic baseline, feature ablations, a shuffled-label negative control, and paired document-bootstrap intervals. See STAGE2_PROTOCOL.md for the fixed protocol and limitations.
python scripts/sample_stage2.py # resume missing samples via API
python scripts/check_stage2_partial_signal.py # optional no-API development signal check
python scripts/run_stage2.py # offline fitting and evaluation
python scripts/render_results.py
python scripts/verify_stage2.py # requires complete samples; fresh database checkN means additional samples: two per calibration document and three per test document, totaling 1,100 calls. Sampling is resumable, and --max-calls sets a per-run limit. sample_stage2.py --offline restores only cached responses. Both Stage 2 commands exit 2 when coverage is incomplete, recording coverage without partial-split metrics. The report includes both matchers and a test N=2 sensitivity check without refitting.
python -m pytest -q
python scripts/check_cache_safe.py
python scripts/verify_stage3.py
cd dashboard
npm run buildscripts/verify_stage3.py checks that the serving export matches the committed
Stage 2 SHA-256, coverage, threshold, and abstention decisions without
recalculating the confidence model. See RESULTS.md for the generated Stage 1
and Stage 2 analysis.