Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DocExtract

Receipt field extraction with calibrated confidence, selective prediction, and an append-only human-review workflow. On the 200-document held-out test split, the frozen score can review 75.7% of normalized field predictions and retain 24.3% at 95.3% observed accuracy. See RESULTS.md for the complete evaluation and its limitations.

The repository contains three layers:

  • Stage 1: structured receipt extraction with source spans.
  • Stage 2: reproducible confidence calibration and held-out evaluation.
  • Stage 3: a FastAPI result service and Next.js review dashboard over a frozen serving export. Stage 3 does not call the model, recompute signals, refit the calibrator, or change the 0.90 threshold.

Run the API and review dashboard

The committed data/stage3_serving.json contains the frozen calibration/test field results needed by the application. Running the website does not require a Gemini key, downloading CORD-v2, or rerunning Stages 1 and 2.

Requirements:

  • Python 3.10 or newer
  • Node.js 20.9 or newer

From the repository root, install the API dependencies:

python -m pip install -r requirements.txt

Start the API in terminal 1:

python -m uvicorn docextract.api:app --host 127.0.0.1 --port 8000

The health/provenance response is at http://127.0.0.1:8000/health; interactive API documentation is at http://127.0.0.1:8000/docs.

Start the dashboard in terminal 2:

cd dashboard
Copy-Item .env.example .env.local -Force
npm install
npm run dev -- -p 3001

Open http://127.0.0.1:3001. Development mode uses a WebSocket only for Next.js hot reloading; it is not a DocExtract or model request.

For a production-mode local run without the development WebSocket:

cd dashboard
Copy-Item .env.example .env.local -Force
npm install
npm run build
npm run start -- -p 3001

The dashboard renders queue filtering, detail pages, and correction submission server-side. Corrections append to the ignored local file data/corrections.sqlite; they never overwrite the committed model output or change its calibrated score or abstention decision.

API surface

  • GET /health: frozen-artifact provenance and test field/document counts.
  • GET /results: paginated fields, filterable by matcher, abstention, correction state, confidence range, and split.
  • GET /results/{document_id}: source text, original values, ground truth, five signals, calibrated scores, decisions, and correction history.
  • GET /review-queue: abstained fields ordered by lowest confidence first.
  • POST /results/{document_id}/fields/{field_id}/correction: append a human correction without rescoring the field.

Reproduce the research pipeline

Use Python 3.10 or newer. From the repository root:

python -m pip install -r requirements.txt
python scripts/build_dataset.py
python scripts/run_stage1.py --offline
python scripts/render_results.py
python scripts/verify_stage1.py

Dataset building downloads public annotations. Stage 1 offline mode uses committed model responses and records missing coverage without making API calls. To resume any missing extractions, configure GEMINI_API_KEY in .env using .env.example, then run python scripts/run_stage1.py (makes model API calls).

Evaluation and project scope

Stage 1 extracts receipt totals and line items, stores their source spans in SQLite, and measures exact and normalized accuracy, ECE, and AUROC. The fixed splits contain 150 development, 250 calibration, and 200 test documents. Only the complete test split is a held-out result. Confidence calibration and discrimination answer separate questions and are reported separately.

python scripts/verify_stage1.py requires complete coverage, creates a fresh temporary SQLite database, rebuilds every extraction from cache, and compares the resulting report with results/stage1.json, excluding timestamps, elapsed time, and call statistics. API access is explicitly blocked during verification.

Stage 2 combines self-reported confidence with self-consistency, source spans, value presence, and type validity in logistic calibration models. It includes an isotonic baseline, feature ablations, a shuffled-label negative control, and paired document-bootstrap intervals. See STAGE2_PROTOCOL.md for the fixed protocol and limitations.

python scripts/sample_stage2.py          # resume missing samples via API
python scripts/check_stage2_partial_signal.py  # optional no-API development signal check
python scripts/run_stage2.py             # offline fitting and evaluation
python scripts/render_results.py
python scripts/verify_stage2.py          # requires complete samples; fresh database check

N means additional samples: two per calibration document and three per test document, totaling 1,100 calls. Sampling is resumable, and --max-calls sets a per-run limit. sample_stage2.py --offline restores only cached responses. Both Stage 2 commands exit 2 when coverage is incomplete, recording coverage without partial-split metrics. The report includes both matchers and a test N=2 sensitivity check without refitting.

Validation

python -m pytest -q
python scripts/check_cache_safe.py
python scripts/verify_stage3.py

cd dashboard
npm run build

scripts/verify_stage3.py checks that the serving export matches the committed Stage 2 SHA-256, coverage, threshold, and abstention decisions without recalculating the confidence model. See RESULTS.md for the generated Stage 1 and Stage 2 analysis.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages