5.7 KiB
RUNBOOK — ky_court_pipeline.ipynb
One notebook to restore, discover, download, extract, sync, and verify the full Kentucky appellate court document corpus (~98.7k documents). It is designed to run in Google Colab, auto-resume after disconnects, and mirror all state up to a Drive folder so nothing is lost across sessions.
What this is
A single generated notebook (ky_court_pipeline.ipynb) whose cells are inlined
from real Python modules in notebook_builder/cells/. The notebook is rebuilt
with notebook_builder/build.py. All stages read/write the same SQLite DB and a
Store abstraction that is Google Drive in Colab and the local filesystem in
tests/smoke.
Stage order: restore → discover → backfill → download → extract → sync/verify. Any stage stops gracefully at the session budget; re-running resumes where it left off.
Files on Drive (MyDrive/ky_court-WIP/)
| Path | Purpose |
|---|---|
ky_court.sqlite |
source of truth (documents, discovery_runs, kv, extraction_log) |
pdf_by_hash/<sha256>.pdf |
deduped PDFs keyed by content hash (3,326 unique from 9,582 backup files) |
pdf_backup/<case>_<n>.pdf |
original named backups — retained, never deleted |
text/<sha256>.txt |
extracted text per PDF |
exports/file_manifest.csv |
sha256 → size manifest |
exports/document_inventory.csv / failed_inventory.csv |
exported summaries from verify |
status.json |
live heartbeat + phase/counts (the "is it alive" check) |
ky_court-WIP-resume-*.tar.gz |
(MyDrive root) resume tarball used to rehydrate a fresh runtime |
The store root is MyDrive/ky_court-WIP (Colab) or <WORKDIR>/ky_court-WIP (local).
First run
- If Drive is empty (no
ky_court-WIP), uploadky_court-WIP-resume-YYYY-MM-DD.tar.gztoMyDrive/. Stage restore will extract it. - Open
ky_court_pipeline.ipynbin Colab. - Edit the first code cell (
CFG) if needed (budget, phases — see Config knobs). - Runtime > Run all. A session mounts Drive, opens the DB (pulling it from the store if absent), heals stale state, then runs the enabled phases.
Expected session timeline (typical):
- setup: ~1 min (Drive mount, dirs, recover)
- discover: 10–30 min (74 shards, paged search, paced)
- backfill: ~20 min (hash
pdf_backup→ dedupe intopdf_by_hash) - download: ~10 h per session ≈ 6–12k documents (budget-capped)
- extract: rides along between batches (thread-pooled PDF text extraction)
- verify: ~5 min (integrity check, inventory exports, status push)
Resuming
Just re-run the notebook. recover() closes any stale running discovery runs,
drops leftover .part files, and each stage skips what's already done:
- backfill: guarded by
kv backfill_done - download: only
discovered/faileddocs withattempts < MAX_ATTEMPTS - extract: only
downloadeddocs with no extracted text yet
Idempotent by design — safe to rerun as often as you like.
What "done" looks like
Verify prints PASS {checks} and, when every queue is empty:
- notebook verify output includes
done: True status.jsonon Drive contains"done": true
A session that merely hit its budget is not "done" — it stops early with
status.json showing the phase and remaining queues. That is normal mid-flight
state, not an error.
Failure modes
- Colab disconnect mid-run → just re-run.
recover()heals stale rows;.partfiles are dropped; partial downloads retry from the queue. - Drive quota exceeded → verify/checkpoint push fails; the run continues locally in the session but the store lags. Check Drive quota, free space, re-run.
- 5xx / connection storms →
api_downloadreturns an error tuple andstage_downloadbacks offmin(60, 5*2^attempts)seconds, retrying up toMAX_ATTEMPTSbefore marking the docfailed. - Search shard hits the 10k ceiling →
api_search_pagesraises; the planner sub-shards that seed (YYYY-CA-N) on the next run. integritycheck not "ok" → verify printsFAIL; the DB is corrupted and should be restored from the last resume tar before trusting anything.
Config knobs (top code cell CFG)
| Key | Default | Meaning |
|---|---|---|
BUDGET_SECONDS |
10 * 3600 | max wall-clock per session (env KY_BUDGET) |
SHUTDOWN_SECONDS |
900 | stop stages this long before the deadline so the session can finish cleanly |
PACING_SEARCH / PACING_DL |
(0.3,0.7) / (0.3,0.5) | random sleep range between API calls (polite throttling) |
MAX_ATTEMPTS |
4 | download retries before failed |
MAX_PDF_BYTES |
200 MiB | size cap; oversized (HTTP 413) fails permanently, no retry/backoff |
MAX_EXTRACT_ATTEMPTS |
3 | PDF-extraction retries before skip |
RETRY_HARD |
0 (env KY_RETRY_HARD) |
on resume, reset failed docs back to discovered |
PHASES |
discover,backfill,download,extract | comma-separated phases to run (env KY_PHASES) |
DRY_RUN |
0 (env KY_DRY) |
1 = fully offline: discovery/download skip all network |
Safety notes
- The pipeline never deletes on Drive automatically.
pdf_backupis retained; pre-restore move-aside only renames (-pre-restore), never deletes. - Set
DRY_RUN=1(KY_DRY=1) to exercise the whole flow offline with zero network calls — the smoke script (scripts/smoke_local.py) does exactly this. - Live API tests are gated behind
KY_LIVE=1and share a ≤10-request budget; keep them there so the ordinary test suite stays offline and fast.
Local / tests
python3 -m venv .venv && .venv/bin/pip install -r <pinned deps>
.venv/bin/pytest tests -q # 18 offline tests (live ones skipped)
.venv/bin/python scripts/smoke_local.py # full pipeline over the real 98,724-row DB
.venv/bin/python notebook_builder/build.py # regenerate the notebook