UNDT Processing Functions
Button Description
Retrieves UNDT English and French metadata for all years and builds the UNDT table entries
Retrieves UNDT English and French metadata for this year and builds the UNDT table entries
Retrieves UNDT English and French metadata for all years
UNDT Process Management Tools
Stop UNDT build process
UNAT Processing Functions
Button Description
Retrieves UNAT English and French metadata for all years and builds the UNAT table entries
Retrieves UNAT English and French metadata for this year and builds the UNAT table entries
Processes UNAT data and Downloads/OCR's as needed(
UNAT Process Management Tools
Stop UNAT build process
ILOAT Processing Functions
Button Description
Processes ILOAT data and Downloads/OCR's as needed
Finds sessions ILO has published that aren't harvested yet, and pulls them (rate-limited) into the scratch area - never touches production
Same as above, capped at 5 documents - a quick sanity check before running a full session's backlog
Reattach to the background harvest (running/complete/failed - runs on the server regardless of this page being open)
Copies sessions/documents already harvested into scratch into the real production database and corpus - only after you've tested them
Read-only cross-check of our ILOSessions ranges against what ILO's site currently lists - HTML pages only, no PDF/document hits
Same as above, capped at 5 sessions - a quick sanity check before checking every session on ILO's site
Reattach to the background consistency check and see any discrepancies found so far
ILOAT Process Management Tools
Stop ILOAT build process
Results:
Ready
System Testing Functions
Button Description
Test collection info update end editing
Display connection pool info
See which files need downloading and ocring
Fetches one ILOAT judgment (EN) into a scratch temp file - confirms scrape.do + ilo.org's URL pattern still work
Fetches one UNAT judgment into a scratch temp file
Fetches one UNDT judgment into a scratch temp file
Re-OCR - corrects documents OCR'd with the wrong tesseract language model before this fix (French documents were OCR'd as English). Overwrites existing text output in place; does not re-download anything.
Runs at low priority in the background - can take hours for the full corpus. Check status any time, even after navigating away.
Test Completeness of Collection - read-only gap report: missing database rows (ILOAT only, via its session case-range), missing PDFs, missing OCR text.
Runs synchronously - just filesystem/database checks, no network calls or OCR
Language Tag Check - flags documents whose actual OCR'd text doesn't match their tagged language (e.g. a French document mistakenly tagged English).
"Flag" resets both language versions of each mismatched case for redownload/re-OCR and clears their current (wrong) search index entries - run "Check Only" first to see the scope

Incremental Update

TBD

Testing

Simple function testing