Community · E2 · artifact verified

Recover OCR marker formatting with typed choices

OCR flattens superscripts: a footnote star, an endnote number and a unit power land in the stream as look-alike tokens. A regex over-finds the suspects, then Jev classifies each - footnote, citation or unit - as a typed choice so formatting can be restored.

01 · Role in the system

What Jev does here

The pipeline is two-stage by design: a regex pass over-finds every token that could be a lost marker - fox*, 2m2, tortoises1 - casting a deliberately wide net, and Jev then classifies each suspect with a typed choice question: footnote, citation or plain text. Because the model only decides among fixed labels, the restoration step is mechanical: strip the false positives, re-superscript the true markers, and leave an audit trail of what was changed and why. The script is deliberately small - one Python file - making the pattern easy to lift into any extraction pipeline that suffers the same flattening.

02 · Control boundary

Where Jev sits

Jev as the disambiguator behind a wide regex net: over-find every marker-shaped token, classify each with one typed choice among fixed labels, and restore formatting mechanically from the calibrated answers.

Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.

03 · Known limits

What this evidence does not prove

  • English-centric marker patterns; other scripts need new regex and criteria.
  • Single-file script; no packaged API or test suite.
  • 0 stars; precision and recall are not formally measured.

04 · Attribution

Public sources

This is a Community record: the project was published by a third-party community author.