Engine · PDF tools

Searchable PDF

Follow a scanned page and its saved text boxes into a new PDF whose text can be selected and searched.

Workflow

  1. Read PDF and sidecarLoad the source PDF and its stored blocks. A missing sidecar or stored failed read is refused. The source is not overwritten; the output must be a new path.
  2. Check text and geometryInspect actual PDF text presence, page count, block bounds and supported text encoding. A sidecar label cannot override the bytes of a born-digital PDF. Multi-page blocks need page numbers that this contract does not yet carry.
  3. REFUSED · no accepted outputDo not guess a text position, flatten a rotation or place all multi-page blocks on page one. Return the named refusal. Already born-digital input is also refused because adding an overlay would duplicate its text.
  4. Place invisible textPosition the sidecar text on the supported page while preserving the visual source. Invisible text supports selection and search; it does not alter the source photograph or improve the OCR itself.
  5. Read the copy backUse the independent text extractor to compare the new reading order with the intended text. If verification fails, the operation must not be reported as successful. Inspect refusal handling before reusing a failed output path.
  6. REFUSED · verification failedA scrambled reading order or unavailable verifier is a named refusal. A file being present is not proof that this operation completed successfully.
  7. SEARCHABLE · new PDFReturn SearchableResult with output path and block and byte counts after successful verification. This is a new artifact associated with the source, not a replacement of the original.

Failure boundaries

  • No stored readRefused: No accepted output
  • Multiple/rotated pagesRefused: No guessed placement
  • Already digital PDFRefused: No duplicated text overlay
  • Extractor absent/order differsVerification refuses: Do not offer file as successful

engine.pdf.searchable