What is OCR (Optical Character Recognition)?

OCR is the technology that converts text inside an image or scanned PDF into machine-readable, searchable, editable text.

Optical Character Recognition (OCR) analyzes the pixels of a scanned document or photo, identifies character shapes, and outputs them as Unicode text. Modern OCR engines combine computer vision with language models to recover words, paragraphs, tables, and reading order. OCR is the bridge between an image-based PDF and a searchable, editable document.

Also known as

  • Optical Character Recognition
  • text recognition

Related PDF tools

Related terms

  • ICR (Intelligent Character Recognition) — ICR is an advanced form of OCR that recognizes handwritten or cursive characters using machine learning.
  • Searchable PDF — A PDF that contains an invisible text layer behind the scanned image, making the document searchable and selectable.
  • Image-Based PDF — A PDF whose pages are scanned images with no underlying text layer — meaning text cannot be searched or selected.
  • OCR Confidence Score — A 0–100 score an OCR engine assigns to each recognized character or word, indicating how certain it is.

Browse all 40 glossary terms

Category: OCR