Data Rooms Providers Find a data room
VDR glossary · Features

What is optical character recognition (OCR)?

Definition

Optical character recognition (OCR): Technology that turns images of text, such as scanned contracts or photographed pages, into machine-readable words so they can be searched, copied for analysis or redacted.

How it works in a data room

When a scanned PDF or image is uploaded, the platform runs it through an OCR engine that recognizes letters and lays an invisible text layer over the picture of each page. The document still looks like the original scan, but its words are now available to full-text search, automated categorization and redaction tools. Quality depends on the scan: clean typed pages convert almost perfectly, while faded faxes, handwriting and stamps produce errors. Some providers process OCR automatically on upload; others charge for it or limit it to certain plans, which is worth asking about.

Why it matters in a deal

Older companies hold years of paper records: leases signed in the 1990s, title deeds, signed board resolutions. In a sale these are scanned in bulk, and without OCR they become blind images that no search can reach. That slows review and raises the chance that a key clause is missed. OCR also matters for privacy work, because tools that find names or account numbers for redaction can only see text, not pixels.

Example

A regional utility in Australia is selling a network of pipelines. Many easement agreements exist only as paper copies stored in a records center. The seller scans about 9,000 pages and uploads them. OCR lets bidders search for landowner names and parcel numbers, and the sell-side team runs a personal data scan before release. For more on how infrastructure deals use rooms, see the energy and infrastructure guide.

Related terms