Guide · 3 min read · Updated 8 October 2026
How to copy text from a scanned PDF or a photo (Georgian, English, Russian)
You can’t select the text in a PDF, or all you have is a photo of a page. To copy, edit or search that text you need OCR — optical character recognition, which reads the letters from the picture. This guide shows how to tell whether you need it, how to get the best result, and what to do with old Georgian files that come out as nonsense.
First: is it a scan or real text?
Open the PDF in your browser or a PDF viewer and try to select a sentence with the mouse.
- You can select the text. The PDF already contains real text, so you don’t need OCR. Use PDF to text for plain text or PDF to Word for an editable document.
- You can’t select anything, or the whole page highlights as one block. It is a picture of text — a scan or a photo — and OCR is the right tool.
Read a photo or scanned PDF with OCR
- Open Image to text and add photos, screenshots or a scanned PDF (up to 50 pages at a time).
- Tick the languages that are really in the text. Georgian and English are on by default; Russian is available too, and mixed pages work.
- Press Extract text. For a PDF, each page is turned into an image and read, and the text is labelled by page.
- Copy the result or download it as a .txt file.
The first time you use it, the OCR engine and the language models are downloaded to your browser (5–10 MB, then cached). After that the recognition runs on your device, so IDs, contracts and other private papers are never uploaded.
How to get a better result
OCR is only as good as the picture it reads. A few habits make a big difference:
- Shoot straight on. Hold the phone parallel to the page, not at an angle, and avoid curved pages near a book’s spine.
- Use even light. Shadows across the text are the most common cause of mistakes. Daylight from a window works well.
- Scan at 300 DPI if you have a scanner. Lower resolutions blur small letters, and far higher ones only make the file slow.
- Crop out the clutter. Remove borders, fingers and the table with Crop image before reading.
- Tick only the languages you need. A page in Georgian and English reads better with exactly those two selected than with every language on.
Printed text reads far better than handwriting or decorative fonts, which are hard for any OCR engine.
Proofread what you copied
Even clean OCR makes small mistakes. Read the result against the original once, and pay most attention to numbers, names, dates and amounts — a swapped digit looks like normal text and is easy to miss. Tables come out as lines of text, not as a table, so you will have to rebuild them.
Old Georgian PDFs that turn into Latin letters
Many Georgian documents made years ago used legacy fonts such as AcadNusx. The PDF looks right on screen, but the characters stored inside are Latin letters, so copying the text gives Latin letters such as “saqarTvelo” instead of საქართველო. This is not a scan: the text is real, just encoded the old way.
Copy it, paste it into the AcadNusx → Unicode converter and you get normal Georgian text back.
AcadNusx → UnicodeOld Georgian fonts → UnicodeConvert text typed in AcadNusx, AcadMtavr or LitNusx fonts to Unicode Georgian (and back). Fix “saqarTvelo” gibberish in old documents.Which tool for which file?
| You have | Use |
|---|---|
| A photo, screenshot or scan | Image to text |
| A PDF where text can be selected | PDF to text or PDF to Word |
| Georgian text that appears as Latin letters | AcadNusx → Unicode |
Frequently asked questions
Does OCR work for Georgian?
Yes. Toola uses Tesseract’s Georgian model, which reads printed Georgian (Mkhedruli) text. Handwriting and decorative fonts are much harder.
Are my files uploaded?
No. Only the OCR engine and language models are downloaded, once. The recognition runs on your device.
How many pages can I read at once?
Up to 50 pages of a scanned PDF at a time. Split larger files first.
Why does my Georgian PDF show Latin letters?
It was made with an old non-Unicode font such as AcadNusx. Use the AcadNusx → Unicode converter on the copied text.