Image to Text
Free OCR that pulls the text out of an image, built on the open-source Tesseract engine, with no signup.
Overview
Image to Text is an OCR tool: it turns a picture of words into text you can copy. The site states it runs on Tesseract, the open-source engine maintained by Google, so the accuracy is Tesseract's. It takes JPG, PNG, GIF, JFIF, HEIC and PDF, and reads more than twenty languages, including Chinese, Japanese and Arabic. Simple OCR returns plain text. Formatted Text keeps tables, lists and headings, and is paid only.
It runs in a browser. Guests get 30 images a day, 2 per submission and 7 MB per file, with ads and a captcha. Registering raises that to 50 a day and 5 per submission. Paid plans are monthly subscriptions that give credits, take files up to 30 MB and remove the ads.
Key features
When to use
Use Image to Text if:
- You photographed a page, a sign, a receipt or a whiteboard and want the text rather than the picture.
- Your document is not in English. It covers more than twenty languages, including non-Latin scripts.
- Your file is a PDF or an iPhone HEIC, both of which it takes directly without converting first.
- You want to extract text without signing up, which guests can do 30 times a day, 2 images per submission.
- You have a large set of images, which the separate batch tool takes 1000 at a time.
- You are happy with a well-known free engine rather than paying for a proprietary one.
When NOT to use
Skip Image to Text if any of these describe you:
- Your document is sensitive. The site makes claims about your data that you cannot verify from outside, so read them before uploading a contract, an ID or medical paperwork.
- You extract text often. Free use stops at 30 images a day for guests and 50 once you register.
- Ads and captchas appear while you work. Both sit on the free tier and paying is what removes them.
- The writing is handwritten and messy. Difficult handwriting still needs checking line by line.
- You want to compare accuracy on your own documents. Other image to text tools such as ImgOCR read the same page differently.
Frequently asked questions
Yes, and you can start as a guest without registering. Guests get 30 images a day, 2 images per submission and 7 MB per file, with ads and a captcha. Registering raises that to 50 images a day and 5 per submission. Paid plans are monthly subscriptions that give you credits: Simple OCR costs one credit per image and Formatted Text costs ten. Paying also takes files up to 30 MB and removes the ads.
Simple OCR gives you the words as plain text. Formatted Text keeps the layout of the page: tables, lists and headings. Formatted Text is a paid mode, and switching to it as a free user brings up the subscription plans. Results save as TXT or DOCX, and the formatted mode can also return HTML or Markdown.
The site states it uses Tesseract, an open-source engine originally developed at Hewlett-Packard and now funded and maintained by Google. It means the recognition quality is Tesseract's, which is used in many products, and it means the same engine is free for anyone to run. What you are paying for on the paid plans is convenience, speed, higher limits and the formatted mode, not a better model.
More than twenty, and the list goes well beyond European ones. It covers Chinese, Japanese, Korean, Arabic, Thai and Georgian among others, so scripts that do not use the Latin alphabet are handled rather than ignored. Setting the language before converting tells the engine which script to expect.
It will try, and the output needs checking line by line. Handwriting is much harder for any OCR engine than printed text, because printed letters are consistent and handwritten ones are not. Neat, clearly separated writing does far better than a hurried scrawl. The output is a first draft to correct rather than a transcript to rely on.
The site states that no data is transmitted or stored. That is its own claim and not something anyone outside can verify. If the document is sensitive, read the current terms yourself first, or use OCR that runs on your own machine.