Read PDFs with RSVP and local OCR
Import selectable or scanned PDFs, check OCR text, correct extraction errors and save your reading position in readbooks.in.
readbooks.in can extract a PDF’s text in your browser and use it in RSVP or normal mode. A PDF can contain selectable text, scanned page images, or both. The right import option depends on what is actually inside the file.
Choose the extraction method
- Selectable text
- If you can select a sentence in a PDF viewer, try Automatic or Selectable text only. readbooks.in uses PDF.js to extract the text without OCR.
- Scanned pages
- Automatic uses Tesseract.js OCR on pages without selectable text. OCR every page is useful when a scanned paragraph sits beside a selectable heading; that heading can otherwise make the page appear to have text.
- Handwritten pages
- The current OCR is intended for printed text. Handwriting may be missed or misread. readbooks.in does not offer dependable handwriting recognition.
Import, review, then read
- Open file import and choose your PDF, JPG or PNG. For an EPUB book, use the same file chooser; EPUB text does not need OCR.
- For OCR, select the language that matches the page. English, Hindi and Spanish models are available. Allow the browser to finish loading the local engine and model.
- Review the extracted page preview and word count. Page numbers and progress show where the import is working. Cancel is available if you picked the wrong file.
- Use edit text to repair errors, remove repeated headers or fix spacing. Compare against the original extraction, then apply your changes.
- Start reading. Use page or chapter navigation for longer documents. Select save to library to keep the text and resume later in the same browser.
When the text looks wrong
Multi-column layouts, tables, footnotes and unusual fonts can produce an unexpected reading order. OCR can confuse similar characters, especially in a blurred, rotated or low-contrast scan. A clean scan of one upright page is easier to inspect than a photograph containing several pages. Review names, dates and numbers before relying on them.
If a page is empty, check the chosen mode. Selectable text only will not recognize an image. If a password is requested, use the document’s password; importing does not remove access restrictions. An extraction failure leaves the current reading available.
What stays on your device?
File contents are processed locally. The site downloads its PDF/OCR/EPUB code and models, but does not upload your document for extraction. Saved text, edits, progress and bookmarks live in browser storage. The original file bytes are not included in a library backup.
The optional offline copy can reopen saved readings. New file extraction and opening a shelf book that has not been saved still need a connection. Keep your original files and export a backup before clearing browser data or moving to a different browser or domain. See Privacy and local data for details.
For open-source implementation details, see PDF.js and Tesseract.js. Continue with How to use an RSVP reader.