Open Source Hacker News (LLM)

OCR It – pull text out of un-copyable documents for your LLM

chrome extensionOCRTesseractLLM tooling

OCR It is a Chrome extension designed to extract text from documents trapped in viewers that don't allow selection — scanned books, slide decks, PDFs, or embedded readers. The user pins a capture region once, then presses a hotkey on each page to screenshot that exact rectangle, run OCR locally, and append the text to a running transcript. An auto-run mode (⌥⇧A) captures, turns the page, and repeats until the document ends, making it possible to convert an entire book into a text file that can be handed to Claude or ChatGPT for summarization, search, or Q&A.

OCR runs entirely locally using a bundled Tesseract build. No API key, no network, no images leave the machine — the extension makes no outbound requests at all. Installation is straightforward: clone the repo, load it as an unpacked extension in developer mode, pin it, and verify the hotkeys in chrome://extensions/shortcuts. There is no build step; npm install is only for tests or re-vendoring Tesseract. The extension asks for no site access at install; single captures use activeTab, while auto-run and cross-origin iframe page-turning require a durable grant via an Allow button in the popup.

Usage involves four steps: (1) pin the region with ⌥⇧R and drag a box, with fine-tuning via arrow keys or handles; (2) capture with ⌥⇧S, which screenshots immediately and runs OCR in the background so captures queue without waiting; (3) set up a next-page control and let ⌥⇧A run fully automatically — capture, turn, repeat until the document ends (Esc stops); (4) export — every page is listed with a thumbnail of what was cropped to catch region drift, text is editable in place, and Copy all / Download .txt emit pages in order with --- page N --- separators. A DUPLICATE marker flags pages with identical text, usually indicating the page didn't actually turn.

The extension can also turn pages automatically if you enable that and pick a control (such as a next-page button or keyboard shortcut) — you hit 'Pick control' and click the element, then it clicks it after each capture. Everything needed is committed to the repo, and the toolbar icon doubles as a page counter.

Read original →

← Back to home