macOS CLI for OCR and searchable PDFs using Apple's Vision framework
-
Updated
Jul 3, 2026 - Swift
macOS CLI for OCR and searchable PDFs using Apple's Vision framework
OCR library to extract text & tables from PDF files and images. Convert any image or PDF to CSV / TXT / JSON / Searchable PDF.
Convert scanned PDFs into searchable text locally using Vision LLMs (olmOCR). 100% private, offline, and free. Features a modern Web UI & CLI.
A powerful and user-friendly tool based on OCRmyPDF, offering a seamless GUI for conversion of image-based PDFs into searchable text.
Perform Optical Character Recognition (OCR) on a scanned PDF file containing Arabic text and output a searchable PDF
A Python script that runs Paddle OCR on a possibly unsearchable PDF to make it searchable.
This batch script creates a searchable PDF of a PDF with one or more scanned pages which contain images.
Business-friendly open source OCR pipeline to produce fully searchable PDFs
Self-hosted GPU-accelerated OCR web app — convert scanned PDFs to searchable PDF, Markdown, or Word. Powered by PaddleOCR. Supports Chinese (Traditional & Simplified) and multilingual documents. Single Docker container deployment.
Create a searchable PDF with ALTO-XML and JP2 files.
Extract tables from searchable as well as non-searchable pdf files
NeuroScan-AI is an advanced document-understanding engine built with modern computer vision and OCR pipelines. It performs smart perspective correction, illumination normalization, and adaptive enhancement to transform raw camera captures into clean, searchable, professional-grade documents.
A lightning-fast, privacy-first web app for offline text extraction. Paste (Ctrl+V) or drop any image to instantly generate plain text and a searchable PDF entirely within your browser using Tesseract.js. No server uploads required.
DocuSEO: A platform to make PDFs 100% searchable with AI and optimize for Google Indexing.
Offline Windows desktop app for converting scanned PDFs into searchable PDFs with English and Vietnamese OCR.
Quick proof of concept to perform OCR on images.
Convert scanned PDF documents into searchable, OCR-processed, and PDF compliant files using ocrmypdf, powered by an interactive Streamlit interface. Supports parallel processing to handle large documents efficiently.
NDLOCR-Liteを使い、PDFからテキストを抽出し、検索可能なPDFを生成するWindows向けスタンドアロン日本語OCRアプリ。ページ指定と図中テキストOCRに対応。
PySide6 app to perform batch image/PDF processing and OCR.
Add a description, image, and links to the searchable-pdf topic page so that developers can more easily learn about it.
To associate your repository with the searchable-pdf topic, visit your repo's landing page and select "manage topics."