OCR AI
PDF Extractor
- gptpdf - 使用 OpenAI API 提取 PDF 內容,輸出為 Markdown 格式。
- omniparse - PDF to Markdown
- PDF-Extract-Kit - Layout Detection, Formula Detection, Formula Recognition
- Marker - Marker converts PDF to markdown quickly and accurately.
- Mathpix (cloud)
- tabled - 提取表格內容
- MarkItDown - Microsoft 開發的各種類型檔案轉換成 Markdown 格式,支援指令與 Python API 兩種方式
- MinerU - 一站式開源高品質資料提取工具,將PDF轉換成Markdown和JSON格式。
- OpenDataLoader PDF - PDF parser for AI data extraction — Extract Markdown, JSON (with bounding boxes), and HTML from any PDF.
OCR
- OCRFlux is a multimodal large language model based toolkit for converting PDFs and images into clean, readable, plain Markdown text.
dots.ocr
是一款強大的多語言文件解析器,能在單一視覺語言模型中整合版面檢測與內容識別功能,同時維持良好的閱讀順序。儘管其基礎模型僅採用精簡的 1.7B參數大型語言模型架構,仍能達到頂尖技術水準(SOTA)的表現。
- GitHub: https://github.com/rednote-hilab/dots.ocr
- YT: https://www.youtube.com/watch?v=t_8ZgUIgnLo
- 🚀重磅开源!本地部署1.7B参数超强OCR大模型dots.ocr!超越GPT-4o和olmOCR!结构化精准提取复杂PDF扫描件!完美识别中英文文档、模糊扫描件与复杂表格!文档解析准确率接近100%! - AI超元域的博客
DeepSeek-OCR
只有3B參數,採用「光學上下文壓縮」技術,將文字視為圖像,利用視覺token進行壓縮和理解,把長文字轉換成圖像進行處理,極大地降低了計算資源消耗。
- GitHub: https://github.com/deepseek-ai/DeepSeek-OCR
- GitHub: https://github.com/deepseek-ai/DeepSeek-OCR-2
- HF: https://huggingface.co/deepseek-ai/DeepSeek-OCR
- HF: https://huggingface.co/deepseek-ai/DeepSeek-OCR-2
- YT: https://www.youtube.com/watch?v=9oICqbApvTg
- 🚀DeepSeek又放大招!这个OCR模型让文档识别效率倍增!本地部署+客观实测DeepSeek-OCR!OCR识别准确率97%,支持100+语言,每天处理3300万页文档的开源大模型! - AI超元域的博客
olmOCR - 支持結構化精准提取復雜PDF文件內容!完美識別中英文文檔、模糊掃描件與復雜表格!本地部署與實際測試全過程!醫療法律行業必備!輕松應對企業級PDF批量轉換需求
- GH: https://github.com/allenai/olmocr
- Demo: https://olmocr.allenai.org/
- 🚀本地部署最强OCR大模型olmOCR!支持结构化精准提取复杂PDF文件内容!完美识别中英文文档、模糊扫描件与复杂表格!本地部署与实际测试全过程!医疗法律行业必备!轻松应对企业级PDF批量转换需求! - AI超元域的博客
- YT: https://www.youtube.com/watch?v=XF3Q_ZjwfaI
- GitHub: https://github.com/PaddlePaddle/PaddleOCR
- HF: https://huggingface.co/spaces/PaddlePaddle/PaddleOCR-VL-1.5_Online_Demo