OCR

Published entries across all sections carrying the “OCR” tag, newest first by publication date on this site.

3 entries

  1. APIsProducts

    CapSolver

    An AI and computer-vision powered automated CAPTCHA recognition service accelerating web data extraction pipelines.

    #OCR#Browser automation

  2. DocumentsGitHub

    MonkeyOCR: a small, multimodal document-parsing model

    MonkeyOCR is an open-source document-parsing tool built on a lightweight multimodal LLM that parses PDFs in three stages—layout, recognition, relation—to reconstruct formulas, tables, and reading order into Markdown, with local GPU inference.

    #OCR#Extraction

  3. DocumentsGitHub

    BabelDOC: A Bilingual Translation Library for PDF Papers

    An open-source translation library that parses and reflows English PDF papers into a bilingual side-by-side document, using an LLM for translation while preserving formulas, tables, and other original layout as closely as possible; usable as a CLI or embedded in other programs.

    #Translation#OCR