<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
    <channel>
        <title>PaddleOCR - Tag - Daily Deep Think</title>
        <link>https://blog.baifan.site/en/tags/paddleocr/</link>
        <description>PaddleOCR - Tag - Daily Deep Think</description>
        <generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>blog@baifan.site (ByF)</managingEditor>
            <webMaster>blog@baifan.site (ByF)</webMaster><lastBuildDate>Tue, 25 Nov 2025 14:00:00 &#43;0800</lastBuildDate><atom:link href="https://blog.baifan.site/en/tags/paddleocr/" rel="self" type="application/rss+xml" /><item>
    <title>PDF2Markdown: Complete Guide to Smart PDF Article Extraction</title>
    <link>https://blog.baifan.site/en/pdf2markdown-intelligent-pdf-extraction-tool/</link>
    <pubDate>Tue, 25 Nov 2025 14:00:00 &#43;0800</pubDate><author>
                    <name>ByF</name>
                </author><guid>https://blog.baifan.site/en/pdf2markdown-intelligent-pdf-extraction-tool/</guid>
    <description><![CDATA[<div class="featured-image">
                <img src="/pictures/note/pdf2markdown-featured.svg" referrerpolicy="no-referrer">
            </div><h1 id="pdf2markdown---an-intelligent-article-extraction-tool-for-large-pdf-documents" class="headerLink">
    <a href="#pdf2markdown---an-intelligent-article-extraction-tool-for-large-pdf-documents" class="header-mark"></a>PDF2Markdown - An Intelligent Article Extraction Tool for Large PDF Documents</h1><p><a href="https://python.org" target="_blank" rel="noopener noreferrer"><img class="tw-inline" loading="lazy" src='/python-3.13+-blue_1423277123478464751.svg'   alt="Python Version"  ></a>
<a href="https://github.com/astral-sh/ruff" target="_blank" rel="noopener noreferrer"><img class="tw-inline" loading="lazy" src='/code%20style-ruff-green_7678052461911854896.svg'   alt="Code Style"  ></a>
<a href="https://mypy.readthedocs.io/" target="_blank" rel="noopener noreferrer"><img class="tw-inline" loading="lazy" src='/type%20checking-mypy-blue_13250137686053437545.svg'   alt="Type Checking"  ></a>
<a href="LICENSE" rel=""><img class="tw-inline" loading="lazy" src='/license-MIT-green_10730612941528799645.svg'   alt="License"  ></a></p>
<h2 id="project-overview" class="headerLink">
    <a href="#project-overview" class="header-mark"></a>Project Overview</h2><p>PDF2Markdown is an intelligent content extraction tool built specifically for large scanned PDF files. It combines traditional OCR technology with modern AI large language models to intelligently extract pure article content from documents, automatically filtering out non-article elements such as images and tables. Mixed Chinese-English documents are fully supported.</p>]]></description>
</item>
</channel>
</rss>
