python-3.x - 从带有图像的扫描pdf中提取文本？-6ren

python-3.x - 从带有图像的扫描pdf中提取文本？

转载作者：行者123 更新时间：2023-12-03 16:26:44

25

4

我尝试从计算机创建的pdf中提取文本，并且可以正常工作，但是我无法从扫描的pdf which you can find here中提取文本，其中包含图像和诸如此类的几页内容:

这是我使用的代码:

# libraries
## split
from PyPDF2 import PdfFileWriter, PdfFileReader
## read 
import sys
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.pdfpage import PDFPage
from pdfminer.converter import XMLConverter, HTMLConverter, TextConverter
from pdfminer.layout import LAParams
import io
# remove files
import os

# split in case there is several pages
def pdfspliter(filename):
    inputpdf = PdfFileReader(open(filename, "rb"))

    for i in range(inputpdf.numPages):
        output = PdfFileWriter()
        output.addPage(inputpdf.getPage(i))
        with open("document-page%s.pdf" % i, "wb") as outputStream:
            output.write(outputStream)
        pdfparser("document-page%s.pdf" % i)
        os.remove("document-page%s.pdf" % i)

# read a given page
def pdfparser(data):

    fp = open(data, 'rb')
    rsrcmgr = PDFResourceManager()
    retstr = io.StringIO()
    codec = 'utf-8'
    laparams = LAParams()
    device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams)
    # Create a PDF interpreter object.
    interpreter = PDFPageInterpreter(rsrcmgr, device)
    # Process each page contained in the document.

    for page in PDFPage.get_pages(fp):
        interpreter.process_page(page)
        data =  retstr.getvalue()

    print(data)

if __name__ == '__main__':
    filename = sys.argv[1]
    pdfspliter(filename)

您可以帮助从此类文件中提取文本吗？

使用Tesseract OCR更新

我尝试使用带有Python的Tesseract OCR，它提取了pdf文本的某些页面，但确实花费时间，而且似乎停在了一点:

# import the necessary packages
from PIL import Image
import pytesseract
import argparse
import cv2
import os
## split
from PyPDF2 import PdfFileWriter, PdfFileReader
# remove
import sys
# 
from pdf2image import convert_from_path
# import all files with a name
import glob

# functions
def pdfspliterimager(filename):
    inputpdf = PdfFileReader(open(filename, "rb"))
    for i in range(inputpdf.numPages):
        output = PdfFileWriter()
        output.addPage(inputpdf.getPage(i))
        with open("document-page%s.pdf" % i, "wb") as outputStream:
            output.write(outputStream)
        pages = convert_from_path("document-page%s.pdf" % i, 500)
        for page in pages:
            page.save('out%s.jpg'%i, 'JPEG')

        os.remove("document-page%s.pdf" % i)

# construct the argument parse and parse the arguments
ap = argparse.ArgumentParser()
ap.add_argument("-i", "--image", required=True,
    help="path to input image to be OCR'd")
ap.add_argument("-p", "--preprocess", type=str, default="thresh",
    help="type of preprocessing to be done")
args = vars(ap.parse_args())

# we test if it is a pdf
image_path = args["image"]
# if it is a pdf we convert it to an image
if image_path.endswith('.pdf'):
    pdfspliterimager(image_path)

# for all files with out in their name
file_names = glob.glob("out*")
for file_name in file_names:
    # load the image and convert it to grayscale
    image = cv2.imread(file_name)
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

    # check to see if we should apply thresholding to preprocess the
    # image
    if args["preprocess"] == "thresh":
        gray = cv2.threshold(gray, 0, 255,
            cv2.THRESH_BINARY | cv2.THRESH_OTSU)[1]

    # make a check to see if median blurring should be done to remove
    # noise
    elif args["preprocess"] == "blur":
        gray = cv2.medianBlur(gray, 3)

    # write the grayscale image to disk as a temporary file so we can
    # apply OCR to it
    filename = "{}.png".format(os.getpid())
    cv2.imwrite(filename, gray)

    # load the image as a PIL/Pillow image, apply OCR, and then delete
    # the temporary file
    text = pytesseract.image_to_string(Image.open(filename))
    os.remove(filename)
    print(text)

    # show the output images
    cv2.imshow("Image", image)
    cv2.imshow("Output", gray)
    cv2.waitKey(0)

最佳答案

使用Python 对PDF文件进行 OCR

import os
import io
from PIL import Image
import pytesseract
from wand.image import Image as wi
import gc

def Get_text_from_image(pdf_path):
 pdf=wi(filename=pdf_path,resolution=300)
 pdfImg=pdf.convert('jpeg')
 imgBlobs=[]
 extracted_text=[]
 for img in pdfImg.sequence:
 page=wi(image=img)
 imgBlobs.append(page.make_blob('jpeg'))
 for imgBlob in imgBlobs:
 im=Image.open(io.BytesIO(imgBlob))
 text=pytesseract.image_to_string(im,lang='eng')
 extracted_text.append(text)
 return ([i.replace("\n","") for i in extracted_text])

我做了一个小修改
下面的代码在破坏图像序列的代码末尾按顺序将PDF的所有页面转换为图像，因为它占用了大量的内存来处理
def Get_text_from_image(pdf_path): import pytesseract,io,gc from PIL import Image from wand.image import Image as wi import gc """ Extracting text content from Image """ pdf=wi(filename=pdf_path,resolution=300) pdfImg=pdf.convert('jpeg') imgBlobs=[] extracted_text=[] try: for img in pdfImg.sequence: page=wi(image=img) imgBlobs.append(page.make_blob('jpeg')) for i in range(0,5): [gc.collect() for i in range(0,10)] for imgBlob in imgBlobs: im=Image.open(io.BytesIO(imgBlob)) text=pytesseract.image_to_string(im,lang='eng') text = text.replace(r"\n", " ") extracted_text.append(text) for i in range(0,5): [gc.collect() for i in range(0,10)] return (''.join([i.replace("\n"," ").replace("\n\n"," ") for i in extracted_text])) [gc.collect() for i in range(0,10)] finally: [gc.collect() for i in range(0,10)] img.destroy()

关于python-3.x - 从带有图像的扫描pdf中提取文本？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/53033077/

25

4

0

文章推荐： r - 如何以 Shiny 的百分比显示 "numericInput"中的值？

文章推荐： batch-file - CMD-批处理文件简单随机

文章推荐： testing - 如何使用 Jest 来测试iFrame的内容

javascript - 图像->JSON->.Net 图像
我正在尝试学习 Knockout 并尝试创建一个照片 uploader 。我已成功将一些图像存储在数组中。现在我想回帖。在我的 knockout 码(Javascript)中，我这样做: 我在 Jav
php - 如何在mysql中添加文本+图像+文本+图像+ ...？
我正在使用 php 编写脚本。我的典型问题是如何在 mysql 中添加一个有很多替代文本和图像的问题。想象一下有机化学中具有苯结构的描述。最有效的方法是什么？据我所知，如果我有一个图像，我可以在数据
html - Bootstrap 图像/按钮/图像
我在两个图像之间有一个按钮，我想将按钮居中到图像高度。有人可以帮帮我吗？ Entrar
javascript - 从中心点旋转全尺寸 Canvas 图像，而不移动其他 Canvas 图像
下面的代码示例可以在这里查看 - http://dev.touch-akl.com/celebtrations/ 我一直在尝试做的是在 Canvas 上绘制 2 个图像(发光，然后耀斑。这些图像的链接
javascript - 为什么相同的 JS 代码不适用于第二个帖子/图像，但适用于第一个帖子/图像？
请检查此https://jsfiddle.net/rhbwpn19/4/ 图像预览对于第一篇帖子工作正常，但对于其他帖子则不然。我应该在这里改变什么？ function readURL(input)
javascript - HTML Canvas 图像 - 在页面中绘制多 Canvas 图像
我对 Canvas 有疑问。我可以用单个图像绘制 Canvas ，但我不能用单独的图像绘制每个 Canvas 。- 如果数据只有一个图像，它工作正常，但数据有多个图像，它不工作你能帮帮我吗？ va
ios - 查找指定文件的扩展文件类型(JPEG 图像、TIFF 图像、...) objective c
我的问题很简单。如何获取 UIImage 的扩展类型？我只能将图像作为 UIImage 而不是它的名称。图像可以是静态的，也可以从手机图库甚至文件路径中获取。如果有人可以为此提供一点帮助，将不胜感激。
image-processing - 如何将 SVG 图像 "paths"转换为单独的 PNG 图像？
我有一个包含 67 个独立路径的 SVG 图像。是否有任何库/教程可以为每个路径创建单独的光栅图像(例如 PNG)，并可能根据路径 ID 命名它们？最佳答案谢谢大家。我最终使用了两个答案的组合。
javascript - 使一个(图像)向右移动并旋转 45 度，同时将鼠标悬停在另一个(图像)上
我想将鼠标悬停在一张图片(音乐专辑)上，然后播放一张唱片，所以我希望它向右移动并旋转一点，当它悬停时我希望它恢复正常动画片。它已经可以向右移动，但我无法让它随之旋转。我喜欢让它尽可能简单，因为我不是编
ios - Retina iOS 设备不显示@2X 图像，它显示 1X 图像
Retina iOS 设备不显示@2X 图像，它显示 1X 图像。我正在使用 Xcode 4.2.1 Build 4D502，该应用程序的目标是 iOS 5。我创建了一个测试应用(主/细节)并添加
javascript - 如何将 HTML CSS 图像 slider 转换为 Angular 图像 slider ？
我正在尝试从头开始以 Angular 实现图像 slider ，并尝试复制 w3school基于图像 slider 。下面我尝试用 Angular 实现，谁能指导我如何使用 Angular 实现？
c++ - 如何从 opencv 中的 RGB 图像(3 channel 图像)访问图像数据
我正在尝试获取图像的图像数据，其中 w= 图像宽度，h = 图像高度 for (int i = x; i imageData[pos]>0) //Taking data (here is the pr
php - HTML (或 JS 图像)与通过 PHP 的内联图像(base64_encoded 图像)
我的网页最初通过在 javascript 中动态创建图像填充了大约 1000 个缩略图。由于权限问题，我迁移到 suPHP。现在不用标准标签本身我正在通过这个 php 脚本进行检索 $file
python - 将 Python Opencv 图像(numpy 数组)转换为 PyQt QPixmap 图像
我正在尝试将 python opencv 图像转换为 QPixmap。我按照指示显示Page Link我的代码附在下面 img = cv2.imread('test.png')[:,:,::1]/2
python - OpenCV 将图像读取为 3 channel 图像，而 PIL 将图像读取为 1 channel 图像
我试图在这个 Repository 中找出语义分割数据集的 NYU-v2 . 我很难理解图像标签是如何存储的。例如，给定以下图像: 对应的标签图片为: 现在，如果我在 OpenCV 中打开标签图像，
java - 如何循环 8*1 svg 图像(由 java 生成)以获得 8*8 svg 图像？
import java.util.Random; class svg{ public static void main(String[] args){ String f="\"
Android自定义画笔图案/图像
我有一张 8x8 的图片。 (位图 - 可以更改) 我想做的是能够绘制一个形状，给定一个 Path 和 Paint 对象到我的 SurfaceView 上。目前我所能做的就是用纯色填充形状。我怎样才
HTML 图像
要在页面上显示图像，你需要使用源属性（src）。src 指 source 。源属性的值是图像的 URL 地址。定义图像的语法是：在浏览器无法载入图像时，替换文本属性告诉读者她们失去的信息。此
图像&视频编辑工具箱MMEditing安装及使用示例(Inpainting)
**MMEditing是基于PyTorch的图像&视频编辑开源工具箱，支持图像和视频超分辨率(super-resolution)、图像修复(inpainting)、图像抠图(matting)、
来自资源文件的 Qt 图像
我正在尝试通过资源文件将图像插入到我的程序中，如下所示: green.png other files 当我尝试使用 QImage 或 QPixm

首页

博学

6Ren·AI

商城

python-3.x - 从带有图像的扫描pdf中提取文本？