python - 在 python 中使用 opencv 检测低对比度图像中的正方形，以便通过 tesseract 读取-6ren

python - 在 python 中使用 opencv 检测低对比度图像中的正方形，以便通过 tesseract 读取

转载作者：太空宇宙更新时间：2023-11-03 22:23:51

26

4

我想检测像这样的图像中的标签，以便使用 tesseract 提取文本。我尝试了各种阈值组合和使用边缘检测。但是我一次最多只能检测大约一半的标签。这些是我一直试图从中读取标签的一些图像:

enter image description here

所有标签都具有相同的纵横比(宽度是高度的 3.5 倍)，因此我试图找到具有相同纵横比的 minAreaRect 的轮廓。困难的部分是在较浅的背景上处理标签。这是我到目前为止的代码:

from PIL import Image
import pytesseract
import numpy as np
import argparse
import cv2
import os

ap = argparse.ArgumentParser()
ap.add_argument("-i", "--image", required=True,
    help="path to input image to be OCR'd")
args = vars(ap.parse_args())

#function to crop an image to a minAreaRect
def crop_minAreaRect(img, rect):
    # rotate img
    angle = rect[2]
    rows,cols = img.shape[0], img.shape[1]
    M = cv2.getRotationMatrix2D((cols/2,rows/2),angle,1)
    img_rot = cv2.warpAffine(img,M,(cols,rows))

    # rotate bounding box
    rect0 = (rect[0], rect[1], 0.0)
    box = cv2.boxPoints(rect)
    pts = np.int0(cv2.transform(np.array([box]), M))[0] 
    pts[pts < 0] = 0

    # crop
    img_crop = img_rot[pts[1][1]:pts[0][1], 
                       pts[1][0]:pts[2][0]]

    return img_crop




# load image and apply threshold
image = cv2.imread(args["image"])
bw = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
#bw = cv2.threshold(bw, 210, 255, cv2.THRESH_BINARY)[1]
bw = cv2.adaptiveThreshold(bw, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY, 27, 20)
#do edge detection
v = np.median(bw)
sigma = 0.5
lower = int(max(0, (1.0 - sigma) * v))
upper = int(min(255, (1.0 + sigma) * v))
bw = cv2.Canny(bw, lower, upper)
kernel = np.ones((5,5), np.uint8)
bw = cv2.dilate(bw,kernel,iterations=1)

#find contours
image2, contours, hierarchy = cv2.findContours(bw,cv2.RETR_TREE,cv2.CHAIN_APPROX_SIMPLE)
bw = cv2.drawContours(bw,contours,0,(0,0,255),2)
cv2.imwrite("edge.png", bw)

#test which contours have the correct aspect ratio
largestarea = 0.0
passes = []
for contour in contours:
    (x,y),(w,h),a = cv2.minAreaRect(contour)
    if h > 20 and w > 20:
        if h > w:
            maxdim = h
            mindim = w
        else:
            maxdim = w
            mindim = h
        ratio = maxdim/mindim
        print("ratio: {}".format(ratio))
        if (ratio > 3.4 and ratio < 3.6):
            passes.append(contour)
if not passes:
    print "no passes"
    exit()

passboxes = []
i = 1

#crop out each label and attemp to extract text
for ps in passes:
    rect = cv2.minAreaRect(ps)
    bw = crop_minAreaRect(image, rect)
    cv2.imwrite("{}.png".format(i), bw)
    i += 1
    h, w = bw.shape[:2]
    print str(h) + "x" + str(w)
    if w and h:
        bw = cv2.cvtColor(bw, cv2.COLOR_BGR2GRAY)
        bw = cv2.threshold(bw, 50, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)[1]
        cv2.imwrite("output.png", bw)
        im = Image.open("output.png")
        w, h = im.size
        print "W:{} H:{}".format(w,h)
        if h > w:
            print ("rotating")
            im.rotate(90)
            im.save("output.png")
        print pytesseract.image_to_string(Image.open("output.png"))
        im.rotate(180)
        im.save("output.png")
        print pytesseract.image_to_string(Image.open("output.png"))
        box = cv2.boxPoints(cv2.minAreaRect(ps))
        passboxes.append(np.int0(box))
        im.close()

cnts = cv2.drawContours(image,passboxes,0,(0,0,255),2)
cnts = cv2.drawContours(cnts,contours,-1,(255,255,0),2)
cnts = cv2.drawContours(cnts, passes, -1, (0,255,0), 3)
cv2.imwrite("output2.png", image)

我相信我遇到的问题可能是阈值参数。或者我可能把这个复杂化了。

最佳答案

只有带有“A-08337”之类的白色标签？以下代码在两张图片上检测到所有这些:

import numpy as np
import cv2

img = cv2.imread('labels.jpg')

#downscale the image because Canny tends to work better on smaller images
w, h, c = img.shape
resize_coeff = 0.25
img = cv2.resize(img, (int(resize_coeff*h), int(resize_coeff*w)))

#find edges, then contours
canny = cv2.Canny(img, 100, 200)
_, contours, _ = cv2.findContours(canny, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)

#draw the contours, do morphological close operation
#to close possible small gaps, then find contours again on the result
w, h, c = img.shape
blank = np.zeros((w, h)).astype(np.uint8)
cv2.drawContours(blank, contours, -1, 1, 1)
blank = cv2.morphologyEx(blank, cv2.MORPH_CLOSE, np.ones((3, 3), np.uint8))
_, contours, _ = cv2.findContours(blank, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)

#keep only contours of more or less correct area and perimeter
contours = [c for c in contours if 800 < cv2.contourArea(c) < 1600]
contours = [c for c in contours if cv2.arcLength(c, True) < 200]
cv2.drawContours(img, contours, -1, (0, 0, 255), 1)

cv2.imwrite("contours.png", img)

可能通过一些额外的凸性检查，您可以摆脱“Verbatim”等高线等(例如，只保留其面积与凸包面积之间几乎为零的差异的等高线)。

关于python - 在 python 中使用 opencv 检测低对比度图像中的正方形，以便通过 tesseract 读取，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/45202770/

26

4

0

文章推荐： html - 启用或禁用输入文件时从 CSS 更改标签颜色

文章推荐： node.js - Webpack 4 和用于长期缓存的哈希

文章推荐： javascript - 定高容器如何分成80%和20%

css - 三个水平正方形|正方形| |正方形| |正方形|使用 CSS
我试图使用显示正方形(如元素符号正方形) .squares{ list-style-type: square; display:inline; } 但我希望它们是水
html - 嵌套 DIV - 4 个 div(正方形)，其中 1 个有 4 个 div(正方形)，其中 1 个有 4 个 div(正方形)
这是关于在作为 4 个 div 的一部分的 1 个 div 中嵌套 4 个 div(正方形)... 我在包装器中使用 display: flex 并用于包装的元素本身，否则它不会工作对我来说，这感觉
python - 有什么方法可以从图像中提取变形的矩形/正方形？
这是图像，我想填充此矩形或正方形的边缘，以便可以使用轮廓对其进行裁剪。到目前为止，我所做的是我使用了canny边缘检测器来找到边缘，然后使用bitwise_or或将这个矩形填充了一点，但没有完全填充。
c# - 在图片框内动态创建点/正方形
我希望能够在图片框内创建 X x Y 数量的框/圆圈/按钮。完全像 Windows 碎片整理工具。我尝试创建一个布局并不断向其添加按钮或图片框，但速度非常慢，在添加 200 个左右的图片框后它崩溃，
python - 使用skimage python从图像中选择具有特定度量的区域(正方形)
我正在尝试从图像(肺部图像)中提取 3 个区域，这些区域在软组织中显示，每个区域都是具有特定高度和宽度的正方形，例如宽高各10mm，如下图，如图所示，该区域也是均匀的，这意味着它只包含相同的颜色(在
java - 尝试用鼠标移动四边形/正方形 - LWJGL
在我左键单击它后，我试图让一个正方形跟随我的鼠标。当我右键单击时，方 block 应该停止跟随我的鼠标。我的程序检测到我在方 block 内单击，但由于某种原因，它没有根据 Mouse.getDX/
c - 将二维数组划分为所有可能的 nxn 正方形
已经花了几个小时在这上面了(因为我还在学习)，所以也许你们可以帮忙。问题是我无法弄清楚如何将二维数组划分为所有可能的 nxn 正方形。我正在随机化二维数组，可以说它是这样的: 1 0 1 0 2
facebook - 正方形 Facebook 图片
使用 Graph API，我可以获得小型、大型、中型图片。或者我可以获得小方形图片。但是我怎样才能得到大方形图片呢？有什么服务可以使用吗？最佳答案很简单，我刚发现这个。例子， https://
html - CSS
正方形
我是 HTML 和 CSS 的新手。尝试创建 3 x 3 正方形“图片”，使用，但无法找到将正方形放在页面中间的简单解决方案，例如中间有九个正方形。如何把所有的方 block 都放在大边框的正方
CSS动画圆-三 Angular -正方形
我正在玩弄 CSS 动画以获得乐趣。我有限的经验阻碍了这一进程。下面的脚本将圆形转换为三 Angular 形，再转换为正方形，然后反转。然而，圆形和三 Angular 形之间的动画有一个小错误。我希
html - 如何在网格中正确排列(不同的)正方形？
我的标准布局(最小宽度 1024 像素)有 4 行。第一个和最后一个有 6 个正方形，中间有两个组合正方形。但是第三行的第一个方 block 不见了。我没有使用不同的 CSS 设置。我试过 clear
javascript - 带正方形的特殊布局 div 正方形
关闭。这个问题需要更多focused .它目前不接受答案。想改进这个问题吗？更新问题，使其只关注一个问题 editing this post . 关闭 7 年前。 Improve this q
html - CSS 图像定位(正方形)
我的网站上有 4 张图片，我试图对其进行定位，以便它们在我的 DIV 中形成一个相等的正方形，但它看起来像一条由 4 张图片组成的垂直线。我希望它看起来像 2 个图像的 2 条垂直线，彼此相邻，使其成
OpenCV 正方形 : filtering output
这是方 block 检测示例的输出我的问题是过滤这个方 block 第一个问题是它为同一区域绘制多条线；第二个是我只需要检测对象而不是所有图像。另一个问题是我必须只取除所有图像之外的最大对象。检
python - 矢量化行进立方体(正方形) - 将直线连接成曲线
我正在绘制一个带有移动立方体(正方形，因为它是 2d)算法的元球。一切都很好，但我想将其作为矢量对象获取。到目前为止，我已经从每个事件方 block 中得到一两条矢量线，将它们保存在列表线中。换句话
android - 如何用手指移动 OpenGL 正方形？
实际上，我有一个适用于 Android 1.5 的应用程序，其中包含一个 GLSurfaceView 类，它在屏幕上显示一个简单的方形多边形。我想学习如何添加一个新功能，即移动用手指触摸的方 blo
java - 制作一个 JPanel 正方形
如果我有一个包含多个子组件的 JPanel，我该如何使 JPanel 保持正方形，而不管其父组件的大小如何调整？我尝试了以下代码的变体，但它不会导致子组件也变成正方形。 public void pai
android - 如何保持 TextView 正方形？
我找到了 this answer ，它确保 ImageView 的宽高比得以保留。我如何使用带有可绘制背景的 TextView 来做到这一点？我有这个 TextView: 这是我的背
html - 基于高度的 CSS 正方形
这个问题在这里已经有了答案: Maintain aspect ratio of a div according to height [duplicate] (1 个回答) 关闭 8 年前。是否可以
html - 连接圆 Angular 正方形
如何创建 div Logo ，如下图所示: 这是我在 JsFiddle 中创建的主要问题是如何将两个形状如下图的盒子连接起来，有人可以提出建议吗？ body,html { width: 100%

首页

博学

6Ren·AI

商城

python - 在 python 中使用 opencv 检测低对比度图像中的正方形，以便通过 tesseract 读取