python - re.findall 和 re.finditer 的区别——Python 2.7 re 模块中的错误？-6ren

python - re.findall 和 re.finditer 的区别——Python 2.7 re 模块中的错误？

转载作者：行者123 更新时间：2023-11-28 22:52:25

24

4

在演示 Python 的正则表达式功能时，我编写了一个小程序来比较 re.search()、re.findall() 和 re 的返回值.finditer()。我知道 re.search() 每行只会找到一个匹配项，而 re.findall() 只会返回匹配的子字符串而不是任何位置信息。然而，令我惊讶的是，匹配的子字符串在这三个函数之间可以不同。

代码(available on GitHub):

#! /usr/bin/env python
# -*- coding: utf-8 -*-

# License: CC-BY-NC-SA 3.0

import re
import codecs

# download kate_chopin_the_awakening_and_other_short_stories.txt
# from Project Gutenberg:
# http://www.gutenberg.org/ebooks/160.txt.utf-8
# with wget:
# wget http://www.gutenberg.org/ebooks/160.txt.utf-8 -O kate_chopin_the_awakening_and_other_short_stories.txt


# match for something o'clock, with valid numerical time or
# any English word with proper capitalization

oclock = re.compile(r"""
                    (
                          [A-Z]?[a-z]+ # word mit max. 1 capital letter
                        | 1[012]       # 10,11,12
                        | [1-9]        # 1,2,3,5,6,7,8,9
                    )
                    \s
                    o'clock""",
                    re.VERBOSE)

path = "kate_chopin_the_awakening_and_other_short_stories.txt"

print
print "re.search()"
print
print u"{:>6} {:>6} {:>6}\t{}".format("Line","Start","End","Match")
print u"{:=>6} {:=>6} {:=>6}\t{}".format('','','','=====')

with  codecs.open(path,mode='r',encoding='utf-8') as f:
    for lineno, line in enumerate(f):
        atime = oclock.search(line)
        if  atime:
            print u"{:>6} {:>6} {:>6}\t{}".format(lineno,
                                            atime.start(),
                                            atime.end(),
                                            atime.group())


print
print "re.findall()"
print
print u"{:>6} {:>6} {:>6}\t{}".format("Line","Start","End","Match")
print u"{:=>6} {:=>6} {:=>6}\t{}".format('','','','=====')
with  codecs.open(path,mode='r',encoding='utf-8') as f:
    for lineno, line in enumerate(f):
        times = oclock.findall(line)
        if times:
            print u"{:>6} {:>6} {:>6}\t{}".format(lineno,
                                            '',
                                            '',
                                            ' '.join(times))


print
print "re.finditer()"
print
print u"{:>6} {:>6} {:>6}\t{}".format("Line","Start","End","Match")
print u"{:=>6} {:=>6} {:=>6}\t{}".format('','','','=====')
with  codecs.open(path,mode='r',encoding='utf-8') as f:
    for lineno, line in enumerate(f):
        times = oclock.finditer(line)
        for m in times:
            print u"{:>6} {:>6} {:>6}\t{}".format(lineno,
                                            m.start(),
                                            m.end(),
                                            m.group())

和输出(在 Python 2.7.3 和 2.7.5 上测试):

re.search()

  Line  Start    End    Match
====== ====== ======    =====
   248      7     21    eleven o'clock
  1520     24     35    one o'clock
  1975     21     33    nine o'clock
  2106      4     16    four o'clock
  4443     19     30    ten o'clock

re.findall()

  Line  Start    End    Match
====== ====== ======    =====
   248                  eleven
  1520                  one
  1975                  nine
  2106                  four
  4443                  ten

re.finditer()

  Line  Start    End    Match
====== ====== ======    =====
   248      7     21    eleven o'clock
  1520     24     35    one o'clock
  1975     21     33    nine o'clock
  2106      4     16    four o'clock
  4443     19     30    ten o'clock

我在这里遗漏了什么？为什么 re.findall() 不返回 o'clock 位？

最佳答案

根据 re.findall documentation :

... If one or more groups are present in the pattern, return a list of groups; this will be a list of tuples if the pattern has more than one group.

pattern 只包含一组； findall 返回组的列表。

>>> import re
>>> re.findall('abc', 'abc')
['abc']
>>> re.findall('a(b)c', 'abc')
['b']
>>> re.findall('a(b)(c)', 'abc')
[('b', 'c')]

使用括号的非捕获版本:

>>> re.findall('a(?:b)c', 'abc')
['abc']

关于python - re.findall 和 re.finditer 的区别——Python 2.7 re 模块中的错误？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/20424661/

24

4

0

文章推荐： python - 仅对大文件 (>3M) 的已关闭文件进行 I/O 操作

文章推荐： objective-c - 按顺序更改 UIView 的属性，但没有动画？

文章推荐： python - 自动启动调试器，没有断点？

文章推荐： Python - 如何使用 ioctl 或 spidev 从设备读取输入？

找不到 Python 模块 "cx_Oracle"模块
我最近在我的机器上安装了 cx_Oracle 模块，以便连接到远程 Oracle 数据库服务器。 (我身边没有 Oracle 客户端)。 Python:版本 2.7 x86 Oracle:版本 11.
python - Timeit 模块 - 将参数传递给 python timeit 模块
我想从 python timeit 模块检查打印以下内容需要多少时间，如何打印， import timeit x = [x for x in range(10000)] timeit.timeit("
javascript - 我该如何修复 --> 文件是 CommonJS 模块；它可以转换为 ES6 模块
我盯着 vs 代码编辑器上的 java 脚本编码，当我尝试将外部模块包含到我的项目中时，代码编辑器提出了这样的建议 -->(文件是 CommonJS 模块；它可能会转换为 ES6 模块。 )..有什么
javascript - 如何在 ES6 模块 Node 应用程序中包含 commonjs 模块？
我有一个 Node 应用程序，我想在标准 ES6 模块格式中使用(即 "type": "module" in the package.json ，并始终使用 import 和 export)而不转译为
css - BlueprintJS 未加载其图标 CSS 模块，但能够加载核心 CSS 模块
我正在学习将 BlueprintJS 合并到我的 React 网络应用程序中，并且在加载某些 CSS 模块时遇到了很多麻烦。我已经安装了 npm install @blueprintjs/core和
javascript - 将 RequireJS 模块 (AMD) 重构为 Webpack 模块 (CommonJS)
我需要重构一堆具有这样的调用的文件 define(['module1','module2','module3' etc...], function(a, b, c etc...) { //bun
javascript - 是否使用 : var app = angular. 模块...或简单地使用 : angular. 模块(
我是 Angular 的新手，正在学习各种教程(Codecademy、thinkster.io 等)，并且已经看到了声明应用程序容器的两种方法。首先: var app = angular.module
unit-testing - 在 OCaml 中使用 OUnit 模块 - 未绑定(bind)模块 OUnit 错误
我正在尝试将 OUnit 与 OCaml 一起使用。单元代码源码(unit.ml)如下: open OUnit let empty_list = [] let list_a = [1;2;3] le
javascript - "Argument ' 模块 ' is not a function, got Object"- 使用 webpack 导入 Angular 模块
我在 Angular 1.x 应用程序中使用 webpack 和 ES6 模块。在我设置的 webpack.config 中: resolve: { alias: { 'angular':
node-modules - 内部/模块/cjs/loader.js :750 return process. dlopen(模块，path.toNamespacedPath(文件名))；
internal/modules/cjs/loader.js:750 return process.dlopen(module, path.toNamespacedPath(filename));
JavaScript 模块
在本教程中，您将借助示例了解 JavaScript 中的模块。随着我们的程序变得越来越大，它可能包含许多行代码。您可以使用模块根据功能将代码分隔在单独的文件中，而不是将所有内容都放在一个文件
JavaScript 模块
我想知道是否可以将此代码更改为仅调用 MyModule.RED 而不是 MyModule.COLORS.RED。我尝试将 mod 设置为变量来存储颜色，但似乎不起作用。难道是我方法不对？ (funct
JavaScript 模块
我有以下代码。它是一个 JavaScript 模块。 (function() { // Object var Cahootsy; Cahootsy = { hello:
angular - 模块 : when and why?
关闭。这个问题是 opinion-based 。它目前不接受答案。想要改进这个问题？更新问题，以便 editing this post 可以用事实和引文来回答它。关闭 2 年前。 Improve
Lua极简入门指南（六）：模块
从用户的角度来看，一个模块能够通过 require 加载并返回一个 table，模块导出的接口都被定义在此 table 中（此 table 被作为一个 namespace）。所有的标准库都是模块。标
ruby 模块
Ruby的模块非常类似类,除了: 模块不可以有实体模块不可以有子类模块由module...end定义. 实际上...模块的'模块类'是'类的类'这个类的父类.搞懂了吗?不懂?让我们继续看
Perl GetOptions 模块
我有一个脚本，它从 CLI 获取 3 个输入变量并将其分别插入到 3 个变量: GetOptions("old_path=s" => \$old_path, "var=s" =
Python 模块，引用同一包中的其他模块？
我有一个简单的 python 包，其目录结构如下: wibble | |-----foo | |----ping.py | |-----bar | |----pong.py 简单的
ocaml - 无法解构仿函数(模块)
这种语法会非常有用——这不起作用有什么原因吗？谢谢! module Foo = { let bar: string = "bar" }; let bar = Foo.bar; /* works *
shell 模块 : < with ansible
我想运行一个命令: - name: install pip shell: "python {"changed": true, "cmd": "python <(curl https://boot

首页

博学

6Ren·AI

商城

python - re.findall 和 re.finditer 的区别——Python 2.7 re 模块中的错误？