Python re.split() 与 nltk word_tokenize 和 sent

Python re.split() 与 nltk word_tokenize 和 sent_tokenize

转载作者：太空狗更新时间：2023-10-29 17:40:42

我正在浏览 this question .

我只是想知道 NLTK 在单词/句子标记化方面是否会比正则表达式更快。</p>

最佳答案

默认的 nltk.word_tokenize() 使用 Treebank tokenizer模拟来自 Penn Treebank tokenizer 的分词器.

请注意，str.split() 并未实现语言学意义上的记号，例如:

>>> sent = "This is a foo, bar sentence."
>>> sent.split()
['This', 'is', 'a', 'foo,', 'bar', 'sentence.']
>>> from nltk import word_tokenize
>>> word_tokenize(sent)
['This', 'is', 'a', 'foo', ',', 'bar', 'sentence', '.']

通常用于以指定的分隔符分隔字符串，例如在制表符分隔的文件中，您可以使用 str.split('\t') 或者当您尝试用换行符拆分字符串时 \n 当您的文本文件每行一个句子。

让我们在 python3 中做一些基准测试:

import time
from nltk import word_tokenize

import urllib.request
url = 'https://raw.githubusercontent.com/Simdiva/DSL-Task/master/data/DSLCC-v2.0/test/test.txt'
response = urllib.request.urlopen(url)
data = response.read().decode('utf8')

for _ in range(10):
    start = time.time()
    for line in data.split('\n'):
        line.split()
    print ('str.split():\t', time.time() - start)

for _ in range(10):
    start = time.time()
    for line in data.split('\n'):
        word_tokenize(line)
    print ('word_tokenize():\t', time.time() - start)

[输出]:

str.split():     0.05451083183288574
str.split():     0.054320573806762695
str.split():     0.05368804931640625
str.split():     0.05416440963745117
str.split():     0.05299568176269531
str.split():     0.05304527282714844
str.split():     0.05356955528259277
str.split():     0.05473494529724121
str.split():     0.053118228912353516
str.split():     0.05236077308654785
word_tokenize():     4.056122779846191
word_tokenize():     4.052812337875366
word_tokenize():     4.042144775390625
word_tokenize():     4.101543664932251
word_tokenize():     4.213029146194458
word_tokenize():     4.411528587341309
word_tokenize():     4.162556886672974
word_tokenize():     4.225975036621094
word_tokenize():     4.22914719581604
word_tokenize():     4.203172445297241

如果我们尝试 another tokenizers in bleeding edge NLTK来自 https://github.com/jonsafari/tok-tok/blob/master/tok-tok.pl :

import time
from nltk.tokenize import ToktokTokenizer

import urllib.request
url = 'https://raw.githubusercontent.com/Simdiva/DSL-Task/master/data/DSLCC-v2.0/test/test.txt'
response = urllib.request.urlopen(url)
data = response.read().decode('utf8')

toktok = ToktokTokenizer().tokenize

for _ in range(10):
    start = time.time()
    for line in data.split('\n'):
        toktok(line)
    print ('toktok:\t', time.time() - start)

[输出]:

toktok:  1.5902607440948486
toktok:  1.5347232818603516
toktok:  1.4993178844451904
toktok:  1.5635688304901123
toktok:  1.5779635906219482
toktok:  1.8177132606506348
toktok:  1.4538452625274658
toktok:  1.5094449520111084
toktok:  1.4871931076049805
toktok:  1.4584410190582275

(注:文本文件来源自https://github.com/Simdiva/DSL-Task)

如果我们查看 native perl 实现，ToktokTokenizer 的 python 与 perl 时间是可比的.但是在 python 实现中，正则表达式是预编译的，而在 perl 中，它不是，但是然后 the proof is still in the pudding :

alvas@ubi:~$ wget https://raw.githubusercontent.com/jonsafari/tok-tok/master/tok-tok.pl
--2016-02-11 20:36:36--  https://raw.githubusercontent.com/jonsafari/tok-tok/master/tok-tok.pl
Resolving raw.githubusercontent.com (raw.githubusercontent.com)... 185.31.17.133
Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|185.31.17.133|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: 2690 (2.6K) [text/plain]
Saving to: ‘tok-tok.pl’

100%[===============================================================================================================================>] 2,690       --.-K/s   in 0s      

2016-02-11 20:36:36 (259 MB/s) - ‘tok-tok.pl’ saved [2690/2690]

alvas@ubi:~$ wget https://raw.githubusercontent.com/Simdiva/DSL-Task/master/data/DSLCC-v2.0/test/test.txt
--2016-02-11 20:36:38--  https://raw.githubusercontent.com/Simdiva/DSL-Task/master/data/DSLCC-v2.0/test/test.txt
Resolving raw.githubusercontent.com (raw.githubusercontent.com)... 185.31.17.133
Connecting to raw.githubusercontent.com (raw.githubusercontent.com)|185.31.17.133|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: 3483550 (3.3M) [text/plain]
Saving to: ‘test.txt’

100%[===============================================================================================================================>] 3,483,550    363KB/s   in 7.4s   

2016-02-11 20:36:46 (459 KB/s) - ‘test.txt’ saved [3483550/3483550]

alvas@ubi:~$ time perl tok-tok.pl < test.txt > /tmp/null

real    0m1.703s
user    0m1.693s
sys 0m0.008s
alvas@ubi:~$ time perl tok-tok.pl < test.txt > /tmp/null

real    0m1.715s
user    0m1.704s
sys 0m0.008s
alvas@ubi:~$ time perl tok-tok.pl < test.txt > /tmp/null

real    0m1.700s
user    0m1.686s
sys 0m0.012s
alvas@ubi:~$ time perl tok-tok.pl < test.txt > /tmp/null

real    0m1.727s
user    0m1.700s
sys 0m0.024s
alvas@ubi:~$ time perl tok-tok.pl < test.txt > /tmp/null

real    0m1.734s
user    0m1.724s
sys 0m0.008s

(注意:在为 tok-tok.pl 计时时，我们必须将输出通过管道传输到文件中，因此这里的计时包括机器输出到文件所花费的时间，而在nltk.tokenize.ToktokTokenizer 时间，它不包括输出到文件的时间)

关于 sent_tokenize()，它有点不同，在不考虑准确性的情况下比较速度基准有点古怪。

考虑一下:

如果正则表达式将文本文件/段落拆分为 1 个句子，那么速度几乎是瞬时的，即完成 0 个工作。但那将是一个可怕的句子分词器......
如果文件中的句子已经被 \n 分隔，那么这只是比较 str.split('\n') 的情况vs re.split('\n') 和 nltk 与句子标记化无关;P

有关 sent_tokenize() 在 NLTK 中如何工作的信息，请参阅:

因此，为了有效地比较 sent_tokenize() 与其他基于正则表达式的方法(不是 str.split('\n'))，还必须评估准确性并拥有一个以标记化格式包含人工评估句子的数据集。

考虑这个任务:https://www.hackerrank.com/challenges/from-paragraphs-to-sentences

给定文本:

In the third category he included those Brothers (the majority) who saw nothing in Freemasonry but the external forms and ceremonies, and prized the strict performance of these forms without troubling about their purport or significance. Such were Willarski and even the Grand Master of the principal lodge. Finally, to the fourth category also a great many Brothers belonged, particularly those who had lately joined. These according to Pierre's observations were men who had no belief in anything, nor desire for anything, but joined the Freemasons merely to associate with the wealthy young Brothers who were influential through their connections or rank, and of whom there were very many in the lodge.Pierre began to feel dissatisfied with what he was doing. Freemasonry, at any rate as he saw it here, sometimes seemed to him based merely on externals. He did not think of doubting Freemasonry itself, but suspected that Russian Masonry had taken a wrong path and deviated from its original principles. And so toward the end of the year he went abroad to be initiated into the higher secrets of the order.What is to be done in these circumstances? To favor revolutions, overthrow everything, repel force by force?No! We are very far from that. Every violent reform deserves censure, for it quite fails to remedy evil while men remain what they are, and also because wisdom needs no violence. "But what is there in running across it like that?" said Ilagin's groom. "Once she had missed it and turned it away, any mongrel could take it," Ilagin was saying at the same time, breathless from his gallop and his excitement.

我们想要得到这个:

In the third category he included those Brothers (the majority) who saw nothing in Freemasonry but the external forms and ceremonies, and prized the strict performance of these forms without troubling about their purport or significance.
Such were Willarski and even the Grand Master of the principal lodge.
Finally, to the fourth category also a great many Brothers belonged, particularly those who had lately joined.
These according to Pierre's observations were men who had no belief in anything, nor desire for anything, but joined the Freemasons merely to associate with the wealthy young Brothers who were influential through their connections or rank, and of whom there were very many in the lodge.
Pierre began to feel dissatisfied with what he was doing.
Freemasonry, at any rate as he saw it here, sometimes seemed to him based merely on externals.
He did not think of doubting Freemasonry itself, but suspected that Russian Masonry had taken a wrong path and deviated from its original principles.
And so toward the end of the year he went abroad to be initiated into the higher secrets of the order.
What is to be done in these circumstances?
To favor revolutions, overthrow everything, repel force by force?
No!
We are very far from that.
Every violent reform deserves censure, for it quite fails to remedy evil while men remain what they are, and also because wisdom needs no violence.
"But what is there in running across it like that?" said Ilagin's groom.
"Once she had missed it and turned it away, any mongrel could take it," Ilagin was saying at the same time, breathless from his gallop and his excitement.

所以简单地执行 str.split('\n') 不会给你任何东西。即使不考虑句子的顺序，你也会得到 0 个肯定结果:

>>> text = """In the third category he included those Brothers (the majority) who saw nothing in Freemasonry but the external forms and ceremonies, and prized the strict performance of these forms without troubling about their purport or significance. Such were Willarski and even the Grand Master of the principal lodge. Finally, to the fourth category also a great many Brothers belonged, particularly those who had lately joined. These according to Pierre's observations were men who had no belief in anything, nor desire for anything, but joined the Freemasons merely to associate with the wealthy young Brothers who were influential through their connections or rank, and of whom there were very many in the lodge.Pierre began to feel dissatisfied with what he was doing. Freemasonry, at any rate as he saw it here, sometimes seemed to him based merely on externals. He did not think of doubting Freemasonry itself, but suspected that Russian Masonry had taken a wrong path and deviated from its original principles. And so toward the end of the year he went abroad to be initiated into the higher secrets of the order.What is to be done in these circumstances? To favor revolutions, overthrow everything, repel force by force?No! We are very far from that. Every violent reform deserves censure, for it quite fails to remedy evil while men remain what they are, and also because wisdom needs no violence. "But what is there in running across it like that?" said Ilagin's groom. "Once she had missed it and turned it away, any mongrel could take it," Ilagin was saying at the same time, breathless from his gallop and his excitement. """
>>> answer = """In the third category he included those Brothers (the majority) who saw nothing in Freemasonry but the external forms and ceremonies, and prized the strict performance of these forms without troubling about their purport or significance.
... Such were Willarski and even the Grand Master of the principal lodge.
... Finally, to the fourth category also a great many Brothers belonged, particularly those who had lately joined.
... These according to Pierre's observations were men who had no belief in anything, nor desire for anything, but joined the Freemasons merely to associate with the wealthy young Brothers who were influential through their connections or rank, and of whom there were very many in the lodge.
... Pierre began to feel dissatisfied with what he was doing.
... Freemasonry, at any rate as he saw it here, sometimes seemed to him based merely on externals.
... He did not think of doubting Freemasonry itself, but suspected that Russian Masonry had taken a wrong path and deviated from its original principles.
... And so toward the end of the year he went abroad to be initiated into the higher secrets of the order.
... What is to be done in these circumstances?
... To favor revolutions, overthrow everything, repel force by force?
... No!
... We are very far from that.
... Every violent reform deserves censure, for it quite fails to remedy evil while men remain what they are, and also because wisdom needs no violence.
... "But what is there in running across it like that?" said Ilagin's groom.
... "Once she had missed it and turned it away, any mongrel could take it," Ilagin was saying at the same time, breathless from his gallop and his excitement."""
>>> 
>>> output = text.split('\n')
>>> sum(1 for sent in text.split('\n') if sent in answer)
0

关于Python re.split() 与 nltk word_tokenize 和 sent_tokenize，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/35345761/

文章推荐： c# - 如何判断三个int是否都相等

文章推荐： c# - 你如何使用 SqlDataRecord 处理 Nullable 类型

文章推荐： angular - 将 Electron 集成到 Angular 项目时 fs 模块失败

regex - R split on delimiter (split) 保留分隔符 (split)
在 R 中，您可以使用 strsplit在分隔符( split )上分割向量的函数如下: x <- "What is this? It's an onion. What! That's| Well
split - 。 split ();让一个函数运行另一个函数时出错
我的 .split(); 方法有问题。我称这个函数为: get_content_ajax("html/settings.html", "#ajax", 1, "Settings page have
split - Elixir中的String.split()在输出中的子字符串列表的末尾放置一个空字符串
我是Elixir的新手。我正在尝试对字符串split的基本操作，如下所示 String.split("Awesome",""); 根据elixir document，它应该根据提供的模式split字符
split - ARRAYFORMULA() 不适用于 SPLIT()
当我使用 =arrayformula(split(input!G2:G, ",")) 时，为什么拆分公式没有扩展到整个列? 我只得到输入的结果!G2 单元格，而不是 G 列中的其余部分。其他公式如 =
javascript - Polymer 1.0 尝试制作一种类似于核心 split 器的 split 器，可以称为铁 split 器
我正在尝试制作一个名为 core-splitter 的元素，该元素在 1.0 中已弃用，因为它在我们的项目中起着关键作用。如果您不知道 core-splitter 的作用，我可以提供一个简短的描述。
split - 具有多个定界符的 ansible string.split()
我很难尝试使用多个定界符将字符串拆分为列表。我可以像下面这样将它拆分两次: myString.split(':')[1].split('.') 然而，这看起来很不优雅。在我的脑海里，我想做这样的事情:
split - AttributeError: 'DatasetAutoFolds' 对象没有属性 'split'
来自使用惊喜模块的推荐引擎的代码，我在任何地方都找不到答案。最佳答案根据您的目标，您可以使用 cross_validation方法，它将自动为您执行拆分。示例:cross_validate(alg
javascript - 尝试在有丝 split 模拟中模拟细胞的 split
我正在制作一个有丝 split 模拟器，我希望它在细胞足够大并 split 时运行有丝 split 功能。当它分割时，我希望它能够将分割从初始 x 值(前一个单元格的 x)动画化为新的 x 值(右侧的
split - Split 函数(jquery)后使用索引吗？
我有一个用于三个按钮的点击处理程序，在这个处理程序中我想提取所点击按钮的 ID。我有一行这样的代码: $('#switch button').click(function(){ var cla
javascript - .split() 保持 split 特征
我需要像这样分割一个字符串 var val = "$cs+55+mod($a)"; 放入数组 arr = val.split( /[+-/*()\s*]/ ); 问题是将分隔符保留为数组元素，如 ar
python - 为什么 split() 在同一字符串上返回的元素多于 split ("")？
我在同一个 string 上使用 split() 和 split("") .但为什么 split("") 返回的元素数量少于 split()？我想知道在什么特定的输入情况下会发生这种情况。最佳答案
javascript - jQuery split() 不是...... split ？
我的代码中某处有错误，但看不到我做错了什么。我拥有的是 facebook 用户 ID 的隐藏输入，它是通过 jQuery UI 自动完成填充的: 然后，我有一个 jQuery 函数，当单击链接将其
Python Split() 和 re.split()
我正在寻找一个程序来读取字符串/文件并显示其中的前三个单词。所以我尝试了: letter= "a,b,c" print(letter.split(',')[0]) 这对获取一个单词有效，但执行 [0
c# - SQL : To split or not to split? 中的邮件表
我有一个存储邮件的表 Mails(谁会想到... ;))。通过 tinyint MailStatus，我决定这是 SentMail、Draft 还是 ReceivedMail。现在我想知道 Tab
Python re.split() 与 split()
在我的优化探索中，我发现内置的 split() 方法比等效的 re.split() 方法快大约 40%。虚拟基准(易于复制粘贴): import re, time, random def rando
perl - Perl `split`不能 `split`到默认数组
我对split有一个奇怪的问题，因为默认情况下它不会将split放入默认数组中。以下是一些玩具代码。 #!/usr/bin/perl $A="A:B:C:D"; split (":",$A); pr
split - 为什么 SPLIT 到属于同一 PDS 的多个成员会产生克隆成员？
我目前正在学习 JCL，并且正在使用 SORT 程序。作为练习，我想将一些输入记录拆分为属于同一 PDS 的多个成员。这是我的 JCL 代码: //FAILJ JOB //STEP1 EX
powershell - Powershell split()vs -split-有什么区别？
在苦苦挣扎了半小时之后，我在使用空格分割字符串时遇到了这种差异，具体取决于您使用的语法。简单字符串: $line = "1: 2: 3: 4: 5: " 拆分示例1 -从1开始注意带有 token
python split 和 re.split 没有捕获字符串中的制表符或空格
我有一个像这样的字符串: 'Agendas / Schedules meetings and speakers 4 F 1928-1209 Box 2' 我正在尝试将其
algorithm - 二次 split 和线性 split 的区别
我试图了解 r-tree 的工作原理，发现有两种类型的拆分:二次拆分和线性拆分。线性和二次实际上有什么区别？在哪种情况下，一个会比另一个更受欢迎？最佳答案原始 R-Tree 论文在 3.5.2

太空狗

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

Python re.split() 与 nltk word_tokenize 和 sent_tokenize