python - 从 difflib 中获取更细粒度的差异(或者通过后处理差异来实现相同目的的方法)-6ren

python - 从 difflib 中获取更细粒度的差异(或者通过后处理差异来实现相同目的的方法)

转载作者：太空宇宙更新时间：2023-11-03 16:56:40

正在下载this页面并对其进行较小的编辑，将本段中的第一个 65 更改为 68:

然后，我使用 BeauifulSoup 解析两个源，并使用 difflib 区分它们。

url = 'https://secure.ssa.gov/apps10/reference.nsf/links/02092016062645AM'
response = urllib2.urlopen(url)
content = response.read()  # get response as list of lines

url2 = 'file:///Users/Pyderman/projects/temp/02092016062645AM-modified.html'
response2 = urllib2.urlopen(url2)
content2 = response2.read()  # get response as list of lines
import difflib
d = difflib.Differ()

diffed = d.compare(content, content)

soup = bs4.BeautifulSoup(content, "lxml")
soup2= bs4.BeautifulSoup(content2, "lxml")
diff = d.compare(list(soup.stripped_strings), list(soup2.stripped_strings))
changes = [change for change in diff if change.startswith('-') or  change.startswith('+')]
for change in changes:
    print change

打印更改会给出:

- The Achieving a Better Life Experience (ABLE) Act, H.R. 5771, legislation passed on December 19, 2014. It contains a Title II provision that changes the age at which workers compensation/public disability offset ends for disability beneficiaries from age 65 to full retirement age (FRA).  This provision will apply to any individual who attains age 65 on or after December 19, 2015 (the one year anniversary of enactment of this bill).  Two new Universal Text Identifiers (UTIs), UTI WCP060 and WCP061 were created to comply with this change.
+ The Achieving a Better Life Experience (ABLE) Act, H.R. 5771, legislation passed on December 19, 2014. It contains a Title II provision that changes the age at which workers compensation/public disability offset ends for disability beneficiaries from age 68 to full retirement age (FRA).  This provision will apply to any individual who attains age 65 on or after December 19, 2015 (the one year anniversary of enactment of this bill).  Two new Universal Text Identifiers (UTIs), UTI WCP060 and WCP061 were created to comply with this change.

因此，尽管变化很小，但它还是打印了整个段落。我认为它通过整个段落而不是句子来显示差异是一件好事，但是我们可以以某种方式使输出更加精细吗？就目前情况而言，如果我想突出显示仅更改的文本，我将必须对这两个几乎相同的字符串进行一些额外的增量比较。

最佳答案

您可以使用nltk.sent_tokenize()将汤字符串拆分成句子:

from nltk import sent_tokenize

sentences = [sentence for string in soup.stripped_strings for sentence in sent_tokenize(string)]
sentences2 = [sentence for string in soup2.stripped_strings for sentence in sent_tokenize(string)]

diff = d.compare(sentences, sentences2)
changes = [change for change in diff if change.startswith('-') or  change.startswith('+')]
for change in changes:
    print(change)

仅打印检测到更改的适当句子:

- It contains a Title II provision that changes the age at which workers compensation/public disability offset ends for disability beneficiaries from age 65 to full retirement age (FRA).
+ It contains a Title II provision that changes the age at which workers compensation/public disability offset ends for disability beneficiaries from age 68 to full retirement age (FRA).

关于python - 从 difflib 中获取更细粒度的差异(或者通过后处理差异来实现相同目的的方法)，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/35375004/

文章推荐： c# - 如何在 C# 中访问 GDI+ 效果类

文章推荐： ruby - Dir globbing 没有完全递归

文章推荐： c# - 通过 C# 通过 NFS 读取 UNIX 文件属性

文章推荐： c# - 将一堆位图转储为 PDF 的最快的 .Net PDF 库是什么？

c# - 将控制台光标设置为粗/细
在命令提示符下，当你按下插入按钮时，光标从细条变为粗条，表明它处于覆盖模式，当你再次按下它时，它又变细表明它处于插入模式有什么办法可以在 C# 中执行此操作吗？编辑:我想知道是否有办法使光标变粗/变
javascript - Rails 3 javascript 细 CoffeeScript 引用错误(类)未定义
RubyRogues 播客上有人曾经说过“学习 CoffeeScript，因为 CoffeeScript 编写的 javascript 比你更好。”抱歉，不记得是谁说的... 所以，我采用了一个非常简

太空宇宙

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 从 difflib 中获取更细粒度的差异(或者通过后处理差异来实现相同目的的方法)