gpt4 book ai didi

python - 如何使用 Python 搜索和替换 XML 文件中的文本?

转载 作者:数据小太阳 更新时间:2023-10-29 01:51:34 26 4
gpt4 key购买 nike

如何在整个 xml 文件中搜索特定的文本模式,然后在 Python 3.5 中用新的文本模式替换每次出现的该文本?

其他所有内容(格式、属性、注释等)都需要保留在原始 xml 文件中。

我在 Windows (win32) 上运行 Python 3.5.1。

具体来说,我想将每次出现的“FEATURE NAME”替换为“THIS WORKED”,并将每次出现的“FEATURE NUMBER”替换为“12345”。

我一直在尝试学习 Python 和 xml.etree.ElementTree,但无法解决这个问题。我已经看过“在 Python 中搜索和替换 .xml 文件中的一行”、“在 Python 中搜索和替换文件中的一行”和“如何使用 Python 搜索和替换文件中的文本?”和本网站上的其他现有问答,但无法弄清楚 - 我不是经验丰富的程序员,所以如果需要更多输入,请告诉我。非常感谢您的帮助!!!

这是我在记事本中打开 xml 代码时的样子的副本(除了我添加空格以缩进每行并在我将其粘贴到此问题时按回车键返回某些行):

<description-topic>
<access-info>
<index-term-set>
<index-term>
<primary>FID FEATURE NUMBER</primary>
</index-term>
<index-term>
<primary>FEATURE NAME</primary>
</index-term>
<index-term>
<primary>Common features</primary>
<secondary>FID FEATURE NUMBER</secondary>
</index-term>
</index-term-set>
</access-info>
<title>FEATURE NUMBER - FEATURE NAME</title>
<block>
<label>Platform</label>
<comment>REVIEWERS: I guessed at the FEATURE NAME</comment>
<para>
This feature applies to the following platforms: FEATURE NAME<!--Check the values--></para>
</block>
<block branch="no">
<label>Feature Benefits</label>
<para>
<comment>REVIEWERS: What do we put here? See template (link given in review email) for more information.</comment>
</para>
</block>
<block branch="no">
<label>Dependencies</label>
<para/>
<subblock>
<label>Features</label>
<comment>What FEATURE NAME do we put here?</comment>
</subblock>
<subblock>
<label>Hardware</label>
<comment>What FEATURE NAME do we put here?</comment>
<para>This feature applies to the following: FEATURE NUMBER and text.</para><?Pub Caret -1?>
</subblock>
<subblock>
<label>Dependencies outside the eNodeB</label>
<comment>What FEATURE NAME do we put here?</comment>
</subblock>
</block>
<block branch="no">
<label>Impacts</label>
<comment>REVIEWERS: What FEATURE NUMBER do we put here?</comment>
<para>
<comment/>
</para>
</block>
</description-topic>

这是我尝试开始工作的最新代码:

from xml.etree import ElementTree as et
tree = et.parse('Atemplate2.xml')
tree.find('description-topic/access-info/index-term-set/index-term/primary/').text = '12345'
tree.write('Atemplate2.xml')

我收到以下错误:追溯(最近一次通话): 文件“ajktest18.py”,第 15 行,位于 tree.find('description-topic/access-info/index-term-set/index-term/primary/').text = '12345'

AttributeError: 'NoneType' 对象没有属性 'text'

我更希望能够搜索和修改整个文件中的任何匹配项,但我不知道如何找到我正在搜索的文本的一个特定匹配项。

这是我试图用来查找路径的代码:

import xml.etree.ElementTree as ET
tree = ET.parse('Atemplate.xml')
root = tree.getroot()

print(root.tag, root.attrib, root.text)

for child in root:
print(child.tag, child.attrib, child.text)
for label in root.iter('label'):
print(label.tag, label.attrib, label.text)
for title in root.iter('title'):
print(title.attrib)

我也试过下面的代码:

with open('Atemplate2.xml') as f:
tree = ET.parse(f)
root = tree.getroot()

for elem in root.getiterator():
try:
elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
elem.text = elem.text.replace('FEATURE NUMBER', '12345')
except AttributeError:
pass

tree.write('output.xml')

但是会出现以下错误:

File "<pyshell#40>", line 2, in <module>
tree = ET.parse(f)
File "C:\MyPath\Python35-32\lib\xml\etree\ElementTree.py", line 1182, in parse
tree.parse(source, parser)
File "C:\ MyPath \Python35-32\lib\xml\etree\ElementTree.py", line 594, in parse
self._root = parser._parse_whole(source)
File "C:\ MyPath \Python35-32\lib\encodings\cp1252.py", line 23, in decode
return codecs.charmap_decode(input,self.errors,decoding_table)[0]

UnicodeDecodeError: 'charmap' 编解码器无法解码位置 1119 中的字节 0x9d:字符映射到

##

最终更新 - 这是最终对我有用的代码(谢谢你,Jarad!):

import lxml.etree as ET
#using lxml instead of xml preserved the comments

#adding the encoding when the file is opened and written is needed to avoid a charmap error
with open('filename.xml', encoding="utf8") as f:
tree = ET.parse(f)
root = tree.getroot()


for elem in root.getiterator():
try:
elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
elem.text = elem.text.replace('FEATURE NUMBER', '123456')
except AttributeError:
pass

#tree.write('output.xml', encoding="utf8")
# Adding the xml_declaration and method helped keep the header info at the top of the file.
tree.write('output.xml', xml_declaration=True, method='xml', encoding="utf8")

最佳答案

注意事项:

  • 我从未使用过 xml.etree.ElementTree
  • 我从未使用过它,因为我从未发现自己在操纵 XML
  • 与对图书馆了如指掌的人相比,我不知道这是否是“最佳”方式
  • 评论者似乎是在评判你,而不是帮助你走出困境

这是对 this excellent answer 的修改。问题是,您需要读取 XML 文件并对其进行解析。

import xml.etree.ElementTree as ET

with open('xmlfile.xml', encoding='latin-1') as f:
tree = ET.parse(f)
root = tree.getroot()

for elem in root.getiterator():
try:
elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
elem.text = elem.text.replace('FEATURE NUMBER', '123456')
except AttributeError:
pass

tree.write('output.xml', encoding='latin-1')

请注意,您可以将 encoding 参数更改为其他内容,例如:utf-8cp1252ISO-8859 -1 等。确实取决于您的系统和文件。

关于python - 如何使用 Python 搜索和替换 XML 文件中的文本?,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/37868881/

26 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com