gpt4 book ai didi

python - 使用 lxml 和 xpath 加速 xml 解析过程

转载 作者:太空宇宙 更新时间:2023-11-03 15:03:19 25 4
gpt4 key购买 nike

<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<document DateTime="2017-06-23T04:27:08.592Z">
<PeakInfo No="1" mz="505.2315648572003965"
Intensity="4531.0000000000000000"
Rel_Intensity="3.2737729673489735"
Resolution="1879.5638812957554364"
SNR="14.0278637770897561"
Area="1348.1007591467391649"
Rel_Area="2.3371194184605959"
Index="238.9999999999976694"/>
<PeakInfo No="2" mz="522.1330917856538463"
Intensity="3382.0000000000000000"
Rel_Intensity="2.4435886505350317"
Resolution="3502.9921209527169594"
SNR="10.4705882352940982"
Area="881.4468100654634100"
Rel_Area="1.5281101521284057"
Index="925.0000000000000000"/>
</document>

上面是我最近使用的 xml 文件的一部分。每个文件包含超过 400 个 PeakInfo,我确实制作了一个 python 脚本来解析每个文件:

from lxml import etree
import pandas as pd
import tkinter.filedialog
import os
import pandas.io.formats.excel

full_path = tkinter.filedialog.askdirectory(initialdir='.')
newfolder = full_path+'\\xls files'
os.chdir(full_path)
os.makedirs(newfolder)

data = {}
for files in os.listdir(full_path):
if os.path.isfile(os.path.join(full_path, files)):
plist = pd.DataFrame()
filename = os.path.basename(files).rpartition('.')[0]

if len(filename) == 2:
filename = filename[:1]+'0'+filename[1:]

xmlp = etree.parse(files)
for p in xmlp.xpath('//PeakInfo'):
data['Exp. m/z'] = p.attrib['mz']
data['Intensity'] = p.attrib['Intensity']
plist = plist.append(data, ignore_index=True)
plist['Exp. m/z'] = plist['Exp. m/z'].astype(float)
plist['Exp. m/z'] = plist['Exp. m/z'].map('{:.4f}'.format)
plist['Intensity'] = plist['Intensity'].astype(float)
plist['Intensity'] = plist['Intensity'].map('{:.0f}'.format)
pandas.io.formats.excel.header_style = None
plist.to_excel(os.path.join(newfolder, filename+'.xls'),index=False)

如果文件名只有两个字符(即 A1 到 A01),此代码会更改文件名,然后提取 mz 和 Intensity 并保存为 xls 文件。问题是解析每个文件花费的时间太长。有什么技巧可以显着加快这个过程吗?

最佳答案

from lxml import etree
import pandas as pd
import tkinter.filedialog
import os
import pandas.io.formats.excel

full_path = tkinter.filedialog.askdirectory(initialdir='.')
newfolder = full_path+'\\xls files'
os.chdir(full_path)
os.makedirs(newfolder)

data = {}
for files in os.listdir(full_path):
if os.path.isfile(os.path.join(full_path, files)):
plist = pd.DataFrame()
filename = os.path.basename(files).rpartition('.')[0]

if len(filename) == 2:
filename = filename[:1]+'0'+filename[1:]

xmlp = etree.parse(files)
for p in xmlp.xpath('//PeakInfo'):
data['Exp. m/z'] = p.attrib['mz']
data['Intensity'] = p.attrib['Intensity']
plist = plist.append(data, ignore_index=True)
plist['Exp. m/z'] = plist['Exp. m/z'].astype(float)
plist['Exp. m/z'] = plist['Exp. m/z'].map('{:.4f}'.format)
plist['Intensity'] = plist['Intensity'].astype(float)
plist['Intensity'] = plist['Intensity'].map('{:.0f}'.format)
pandas.io.formats.excel.header_style = None
plist.to_excel(os.path.join(newfolder, filename+'.xls'),index=False)

只要改变空格,你像to_excel这样的代码执行太多次,而且速度很慢,而且“astype”会复制元素,占用太多内存,速度变慢。

关于python - 使用 lxml 和 xpath 加速 xml 解析过程,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/44880509/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com