gpt4 book ai didi

Python HTML 解析器 : UnicodeDecodeError

转载 作者:太空狗 更新时间:2023-10-29 17:43:37 26 4
gpt4 key购买 nike

我正在使用 HTMLParser 来解析我使用 urllib 提取的页面,并且在将某些内容传递给 HTMLParser 时遇到了 UnicodeDecodeError 异常。

我尝试使用 chardet 检测编码并转换为 asciiutf-8(docs 不似乎在说它应该是什么)。有损是可以接受的,但是虽然解码/编码行工作得很好,但我总是在 self.feed() 之后得到错误。

如果我只是打印,信息就在那里。

from HTMLParser import HTMLParser
import urllib
import chardet

class search_youtube(HTMLParser):

def __init__(self, search_terms):
HTMLParser.__init__(self)
self.track_ids = []
for search in search_terms:
self.__in_result = False
search = urllib.quote_plus(search)
query = 'http://youtube.com/results?search_query='
page = urllib.urlopen(query + search).read()
try:
self.feed(page)
except UnicodeDecodeError:
encoding = chardet.detect(page)['encoding']
if encoding != 'unicode':
page = page.decode(encoding)
page = page.encode('ascii', 'ignore')
self.feed(page)
print 'success'

searches = ['telepopmusik breathe']
results = search_youtube(searches)
print results.track_ids

这是输出:

Traceback (most recent call last):
File "test.py", line 27, in <module>
results = search_youtube(searches)
File "test.py", line 23, in __init__
self.feed(page)
File "/usr/lib/python2.6/HTMLParser.py", line 108, in feed
self.goahead(0)
File "/usr/lib/python2.6/HTMLParser.py", line 148, in goahead
k = self.parse_starttag(i)
File "/usr/lib/python2.6/HTMLParser.py", line 252, in parse_starttag
attrvalue = self.unescape(attrvalue)
File "/usr/lib/python2.6/HTMLParser.py", line 390, in unescape
return re.sub(r"&(#?[xX]?(?:[0-9a-fA-F]+|\w{1,8}));", replaceEntities, s)
File "/usr/lib/python2.6/re.py", line 151, in sub
return _compile(pattern, 0).sub(repl, string, count)
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 1: ordinal not in range(128)

最佳答案

确实是 UTF-8。这有效:

from HTMLParser import HTMLParser
import urllib

class search_youtube(HTMLParser):

def __init__(self, search_terms):
HTMLParser.__init__(self)
self.track_ids = []
for search in search_terms:
self.__in_result = False
search = urllib.quote_plus(search)
query = 'http://youtube.com/results?search_query='
connection = urllib.urlopen(query + search)
encoding = connection.headers.getparam('charset')
page = connection.read().decode(encoding)
self.feed(page)
print 'success'

searches = ['telepopmusik breathe']
results = search_youtube(searches)
print results.track_ids

你不需要 chardet,Youtube 不是白痴,他们实际上在标题中发送了正确的编码。

关于Python HTML 解析器 : UnicodeDecodeError,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/4790078/

26 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com