Python Mechanize 在处理网站 html 时遇到问题-6ren

Python Mechanize 在处理网站 html 时遇到问题

转载作者：行者123 更新时间：2023-12-04 16:19:23

24

4

我正在尝试使用 python 模块 mechanize 在网页上填写表单，然后下载生成的 html。网站如下:

“xxx”

但首先，我只想列出表格。我的代码如下:

import mechanize

br = mechanize.Browser()
br.set_handle_robots(False)   # ignore robots
br.set_handle_refresh(False)  # can sometimes hang without this

url = "xxx"
response = br.open(url)
a = response.read()      # the text of the page

for form in br.forms():
    print "Form name:", form.name
    print form

给出的错误列表相当大:

In [75]: % run -i form_filler.py
---------------------------------------------------------------------------
ParseError                                Traceback (most recent call last)
/home/blake/Python/form_filler.py in <module>()
     19 
     20 
---> 21 for form in br.forms():
     22         print "Form name:", form.name
     23         print form

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_mechanize.pyc in forms(self)
    418         if not self.viewing_html():
    419             raise BrowserStateError("not viewing HTML")
--> 420         return self._factory.forms()
    421 
    422     def global_form(self):

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_html.pyc in forms(self)
    555             try:
    556                 self._forms_genf = CachingGeneratorFunction(
--> 557                     self._forms_factory.forms())
    558             except:  # XXXX define exception!
    559                 self.set_response(self._response)

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_html.pyc in forms(self)
    235             _urljoin=_rfc3986.urljoin,
    236             _urlparse=_rfc3986.urlsplit,
--> 237             _urlunparse=_rfc3986.urlunsplit,
    238             )
    239         self.global_form = forms[0]

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_form.pyc in ParseResponseEx(response, select_default, form_parser_class, request_class, entitydefs, encoding, _urljoin, _urlparse, _urlunparse)
    842                         _urljoin=_urljoin,
    843                         _urlparse=_urlparse,
--> 844                         _urlunparse=_urlunparse,
    845                         )
    846 

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_form.pyc in _ParseFileEx(file, base_uri, select_default, ignore_errors, form_parser_class, request_class, entitydefs, backwards_compat, encoding, _urljoin, _urlparse, _urlunparse)
    979         data = file.read(CHUNK)
    980         try:
--> 981             fp.feed(data)
    982         except ParseError, e:
    983             e.base_uri = base_uri

/usr/local/lib/python2.7/dist-packages/mechanize-0.2.5-py2.7.egg/mechanize/_form.pyc in feed(self, data)
    758             _sgmllib_copy.SGMLParser.feed(self, data)
    759         except _sgmllib_copy.SGMLParseError, exc:
--> 760             raise ParseError(exc)
    761 
    762     def close(self):

ParseError: expected name token at '<!!--end footer-->\r\n'

我的理解是 html 以某种方式“写得不好”。我已经在 google 等示例网站上尝试了上述代码，并且运行良好。怀疑是回车有关系，但是怎么绕过这个问题呢？

最佳答案

显然，如果我使用

br = mechanize.Browser(factory=mechanize.RobustFactory())

代替

br = mechanize.Browser()

它给出了正确的输出:

In [18]: % run -i form_filler.py
Form name: qSearchForm
<qSearchForm GET xxx
  <TextControl(q=Search)>
  <SubmitControl(qSearchBtn=) (readonly)>>
Form name: form1
<form1 POST xxx
  <TextControl(name=)>
  <SelectControl(code=[*LER, Lerwick, Scotland, 29.9 / 358.8, ESK, Eskdalemuir, Scotland, 34.7 / 356.8, HAD, Hartland, England, 39.0 / 355.5, ABN, Abinger, England, 38.8 / 359.6, GRW, Greenwich, England, 38.5 / 0.0])>
  <TextControl(month=)>
  <TextControl(year=)>
  <SubmitControl(<None>=Submit Query) (readonly)>
  <IgnoreControl(<None>=<None>)>>

关于Python Mechanize 在处理网站 html 时遇到问题，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/25407738/

24

4

0

文章推荐： ruby - Mechanize 和 NTLM 身份验证

文章推荐： python - 编译网页表单并使用 Mechanize 检索文件

文章推荐： python - 找不到匹配的表格

文章推荐： perl - 使用 www::mechanize 的爬虫

mechanize - Mechanize python中的模块错误
我正在使用 mechanize python 登录网站 combochat2.us用户名 mask3和密码findnext ，但它显示了“没有找到 Mechanize 模块”之类的错误 import
mechanize - 如何避免 Mechanize 解析文件或图像的 url？
我在我的 rails 应用程序中使用 gem mechanize 来抓取网页数据。我这样使用它: agent = Mechanize.new document = agent.get("http:/
python, mechanize - 使用 mechanize 打开文本文件
我正在学习机械。我正在尝试打开一个文本文件，您点击的链接显示文本 (.prn)我遇到的一个问题是此页面上只有 1 个表单，并且该文件不在表单中。对我来说另一个问题是此页面上有几个文本文件，但它们都具
mechanize - Beautiful Soup 在解析 Mechanize 输出时遇到问题
def return_with_soup(url): #uses mechanize to tell the browser we aren't a bot #and to retri
python - 如何通过python中的 Mechanize 来 Mechanize 返回页面内的网址？
我正在开发一个项目，使用 python 和 Mechanize 。我有个问题 : Mechanize 返回的页面，有不是的 URLS Mechanize ，如果用户点击它，他们将通过链接他们自己计算
html - Ruby Mechanize - 如何在 Mechanize 解析站点响应之前解析它？
问题: 解析网站时，有些字符会导致 Mechanize 无法正确解析。提出的解决方案解析来自网站的响应以删除这些字符在 Mechanize 之前尝试解析它。或者，在 Mechanize 解析网
ruby - 如何将新字段添加到 Mechanize 表单( ruby /Mechanize )
有一个public class method将字段添加到 Mechanize 表单我试过了.. #login_form.field.new('auth_login','Login') #login_
ruby-on-rails - Mechanize 重定向/Nokogiri(菜鸟使用 Mechanize )
我有一些看起来像这样的东西: def self.foo agent = Mechanize.new form = agent.get("link/to/form/url") form.f
ruby - 是否可以将 Mechanize::File 转换为 Mechanize::Page
我在使用 Mechanize gem 时遇到问题，如何转换 Mechanize::文件进入 Mechanize::页面 , 这是我的一段代码: **link** = page.link_with(:
ruby - 如何从 Mechanize::Page 的搜索方法中获取 Mechanize 对象？
我正在尝试抓取一个只能依靠类和元素层次结构来找到正确节点的站点。但是使用 Mechanize::Page#search 返回 Nokogiri::XML::Element，我不能用它来填写和提交表单等
ruby - Mechanize 的局限性是什么？ mechanize 和 watir 之间的区别是什么
我正在使用 mechanize 来抓取一些网页。我需要知道什么是 Mechanize 限制？ Mechanize 不能做什么？它可以执行网页中嵌入的javascripts吗？我可以用它来调用 j
perl - WWW::Mechanize 中的基本表单方法在 WWW::Mechanize::PhantomJS 中不起作用
在 WWW::Mechanize 中使用表单方法 my @form = $mech->form_number(1); foreach my $sum_form ( @form ) {
mechanize - 如何使用 Mechanize 单击没有 id 和 name 的提交按钮？
找到以下 HTML 代码: 如何使用 Mechanize 单击没有 id 和 name 的提交按钮？最佳答案我已经找到了此类场景的答案，代码如下: agent = Mechanize.new
ruby-on-rails - Rails rake Mechanize - 错误 - 没有要加载的文件 - Mechanize
这个问题不太可能对 future 的访客有帮助；它只与一个小的地理区域、一个特定的时刻或一个非常狭窄的情况相关，通常不适用于互联网的全局受众。如需帮助使这个问题更广泛地适用，visit the hel
ruby - Mechanize/ ruby : `require' : cannot load such file -- mechanize (LoadError)
我一直在尝试使用以下方法从终端运行 ruby 文件: ruby file_cleanse_auto.rb 但是我从 mechanize 得到一个错误: /Library/Ruby/Site/2.0
Ruby Mechanize gem，从本地 html 副本恢复 Mechanize::Page 对象
这是我拥有的代码: agent = Mechanize.new page = agent.get 'http://google.com' page.save 'google_index.htm' 我怎
python - Python Mechanize 错误 - "mechanize._mechanize.BrowserStateError: not viewing HTML"
for link in br.links(url_regex="inquiry-results.jsp"): cb[link.url] = link for page_link in cb.v
ruby-on-rails - 如何从 Mechanize::File 对象转换为 Mechanize::Page 对象？
我有一个登录表单的页面。登录后有一些重定向。第一个看起来像这样: #"no-cache=\"set-cookie\"", "content-length"=>"114", "set-cookie"=>
mechanize - 获取 Mechanize::UnauthorizedError: 401 => Net::HTTPUnauthorized 使用基本身份验证访问 API 时
我正在尝试使用基本身份验证访问 API。它适用于 HTTParty，但不适用于 2.7.6 Mechanize。这是我尝试过的: agent = Mechanize.new agent.log =
Ruby 的 Mechanize 在选择单选按钮时会犹豫(可能是因为它的名称大写)但 Perl 的 WWW::Mechanize 工作正常
我正在尝试使用 Ruby 的 Mechanize gem 提交表单。此表单有一组名为“KeywordType”的单选按钮。各个按钮的名称类似于 rdoAny、rdoAll 和 rdoPhrase。使用

首页

博学

6Ren·AI

商城

Python Mechanize 在处理网站 html 时遇到问题