python - Scrapy Authenticated Spider 获取内部服务器错误-6ren

python - Scrapy Authenticated Spider 获取内部服务器错误

转载作者：太空宇宙更新时间：2023-11-04 05:31:42

我正在尝试制作经过身份验证的蜘蛛。我在这里提到了几乎所有与 Scrapy 身份验证蜘蛛相关的帖子，我找不到我的问题的任何答案。我使用了以下代码:

import scrapy

from scrapy.spider import BaseSpider
from scrapy.selector import  Selector
from scrapy.http import FormRequest, Request
import  logging
from PWC.items import PwcItem


class PwcmoneySpider(scrapy.Spider):
    name = "PWCMoney"
    allowed_domains = ["pwcmoneytree.com"]
    start_urls = (
        'https://www.pwcmoneytree.com/SingleEntry/singleComp?compName=Addicaid',
    )

    def parse(self, response):
        return [scrapy.FormRequest("https://www.pwcmoneytree.com/Account/Login",
                                   formdata={'UserName': 'user', 'Password': 'pswd'},
                                   callback=self.after_login)]

    def after_login(self, response):
      if "authentication failed" in response.body:
        self.log("Login failed", level=logging.ERROR)
        return
    # We've successfully authenticated, let's have some fun!
    print("Login Successful!!")
    return Request(url="https://www.pwcmoneytree.com/SingleEntry/singleComp?compName=Addicaid",
               callback=self.parse_tastypage)


    def parse_tastypage(self, response):
      for sel in response.xpath('//div[@id="MainDivParallel"]'):
                                      item = PwcItem()
                                      item['name'] = sel.xpath('div[@id="CompDiv"]/h2/text()').extract()
                                      item['location'] = sel.xpath('div[@id="CompDiv"]/div[@id="infoPane"]/div[@class="infoSlot"]/div/a/text()').extract()
                                      item['region'] = sel.xpath('div[@id="CompDiv"]/div[@id="infoPane"]/div[@id="contactInfoDiv"]/div[1]/a[2]/text()').extract()
                                      yield item

我得到了以下输出:

Microsoft Windows [Version 6.1.7601]
Copyright (c) 2009 Microsoft Corporation.  All rights reserved.

C:\Python27\PWC>scrapy crawl PWCMoney -o test.csv
2016-04-29 11:37:35 [scrapy] INFO: Scrapy 1.0.5 started (bot: PWC)
2016-04-29 11:37:35 [scrapy] INFO: Optional features available: ssl, http11
2016-04-29 11:37:35 [scrapy] INFO: Overridden settings: {'NEWSPIDER_MODULE': 'PW
C.spiders', 'FEED_FORMAT': 'csv', 'SPIDER_MODULES': ['PWC.spiders'], 'FEED_URI':
 'test.csv', 'BOT_NAME': 'PWC'}
2016-04-29 11:37:35 [scrapy] INFO: Enabled extensions: CloseSpider, FeedExporter
, TelnetConsole, LogStats, CoreStats, SpiderState
2016-04-29 11:37:36 [scrapy] INFO: Enabled downloader middlewares: HttpAuthMiddl
eware, DownloadTimeoutMiddleware, UserAgentMiddleware, RetryMiddleware, DefaultH
eadersMiddleware, MetaRefreshMiddleware, HttpCompressionMiddleware, RedirectMidd
leware, CookiesMiddleware, ChunkedTransferMiddleware, DownloaderStats
2016-04-29 11:37:36 [scrapy] INFO: Enabled spider middlewares: HttpErrorMiddlewa
re, OffsiteMiddleware, RefererMiddleware, UrlLengthMiddleware, DepthMiddleware
2016-04-29 11:37:36 [scrapy] INFO: Enabled item pipelines:
2016-04-29 11:37:36 [scrapy] INFO: Spider opened
2016-04-29 11:37:36 [scrapy] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 i
tems (at 0 items/min)
2016-04-29 11:37:36 [scrapy] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-04-29 11:37:37 [scrapy] DEBUG: Retrying <POST https://www.pwcmoneytree.com/
Account/Login> (failed 1 times): 500 Internal Server Error
2016-04-29 11:37:38 [scrapy] DEBUG: Retrying <POST https://www.pwcmoneytree.com/
Account/Login> (failed 2 times): 500 Internal Server Error
2016-04-29 11:37:38 [scrapy] DEBUG: Gave up retrying <POST https://www.pwcmoneyt
ree.com/Account/Login> (failed 3 times): 500 Internal Server Error
2016-04-29 11:37:38 [scrapy] DEBUG: Crawled (500) <POST https://www.pwcmoneytree
.com/Account/Login> (referer: None)
2016-04-29 11:37:38 [scrapy] DEBUG: Ignoring response <500 https://www.pwcmoneyt
ree.com/Account/Login>: HTTP status code is not handled or not allowed
2016-04-29 11:37:38 [scrapy] INFO: Closing spider (finished)
2016-04-29 11:37:38 [scrapy] INFO: Dumping Scrapy stats:
{'downloader/request_bytes': 954,
 'downloader/request_count': 3,
 'downloader/request_method_count/POST': 3,
 'downloader/response_bytes': 30177,
 'downloader/response_count': 3,
 'downloader/response_status_count/500': 3,
 'finish_reason': 'finished',
 'finish_time': datetime.datetime(2016, 4, 29, 6, 7, 38, 674000),
 'log_count/DEBUG': 6,
 'log_count/INFO': 7,
 'response_received_count': 1,
 'scheduler/dequeued': 3,
 'scheduler/dequeued/memory': 3,
 'scheduler/enqueued': 3,
 'scheduler/enqueued/memory': 3,
 'start_time': datetime.datetime(2016, 4, 29, 6, 7, 36, 193000)}
2016-04-29 11:37:38 [scrapy] INFO: Spider closed (finished)

由于我是 python 和 Scrapy 的新手，我似乎无法理解错误，希望这里的人能帮助我。

所以，我修改了这样的代码 Rejected的建议，只显示修改的部分:

allowed_domains = ["pwcmoneytree.com"]
    start_urls = (
        'https://www.pwcmoneytree.com/Account/Login',
    )

    def start_requests(self):
        return [scrapy.FormRequest.from_response("https://www.pwcmoneytree.com/Account/Login",
                                   formdata={'UserName': 'user', 'Password': 'pswd'},
                                   callback=self.logged_in)]

出现如下错误:

C:\Python27\PWC>scrapy crawl PWCMoney -o test.csv
2016-04-30 11:04:47 [scrapy] INFO: Scrapy 1.0.5 started (bot: PWC)
2016-04-30 11:04:47 [scrapy] INFO: Optional features available: ssl, http11
2016-04-30 11:04:47 [scrapy] INFO: Overridden settings: {'NEWSPIDER_MODULE': 'PW
C.spiders', 'FEED_FORMAT': 'csv', 'SPIDER_MODULES': ['PWC.spiders'], 'FEED_URI':
 'test.csv', 'BOT_NAME': 'PWC'}
2016-04-30 11:04:50 [scrapy] INFO: Enabled extensions: CloseSpider, FeedExporter
, TelnetConsole, LogStats, CoreStats, SpiderState
2016-04-30 11:04:54 [scrapy] INFO: Enabled downloader middlewares: HttpAuthMiddl
eware, DownloadTimeoutMiddleware, UserAgentMiddleware, RetryMiddleware, DefaultH
eadersMiddleware, MetaRefreshMiddleware, HttpCompressionMiddleware, RedirectMidd
leware, CookiesMiddleware, ChunkedTransferMiddleware, DownloaderStats
2016-04-30 11:04:54 [scrapy] INFO: Enabled spider middlewares: HttpErrorMiddlewa
re, OffsiteMiddleware, RefererMiddleware, UrlLengthMiddleware, DepthMiddleware
2016-04-30 11:04:54 [scrapy] INFO: Enabled item pipelines:
Unhandled error in Deferred:
2016-04-30 11:04:54 [twisted] CRITICAL: Unhandled error in Deferred:


Traceback (most recent call last):
  File "c:\python27\lib\site-packages\scrapy\cmdline.py", line 150, in _run_comm
and
    cmd.run(args, opts)
  File "c:\python27\lib\site-packages\scrapy\commands\crawl.py", line 57, in run

    self.crawler_process.crawl(spname, **opts.spargs)
  File "c:\python27\lib\site-packages\scrapy\crawler.py", line 153, in crawl
    d = crawler.crawl(*args, **kwargs)
  File "c:\python27\lib\site-packages\twisted\internet\defer.py", line 1274, in
unwindGenerator
    return _inlineCallbacks(None, gen, Deferred())
--- <exception caught here> ---
  File "c:\python27\lib\site-packages\twisted\internet\defer.py", line 1128, in
_inlineCallbacks
    result = g.send(result)
  File "c:\python27\lib\site-packages\scrapy\crawler.py", line 72, in crawl
    start_requests = iter(self.spider.start_requests())
  File "C:\Python27\PWC\PWC\spiders\PWCMoney.py", line 16, in start_requests
    callback=self.logged_in)]
  File "c:\python27\lib\site-packages\scrapy\http\request\form.py", line 36, in
from_response
    kwargs.setdefault('encoding', response.encoding)
exceptions.AttributeError: 'str' object has no attribute 'encoding'
2016-04-30 11:04:54 [twisted] CRITICAL:

最佳答案

从您的错误日志中可以看出，是对 https://www.pwcmoneytree.com/Account/Login 的 POST 请求给您带来了 500 错误。

我尝试使用 POSTman 手动发出相同的 POST 请求.它给出了 500 错误代码和包含此错误消息的 HTML 页面:

The required anti-forgery cookie "__RequestVerificationToken" is not present.

这是许多 API 和网站用来防止 CSRF 攻击的功能。如果您仍想抓取网站，则必须先访问登录表单并在登录前获取正确的 cookie。

关于python - Scrapy Authenticated Spider 获取内部服务器错误，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/36930865/

文章推荐： c - C 中的 -mno-sse 标志和 gettimeofday() 出错

文章推荐： Python 过滤器脚本

文章推荐： python - 在 linux 中执行远程命令时如何处理错误

文章推荐： python - 将 curl 调用转换为 python 请求

authentication - OKHttp Authenticator 401和407以外的自定义http代码
我在服务器端实现了 oauth token ，但在无效 token 或 token 过期时，我收到 200 http 状态代码，但在响应正文中我有{"code":"4XX", "data":{"som
authentication - 未找到 sinatra-authentication 路线
我正在尝试将 sinatra-authentication gem 添加到 Sinatra 应用程序中，虽然它在那里并完成了它的一部分工作，但由于某种原因，路由似乎没有被添加。代码基础: requir
authentication - 移动应用程序 : how do I provide client authentication
我有一个健身移动应用程序的想法，我一直在为 iPhone(基于 Obj-C)、Android(基于 Java)、WebOS(基于 html5)和诺基亚 Qt 开发基于这个想法的应用程序。我现在需要向
authentication - SessionId/Authentication Token 生成最佳实践
我见过有人使用 UUID 生成身份验证 token 。然而，在 RFC 4122据说 Do not assume that UUIDs are hard to guess; they should n
authentication - 如何使用 pouchdb-authentication 检查用户名是否已存在？
上下文如下。 pouchdb-authentication API没有为此提供明确的方法。我考虑过使用db.getUser(username [, opts][, callback]) 。然而，该方法
authentication - Edge 浏览器中的 "Basic authentication"
Edge 浏览器中的“基本身份验证”没有保存密码的选项。当浏览器关闭并重新打开时，用户必须重新输入密码。有没有人解决这个问题？最佳答案它仍然存在并且仍在工作，他们只是从那些对话框窗口中删除了复选
authentication - oAuth Authentication Twitter - 如何自动登录
嗨，我需要知道如何在 iPhone 上使用 oAuth for twitter 自动登录帐户。该应用程序应登录并向用户显示该帐户的提要。最佳答案 OAuthentication 需要几个阶段，您可以
authentication - Edge 浏览器中的 "Basic authentication"
Edge 浏览器中的“基本身份验证”没有保存密码的选项。当浏览器关闭并重新打开时，用户必须重新输入密码。有没有人解决这个问题？最佳答案它仍然存在并且仍在工作，他们只是从那些对话框窗口中删除了复选
authentication - 信箱不可用。服务器响应是 : Authentication Required
关闭。这个问题是off-topic .它目前不接受答案。想改进这个问题吗？ Update the question所以它是on-topic用于堆栈溢出。关闭 10 年前。 Improve thi
authentication - OpenVAS 命令获取 Failed Authentication 错误
在尝试运行一些 OpenVAS CLI 命令时，我收到 Failed Authentication 错误消息。 OpenVAS 安装在 CentOS 机器上。我尝试使用用户帐户凭据，但仍然收到相同的错
restful-authentication - 设计一个 web api : How to authenticate?
我正在设计一个 web api。我需要让用户对自己进行身份验证。我有点犹豫让用户以明文形式传递他们的用户名/密码.. 类似于:api.mysite.com/auth.php?user=x&pass=y
authentication - Spring 安全 : Authentication user manually
我尝试通过 oAuth 在 Spring Security 应用程序中验证用户。我已经收到 token 和用户数据。如何在没有密码和经典登录表单的情况下手动验证用户？谢谢你。最佳答案像这样的东
authentication - Symfony 4 中的 Guard Authenticator
我正在 Symfony 4 中创建一个简单的登录身份验证系统并使用安全组件 Guard。我的 FormLoginAuthenticator 如下: router = $router;
authentication - react 路由器 : Handling role authentication
我正在开发一个具有多个角色的网络应用程序。我想到了一种方法，可以使用 React Router 通过 onEnter 触发器来限制对某些路由的访问。现在我想知道这是否是防止访问未经授权的页面的可靠方
http-authentication - 多个方案的 WWW-Authenticate 的分隔符是什么？
我已通读 RFC 2617如果支持多种方案，则无法在那里或其他任何地方找到分隔符。例如，假设支持 Basic 和 Digest。我知道它可能会以这种方式出现: HTTP/1.1 401 Unautho
authentication - OWIN - Authentication.SignOut() 似乎没有删除 cookie
我在 OWIN Cookie 身份验证方面遇到了一些问题。我有一个 .Net 站点，它有一些 MVC 页面，这些页面使用 cookie 身份验证和受不记名 token 保护的 WebAPI 资源。当
authentication - Telnet 自动化与 Expect : Slow authentication?
我正在使用 Telnet 向 Mikrotik 路由器发送命令。 telnet 192.168.100.100 -l admin Password: pass1234 [admin@ZYMMA] >
authentication - 远程图像嵌入 : how to handle ones that require authentication?
我管理着一个庞大而活跃的论坛，但我们正被一个非常严重的问题所困扰。我们允许用户嵌入远程图像，就像 stackoverflow 处理图像 (imgur) 的方式一样，但是我们没有一组特定的主机，可以使用
authentication - Drools KnowledgeAgent 和 Guvnor : authentication fails
这个真的让我抓狂。我在 JBoss AS 中有一个 Guvnor(稍后会详细介绍)。我编辑了 components.xml 以启用身份验证(使用 JAAS，我已经很好地设置了用户和密码)和基于角色的
authentication - IIS/冷融合 : authentication request on anonymous content
我们有一个管理站点，需要身份验证才能访问。站点上的页面包裹在 Coldfusion 自定义标签中，其中包括所有样式和 JS，以及一些其他信息。我最近制作了一份自定义标签包装器的副本。我将副本放在与原

太空宇宙

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - Scrapy Authenticated Spider 获取内部服务器错误