javascript - 跟踪页面的每个链接并抓取内容，Scrapy + Selenium-6ren

javascript - 跟踪页面的每个链接并抓取内容，Scrapy + Selenium

转载作者：行者123 更新时间：2023-11-27 23:25:52

This是我正在开发的网站。每页有一个表格，共 18 个帖子。我想访问每篇文章并抓取其内容，并对前 5 页重复此操作。

我的方法是让我的蜘蛛抓取 5 个页面中的所有链接并迭代它们以获取内容。因为“下一页”按钮和每篇文章中的某些文本是由 JavaScript 编写的，所以我使用 Selenium 和 Scrapy。我运行我的蜘蛛，可以看到 Firefox webdriver 显示前 5 页，但随后蜘蛛停止了，没有抓取任何内容。 Scrapy 也不会返回错误消息。

现在我怀疑失败的原因可能是:

1) 没有链接存储到 all_links 中。

2) 不知何故，parse_content 没有运行。

我的诊断可能是错误的，我需要帮助来查找问题。非常感谢!

这是我的蜘蛛:

import scrapy
from bjdaxing.items_bjdaxing import BjdaxingItem
from selenium import webdriver
from scrapy.http import TextResponse 
import time

all_links = [] # a global variable to store post links


class Bjdaxing(scrapy.Spider):
    name = "daxing"

    allowed_domains = ["bjdx.gov.cn"] # DO NOT use www in allowed domains
    start_urls = ["http://app.bjdx.gov.cn/cms/daxing/lookliuyan_bjdx.jsp"] # This has to start with http

    def __init__(self):
        self.driver = webdriver.Firefox()

    def parse(self, response):
        self.driver.get(response.url) # request the start url in the browser         

        i = 1

        while i <= 5: # The number of pages to be scraped in this session

            response = TextResponse(url = response.url, body = self.driver.page_source, encoding='utf-8') # Assign page source to response. I can treat response as if it's a normal scrapy project.           

            global all_links
            all_links.extend(response.xpath("//a/@href").extract()[0:18])

            next = self.driver.find_element_by_xpath(u'//a[text()="\u4e0b\u9875\xa0"]') # locate "next" button
            next.click() # Click next page            
            time.sleep(2) # Wait a few seconds for next page to load. 

            i += 1


    def parse_content(self, response):
        item = BjdaxingItem()
        global all_links
        for link in all_links: 
            self.driver.get("http://app.bjdx.gov.cn/cms/daxing/") + link

            response = TextResponse(url = response.url, body = self.driver.page_source, encoding = 'utf-8')

            if len(response.xpath("//table/tbody/tr[1]/td[2]/text()").extract() > 0):
                item['title'] =     response.xpath("//table/tbody/tr[1]/td[2]/text()").extract()
            else: 
                item['title'] = ""    

            if len(response.xpath("//table/tbody/tr[3]/td[2]/text()").extract() > 0):
                item['netizen'] =    response.xpath("//table/tbody/tr[3]/td[2]/text()").extract()
            else: 
                item['netizen'] = ""    

            if len(response.xpath("//table/tbody/tr[3]/td[4]/text()").extract() > 0):
                item['sex'] = response.xpath("//table/tbody/tr[3]/td[4]/text()").extract()
            else: 
                item['sex'] = ""   

            if len(response.xpath("//table/tbody/tr[5]/td[2]/text()").extract() > 0):
                item['time1'] = response.xpath("//table/tbody/tr[5]/td[2]/text()").extract()
            else: 
                item['time1'] = ""

            if len(response.xpath("//table/tbody/tr[11]/td[2]/text()").extract() > 0):
                item['time2'] =   response.xpath("//table/tbody/tr[11]/td[2]/text()").extract()
            else: 
                item['time2'] = "" 

            if len(response.xpath("//table/tbody/tr[7]/td[2]/text()").extract()) > 0:
                question = "".join(response.xpath("//table/tbody/tr[7]/td[2]/text()").extract())
                item['question'] = "".join(map(unicode.strip, question))
            else: item['question'] = ""  

            if len(response.xpath("//table/tbody/tr[9]/td[2]/text()").extract()) > 0:
                reply = "".join(response.xpath("//table/tbody/tr[9]/td[2]/text()").extract()) 
                item['reply'] = "".join(map(unicode.strip, reply))
            else: item['reply'] = ""    

            if len(response.xpath("//table/tbody/tr[13]/td[2]/text()").extract()) > 0:
                agency = "".join(response.xpath("//table/tbody/tr[13]/td[2]/text()").extract())
                item['agency'] = "".join(map(unicode.strip, agency))
            else: item['agency'] = ""    

            yield item

最佳答案

此处存在多个问题和可能的改进:

parse() 和 parse_content() 方法之间没有任何“链接”
使用全局变量通常是一种不好的做法
这里根本不需要selenium。要跟踪分页，您只需向同一网址发出 POST 请求并提供 currPage 参数

这个想法是使用.start_requests()并创建请求列表/队列来处理分页。按照分页并从表中收集链接。一旦请求队列为空，请切换到之前收集的链接。实现:

import json
from urlparse import urljoin

import scrapy


NUM_PAGES = 5

class Bjdaxing(scrapy.Spider):
    name = "daxing"

    allowed_domains = ["bjdx.gov.cn"] # DO NOT use www in allowed domains

    def __init__(self):
        self.pages = []
        self.links = []

    def start_requests(self):
        self.pages = [scrapy.Request("http://app.bjdx.gov.cn/cms/daxing/lookliuyan_bjdx.jsp",
                                     body=json.dumps({"currPage": str(page)}),
                                     method="POST",
                                     callback=self.parse_page,
                                     dont_filter=True)
                      for page in range(1, NUM_PAGES + 1)]

        yield self.pages.pop()

    def parse_page(self, response):
        base_url = response.url
        self.links += [urljoin(base_url, link) for link in response.css("table tr td a::attr(href)").extract()]

        try:
            yield self.pages.pop()
        except IndexError:  # no more pages to follow, going over the gathered links
            for link in self.links:
                yield scrapy.Request(link, callback=self.parse_content)

    def parse_content(self, response):
        # your parse_content method here

关于javascript - 跟踪页面的每个链接并抓取内容，Scrapy + Selenium，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/34968262/

文章推荐： c++ - 成员变量的通用声明

文章推荐： javascript - 模态显示不正确

selenium - Selenium IDE、Selenium RC 和 Selenium WebDriver 之间有什么区别？
Selenium IDE、Selenium RC 和 Selenium WebDriver 有什么区别；我们可以在什么样的项目中使用它们？任何建议将不胜感激。最佳答案 Selenium IDE 是一
selenium - 如何压缩 Selenium 客户端和 Selenium 服务器之间的传输
我的 Selenium 服务器在远程服务器上运行。我从我的本地 PC 启动我的 Selenium 脚本，它从网站获取数据。例如，我的 Selenium 脚本执行这段 JS 代码: JSON.stri
selenium - "//div[.//a[text()=' SELENIUM'] ]"and "//div[//a[text() ='SELENIUM' ]]"在 Selenium xpath中有什么区别
Selenium 中“//div[.//a[text()='SELENIUM']]”和“//div[//a[text()='SELENIUM']]”有什么区别xpath。有人可以澄清我在 xpath
selenium - Selenium 中每个测试的多个断言与单个断言？
我正在创建自动冒烟测试。我读到在单元测试中使用多个断言不是一个好的做法，这条规则是否也适用于使用 selenium 的 webdriver 测试？在我的冒烟测试中，有时我会使用 20 多个断言来验证
selenium - selenium IDE中添加两个变量
我在一个变量中存储了一个值，在另一个变量中存储了第二个值，现在我想将这两个数字相加。我无法做到这一点，我尝试过下面的代码，但它不起作用 store 6 w sto
selenium - Selenium 中回车键和回车键的区别
Selenium 中的回车键和回车键有什么区别？ This related SO answer并且提供的链接说明它们是不同的。我还注意到，在使用 Firefox 24.2 时，回车键将发送一个 HTM
selenium - 如何使用 Selenium 3 设置 Selenium Grid
以下是我遇到异常的详细信息: 当我使用以下命令启动节点时，出现如下错误: F:\SeleniumGrid\Jars>java -jar selenium-server-standalone-3.0.0
selenium - 是否有 Selenium 2 版本的 Selenium IDE？
我是的新手 Selenium 我对版本号有点困惑。 Selenium 2.0 2011年发布。我刚刚下载了 Selenium IDE Firefox 扩展，版本为 1.7.2 .是否还有 IDE 的
selenium - 我如何停止断言失败时关闭浏览器窗口的代码接收/ Selenium ？
我正在使用 Selenium 运行Codeception 2。我可以看到 Selenium 打开了浏览器并运行了测试。然后，我从代码接收中得到一个错误，即存在失败的断言。我知道有一个HTML文件可以
selenium - Selenium 3的新功能是什么
Closed. This question needs to be more focused。它当前不接受答案。想要改善这个问题吗？更新问题，使它仅关注editing this post的一个问题。
selenium - Selenium 运行中如何关闭弹出窗口？
我想关闭弹出窗口(已知的窗口名称)，然后返回到原始窗口。我该怎么办？如果我无法获得窗口中关闭按钮的常量。那么有没有达到目标的一般行为？最佳答案你有没有尝试过: selenium.Close()
selenium - 错误后如何继续使用webdriver/selenium
我正在用webdriver做一个测试机器人。我有一个场景，它单击一个按钮，打开一个新窗口，并且它通过特定的xpath搜索元素，但是有时没有这样的元素，因为可以将其禁用，并且出现此错误：org.open
selenium - Selenium :如何等待选择中的选项被填充？
我是第一次使用Selenium，对这些选项不知所措。我在Firefox中使用IDE。当我的页面加载时，它随后通过JSONP请求获取值，并在其中填充选择中的选项。我如何让Selenium等待选择中的
selenium-webdriver - 如何在运行 Selenium Selenium nightwatch.js测试时保持打开的开发人员工具？
我开始使用nightwatch.js编写e2e测试，我注意到我想在目标浏览器的控制台（开发人员工具）中手动检查一些错误。但总是在我打开开发者控制台时，浏览器会自动关闭它。这是selenium还是nig
selenium - Selenium 没有这种元素异常
我正在尝试使用以下方式刮除Glassdoor的评论: https://github.com/MatthewChatham/glassdoor-review-scraper 但是我得到了错误并且不知道如
selenium - Selenium Grid总是执行我的测试的多余实例
背景我设置了一个Selenium Grid项目，以在两种不同的浏览器Chrome和Firefox中执行测试。我正在使用Gradle执行测试。该测试将成功执行两次，一次按预期在Chrome中执行，一次
selenium - 使用 phpunit/selenium 保持 selenium 浏览器打开
当测试失败时，运行 selenium 测试的浏览器将关闭。这在尝试调试时没有帮助。我知道我可以在失败时选择屏幕截图，但如果没有整个上下文，这并没有帮助。在浏览器仍然可用的情况下，我可以回击并检查发生了
selenium - 使用 Selenium Web 驱动程序或 selenium RC
使用 Selenium Web 驱动程序而不是 Selenium RC 启动新的测试框架是个好主意吗？对于 Selenium Web 驱动程序，并非所有 Selenium 方法都已实现。那么使用 Se
selenium - Selenium 标识符的建议命名约定
我使用 selenium 页面对象模型来定义所有页面元素。我对元素命名所遵循的命名约定不太相信，并且感觉太长了。请对此提出建议。 @FindBy(xpath = "//tbody[@id='tabvi
selenium - 等待处理程序注册 - selenium
有一个带有按钮的 html 页面，我的 Selenium 测试正在测试，当单击按钮时，会执行一个操作。问题是，看起来点击发生在 javascript 执行之前 - 在处理程序绑定(bind)到页面之

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

javascript - 跟踪页面的每个链接并抓取内容，Scrapy + Selenium