python - Scrapy python csv 输出在每行之间有空行-6ren

python - Scrapy python csv 输出在每行之间有空行

转载作者：太空狗更新时间：2023-10-30 01:32:16

我在生成的 csv 输出文件中的每一行 scrapy 输出之间得到不需要的空行。

我已经从 python2 迁移到 python 3，并且我使用的是 Windows 10。因此，我正在为 python3 调整我的 scrapy 项目。

我当前(也是目前唯一的)问题是，当我将 scrapy 输出写入 CSV 文件时，每行之间有一个空行。这已在此处的多个帖子中突出显示(与 Windows 相关)，但我无法找到有效的解决方案。

碰巧的是，我还在 piplines.py 文件中添加了一些代码，以确保 csv 输出按给定的列顺序而不是随机顺序。因此，我可以使用普通的 scrapy crawl charleschurch 来运行这段代码，而不是 scrapy crawl charleschurch -o charleschurch2017xxxx.csv

有谁知道如何在 CSV 输出中跳过/省略此空白行？

我的 pipelines.py 代码如下(我可能不需要 import csv 行，但我想我可能会为最终答案做):

# -*- coding: utf-8 -*-

# Define your item pipelines here
#
# Don't forget to add your pipeline to the ITEM_PIPELINES setting
# See: http://doc.scrapy.org/en/latest/topics/item-pipeline.html

import csv
from scrapy import signals
from scrapy.exporters import CsvItemExporter

class CSVPipeline(object):

  def __init__(self):
    self.files = {}

  @classmethod
  def from_crawler(cls, crawler):
    pipeline = cls()
    crawler.signals.connect(pipeline.spider_opened, signals.spider_opened)
    crawler.signals.connect(pipeline.spider_closed, signals.spider_closed)
    return pipeline

  def spider_opened(self, spider):
    file = open('%s_items.csv' % spider.name, 'w+b')
    self.files[spider] = file
    self.exporter = CsvItemExporter(file)
    self.exporter.fields_to_export = ["plotid","plotprice","plotname","name","address"]
    self.exporter.start_exporting()

  def spider_closed(self, spider):
    self.exporter.finish_exporting()
    file = self.files.pop(spider)
    file.close()

  def process_item(self, item, spider):
    self.exporter.export_item(item)
    return item

我将这一行添加到 settings.py 文件中(不确定 300 的相关性):

ITEM_PIPELINES = {'CharlesChurch.pipelines.CSVPipeline': 300 }

我的 scrapy 代码如下:

import scrapy
from urllib.parse import urljoin

from CharlesChurch.items import CharleschurchItem

class charleschurchSpider(scrapy.Spider):
    name = "charleschurch"
    allowed_domains = ["charleschurch.com"]    
    start_urls = ["https://www.charleschurch.com/county-durham_willington/the-ridings-1111"]


    def parse(self, response):

        for sel in response.xpath('//*[@id="aspnetForm"]/div[4]'):
           item = CharleschurchItem()
           item['name'] = sel.xpath('//*[@id="XplodePage_ctl12_dsDetailsSnippet_pDetailsContainer"]/span[1]/b/text()').extract()
           item['address'] = sel.xpath('//*[@id="XplodePage_ctl12_dsDetailsSnippet_pDetailsContainer"]/div/*[@itemprop="postalCode"]/text()').extract()
           plotnames = sel.xpath('//div[@class="housetype js-filter-housetype"]/div[@class="housetype__col-2"]/div[@class="housetype__plots"]/div[not(contains(@data-status,"Sold"))]/div[@class="plot__name"]/a/text()').extract()
           plotnames = [plotname.strip() for plotname in plotnames]
           plotids = sel.xpath('//div[@class="housetype js-filter-housetype"]/div[@class="housetype__col-2"]/div[@class="housetype__plots"]/div[not(contains(@data-status,"Sold"))]/div[@class="plot__name"]/a/@href').extract()
           plotids = [plotid.strip() for plotid in plotids]
           plotprices = sel.xpath('//div[@class="housetype js-filter-housetype"]/div[@class="housetype__col-2"]/div[@class="housetype__plots"]/div[not(contains(@data-status,"Sold"))]/div[@class="plot__price"]/text()').extract()
           plotprices = [plotprice.strip() for plotprice in plotprices]
           result = zip(plotnames, plotids, plotprices)
           for plotname, plotid, plotprice in result:
               item['plotname'] = plotname
               item['plotid'] = plotid
               item['plotprice'] = plotprice
               yield item

最佳答案

我怀疑这不太理想，但我找到了解决此问题的方法。在 pipelines.py 文件中，我添加了更多代码，这些代码实质上是将带有空行的 csv 文件读取到列表中，因此删除空行，然后将清理后的列表写入新文件。

我添加的代码是:

with open('%s_items.csv' % spider.name, 'r') as f:
  reader = csv.reader(f)
  original_list = list(reader)
  cleaned_list = list(filter(None,original_list))

with open('%s_items_cleaned.csv' % spider.name, 'w', newline='') as output_file:
    wr = csv.writer(output_file, dialect='excel')
    for data in cleaned_list:
      wr.writerow(data)

所以整个 pipelines.py 文件是:

# -*- coding: utf-8 -*-

# Define your item pipelines here
#
# Don't forget to add your pipeline to the ITEM_PIPELINES setting
# See: http://doc.scrapy.org/en/latest/topics/item-pipeline.html

import csv
from scrapy import signals
from scrapy.exporters import CsvItemExporter

class CSVPipeline(object):

  def __init__(self):
    self.files = {}

  @classmethod
  def from_crawler(cls, crawler):
    pipeline = cls()
    crawler.signals.connect(pipeline.spider_opened, signals.spider_opened)
    crawler.signals.connect(pipeline.spider_closed, signals.spider_closed)
    return pipeline

  def spider_opened(self, spider):
    file = open('%s_items.csv' % spider.name, 'w+b')
    self.files[spider] = file
    self.exporter = CsvItemExporter(file)
    self.exporter.fields_to_export = ["plotid","plotprice","plotname","name","address"]
    self.exporter.start_exporting()

  def spider_closed(self, spider):
    self.exporter.finish_exporting()
    file = self.files.pop(spider)
    file.close()

    #given I am using Windows i need to elimate the blank lines in the csv file
    print("Starting csv blank line cleaning")
    with open('%s_items.csv' % spider.name, 'r') as f:
      reader = csv.reader(f)
      original_list = list(reader)
      cleaned_list = list(filter(None,original_list))

    with open('%s_items_cleaned.csv' % spider.name, 'w', newline='') as output_file:
        wr = csv.writer(output_file, dialect='excel')
        for data in cleaned_list:
          wr.writerow(data)

  def process_item(self, item, spider):
    self.exporter.export_item(item)
    return item


class CharleschurchPipeline(object):
    def process_item(self, item, spider):
        return item

不理想，但暂时解决了问题。

关于python - Scrapy python csv 输出在每行之间有空行，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/43472847/

文章推荐： python - 是否可以更改本地人的口述？

文章推荐： python - 如何检测哪些圆圈被填充 + OpenCV + Python

文章推荐： python - Doc2Vec 句子聚类

文章推荐： python - KeyError 读取 csv 文件并传输到数组

scrapy - 如何在 scrapy Shell 中使用 scrapy 中间件？
在一个 scrapy 项目中，人们经常使用中间件。在交互式 session 期间是否也有一种通用方法可以在 scrapy shell 中启用中间件？最佳答案尽管如此，在 setting.py 中设
scrapy - scrapy-splash如何处理无限滚动？
我想对网页中向下滚动生成的内容进行反向工程。问题出在url https://www.crowdfunder.com/user/following_page/80159?user_id=80159&li
scrapy - 相对URL到绝对URL Scrapy
我需要帮助将相对URL转换为Scrapy Spider中的绝对URL。我需要将起始页面上的链接转换为绝对URL，以获取起始页面上已草稿的项目的图像。我尝试使用不同的方法来实现此目标失败，但是我陷入了
scrapy - Python Scrapy 错误。不再支持使用多个蜘蛛运行 'scrapy crawl'
我在 Scrapy Python 中制作了一个脚本，它在几个月内一直运行良好(没有更改)。最近，当我在 Windows Powershell 中执行脚本时，它引发了下一个错误: scrapy craw
scrapy - 飞溅内存限制(scrapy)
我已经从 docker 启动了 splash。我为 splash 和 scrapy 创建了大的 lua 脚本，然后它运行我看到了问题: Lua error: error in __gc metamet
scrapy - 让 Scrapy 从上一个断点继续爬行
我正在使用scrapy 来抓取网站，但发生了不好的事情(断电等)。我想知道我怎样才能从它坏了的地方继续爬行。我不想从种子开始。最佳答案这可以通过将预定的请求持久化到磁盘来完成。 scrapy c
scrapy - Scrapy 暂停/恢复如何工作？
有人可以向我解释一下 Scrapy 中的暂停/恢复功能是如何实现的吗？作品？ scrapy的版本我正在使用的是 0.24.5 documentation没有提供太多细节。我有以下简单的蜘蛛: cla
scrapy - Apscheduler+scrapy 信号仅适用于主线程
我想将 apscheduler 与 scrapy.but 我的代码是错误的。我应该如何修改它？ settings = get_project_settings() configure_logging
scrapy - 为什么 Scrapy 很慢？
我正在抓取一个网站并解析一些内容和图像，但即使对于 100 页左右的简单网站，完成这项工作也需要数小时。我正在使用以下设置。任何帮助将不胜感激。我已经看过这个问题- Scrapy 's Scrapyd
scrapy - 为什么 Scrapy 很慢？
我正在抓取一个网站并解析一些内容和图像，但即使对于 100 页左右的简单网站，完成这项工作也需要数小时。我正在使用以下设置。任何帮助将不胜感激。我已经看过这个问题- Scrapy 's Scrapyd
scrapy - 使用 Scrapy 增量爬取网站
我是爬行新手，想知道是否可以使用 Scrapy 逐步爬行网站，例如 CNBC.com？例如，如果今天我从一个站点抓取所有页面，那么从明天开始我只想收集新发布到该站点的页面，以避免抓取所有旧页面。感谢
scrapy - 如何使用 Scrapy 下载图片？
我是scrapy的新手。我正在尝试从 here 下载图像.我在关注 Official-Doc和 this article . 我的 settings.py 看起来像: BOT_NAME = 'shop
python - Scrapy:已抓取 0 页(可在 scrapy shell 中使用，但不适用于 scrapy crawl spider 命令)
我在使用 scrapy 时遇到了一些问题。它没有返回任何结果。我试图将以下蜘蛛复制并粘贴到 scrapy shell 中，它确实有效。真的不确定问题出在哪里，但是当我用“scrapy crawl rx
scrapy - 使用 Scrapy 抓取多个 URL
如何使用 Scrapy 抓取多个 URL？我是否被迫制作多个爬虫？ class TravelSpider(BaseSpider): name = "speedy" allowed_d
scrapy - 如何确保 scrapy-splash 已成功渲染整个页面
当我使用splash渲染整个目标页面来爬取整个网站时出现问题。某些页面不是随机成功的，所以我错误地获取了支持渲染工作完成后出现的信息。这意味着我尽管我可以从其他渲染结果中获取全部信息，但仅从渲染结果中
scrapy - 使用 Scrapy 抓取多个 URL
如何使用 Scrapy 抓取多个 URL？我是否被迫制作多个爬虫？ class TravelSpider(BaseSpider): name = "speedy" allowed_d
scrapy - 如何将所有 CPU 内核用于 Scrapy
我的scrapy程序无论如何只使用一个CPU内核CONCURRENT_REQUESTS我做。 scrapy中的某些方法是否可以在一个scrapy爬虫中使用所有cpu核心？ ps:好像有争论max_pr
python - Scrapy - 动态等待页面加载 - selenium + scrapy
我最近用 python 和 Selenium 做了一个网络爬虫，我发现它做起来非常简单。该页面使用 ajax 调用来加载数据，最初我等待固定的 time_out 来加载页面。这工作了一段时间。之后，我
python - Scrapy:scrapy server需要一个项目，为什么？
我想用这个命令运行 scrapy 服务器: scrapy server 它失败了，因为没有项目。然后我创建一个空项目来运行服务器，并成功部署另一个项目。但是，scrapy 服务器无法处理这个项目，并告
python - Scrapy - 在一个 scrapy 脚本中抓取不同的网页
我正在创建一个网络应用程序，用于从不同网站抓取一长串鞋子。这是我的两个单独的 scrapy 脚本: http://store.nike.com/us/en_us/pw/mens-clearance-s

太空狗

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - Scrapy python csv 输出在每行之间有空行