python - 不可 JSON 序列化的项目的管道-6ren

python - 不可 JSON 序列化的项目的管道

转载作者：行者123 更新时间：2023-12-01 04:10:54

26

4

我正在尝试将抓取的 xml 输出写入 json。由于项目不可序列化，抓取失败。

从这个问题来看，它建议您需要构建一个管道，未提供的答案超出了问题 SO scrapy serializer 的范围。

所以指的是scrapy docs它举例说明了一个示例，但是文档建议不要使用此示例

The purpose of JsonWriterPipeline is just to introduce how to write item pipelines. If you really want to store all scraped items into a JSON file you should use the Feed exports.

如果我去 feed 导出，就会显示

JSON

FEED_FORMAT: json Exporter used: JsonItemExporter See this warning if you’re using JSON with large feeds.

我的问题仍然存在，因为据我所知，是从命令行执行的。

scrapy runspider myxml.py -o ~/items.json -t json

但是，这会产生我打算使用管道来解决的错误。

TypeError: <bound method SelectorList.extract of [<Selector xpath='.//@venue' data=u'Royal Randwick'>]> is not JSON serializable

如何创建 json 管道来纠正 json 序列化错误？

这是我的代码。

# -*- coding: utf-8 -*-
import scrapy
from scrapy.selector import Selector
from scrapy.http import HtmlResponse
from scrapy.selector import XmlXPathSelector
from conv_xml.items import ConvXmlItem
# https://stackoverflow.com/a/27391649/461887
import json


class MyxmlSpider(scrapy.Spider):
    name = "myxml"

    start_urls = (
        ["file:///home/sayth/Downloads/20160123RAND0.xml"]
    )

    def parse(self, response):
        sel = Selector(response)
        sites = sel.xpath('//meeting')
        items = []

        for site in sites:
            item = ConvXmlItem()
            item['venue'] = site.xpath('.//@venue').extract
            item['name'] = site.xpath('.//race/@id').extract()
            item['url'] = site.xpath('.//race/@number').extract()
            item['description'] = site.xpath('.//race/@distance').extract()
            items.append(item)

        return items


        # class JsonWriterPipeline(object):
        #
        #     def __init__(self):
        #         self.file = open('items.jl', 'wb')
        #
        #     def process_item(self, item, spider):
        #         line = json.dumps(dict(item)) + "\n"
        #         self.file.write(line)
        #         return item

最佳答案

问题出在这里:

item['venue'] = site.xpath('.//@venue').extract

您刚刚忘记调用extract。将其替换为:

item['venue'] = site.xpath('.//@venue').extract()

关于python - 不可 JSON 序列化的项目的管道，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/34972368/

26

4

0

文章推荐： java - JFace:Setgrayed 在树查看器中不起作用

文章推荐： java - 基于用户输入的结果

文章推荐： java - 动态JNLP从服务器获取文件

文章推荐： java - ConnectTimeout 并不总是阻止从输入流读取？

Grails 3 Assets 管道/咖啡 Assets 管道
我正在使用 Assets 管道来管理我的 Grails 3.0 应用程序的前端资源。但是，似乎没有创建 CoffeeScript 文件的源映射。有什么办法可以启用它吗？我的 build.gradle
jenkins-pipeline - 失败后继续 Tekton 管道(类似于 jenkins 管道 catchError 行为)
我有一个我想要的管道: 提供一些资源，运行一些测试，拆资源。我希望第 3 步中的拆卸任务运行不管测试是否通过或失败，在第 2 步。据我所知 runAfter如果前一个任务成功，则只运行一个任
PowerShell 管道
如果我运行以下命令: Measure-Command -Expression {gci -Path C:\ -Recurse -ea SilentlyContinue | where Extensio
Java输入解析与分隔符| (管道)
我知道管道是一个特殊字符，我需要使用: Scanner input = new Scanner(System.in); String line = input.next
Powershell 管道 - 返回一个在管道内创建的新对象
我再次遇到同样的问题，我有我的默认处理方式，但它一直困扰着我。有没有更好的办法？所以基本上我有一个运行的管道，在管道内做一些事情，并想从管道内返回一个键/值对。我希望整个管道返回一个类型为 ps
Azure 管道 - 阶段条件取决于
我有三个环境:dev、hml 和 qa。在我的管道中，根据分支，阶段有一个条件来检查它是否会运行: - stage: Project_Deploy_DEV condition: eq(varia
Jenkins 管道 - 为什么管道选项不显示
我有 Jenkins Jenkins ver. 2.82 正在运行并想在创建新作业时使用 Pipeline 功能。但我没有看到这个列为选项。我只能在自由式项目、maven 项目、外部项目和多配置之间进
haskell - 管道:产生内存泄漏
在对上一个问题 (haskell-data-hashset-from-unordered-container-performance-for-large-sets) 进行一些观察时，我偶然发现了一个奇
命令参数的 Unix 管道
我正在寻找有关如何使用管道将标准输出作为其他命令的参数传递的见解。例如，考虑这种情况: ls | grep Hello grep 的结构遵循以下模式:grep SearchTerm PathOfFi
Jenkinsfile 管道，返回警告但不会失败
有没有办法不因声明性管道步骤而失败，而是显示警告？目前我正在通过添加 || exit 0 来规避它到 sh 命令行的末尾，所以它总是可以正常退出。当前示例: sh 'vendor/bin/phpcs
Jenkins 管道 - 手动清除工作区？
我们正在从旧的 Jenkins 设置迁移到所有计划都是声明性 jenkinsfile 管道的新服务器……但是，通过使用管道，我们无法再手动清除工作区。我如何设置 Jenkins 以允许手动点播清理工
python - 管道:多个分类器？
我在 Python 中阅读了有关 Pipelines 和 GridSearchCV 的以下示例: http://www.davidsbatista.net/blog/2017/04/01/docume
Jenkins 管道 - 无法在空对象上调用方法阶段()
我有一个这样的管道脚本: node('linux'){ stage('Setup'){ echo "Build Stage" } stage('Build'){ echo
Bitbucket 管道 - 无法从远程存储库中读取？
我正在使用 bitbucket 管道进行培训这是我的 bitbucket-pipelines.yml: image: php:7.2.9 pipelines: default:
haskell - 管道 - 管道内的多个输出文件
我正在编写一个程序，其中输入文件被拆分为多个文件(Shamir 的 secret 共享方案)。这是我想象的管道: 来源:使用 Conduit.Binary.sourceFile 从输入中读取导管:
Jenkins 管道 - 阶段与时间和输入
我创建了一个管道，它有一个应该只在开发分支上执行的阶段。该阶段还需要用户输入。即使我在不同的分支上，为什么它会卡在这些步骤的用户输入上？当我提供输入时，它们会被正确跳过。 stage('Deplo
R 管道 (%>%) 不适用于复制功能
我正在尝试学习管道功能(％>％)。当试图从这行代码转换到另一行时，它不起作用。 ---- R代码--原版----- set.seed(1014) replicate(6,sample(1:8))
Jenkins 管道，如何将工件从以前的构建复制到当前构建？
在 Jenkins Pipeline 中，如何将工件从以前的构建复制到当前构建？即使之前的构建失败，我也想这样做。最佳答案 Stuart Rowe 还在 Pipeline Authoring Si
Jenkins 管道 - 使用参数构建
我正在尝试使用执行已定义的作业构建使用 Jenkins 管道的方法。这是一个简单的例子: build('jenkins-test-project-build', param1 : 'some-
Powershell 管道，其表现不符合预期
当我使用 where 过滤器通过管道命令排除对象时，它没有给我正确的输出。 PS C:\Users\Administrator> $proall = Get-ADComputer -filter *

首页

博学

6Ren·AI

商城

python - 不可 JSON 序列化的项目的管道