python - 将两个 python 数据处理脚本组合成一个工作流-6ren

python - 将两个 python 数据处理脚本组合成一个工作流

转载作者：太空宇宙更新时间：2023-11-04 05:19:12

25

4

我目前正在处理一项数据处理任务。

我有两个 python 脚本，每个脚本都实现了一个单独的功能，但它们对相同的数据进行操作，我认为它们可以组合成一个单一的工作流程，但我想不出最合乎逻辑的方法来实现这一点。

数据文件是here ，它是 JSON，但它有两个不同的组件。

第一部分是这样的:

{
    "links": {
        "self": "http://localhost:2510/api/v2/jobs?skills=data%20science"
    },
    "data": [
        {
            "id": 121,
            "type": "job",
            "attributes": {
                "title": "Data Scientist",
                "date": "2014-01-22T15:25:00.000Z",
                "description": "Data scientists are in increasingly high demand amongst tech companies in London. Generally a combination of business acumen and technical skills are sought. Big data experience ..."
            },
            "relationships": {
                "location": {
                    "links": {
                        "self": "http://localhost:2510/api/v2/jobs/121/location"
                    },
                    "data": {
                        "type": "location",
                        "id": 3
                    }
                },
                "country": {
                    "links": {
                        "self": "http://localhost:2510/api/v2/jobs/121/country"
                    },
                    "data": {
                        "type": "country",
                        "id": 1
                    }
                },

它是由第一个 python 脚本处理的，在这里:

import json
from collections import defaultdict
from pprint import pprint

with open('data-science.txt') as data_file:
    data = json.load(data_file)

locations = defaultdict(int)

for item in data['data']:
    location = item['relationships']['location']['data']['id']
    locations[location] += 1

pprint(locations)

呈现这种形式的数据:

这些是位置 “id” 和分配给该位置的记录数。

JSON 对象的另一部分如下所示:

"included": [
    {
        "id": 3,
        "type": "location",
        "attributes": {
            "name": "Victoria",
            "coord": [
                51.503378,
                -0.139134
            ]
        }
    },

并由此 python 文件处理:

import json
from collections import defaultdict
from pprint import pprint

with open('data-science.txt') as data_file:
    data = json.load(data_file)

locations = defaultdict(int)

for record in data['included']:
    id = record.get('id', None)
    name = record.get('attributes', {}).get('name', None)
    coord = record.get('attributes', {}).get('coord', None)
    print(id, name, coord)

它以这种格式输出数据:

3 Victoria [51.503378, -0.139134]
1 United Kingdom None
71 data science None
32 None None
3 Victoria [51.503378, -0.139134]
1 United Kingdom None
1 data mining None
22 data analysis None
33 sdlc None
38 artificial intelligence None
39 machine learning None
40 software development None
71 data science None
93 devops None
63 None None
52 Cubitt Town [51.505199, -0.018848]

我真正想要的是最终输出看起来像这样:

3, Victoria, [51.503378, -0.139134], 2673

其中 2673 引用第一个脚本中的作业计数。

如果它没有任何坐标，例如[51.503378, -0.139134] 我可以把它扔掉。

我确信可以将这些脚本组合在一起并获得该输出，但我不是一个全面的思考者，我不知道该怎么做。

所有真实项目文件live here .

最佳答案

使用函数 是结合这两个脚本的一种方法，毕竟它们处理相同的数据。因此，您应该为每个处理逻辑 block 创建一个函数，然后最后合并结果:

import json
from collections import defaultdict
from pprint import pprint

def process_locations_data(data):
    # processes the 'data' block
    locations = defaultdict(int)
    for item in data['data']:
        location = item['relationships']['location']['data']['id']
        locations[location] += 1
    return locations

def process_locations_included(data):
    # processes the 'included' block
    return_list = []
    for record in data['included']:
        id = record.get('id', None)
        name = record.get('attributes', {}).get('name', None)
        coord = record.get('attributes', {}).get('coord', None)
        return_list.append((id, name, coord))
    return return_list    # return list of tuples

# load the data from file once
with open('data-science.txt') as data_file:
    data = json.load(data_file)

# use the two functions on same data
locations = process_locations_data(data)
records = process_locations_included(data)

# combine the data for printing
for record in records:
    id, name, coord = record
    references = locations[id]   # lookup the references in the dict
    print id, name, coord, references

该函数可以有更好的名称，但这应该可以实现您正在寻找的统一。

关于python - 将两个 python 数据处理脚本组合成一个工作流，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/40896330/

25

4

0

文章推荐： python - 需要更好的逻辑来查找范围内的回文数

文章推荐： linux - iftop - 如何找到这些端口相关的进程

文章推荐： python - 在python中解析sql查询

组合/值的mysql分布
我有一个 mysql 表，其中包含一些随机数字组合。为简单起见，以下表为例: index|n1|n2|n3 1 1 2 3 2 4 10 32 3 3 10 4 4
SQL - 组合 AND & OR
我有以下代码: SELECT sdd.sd_doc_classification, sdd.sd_title, sdd.sd_desc, sdr.sd_upl
组合 2 个数据帧时按日期重复变量
如果我有两个要合并的数据框 Date RollingSTD 01/06/2012 0.16 01/07/2012 0.18 01/08/2012 0.17 01/09/20
clojure - 没有码头的环/组合
我知道可以使用 lein ring war 创建一个 war 文件，但它似乎仍然包含码头依赖项。当我构建 war (并在 tomcat 上部署)时，有没有办法排除码头依赖项？如果我根本不能做这件事，
clone - 封装聚合/组合
维基百科关于封装的文章指出: “封装还通过防止用户将组件的内部数据设置为无效或不一致的状态来保护组件的完整性” 我在一个论坛上开始讨论封装，在那里我问你是否应该始终在 setter 和/或 gette
带有复选框的 ExtJS 组合
对于我使用的组合框内的复选框: AOEDComboAssociationName = new Ext.form.ComboBox({ id: 'AOEDComboAssociationName',
c# - 组合 Where 语句的表达式
这个问题在这里已经有了答案: 关闭 10 年前。 Possible Duplicate: How do I combine LINQ expressions into one? public boo
rust - 组合/排列的数量
如何在 rust 中找到排列或组合的数量？例如C(10,6) = 210 我在标准库中找不到这个函数，也找不到那里的阶乘运算符(这就足够了)。最佳答案以@vallentin 的回答为基础，可以进
泛型类型的 Scala 组合
我有一个复杂的泛型类型用例，已在下面进行了简化 trait A class AB extends A{ val v = 10 } trait X[T<:A]{ def request: T }
Hibernate 标准限制 AND/OR 组合
如何使用 Hibernate 限制来实现此目的？ (((A='X') and (B in('X',Y))) or ((A='Y') and (B='Z'))) 最佳答案思考有效 Criteria c
javascript - 在谷歌条形图上绘制直线(组合)
我一定会在我的一个项目中使用谷歌图表。我需要的是，显示一个条形图，并且在条形图中，与每个条形相交的线代表另一个值。如果您查看下面的 jsfiddle，您会发现折线图仅与中间的条形图相交，并继续向其他条
javascript - 组合/匹配数组
只是一个简单的问题，我也很想得到答案，因为我不能百分百理解 Javascript 示例:假设您提示用户输入名称。够简单吧？但是你有一个数组，上面写着一些名字(其中之一就是)，基本上就是我到目前为止所说
具有两个参数的 Haskell 组合
我试图通过 Haskell 理解函数式编程，但在处理函数组合时遇到了很多麻烦。其实我有这两个功能: add:: Integer -> Integer -> Integer add x y = x
Realm :组合 "or"和 "and"
我正在寻找一种在 Realm 查询中组合 AND 和 OR 的方法。这是我的课: class Event extends RealmObject { String id; String
Ruby - 哈希 - 组合
例如，我有一个包含 5 个元素的哈希: my_hash = {a: 'qwe', b: 'zcx', c: 'dss', d: 'ccc', e: 'www' } 我的目标是每次循环哈希时都返回，但没
ios - 组合:以一定的延迟发布序列的元素
我是Combine 的新手，我想得到一个看似简单的东西。假设我有一个整数集合，例如: let myCollection = [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] 我想以例如 0
java - 组合、转发和包装
关于“优先组合而不是继承”的问题，我的老师是这样说的: 组合:现有类成为新类的组件转发:新类中的每个实例方法，在现有类的包含实例上调用相应的方法并返回结果包装器:新类封装了现有的这三个概念我不是
java - 组合 if-then 语句
我正在尝试将单个整数从 ASCII 值转换为 0 和 1。相关代码如下所示: int num1 = bin.charAt(0); int num2 = bin.charAt(1);
java - 组合:如何使用点表示法访问非静态变量而不出现空指针异常？
这个问题已经有答案了: What is a NullPointerException, and how do I fix it? (12 个回答) 已关闭 7 年前。我经常看到“嵌套”类中的非静态变
python - 组合/合并具有重复名称的两个数据集
我尝试合并两个数据集(DataFrame)，如下所示: D1 = pd.DataFrame({'Village':['Ampil','Ampil','Ampil','Bachey','Bachey',

首页

博学

6Ren·AI

商城

python - 将两个 python 数据处理脚本组合成一个工作流