gpt4 book ai didi

python-3.x - Pandas 彼此更改日期

转载 作者:行者123 更新时间:2023-12-03 16:40:22 25 4
gpt4 key购买 nike

我有一个带有日期和用户的 pandas 数据框,看起来像这样-

Initial Dataframe

date = ['1/2/2020','1/9/2020','1/10/2020','1/17/2020','1/18/2020','1/24/2020','1/25/2020','5/17/2019','5/18/2019','5/24/2019','5/29/2019']
user =['A','B','C','B','A','A','B','C','A','A','B']
df = pd.DataFrame(data={"Date":date, "User":user})

我正在尝试查找彼此相邻的所有日期(1 月 1 日和 1 月 2 日)并将它们转换为单个日期,这样两个条目就会成为两者中较低的日期。条目数量超过一百万。此数据是根据夜间触发的扫描结果创建的,有时会流入前几天。更新-我想合并扫描日期,以便我可以正确显示可视化。就目前而言,结果在扫描开始的那一天会有更多的条目,但在扫描溢出的那一天只有很少的条目。存储了主要日期和时间,因此我不会丢失数据。用户列在扫描包含所有用户名的文件时显示,日期存储扫描日期。

到目前为止,我能够读取数据框,然后根据日期对其进行排序,使条目一个接一个地出现。

输出应该如下所示-

Output

有没有一种 pytonic 方法可以做到这一点?

最佳答案

要考虑的一个问题是连续多天的情况以及您希望如何处理这些情况。以下代码将日期设置为每个 block 中连续日期的第一天:

import pandas as pd
from datetime import timedelta

# prepend two dates to show multiple consecutive days "use-case"
date = ['12/31/2019','1/1/2020','1/2/2020','1/9/2020','1/10/2020','1/17/2020','1/18/2020','1/24/2020','1/25/2020','5/17/2019','5/18/2019','5/24/2019','5/29/2019']
user = ['Z','Z','A','B','C','B','A','A','B','C','A','A','B']
df = pd.DataFrame(data={"Date":date, "User":user})

# first convert to datetime to allow date operations
df.Date = pd.to_datetime(df.Date)

# check if the the date is one day after the row before (by shifting the Date column)
df['isConsecutive'] = (df.Date == df.Date.shift()+pd.DateOffset(1))

# get number of consecutive days in each block
df['numConsecutive'] = df.isConsecutive.groupby((~df.isConsecutive).cumsum()).cumsum()

# convert to timedelta
df.numConsecutive = df.numConsecutive.apply(lambda x: timedelta(days=x))

# take this as differnce to Date
df['NewDate'] = df.Date - df.numConsecutive

print(df)

返回:

         Date User  isConsecutive numConsecutive    NewDate
0 2019-12-31 Z False 0 days 2019-12-31
1 2020-01-01 Z True 1 days 2019-12-31
2 2020-01-02 A True 2 days 2019-12-31
3 2020-01-09 B False 0 days 2020-01-09
4 2020-01-10 C True 1 days 2020-01-09
5 2020-01-17 B False 0 days 2020-01-17
6 2020-01-18 A True 1 days 2020-01-17
7 2020-01-24 A False 0 days 2020-01-24
8 2020-01-25 B True 1 days 2020-01-24
9 2019-05-17 C False 0 days 2019-05-17
10 2019-05-18 A True 1 days 2019-05-17
11 2019-05-24 A False 0 days 2019-05-24
12 2019-05-29 B False 0 days 2019-05-29

关于python-3.x - Pandas 彼此更改日期,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/60155468/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com