gpt4 book ai didi

python - 在 Python 中找到最早的出现

转载 作者:太空宇宙 更新时间:2023-11-03 14:09:01 24 4
gpt4 key购买 nike

我在这方面遇到了麻烦:我需要找到用户第一次点击电子邮件(变量发送)的时间,并在它发生时在相应的行中放置一个。

该数据集有数千名用户(散列),他们点击了时事通讯中电子邮件的一部分。我试图通过发送、散列对它们进行分组,然后找到最早的日期,但无法正常工作。

所以我寻求了一个有点讨厌的解决方案,然而它返回了奇怪的东西:

我的数据集(相关变量):

>>> clicks[['datetime','hash','sending']].head()

datetime hash sending
0 2016-11-01 19:13:34 0b1f4745df5925dfb1c8f53a56c43995 5
1 2016-11-01 10:47:14 0a73d5953ebf5826fbb7f3935bad026d 5
2 2016-10-31 19:09:21 605cebbabe0ba1b4248b3c54c280b477 5
3 2016-10-31 13:42:36 d26d61fb10c834292803b247a05b6cb7 5
4 2016-10-31 10:46:30 48f8ab83e8790d80af628e391f3325ad 5

共有6轮发送,datetimedatetime64[ns]

我的做法是这样的:

clicks['first'] = 0

for hash in clicks['hash'].unique():
t = clicks.ix[clicks.hash==hash, ['hash','datetime','sending']]
part = t['sending'].unique()

for i in part:
temp = t.ix[t.sending == i,'datetime']
clicks.ix[t[t.datetime == np.min(temp)].index.values,'first']=1

首先,我不认为它非常 pythonic,而且速度很慢。但主要是它返回一个奇怪的类型!有 0.01.0 值,但我无法使用它们:

    >>> type(clicks.first)
<type 'instancemethod'>

>>> clicks.loc[clicks.first==1]
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/Users/air/anaconda/lib/python2.7/site-packages/pandas/core/indexing.py", line 1296, in __getitem__
return self._getitem_axis(key, axis=0)
File "/Users/air/anaconda/lib/python2.7/site-packages/pandas/core/indexing.py", line 1467, in _getitem_axis
return self._get_label(key, axis=axis)
File "/Users/air/anaconda/lib/python2.7/site-packages/pandas/core/indexing.py", line 93, in _get_label
return self.obj._xs(label, axis=axis)
File "/Users/air/anaconda/lib/python2.7/site-packages/pandas/core/generic.py", line 1749, in xs
loc = self.index.get_loc(key)
File "/Users/air/anaconda/lib/python2.7/site-packages/pandas/indexes/base.py", line 1947, in get_loc
return self._engine.get_loc(self._maybe_cast_indexer(key))
File "pandas/index.pyx", line 137, in pandas.index.IndexEngine.get_loc (pandas/index.c:4154)
File "pandas/index.pyx", line 156, in pandas.index.IndexEngine.get_loc (pandas/index.c:3977)
File "pandas/index.pyx", line 373, in pandas.index.Int64Engine._check_type (pandas/index.c:7634)
KeyError: False

----- 更新:------

  INSTALLED VERSIONS
------------------
commit: None
python: 2.7.12.final.0
python-bits: 64
OS: Darwin
OS-release: 15.6.0
machine: x86_64
processor: i386
byteorder: little
LC_ALL: None
LANG: en_US.UTF-8

pandas: 0.18.1

最佳答案

我想你需要groupbyapply其中将值与 minimal 进行比较,输出为 bool 值 - 需要通过 astype 转换为 int 01 :

clicks = pd.DataFrame({'hash': {0: '0b1f4745df5925dfb1c8f53a56c43995', 1: '0a73d5953ebf5826fbb7f3935bad026d', 2: '605cebbabe0ba1b4248b3c54c280b477', 3: '0b1f4745df5925dfb1c8f53a56c43995', 4: '0a73d5953ebf5826fbb7f3935bad026d', 5: '605cebbabe0ba1b4248b3c54c280b477', 6: 'd26d61fb10c834292803b247a05b6cb7', 7: '48f8ab83e8790d80af628e391f3325ad'}, 'sending': {0: 5, 1: 5, 2: 5, 3: 5, 4: 5, 5: 5, 6: 5, 7: 5}, 'datetime': {0: pd.Timestamp('2016-11-01 19:13:34'), 1: pd.Timestamp('2016-11-01 10:47:14'), 2: pd.Timestamp('2016-10-31 19:09:21'), 3: pd.Timestamp('2016-11-01 19:13:34'), 4: pd.Timestamp('2016-11-01 11:47:14'), 5: pd.Timestamp('2016-10-31 19:09:20'), 6: pd.Timestamp('2016-10-31 13:42:36'), 7: pd.Timestamp('2016-10-31 10:46:30')}})
print (clicks)
datetime hash sending
0 2016-11-01 19:13:34 0b1f4745df5925dfb1c8f53a56c43995 5
1 2016-11-01 10:47:14 0a73d5953ebf5826fbb7f3935bad026d 5
2 2016-10-31 19:09:21 605cebbabe0ba1b4248b3c54c280b477 5
3 2016-11-01 19:13:34 0b1f4745df5925dfb1c8f53a56c43995 5
4 2016-11-01 11:47:14 0a73d5953ebf5826fbb7f3935bad026d 5
5 2016-10-31 19:09:20 605cebbabe0ba1b4248b3c54c280b477 5
6 2016-10-31 13:42:36 d26d61fb10c834292803b247a05b6cb7 5
7 2016-10-31 10:46:30 48f8ab83e8790d80af628e391f3325ad 5
#if column dtype of column datetime is not datetime (with this sample not necessary)
clicks.datetime = pd.to_datetime(clicks.datetime)
clicks['first'] = clicks.groupby(['hash','sending'])['datetime'] \
.apply(lambda x: x == x.min()) \
.astype(int)
print (clicks)
datetime hash sending first
0 2016-11-01 19:13:34 0b1f4745df5925dfb1c8f53a56c43995 5 1
1 2016-11-01 10:47:14 0a73d5953ebf5826fbb7f3935bad026d 5 1
2 2016-10-31 19:09:21 605cebbabe0ba1b4248b3c54c280b477 5 0
3 2016-11-01 19:13:34 0b1f4745df5925dfb1c8f53a56c43995 5 1
4 2016-11-01 11:47:14 0a73d5953ebf5826fbb7f3935bad026d 5 0
5 2016-10-31 19:09:20 605cebbabe0ba1b4248b3c54c280b477 5 1
6 2016-10-31 13:42:36 d26d61fb10c834292803b247a05b6cb7 5 1
7 2016-10-31 10:46:30 48f8ab83e8790d80af628e391f3325ad 5 1

----- 更新:------

INSTALLED VERSIONS
------------------
commit: None
python: 2.7.12.final.0
python-bits: 64
OS: Darwin
OS-release: 15.6.0
machine: x86_64
processor: i386
byteorder: little
LC_ALL: None
LANG: en_US.UTF-8

pandas: 0.18.1

关于python - 在 Python 中找到最早的出现,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/40533617/

24 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com