python - 为什么 pandas "rank"百分位数不介于 0 和 1 之间？-6ren

python - 为什么 pandas "rank"百分位数不介于 0 和 1 之间？

转载作者：太空宇宙更新时间：2023-11-03 16:48:45

25

4

我经常使用 pandas，并且经常执行与以下类似的代码:

df['var_rank'] = df['var'].rank(pct=True)
print( df.var_rank.max() )

并且经常会得到大于 1 的值。无论我保留还是删除“na”值，这种情况仍然会发生。这显然很容易修复(只需除以最大排名的值)，所以我不要求解决方法。我只是好奇为什么会发生这种情况，并且在网上没有找到任何线索。

有人知道为什么会发生这种情况吗？

一些非常简单的示例数据here (dropbox 链接 - 腌制的 Pandas 系列)。

我从 df.rank(pct=True).max() 得到的值为 1.0156 。我的其他数据的值高达 4 或 5。我通常使用非常困惑的数据。

最佳答案

您的数据不正确。

>>> s.rank(pct=True).max()
1.015625

s.sort(inplace=True)
>>> s.tail(7)
8      202512882
6      253661077
102            -
101            -
99             -
58             -
116            -
Name: Total Assets, dtype: object

>>> s[s != u'-'].rank(pct=True).max()
1.0

在 Pandas 0.18.0(上周发布)中，您可以指定 numeric only :

s.rank(pct=True, numeric_only=True)

我已经在 0.18.0 中尝试过上述方法，但似乎无法使其工作，因此您也可以执行此操作来对所有 float 和 int 值进行排名:

>>> s[s.apply(lambda x: isinstance(x, (int, float)))].rank(pct=True).max()
1.0

它创建一个 bool 掩码，确保每个值都是 int 或 float，然后对过滤结果进行排名。

关于python - 为什么 pandas "rank"百分位数不介于 0 和 1 之间？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/36070946/

25

4

0

文章推荐： ruby-on-rails - rake 数据库 :create gives error on server

文章推荐： ubuntu - LINKERD:Ubuntu 上 Kubernetes 中的待定外部 IP

文章推荐： Ruby Tempfile 不在磁盘上创建文件

ranking - 投票算法: how to calculate rank?
我正在尝试找出一种计算排名的方法。现在它只需要每个条目的赢/输的比率，所以例如100 次中，有 99 次获胜，则胜率达到 99%。但如果一个参赛作品在 1 票中赢得 1 票，那么它的获胜排名将是 10
Mysql RANK 没有给出每个类别的 RANK
我尝试了以下操作，但它没有对每个类别进行明智的排名。相反，在不考虑类别的情况下对所有记录进行排名。我希望每个类别重新出现排名 select rs.Section,rs.Field1,rs.Field
sql - RANK() OVER PARTITION 并重置 RANK
如何获得在分区更改时重新启动的 RANK？我有这张表: ID Date Value 1 2015-01-01 1 2 2015-01-02 1 1; 关于
sql - 何时选择 rank() 而不是密集的 rank() 或 row_number()
由于我们可以使用 row_number() 获得分配的行号如果我们想使用 dense_rank() 在不跳过分区内的任何数字的情况下找到每一行的排名，我们为什么需要rank()功能，我想不出任何用例
python - 张量形状错误 : Must be rank 2 but is rank 3
我很难搜索可以帮助我构建文本序列(特征)分类器的文档、研究或博客。我拥有的文本序列包含网络日志。我正在使用 TensorFlow 构建 GRU 模型，并将 SVM 作为分类函数。我在处理张量形状时遇
fortran - 错误 : Rank mismatch in argument (rank-1 and scalar)
我遇到了这类错误。 colsys.f:1367.51: 1 NOLD, ALDIF, K, NCOMP, M, MSTAR, 3,DUMM,0)
tensorflow 值错误 : Shape must be rank 1 but is rank 2
import tensorflow as tf x = [[1,2,3],[4,5,6]] y = [0,1] z = [1,2] x = tf.constant(x) y = tf.constant
python - 获取与 SQL rank 不同的 pandas dataframe rank answer
我在学习 SQL 中的排名函数，发现它使用的排名与 pandas 方法不同。如何得到相同的答案？提问链接:https://www.windowfunctions.com/questions/rank
sql - 在 SQL 中使用 RANK() OVER 将 rank 设置为 NULL
在 SQL Server 数据库中，我有一个我对排名感兴趣的值表。当我执行 RANK() OVER (ORDER BY VALUE DESC) 作为 RANK 时，我得到以下结果(在假设表中): R
sql-server - SQL Server : Rank by sum of points and order by ranking
我有一个包含以下字段的游戏 table : ID Name Email Points ---------------------------------- 1 Jo
python - 值错误 : Shape must be rank 2 but is rank 3 for 'MatMul'
我有以下 TensorFlow 代码: layer_1 = tf.add(tf.matmul(tf.cast(x, tf.float32), weights['h1']), biases['b1'])
Python TensorFlow 值错误 : Shape must be rank 1 but is rank 0
我是 Sentdex 教程的神经网络新手。我尝试运行该代码: import tensorflow as tf from tensorflow.examples.tutorials.mnist i
python - tensorflow : ValueError: Shape must be rank 2 but is rank 3
我是 tensorflow 的新手，我正在尝试将双向 LSTM 的一些代码从旧版本的 tensorflow 更新到最新版本 (1.0)，但我收到此错误: Shape must be rank 2 bu
SAS 9.3 Proc Rank 问题(Rank/Sort Road Block)
我正在使用以下格式的数据集: Column 1 (What I Have), Column 2 (What I need to see) 8 1 8 1 8 1 9 2 9
python - 凯拉斯错误 : "BatchNormalization Shape must be rank 1 but is rank 4 for batch_normalization"
我有一个 Keras 函数模型(具有卷积层的神经网络)，它可以很好地与 tensorflow 配合使用。我可以运行它，我可以适应它。但是，使用tensorflow gpu时无法建立模型。这是构建模
c - mpi 中的进程以什么顺序执行...我的意思是排名顺序？例如 : rank==0 first and rank==1 next?
MPI 中的进程以什么顺序执行？我的意思是排名明智的顺序？例如:rank == 0 首先，rank == 1 接下来？我通过在运行时给出以下命令来考虑两个过程: mpirun -np 2 示例。
python - ArithmeticError 导致 cvxpy 出现 "Rank(A) < p or Rank([G; A]) < n"错误
我正在尝试使用 cvxpy(因此使用 cvxopt)在具有 28 个节点和 37 条线路的相对简单的网络中对最佳功率流进行建模，但得到的是“Rank(A) < p or Rank([G; A] ) <
python - tensorflow 错误 : Shape must be rank 0 but is rank 1 for 'cond_1/Switch'
我是 tensorflow 的新手，我正在做一些在线练习以熟悉 tensorflow。我要执行以下任务: Create two tensors x and y of shape 300 from an
python - Tensorflow - 值错误 : Shape must be rank 1 but is rank 0 for 'ParseExample/ParseExample'
我有一个 Ubuntu 对话语料库的 .tfrecords 文件。我正在尝试读取整个数据集，以便我可以将上下文和话语分成几批。使用 tf.parse_single_example 我能够阅读一个示例。
tensorflow - 值错误: Shape must be rank 0 but is rank 1 for 'cond_11/Switch' (op: 'Switch' )
实际上我们不能在 if 语句中使用 tf.var 作为 bool 来代替使用 tf.cond。我为规范化输入数据编写了这段代码，但出现了令人困惑的错误，我哪里做错了？ def global_co

首页

博学

6Ren·AI

商城

python - 为什么 pandas "rank"百分位数不介于 0 和 1 之间？