python - 在 Python 中运行 Theil-Sen 回归时出错-6ren

python - 在 Python 中运行 Theil-Sen 回归时出错

转载作者：行者123 更新时间：2023-12-01 07:45:10

我有一个类似于以下内容的数据框，我们将其称为“df”:

id    value    time
a      1        1
a      1.5      2
a      2        3
a      2.5      4
b      1        1
b      1.5      2
b      2        3
b      2.5      4

我正在这个数据帧上通过Python中的“id”运行各种回归。一般来说，这需要按“id”进行分组，然后对这些分组应用一个函数来计算回归。

我正在 Scipy 的统计库中使用 2 种类似的回归技术:

Theil-Sen 估计器:
( https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/scipy.stats.mstats.theilslopes.html )
西格尔估计器:
(https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.siegelslopes.html)。

这两者都接收相同类型的数据。因此，除了实际使用的技术之外，计算它们的函数应该是相同的。

对于 Theil-Sen，我编写了以下函数以及将应用于该函数的 groupby 语句:

def theil_reg(df, xcol, ycol):
   model = stats.theilslopes(ycol,xcol)
   return pd.Series(model)

out = df.groupby('id').apply(theil_reg, xcol='time', ycol='value')

但是，我收到以下错误，我一直很难理解如何解决该错误:

ValueError: could not convert string to float: 'time'

实际变量time是一个numpy浮点对象，所以它不是一个字符串。这让我相信 stats.theilslopes 函数无法识别 time 是数据帧中的一列，而是使用“time”作为函数的字符串输入。

但是，如果是这种情况，那么这似乎是 stats.theilslopes 包中的一个错误，需要由 Scipy 解决。我相信这种情况的原因是因为与上面完全相同的函数，但使用 siegelslopes 包，工作得很好，并提供了我期望的输出，它们本质上是使用相同的输入进行相同的估计。

在 Siegel 上执行以下操作:

def siegel_reg(df, xcol, ycol):
   model = stats.siegelslopes(ycol,xcol)
   return pd.Series(model)

out = df.groupby('id').apply(siegel_reg, xcol='time',ycol='value')

不会产生任何关于时间变量的错误，并根据需要进行回归。

有人知道我是否在这里遗漏了一些东西吗？如果是这样，我将不胜感激任何想法，或者如果不是，任何关于如何使用 Scipy 解决这个问题的想法。

编辑:这是我运行此脚本时显示的完整错误消息:

ValueError Traceback (most recent call last)
C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in apply(self, func, *args, **kwargs)
    688 try:
--> 689 result = self._python_apply_general(f)
    690 except Exception:

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in _python_apply_general(self, f)
    706 keys, values, mutated = self.grouper.apply(f, self._selected_obj,
--> 707                                                    self.axis)
    708 

C:\Anaconda\lib\site-packages\pandas\core\groupby\ops.py in apply(self, f, data, axis)
    189             group_axes = _get_axes(group)
--> 190             res = f(group)
    191             if not _is_indexed_like(res, group_axes):

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in f(g)
    678                     with np.errstate(all='ignore'):
--> 679                         return func(g, *args, **kwargs)
    680             else:

<ipython-input-506-0a1696f0aecd> in theil_reg(df, xcol, ycol)
      1 def theil_reg(df, xcol, ycol):
----> 2     model = stats.theilslopes(ycol,xcol)
      3     return pd.Series(model)

C:\Anaconda\lib\site-packages\scipy\stats\_stats_mstats_common.py in 
theilslopes(y, x, alpha)
    221     else:
--> 222         x = np.array(x, dtype=float).flatten()
    223         if len(x) != len(y):

ValueError: could not convert string to float: 'time'

During handling of the above exception, another exception occurred:

ValueError Traceback (most recent call last)
<ipython-input-507-9a199e0ce924> in <module>
----> 1 df_accel_correct.groupby('chart').apply(theil_reg, xcol='time', 
ycol='value')

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in apply(self, func, *args, **kwargs)
    699 
    700                 with _group_selection_context(self):
--> 701                     return self._python_apply_general(f)
    702 
    703         return result

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in _python_apply_general(self, f)
    705     def _python_apply_general(self, f):
    706         keys, values, mutated = self.grouper.apply(f, 
self._selected_obj,
--> 707                                                    self.axis)
    708 
    709         return self._wrap_applied_output(

C:\Anaconda\lib\site-packages\pandas\core\groupby\ops.py in apply(self, f, data, axis)
    188             # group might be modified
    189             group_axes = _get_axes(group)
--> 190             res = f(group)
    191             if not _is_indexed_like(res, group_axes):
    192                 mutated = True

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in f(g)
    677                 def f(g):
    678                     with np.errstate(all='ignore'):
--> 679                         return func(g, *args, **kwargs)
    680             else:
    681                 raise ValueError('func must be a callable if args or '

<ipython-input-506-0a1696f0aecd> in theil_reg(df, xcol, ycol)
      1 def theil_reg(df, xcol, ycol):
----> 2     model = stats.theilslopes(ycol,xcol)
      3     return pd.Series(model)

C:\Anaconda\lib\site-packages\scipy\stats\_stats_mstats_common.py in theilslopes(y, x, alpha)
    220         x = np.arange(len(y), dtype=float)
    221     else:
--> 222         x = np.array(x, dtype=float).flatten()
    223         if len(x) != len(y):
    224             raise ValueError("Incompatible lengths ! (%s<>%s)" % (len(y), len(x)))

ValueError: could not convert string to float: 'time'

更新2:在函数中调用df后，我收到以下错误消息:

ValueError                                Traceback (most recent call last)
C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in apply(self, func, *args, **kwargs)
    688             try:
--> 689                 result = self._python_apply_general(f)
    690             except Exception:

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in _python_apply_general(self, f)
    706         keys, values, mutated = self.grouper.apply(f, self._selected_obj,
--> 707                                                    self.axis)
    708 

C:\Anaconda\lib\site-packages\pandas\core\groupby\ops.py in apply(self, f, data, axis)
    189             group_axes = _get_axes(group)
--> 190             res = f(group)
    191             if not _is_indexed_like(res, group_axes):

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in f(g)
    678                     with np.errstate(all='ignore'):
--> 679                         return func(g, *args, **kwargs)
    680             else:

<ipython-input-563-5db69048f347> in theil_reg(df, xcol, ycol)
      1 def theil_reg(df, xcol, ycol):
----> 2     model = stats.theilslopes(df[ycol],df[xcol])
      3     return pd.Series(model)

C:\Anaconda\lib\site-packages\scipy\stats\_stats_mstats_common.py in theilslopes(y, x, alpha)
    248     sigma = np.sqrt(sigsq)
--> 249     Ru = min(int(np.round((nt - z*sigma)/2.)), len(slopes)-1)
    250     Rl = max(int(np.round((nt + z*sigma)/2.)) - 1, 0)

ValueError: cannot convert float NaN to integer

During handling of the above exception, another exception occurred:

ValueError                                Traceback (most recent call last)
<ipython-input-564-d7794bd1d495> in <module>
----> 1 correct_theil = df_accel_correct.groupby('chart').apply(theil_reg, xcol='time', ycol='value')

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in apply(self, func, *args, **kwargs)
    699 
    700                 with _group_selection_context(self):
--> 701                     return self._python_apply_general(f)
    702 
    703         return result

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in _python_apply_general(self, f)
    705     def _python_apply_general(self, f):
    706         keys, values, mutated = self.grouper.apply(f, self._selected_obj,
--> 707                                                    self.axis)
    708 
    709         return self._wrap_applied_output(

C:\Anaconda\lib\site-packages\pandas\core\groupby\ops.py in apply(self, f, data, axis)
    188             # group might be modified
    189             group_axes = _get_axes(group)
--> 190             res = f(group)
    191             if not _is_indexed_like(res, group_axes):
    192                 mutated = True

C:\Anaconda\lib\site-packages\pandas\core\groupby\groupby.py in f(g)
    677                 def f(g):
    678                     with np.errstate(all='ignore'):
--> 679                         return func(g, *args, **kwargs)
    680             else:
    681                 raise ValueError('func must be a callable if args or '

<ipython-input-563-5db69048f347> in theil_reg(df, xcol, ycol)
       1 def theil_reg(df, xcol, ycol):
 ----> 2     model = stats.theilslopes(df[ycol],df[xcol])
       3     return pd.Series(model)

C:\Anaconda\lib\site-packages\scipy\stats\_stats_mstats_common.py in theilslopes(y, x, alpha)
    247     # Find the confidence interval indices in `slopes`
    248     sigma = np.sqrt(sigsq)
--> 249     Ru = min(int(np.round((nt - z*sigma)/2.)), len(slopes)-1)
    250     Rl = max(int(np.round((nt + z*sigma)/2.)) - 1, 0)
    251     delta = slopes[[Rl, Ru]]

ValueError: cannot convert float NaN to integer

但是，两列中都没有空值，并且两列都是 float 。对这个错误有什么建议吗？

最佳答案

本质上，您将列名称的字符串值(不是任何值实体)传递到方法中，但 slopes 调用需要 numpy 数组(或可以强制转换为数组的 pandas 系列)。具体来说，您尝试此调用时未引用 df，因此出现错误:

model = stats.theilslopes('value', 'time')

只需在调用中引用df:

model = stats.theilslopes(df['value'], df['time'])

model = stats.theilslopes(df[ycol], df[xcol])

<小时/>

关于跨包的不同结果并不意味着 Scipy 存在错误。包运行不同的实现。仔细阅读文档以了解如何调用方法。可能，您引用的另一个包允许数据输入作为调用内的参数，并且命名字符串引用如下所示的列:

slopes_call(y='y_string', x='x_string', data=df)

一般来说，Python 对象模型始终需要对调用和对象的显式命名引用，并且不假定上下文。

关于python - 在 Python 中运行 Theil-Sen 回归时出错，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/56500036/

文章推荐： python - 不断收到 Kattis 问题的运行时错误

文章推荐： swift3 - 像 SnapChat 一样的 UICollectionView 布局？

文章推荐： java - 将 6 个月前的日期转换为字符串

javascript - npm test 出错，但 mocha test node.js 出错
我正在使用 node.js 和 mocha 单元测试，并且希望能够通过 npm 运行测试命令。当我在测试文件夹中运行 Mocha 测试时，测试运行成功。但是，当我运行 npm test 时，测试给出了
java - java方法replaceAll()出错
我的文本区域中有这些标签 ..... 我正在尝试使用 replaceAll() String 方法替换它们 text.replaceAll("", ""); text.replaceAll("", "
java - ZXing 出错
早上好，我是 ZXing 的新手，当我运行我的应用程序时出现以下错误: 异常Ljava/lang/NoClassDefFoundError；初始化 ICOM/google/zxing/client/a
C - free() 出错？
我正在制作一些哈希函数。它的源代码是... #include #include #include int m_hash(char *input, size_t in_length, char
swift - SKPhysicsContactDelegate 出错？
我正在尝试使用 Spritekit 在 Swift 中编写游戏。目的是带着他的角色迎面而来的矩形逃跑。现在我在 SKPhysicsContactDelegate (didBegin ()) 方法中犯了
java - actionPerformed 出错？
我正在尝试创建一个用于导入 CSV 文件的按钮，但出现此错误: actionPerformed(java.awt.event.ActionEvent) in cannot implement
java - NULLIF() 出错？
请看下面的代码 public List getNames() { List names = new ArrayList(); try { createConnection(); Sta
创建计划事件时 MySQL 出错
我正在尝试添加一个事件以在“dealsArchive”表中创建一个条目，然后从“deals”表中删除该条目。它需要在特定时间执行。这是我正在尝试使用的: DELIMITER $$ CREATE EV
尝试将存储过程结果插入表时 PHPmyadmin 出错
我试图将两个存储过程的表结果存储到 phpmyadmin 例程窗口中的单个表中，这给了我 mariadb 语法错误。单独调用存储过程给出了结果。存储过程代码 BEGIN CREATE TABLE t
android - videoview的onpreparedlistener()出错
我想在 videoview 中加载视频之前有一个进度条。但是我收到以下错误。我还添加了所有必要的导入。我在 ANDROID 中使用 AIDE 这是我的代码 public class MainActi
android - AsyncTask 出错
我已经使用了 AsyncTask，但我不明白为什么在我的设备 (OS 4.0) 上测试时仍然出现错误。我的 apk 构建于 2.3.3 中。我想我把代码弄错了，但我不知道我的错误在哪里。任何人都请帮助
MySQL注入(inject)出错
我在测试 friend 网站的安全性时，通过在 URL 末尾添加 ' 发现了 SQL 注入(inject)漏洞该网站是用zend框架构建的我遇到的问题是 MySQL -- 中的注释语法不起作用，因此页
java - getMap() 出错？
我正在尝试使用堆栈溢出答案之一的交互式信息窗口。链接如下: interactive infowindow 但是我在代码中使用 getMap() 时遇到错误。虽然我尝试使用 getMapAsync 但
java - addMouseListener 出错
当我编译以下代码时出现错误: The method addMouseListener(Player) is undefined for the type Player 代码: import java.
mysql - Async_Http_Response_Handler 出错
我是 Android 开发的初学者。我正在开发一个接收 MySql 数据然后将其保存在 SQLite 中的应用程序。我将 Json 用于同步状态，以便我可以将未同步数据的数量显示为要同步的待处理数据
Java - 转换文件名 - 出错？
(这里是Hello world级别的自动化测试人员) 我正在尝试下载一个文件并将其重命名以便于查找。我收到一个错误....这是代码 @Test public void allDownload(
c++ - while(cin) 出错
我只是在写另一个程序。并使用: while (cin) words.push_back(s); words是string的vector，s是string。我的 RAM 使用量在 4 或 5
javascript - AngularJS 出错
我是 AngularJS 的新手，我遇到了一个问题。我有一个带有提交按钮的页面，当我单击提交模式时必须打开并且来自 URL 的数据必须存在于模式中。现在，模式打开但它是空的并且没有从 URL 获取数据
c++ - 尝试将文件写入数组时运算符 << 出错
我正在尝试读取一个文件(它可以包含任意数量的随机数字，但不会超过 500 个)并将其放入一个数组中。稍后我将需要使用数组来做很多事情。但到目前为止，这一小段代码给了我 no match for o
c++ - benderrmq 出错
有些人在使用 make 命令进行编译时遇到了问题，所以我想我应该在这里尝试一下，我已经在以下操作系统的 ubuntu 32 位和挤压 64 位上尝试过我克隆了 git 项目 https://gith

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 在 Python 中运行 Theil-Sen 回归时出错