r - 用于特征选择的 t-stat-6ren

r - 用于特征选择的 t-stat

转载作者：行者123 更新时间：2023-11-30 09:20:04

25

4

我想用 for 循环计算 R 中特征选择的 t-Statistic。数据有 155 列，因变量是二元的(诱变剂 - 非诱变剂)。我想为每一列分配一个 t-stat。问题是我不知道如何写它。

这是我尝试在 R 中实现的公式:

我还写了一个代码，但我不确定它，它只是用于第一列。我需要将其写入所有列的 for 循环中。

abs(diff(tapply(train_df[,1], train_df$Activity, mean))) / sqrt(sd((train_df$NEG_01_NEG[train_df$Activity == "mutagen"])^2) / (length(train_df$NEG_01_NEG[train_df$Activity == "mutagen"])) + 
   sd((train_df$NEG_01_NEG[train_df$Activity != "mutagen"])^2) / (length(train_df$NEG_01_NEG[train_df$Activity != "mutagen"])))

提前致谢!

最佳答案

如果您不想担心速度(并且您可能不关心 155 列)，您可以使用 t.test 函数并将其应用于每一列。

先模拟一些数据

set.seed(1)
DF <- data.frame(y=rep(1:2, 50), x1=rnorm(100), x2=rnorm(100), x3=rnorm(100))
head(DF)

  y         x1          x2         x3
1 1 -0.6264538 -0.62036668  0.4094018
2 2  0.1836433  0.04211587  1.6888733
3 1 -0.8356286 -0.91092165  1.5865884
4 2  1.5952808  0.15802877 -0.3309078
5 1  0.3295078 -0.65458464 -2.2852355
6 2 -0.8204684  1.76728727  2.4976616

然后我们可以使用公式参数将 t.test 函数应用于除第一列之外的所有列。

group <- DF$y
lapply(DF[,-1], function(x) { t.test(x ~ group)$statistic })

返回每列的检验统计量。

t.test 计算大量您不需要的额外信息，因此您可以通过直接进行计算来大大加快速度，但这里确实没有必要

关于r - 用于特征选择的 t-stat，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/43017689/

25

4

0

文章推荐： java - JProfiler 报告 Long 在 Object.wait() 中的分配

文章推荐： javascript - 将复选框更改为开关样式按钮

文章推荐： javascript - 如何在双击时附加和删除div

文章推荐： java - 在小程序中将图像转换为缓冲图像

c - struct stat 和 stat 函数失败
所以，我正在尝试创建一种 ls 函数。这是我对每个文件的描述的代码 struct stat fileStat; struct dirent **files; num_entries = scandir
C sys/stat.h 并非 stat 结构的每个字段都被初始化
我最近一直在尝试实现我自己的 linux ls 命令版本。一切都很好，但是当我尝试使用 ls -l 功能时，struct stat 的某些字段未初始化 - 我得到 NULL 指针或垃圾值，尽管它似乎只
php - 相关模型的 STAT 关系的 STAT 关系，Yii
我在 Yii 中遇到 STAT 关系问题。我不确定我正在寻找的东西是否可以通过本地 Yii 关系实现。我会尽力描述我的问题，如果不清楚，请询问任何具体细节。我有三个表，因此有三个模型 | table
python - 部署后在 Django 中使用 scipy.stats.stats
我正在为一个严重依赖 scipy.stats.stats(scipy 版本 0.9.0)的包创建一个 django-powered (1.3) 接口(interface)，称为 ovl 。在早期开发阶
c++ - 我想重置 C++ struct stat，我可以以某种方式使用 stat() 语法吗？
为了安全起见，我喜欢显式初始化我的变量(当您编写大量代码时，它通常会使它更安全，因为您的代码最终不会崩溃那么多。) 对于大多数类型，无论是结构还是整数等基本 C++ 类型，我都可以编写以下内容: ti
c - 是否有 stat() 的宽字符版本(来自 sys/stat.h)？
我一直在使用 stat() 检查文件是否存在，据我所知，这比尝试打开文件更好。但是，stat() 不适用于包含其他语言的 unicode 字符的文件名。是否有 stat() 的宽字符版本或我可以使用的
python - 无法从 scipy.stats.stats 导入 ss 函数
错误: File "/usr/lib/python2.7/dist-packages/statsmodels/regression/linear_model.py", line 36, in
linux - 在一行中使用 awk 和 stat 来打印 stat 的值
下面是我要运行的脚本。我不能在 awk 中使用 stat。 cat /etc/passwd | awk 'BEGIN{FS=":"}{print $6 }' | (stat $6 | sed -n '
python - Seaborn regplot 拟合线与来自 stats.linregress 或 stats 模型的计算拟合不匹配
我正在尝试拟合 xlog 线性回归。我使用 Seaborn regplot 来绘制拟合，看起来很合适(绿线)。然后，因为 regplot 不提供系数。我使用 stats.linregress 来查找系
c - libc_nonshared.a(stat.oS) 中的隐藏符号 `stat' 被 DSO 引用
我正在尝试使用共享库 (libscplugin.so) 中包含的方法。我已经满足了库的所有要求: libc.so 带有指向 libc.so.6 的符号链接(symbolic link) libz.s
c - C 编程中的错误 : stat: No Such File Or Directory, opendir() 和 stat()
嘿，感谢阅读。我正在制作一个程序，它接受 1 个参数(目录)并使用 opendir()/readdir() 读取目录中的所有文件，并使用 stat 显示文件类型(reg、链接、目录等)。当我在 sh
c - 非设备文件上的 major(stat.st_rdev) 和 minor(stat.st_rdev)
简单问题:在 Linux 中，我 stat() 一个不是设备的文件。 st_rdev 字段的期望值是多少？我可以运行 major(stat.st_rdev) 和 minor(stat.st_rdev)
使用 --stats-json 构建 Angular 6 不生成 stats.json 文件
我正在尝试为我的 Angular 6 应用程序生成 stats.json 文件。下面的事情我已经尝试过，但根本没有生成文件。我的系统需要有 “npm 运行”在每个 angular cli 命令之前。
c - [C][Stat][Fileinfo] 当我使用 stat() 调用返回结构时，为什么 st_mode 定义为不在结构中的内容？
我正在尝试使用返回的 stat 结构中的 st_mode，该结构是我通过以下方式从 stat() 调用获得的； char *fn = "test.c" struct s
c++ - stat(char[], stat) 当 string.c_str() = "c:/"时返回 -1
关闭。这个问题需要debugging details .它目前不接受答案。编辑问题以包含 desired behavior, a specific problem or error, and th
c - 系统/stat.h :456: error: nested function 'stat' declared 'extern'
我有一个程序，是我通过修改原始暗网(深度学习图像识别，Yolov2)的许多地方而制作的。几个月前我一直在使用它，但是今天当我编译它时，它给了我一个错误: gcc -DSAVE_LAYER_INPUT
python - 为什么 scipy.stats.mstats.pearsonr 结果与 scipy.stats.pearsonr 不一致？
我预计 scipy.stats.mstats.pearsonr 对于屏蔽数组输入的结果将与 scipy.stats.pearsonr 对于输入数据的 unmasked 值给出相同的结果，但它不会't:
从调用 stat 失败并出现 "Value too large for defined data type"错误
给定 tmp.c: #include #include #include int main(int argc, const char *argv[]) { struct stat st;
python - 为什么 scipy.stats.entropy(a, b) 返回 inf 而 scipy.stats.entropy(b, a) 不返回？
In [15]: a = np.array([0.5, 0.5, 0, 0, 0]) In [16]: b = np.array([1, 0, 0, 0, 0]) In [17]: entropy(a
linux - Stat 命令将具有更改日期的文件列入候选名单
当我们运行 stat filename我们得到 Access: 2021-06-25 15:40:18.532621916 +0530 Modify: 2020-08-13 15:57:30.0000

首页

博学

6Ren·AI

商城

r - 用于特征选择的 t-stat