r - 将 vline 添加到 geom_density 和均值 R 的阴影置信区间-6ren

r - 将 vline 添加到 geom_density 和均值 R 的阴影置信区间

转载作者：行者123 更新时间：2023-12-01 01:46:58

在阅读了不同的帖子后，我发现了如何向密度图添加一条均值的 vline，如图所示 here .
使用上述链接中提供的数据:

1) 如何使用 geom_ribbon 在平均值周围添加 95% 的置信区间？
CI 可以计算为

#computation of the standard error of the mean
sem<-sd(x)/sqrt(length(x))
#95% confidence intervals of the mean
c(mean(x)-2*sem,mean(x)+2*sem)

2)如何将vline限制在曲线下的区域？您将在下图中看到曲线外的 vline 图。

可以在 https://www.dropbox.com/s/bvvfdpgekbjyjh0/test.csv?dl=0 找到与我的实际问题非常接近的示例数据

更新

使用上面链接中的真实数据，我使用@beetroot 的答案尝试了以下操作。

# Find the mean of each group
dat=me
library(dplyr)
library(plyr)
cdat <- ddply(data,.(direction,cond), summarise, rating.mean=mean(rating,na.rm=T))# summarize by season and variable
cdat

#ggplot
p=ggplot(data,aes(x = rating)) + 
  geom_density(aes(colour = cond),size=1.3,adjust=4)+
  facet_grid(.~direction, scales="free")+
  xlab(NULL) + ylab("Density")
p=p+coord_cartesian(xlim = c(0, 130))+scale_color_manual(name="",values=c("blue","#00BA38","#F8766D"))+
  scale_fill_manual(values=c("blue", "#00BA38", "#F8766D"))+
  theme(legend.title = element_text(colour="black", size=15, face="plain"))+
  theme(legend.text = element_text(colour="black", size = 15, face = "plain"))+
  theme(title = red.bold.italic.text, axis.title = red.bold.italic.text)+
  theme(strip.text.x = element_text(size=20, color="black",face="plain"))+ # facet labels
  ggtitle("SAMPLE A") +theme(plot.title = element_text(size = 20, face = "bold"))+
    theme(axis.text = blue.bold.italic.16.text)+ theme(legend.position = "none")+
  geom_vline(data=cdat, aes(xintercept=rating.mean, color=cond),linetype="dotted",size=1)
p

## implementing @beetroot's code to restrict lines under the curve and shade CIs around the mean
# I will use ddply for mean and CIs
cdat <- ddply(data,.(direction,cond), summarise, rating.mean=mean(rating,na.rm=T),
              sem = sd(rating,na.rm=T)/sqrt(length(rating)),
              ci.low = mean(rating,na.rm=T) - 2*sem,
              ci.upp = mean(rating,na.rm=T) + 2*sem)# summarize by direction and variable


#In order to limit the lines to the outline of the curves you first need to find out which y values
#of the curves correspond to the means, e.g. by accessing the density values with ggplot_build and 
#using approx:

   cdat.dens <- ggplot_build(ggplot(data, aes(x=rating, colour=cond)) +
                              facet_grid(.~direction, scales="free")+
                              geom_density(aes(colour = cond),size=1.3,adjust=4))$data[[1]] %>%
  mutate(cond = ifelse(group==1, "A",
                       ifelse(group==2, "B","C"))) %>%
  left_join(cdat) %>%
  select(y, x, cond, rating.mean, sem, ci.low, ci.upp) %>%
  group_by(cond) %>%
  mutate(dens.mean = approx(x, y, xout = rating.mean)[[2]],
         dens.cilow = approx(x, y, xout = ci.low)[[2]],
         dens.ciupp = approx(x, y, xout = ci.upp)[[2]]) %>%
  select(-y, -x) %>%
  slice(1)

 cdat.dens

#---
 #You can then combine everything with various geom_segments:

   ggplot(data, aes(x=rating, colour=cond)) +
   geom_density(data = data, aes(x = rating, colour = cond),size=1.3,adjust=4) +facet_grid(.~direction, scales="free")+
   geom_segment(data = cdat.dens, aes(x = rating.mean, xend = rating.mean, y = 0, yend = dens.mean, colour = cond),
                linetype = "dashed", size = 1) +
   geom_segment(data = cdat.dens, aes(x = ci.low, xend = ci.low, y = 0, yend = dens.cilow, colour = cond),
                linetype = "dotted", size = 1) +
   geom_segment(data = cdat.dens, aes(x = ci.upp, xend = ci.upp, y = 0, yend = dens.ciupp, colour = cond),
                linetype = "dotted", size = 1)

给出了这个:

您会注意到均值和 CI 没有像原始图中那样对齐。我做错了什么@beetroot？

最佳答案

使用来自链接的数据，您可以像这样计算均值、se 和 ci(我建议使用 dplyr ， plyr 的后继者):

set.seed(1234)
dat <- data.frame(cond = factor(rep(c("A","B"), each=200)), 
                  rating = c(rnorm(200),rnorm(200, mean=.8)))

library(ggplot2)
library(dplyr)
cdat <- dat %>%
  group_by(cond) %>%
  summarise(rating.mean = mean(rating),
            sem = sd(rating)/sqrt(length(rating)),
            ci.low = mean(rating) - 2*sem,
            ci.upp = mean(rating) + 2*sem)

为了将线条限制为曲线的轮廓，您首先需要找出曲线的哪些 y 值对应于均值，例如通过使用 ggplot_build 访问密度值并使用 approx :

cdat.dens <- ggplot_build(ggplot(dat, aes(x=rating, colour=cond)) + geom_density())$data[[1]] %>%
  mutate(cond = ifelse(group == 1, "A", "B")) %>%
  left_join(cdat) %>%
  select(y, x, cond, rating.mean, sem, ci.low, ci.upp) %>%
  group_by(cond) %>%
  mutate(dens.mean = approx(x, y, xout = rating.mean)[[2]],
         dens.cilow = approx(x, y, xout = ci.low)[[2]],
         dens.ciupp = approx(x, y, xout = ci.upp)[[2]]) %>%
  select(-y, -x) %>%
  slice(1)

> cdat.dens
Source: local data frame [2 x 8]
Groups: cond [2]

   cond rating.mean        sem     ci.low     ci.upp dens.mean dens.cilow dens.ciupp
  <chr>       <dbl>      <dbl>      <dbl>      <dbl>     <dbl>      <dbl>      <dbl>
1     A -0.05775928 0.07217200 -0.2021033 0.08658471 0.3865929   0.403623  0.3643583
2     B  0.87324927 0.07120697  0.7308353 1.01566320 0.3979347   0.381683  0.4096153

然后，您可以将所有内容与各种 geom_segment 结合起来。 s:

ggplot() +
  geom_density(data = dat, aes(x = rating, colour = cond)) +
  geom_segment(data = cdat.dens, aes(x = rating.mean, xend = rating.mean, y = 0, yend = dens.mean, colour = cond),
             linetype = "dashed", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.low, xend = ci.low, y = 0, yend = dens.cilow, colour = cond),
             linetype = "dotted", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.upp, xend = ci.upp, y = 0, yend = dens.ciupp, colour = cond),
               linetype = "dotted", size = 1)

正如 Axeman 指出的，您可以根据 this answer 中所述的带区创建多边形。 .

因此，对于您的数据，您可以子集并添加额外的行，如下所示:

ribbon <- ggplot_build(ggplot(dat, aes(x=rating, colour=cond)) + geom_density())$data[[1]] %>%
  mutate(cond = ifelse(group == 1, "A", "B")) %>%
  left_join(cdat.dens) %>%
  group_by(cond) %>%
  filter(x >= ci.low & x <= ci.upp) %>%
  select(cond, x, y)

ribbon <- rbind(data.frame(cond = c("A", "B"), x = c(-0.2021033, 0.7308353), y = c(0, 0)), 
                as.data.frame(ribbon), 
                data.frame(cond = c("A", "B"), x = c(0.08658471, 1.01566320), y = c(0, 0)))

并添加 geom_polygon情节:

ggplot() +
  geom_polygon(data = ribbon, aes(x = x, y = y, fill = cond), alpha = .5) +
  geom_density(data = dat, aes(x = rating, colour = cond)) +
  geom_segment(data = cdat.dens, aes(x = rating.mean, xend = rating.mean, y = 0, yend = dens.mean, colour = cond),
             linetype = "dashed", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.low, xend = ci.low, y = 0, yend = dens.cilow, colour = cond),
             linetype = "dotted", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.upp, xend = ci.upp, y = 0, yend = dens.ciupp, colour = cond),
               linetype = "dotted", size = 1)

这是您的真实数据的改编代码。合并两个组而不是一个组有点棘手:

cdat <- dat %>%
  group_by(direction, cond) %>%
  summarise(rating.mean = mean(rating, na.rm = TRUE),
            sem = sd(rating, na.rm = TRUE)/sqrt(length(rating)),
            ci.low = mean(rating, na.rm = TRUE) - 2*sem,
            ci.upp = mean(rating, na.rm = TRUE) + 2*sem)

cdat.dens <- ggplot_build(ggplot(dat, aes(x=rating, colour=interaction(direction, cond))) + geom_density())$data[[1]] %>%
  mutate(cond = ifelse((group == 1 | group == 2 | group == 3 | group == 4), "A",
                        ifelse((group == 5 | group == 6 | group == 7 | group == 8), "B", "C")),
         direction = ifelse((group == 1 | group == 5 | group == 9), "EAST",
                            ifelse((group == 2 | group == 6 | group == 10), "NORTH",
                                   ifelse((group == 3 | group == 7 | group == 11), "SOUTH", "WEST")))) %>%
  left_join(cdat) %>%
  select(y, x, cond, direction, rating.mean, sem, ci.low, ci.upp) %>%
  group_by(cond, direction) %>%
  mutate(dens.mean = approx(x, y, xout = rating.mean)[[2]],
         dens.cilow = approx(x, y, xout = ci.low)[[2]],
         dens.ciupp = approx(x, y, xout = ci.upp)[[2]]) %>%
  select(-y, -x) %>%
  slice(1)

ggplot() +
  geom_density(data = dat, aes(x = rating, colour = cond)) +
  geom_segment(data = cdat.dens, aes(x = rating.mean, xend = rating.mean, y = 0, yend = dens.mean, colour = cond),
               linetype = "dashed", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.low, xend = ci.low, y = 0, yend = dens.cilow, colour = cond),
               linetype = "dotted", size = 1) +
  geom_segment(data = cdat.dens, aes(x = ci.upp, xend = ci.upp, y = 0, yend = dens.ciupp, colour = cond),
               linetype = "dotted", size = 1) +
  facet_wrap(~direction)

关于r - 将 vline 添加到 geom_density 和均值 R 的阴影置信区间，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/41971150/

文章推荐： javascript - Rails - 使链接与 ajax 一起工作

r - 如何获得所选列的平均值(均值)
我想获取每一行某些列的平均值。我有此数据: w=c(5,6,7,8) x=c(1,2,3,4) y=c(1,2,3) length(y)=4 z=data.frame(w,x,y) 哪个返回:
python - 带条件的向量化 numpy 均值
类似于Numpy mean with condition我的问题将其扩展到对矩阵进行操作:计算矩阵 rdat 的行均值，跳过某些单元格 - 在本例中我使用 0 作为要跳过的单元格 - 就好像这些值从一
python - 如何对产品推荐数据集使用 k 均值
我有一个数据集，其中的列标题为产品名称、品牌、评级(1:5)、评论文本、评论有用性。我需要的是提出一个使用评论的推荐算法。我这里必须使用 python 进行编码。数据集采用.csv 格式。为了识别数
statistics - 椭圆体的 k 均值
我在 R^3 中有 n 个点，我想用 k 个椭球体或圆柱体覆盖它们(我不在乎；以更容易的为准)。我想大约最小化卷的并集。假设 n 是数万，k 是少数。开发时间(即简单性)比运行时更重要。显然我可以运
java - 均值、中值、方差计算器
我创建了一个计算均值、中位数和方差的程序。该程序最多接受 500 个输入。当有 500 个输入(我的数组的最大大小)时，我的所有方法都能完美运行。当输入较少时，只有“平均值”计算器起作用。这是整个程序
c++ - 使用推力库获取最近的质心？ (K-均值)
我已经完成了距离的计算并存储在推力 vector 中，例如，我有 2 个质心和 5 个数据点，我计算距离的方法是，对于每个质心，我首先计算 5 个数据点的距离并存储在阵列，然后与距离一维阵列中的另一个
python - Pandas GroupBy 均值
下面的代码适用于每一列的总数，但我想计算出每个物种的平均值。 # Read data file into array data = numpy.genfromtxt('data/iris.csv',
python - 仅在相似列上跨两个数据框的 Pandas 均值
我有一个独特的要求，我需要两个数据帧的公共(public)列(每行)的平均值。我想不出这样做的 pythonic 方式。我知道我可以遍历两个数据框并找到公共(public)列，然后获取键匹配的行的平
OpenCV 均值/SD 过滤器
我把它扔在那里，希望有人会尝试过这种荒谬的事情。我的目标是获取输入图像，并根据每个像素周围小窗口的标准差对其进行分割。基本上，这在数学上应该类似于高斯或盒式过滤器，因为它将应用于编译时(甚至运行时)用
python - 跨数组切片向量化 numpy 均值
有没有一种方法可以对函数进行向量化处理，使输出成为均值数组，其中每个均值代表输入数组的 0 索引值的均值？循环这个非常简单，但我正在努力尽可能高效。例如0 = 均值(0)，1 = 均值(0-1)，N
c++ - 如何生成具有指数分布(均值)的随机数？
我正在尝试生成均值为 1 的指数分布随机数。我知道如何获取具有均值和标准差的正态分布随机数。我们可以通过normal(mean, standard_deviation)得到它，但是我不知道如何得到指数
python - 参数中带有比较运算符的 numpy 均值
我遇到了一段 Python 代码，它的内容类似于以下内容: a = np.array([1,2,3,4,5,6,7]) a array([1, 2, 3, 4, 5, 6, 7]) np.mean(a
python - 计算python中分布的矩(均值，方差)
我有两个数组。 x 是独立变量，counts 是 x 出现的次数，就像直方图一样。我知道我可以通过定义一个函数来计算平均值: def mean(x,counts): return np.sum
python - 有条件的 Numpy 均值
我有在纯 python 中计算平均速度的算法: speed = [...] avg_speed = 0.0 speed_count = 0 for i in speed: if i > 0:
r - 按组计算的累积(扩展窗口)均值，对每个计算进行重复检查
我正在尝试计算扩展窗口的平均值，但是数据结构使得之前的答案至少缺少一点所需的内容(最接近的是:link)。我的数据看起来像这样: Company TimePeriod IndividualID
python - 使用具有余弦相似度的 K 均值 - Python
我正在尝试实现 Kmeans python中的算法将使用cosine distance而不是欧几里得距离作为距离度量。我知道使用不同的距离函数可能是致命的，应该小心使用。使用余弦距离作为度量迫使我改
k-means - 自组织映射与 k 均值
有谁知道自组织映射 (SOM) 与 k 均值相比效果如何？我相信通常在颜色空间(例如 RGB)中，SOM 是将颜色聚类在一起的更好方法，因为视觉上不同的颜色之间的颜色空间存在重叠( http://ww
c++ - 无分支 K 均值(或其他优化)
注意:我希望能得到更多有关如何处理和提出此类解决方案的指南，而不是解决方案本身。我的系统中有一个非常关键的功能，它在特定上下文中显示为排名第一的分析热点。它处于 k-means 迭代的中间(已经是多
python - 如何用python描述矩阵中的所有二因子列组合(均值、中位数、计数等)？
我有一个 pandas 数据框，看起来像这样: 给定行中的每个值要么是相同的数字，要么是 NaN。我想计算数据框中所有两列组合的平均值、中位数和获取计数，其中两列都不是 NaN。例如，上述数据帧的结
machine-learning - 改进某些数据集上的 K 均值
任何人都知道如何调整简单的 K 均值算法来处理 this form 的数据集. 最佳答案在仍然使用 k-means 的同时处理该形式的数据的最直接方法是使用 k-means 的内核化版本。 JSAT

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

r - 将 vline 添加到 geom_density 和均值 R 的阴影置信区间