r - K-均值算法，R-6ren

r - K-均值算法，R

转载作者：行者123 更新时间：2023-11-30 08:49:39

24

4

大家好!我被要求在 R 上创建一个 K 均值算法，但我并不真正了解这门语言，所以我在互联网上找到了一些示例代码，并决定使用。我研究了它，了解了其中使用的功能，并对其进行了一些修正，因为它运行得不太好。代码如下:

# Creating a sample of data
y=rnorm(500,1.65)
x=rnorm(500,1.15)
x=cbind(x,y)
centers <- x[sample(nrow(x),5),]

# A function for calculating the distance between centers and the rest of the dots
euclid <- function(points1, points2) {
  distanceMatrix <- matrix(NA, nrow=dim(points1)[1], ncol=dim(points2)[1])
  for(i in 1:nrow(points2)) {
    distanceMatrix[,i] <- sqrt(rowSums(t(t(points1)-points2[i,])^2))
  }
  distanceMatrix
}


# A method function
K_means <- function(x, centers, euclid, nItter) {
  clusterHistory <- vector(nItter, mode="list")
  centerHistory <- vector(nItter, mode="list")

  for(i in 1:nItter) {
    distsToCenters <- euclid(x, centers)
    clusters <- apply(distsToCenters, 1, which.min)
    centers <- apply(x, 2, tapply, clusters, mean)
    # Saving history
    clusterHistory[[i]] <- clusters
    centerHistory[[i]] <- centers
  }

  structure(list(clusters = clusterHistory, centers = centerHistory))

}


res <- K_means(x, centers, euclid, 5)
#To use the same plot operations I had to use unlist, since the resulting object in my function is a list of lists,
#and default object is just a list. And also i store the history of each iteration in that object.
res <- unlist(res, recursive = FALSE)
plot(x, col = res$clusters5)
points(res$centers5, col = 1:5, pch = 8, cex = 2)

它在这个简单的矩阵上运行良好。但有人要求我在 iris 上使用它:

head(iris)
a <-data.frame(iris$Sepal.Length, iris$Sepal.Width, iris$Petal.Length, iris$Petal.Width)
centers <- a[sample(nrow(a),3),]
iris_clusters <- K_means(a, centers, euclid, 3)
iris_clusters <- unlist(iris_clusters, recursive = FALSE)
head(iris_clusters)

问题是它不起作用。错误是:

Error in distanceMatrix[, i] <- sqrt(rowSums(t(t(points1) - points2[i,  : 
  number of items to replace is not a multiple of replacement length

我知道物体的尺寸不匹配，但我不明白为什么。这就是我寻求帮助的原因。我提前对这段代码中可能存在的所有愚蠢之处表示歉意，但我还不太熟悉这门语言，所以不要对我评价太严厉。谢谢!

最佳答案

您的实现应该适用于简单的类型转换

iris_clusters <- K_means(as.matrix(a), as.matrix(centers), euclid, 3) # 3 iterations

iris_clusters <- unlist(iris_clusters, recursive = FALSE)

# plotting the clusters obtained on the first two dimensions at the end of 3rd iteration

plot(a[,1:2], col = iris_clusters$clusters3, pch=19) 
points(iris_clusters$centers3, col = 1:5, pch = 8, cex = 2)

head(iris_clusters)

# cluster assignments and centroids computed at different iterations

$clusters1
  [1] 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 2 2 3 2 3 2 3 2 3 3 3 3 2 3 3 3 3 3 3 2 3 2 2 3 3
 [77] 2 2 3 3 3 3 3 2 3 3 2 3 3 3 3 2 3 3 3 3 3 3 3 3 1 2 1 2 1 1 3 1 1 1 2 2 2 2 2 2 2 1 1 2 1 2 1 2 1 1 2 2 2 1 1 1 2 2 2 1 2 2 2 2 1 2 2 1 1 2 2 2 2 2

$clusters2
  [1] 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 2 2 2 2 2 2 3 2 3 3 2 2 2 3 2 2 2 2 3 2 2 2 2 2 2
 [77] 2 2 2 3 3 3 2 2 2 2 2 2 2 2 2 2 2 3 2 2 2 2 3 2 1 2 1 2 1 1 2 1 1 1 2 2 1 2 2 2 2 1 1 2 1 2 1 2 1 1 2 2 2 1 1 1 2 2 2 1 2 2 2 1 1 2 2 1 1 2 2 2 2 2

$clusters3
  [1] 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 2 2 2 2 2 2 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2
 [77] 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 3 2 2 2 2 3 2 1 2 1 2 1 1 2 1 1 1 2 2 1 2 2 2 2 1 1 2 1 2 1 2 1 1 2 2 1 1 1 1 1 2 2 1 1 2 2 1 1 1 2 1 1 1 2 2 2 2

$centers1
  iris.Sepal.Length iris.Sepal.Width iris.Petal.Length iris.Petal.Width
1          7.150000         3.120000          6.090000        2.1350000
2          6.315909         2.915909          5.059091        1.8000000
3          5.297674         3.115116          2.550000        0.6744186

$centers2
  iris.Sepal.Length iris.Sepal.Width iris.Petal.Length iris.Petal.Width
1          7.122727         3.113636          6.031818        2.1318182
2          6.123529         2.852941          4.741176        1.6132353
3          5.056667         3.268333          1.810000        0.3883333

$centers3
  iris.Sepal.Length iris.Sepal.Width iris.Petal.Length iris.Petal.Width
1          7.014815         3.096296          5.918519         2.155556
2          6.025714         2.805714          4.588571         1.518571
3          5.005660         3.369811          1.560377         0.290566

关于r - K-均值算法，R，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/40979777/

24

4

0

文章推荐： java - db2 - 查询结果到新表

文章推荐：传递的 Javascript 参数小于调用的参数

文章推荐： java - dependsOnGroups顺序-Testng

文章推荐： java - OGNL 调用构造函数失败

r - 如何获得所选列的平均值(均值)
我想获取每一行某些列的平均值。我有此数据: w=c(5,6,7,8) x=c(1,2,3,4) y=c(1,2,3) length(y)=4 z=data.frame(w,x,y) 哪个返回:
python - 带条件的向量化 numpy 均值
类似于Numpy mean with condition我的问题将其扩展到对矩阵进行操作:计算矩阵 rdat 的行均值，跳过某些单元格 - 在本例中我使用 0 作为要跳过的单元格 - 就好像这些值从一
python - 如何对产品推荐数据集使用 k 均值
我有一个数据集，其中的列标题为产品名称、品牌、评级(1:5)、评论文本、评论有用性。我需要的是提出一个使用评论的推荐算法。我这里必须使用 python 进行编码。数据集采用.csv 格式。为了识别数
statistics - 椭圆体的 k 均值
我在 R^3 中有 n 个点，我想用 k 个椭球体或圆柱体覆盖它们(我不在乎；以更容易的为准)。我想大约最小化卷的并集。假设 n 是数万，k 是少数。开发时间(即简单性)比运行时更重要。显然我可以运
java - 均值、中值、方差计算器
我创建了一个计算均值、中位数和方差的程序。该程序最多接受 500 个输入。当有 500 个输入(我的数组的最大大小)时，我的所有方法都能完美运行。当输入较少时，只有“平均值”计算器起作用。这是整个程序
c++ - 使用推力库获取最近的质心？ (K-均值)
我已经完成了距离的计算并存储在推力 vector 中，例如，我有 2 个质心和 5 个数据点，我计算距离的方法是，对于每个质心，我首先计算 5 个数据点的距离并存储在阵列，然后与距离一维阵列中的另一个
python - Pandas GroupBy 均值
下面的代码适用于每一列的总数，但我想计算出每个物种的平均值。 # Read data file into array data = numpy.genfromtxt('data/iris.csv',
python - 仅在相似列上跨两个数据框的 Pandas 均值
我有一个独特的要求，我需要两个数据帧的公共(public)列(每行)的平均值。我想不出这样做的 pythonic 方式。我知道我可以遍历两个数据框并找到公共(public)列，然后获取键匹配的行的平
OpenCV 均值/SD 过滤器
我把它扔在那里，希望有人会尝试过这种荒谬的事情。我的目标是获取输入图像，并根据每个像素周围小窗口的标准差对其进行分割。基本上，这在数学上应该类似于高斯或盒式过滤器，因为它将应用于编译时(甚至运行时)用
python - 跨数组切片向量化 numpy 均值
有没有一种方法可以对函数进行向量化处理，使输出成为均值数组，其中每个均值代表输入数组的 0 索引值的均值？循环这个非常简单，但我正在努力尽可能高效。例如0 = 均值(0)，1 = 均值(0-1)，N
c++ - 如何生成具有指数分布(均值)的随机数？
我正在尝试生成均值为 1 的指数分布随机数。我知道如何获取具有均值和标准差的正态分布随机数。我们可以通过normal(mean, standard_deviation)得到它，但是我不知道如何得到指数
python - 参数中带有比较运算符的 numpy 均值
我遇到了一段 Python 代码，它的内容类似于以下内容: a = np.array([1,2,3,4,5,6,7]) a array([1, 2, 3, 4, 5, 6, 7]) np.mean(a
python - 计算python中分布的矩(均值，方差)
我有两个数组。 x 是独立变量，counts 是 x 出现的次数，就像直方图一样。我知道我可以通过定义一个函数来计算平均值: def mean(x,counts): return np.sum
python - 有条件的 Numpy 均值
我有在纯 python 中计算平均速度的算法: speed = [...] avg_speed = 0.0 speed_count = 0 for i in speed: if i > 0:
r - 按组计算的累积(扩展窗口)均值，对每个计算进行重复检查
我正在尝试计算扩展窗口的平均值，但是数据结构使得之前的答案至少缺少一点所需的内容(最接近的是:link)。我的数据看起来像这样: Company TimePeriod IndividualID
python - 使用具有余弦相似度的 K 均值 - Python
我正在尝试实现 Kmeans python中的算法将使用cosine distance而不是欧几里得距离作为距离度量。我知道使用不同的距离函数可能是致命的，应该小心使用。使用余弦距离作为度量迫使我改
k-means - 自组织映射与 k 均值
有谁知道自组织映射 (SOM) 与 k 均值相比效果如何？我相信通常在颜色空间(例如 RGB)中，SOM 是将颜色聚类在一起的更好方法，因为视觉上不同的颜色之间的颜色空间存在重叠( http://ww
c++ - 无分支 K 均值(或其他优化)
注意:我希望能得到更多有关如何处理和提出此类解决方案的指南，而不是解决方案本身。我的系统中有一个非常关键的功能，它在特定上下文中显示为排名第一的分析热点。它处于 k-means 迭代的中间(已经是多
python - 如何用python描述矩阵中的所有二因子列组合(均值、中位数、计数等)？
我有一个 pandas 数据框，看起来像这样: 给定行中的每个值要么是相同的数字，要么是 NaN。我想计算数据框中所有两列组合的平均值、中位数和获取计数，其中两列都不是 NaN。例如，上述数据帧的结
machine-learning - 改进某些数据集上的 K 均值
任何人都知道如何调整简单的 K 均值算法来处理 this form 的数据集. 最佳答案在仍然使用 k-means 的同时处理该形式的数据的最直接方法是使用 k-means 的内核化版本。 JSAT

首页

博学

6Ren·AI

商城

r - K-均值算法，R