gpt4 book ai didi

r - Rhadoop-使用RMR的字数统计

转载 作者:行者123 更新时间:2023-12-02 21:38:50 25 4
gpt4 key购买 nike

我正在尝试使用Rhadoop软件包运行一个简单的rmr作业,但是它不起作用。这是我的R脚本

print("Initializing variable.....")
Sys.setenv(HADOOP_HOME="/usr/hdp/2.2.4.2-2/hadoop")
Sys.setenv(HADOOP_CMD="/usr/hdp/2.2.4.2-2/hadoop/bin/hadoop")
print("Invoking functions.......")
#Referece taken from Revolution Analytics
wordcount = function( input, output = NULL, pattern = " ")
{
mapreduce(
input = input ,
output = output,
input.format = "text",
map = wc.map,
reduce = wc.reduce,
combine = T)
}

wc.map =
function(., lines) {
keyval(
unlist(
strsplit(
x = lines,
split = pattern)),
1)}

wc.reduce =
function(word, counts ) {
keyval(word, sum(counts))}

#Function Invoke

wordcount('/user/hduser/rmr/wcinput.txt')

我正在上面的脚本中运行
Rscript wordcount.r

我正在错误以下。
[1] "Initializing variable....."
[1] "Invoking functions......."
Error in wordcount("/user/hduser/rmr/wcinput.txt") :
could not find function "mapreduce"
Execution halted

请让我知道是什么问题。

最佳答案

首先,您必须在代码中设置HADOOP_STREAMING环境变量。

尝试下面的代码,并注意该代码假定您已将文本文件复制到hdfs文件夹examples/wordcount/data
R代码:

Sys.setenv("HADOOP_CMD"="/usr/local/hadoop/bin/hadoop")
Sys.setenv("HADOOP_STREAMING"="/usr/local/hadoop/share/hadoop/tools/lib/hadoop-streaming-2.4.0.jar")

# load librarys
library(rmr2)
library(rhdfs)

# initiate rhdfs package
hdfs.init()

map <- function(k,lines) {
words.list <- strsplit(lines, '\\s')
words <- unlist(words.list)
return( keyval(words, 1) )
}

reduce <- function(word, counts) {
keyval(word, sum(counts))
}

wordcount <- function (input, output=NULL) {
mapreduce(input=input, output=output, input.format="text", map=map, reduce=reduce)
}

## read text files from folder example/wordcount/data
hdfs.root <- 'example/wordcount'
hdfs.data <- file.path(hdfs.root, 'data')

## save result in folder example/wordcount/out
hdfs.out <- file.path(hdfs.root, 'out')

## Submit job
out <- wordcount(hdfs.data, hdfs.out)

## Fetch results from HDFS
results <- from.dfs(out)
results.df <- as.data.frame(results, stringsAsFactors=F)
colnames(results.df) <- c('word', 'count')

head(results.df)

输出:
word count
AS 16
As 5
B. 1
BE 13
BY 23
By 7

供您引用, here是运行R字数映射减少程序的另一个示例。

希望这可以帮助。

关于r - Rhadoop-使用RMR的字数统计,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/30261383/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com