r - 将 plm 拟合值合并到数据集-6ren

r - 将 plm 拟合值合并到数据集

转载作者：行者123 更新时间：2023-12-04 02:07:32

36

4

我正在使用 plm 处理固定效应回归模型。

模型如下所示:

FE.model <-plm(fml, data = data.reg2,
           index=c('Site.ID','date.hour'), # cross section ID and time series ID
           model='within', #coefficients are fixed
           effect='individual')
summary(FE.model)

“fml”是我之前定义的公式。我有很多自变量，所以这使它更有效率。

我想要做的是获取我的拟合值(我的 yhats)并将它们加入我的基础数据集；数据.reg2

我能够使用此代码获得拟合值:

 Fe.model.fitted <- FE.model$model[[1]] - FE.model$residuals

但是，这只给了我一个仅包含拟合值的列向量 - 我无法将它加入我的基础数据集。

或者，我尝试过这样的事情:

 Fe.model.fitted <- cbind(data.reg2, resid=resid(FE.model), fitted=fitted(FE.model))

但是，我得到了这个错误:

 Error in as.data.frame.default(x[[i]], optional = TRUE) : cannot coerce class ""pseries"" to a data.frame

还有其他方法可以在我的基础数据集中获取我的拟合值吗？或者有人可以解释我遇到的错误以及修复它的方法吗？

我应该注意，我不想根据我的 beta 手动计算 yhats。我对该选项有太多的自变量，并且我定义的公式(fml)可能会改变，因此该选项不会有效。

非常感谢!!

最佳答案

将 plm 拟合值合并回原始数据集需要一些中间步骤 - plm 删除任何缺少数据的行，据我所知，plm 对象不包含索引信息。数据的顺序不保留——请参阅 plm 的作者之一 Giovanni Millo 在 this thread 中的评论:

"...the input order is not always preserved: observations are always reordered by (individual, time) internally, so that the output you get is ordered accordingly..."

简单的步骤:

从估计的 plm 对象中获取拟合值。它是单个向量，但条目已命名。名称与索引中的位置相对应。
使用 index() 函数获取索引。它可以返回单个索引和时间索引。请注意，索引可能包含比拟合值更多的行，以防因缺失数据而删除行。 (也可以直接从原始数据生成索引，但我没有看到 plm 返回的内容中保留数据原始顺序的 promise 。)
合并到原始数据中，从索引中查找 id 和 time 值。

下面提供了示例代码。有点长，但我试图发表评论。代码没有优化，我的意图是明确列出这些步骤。另外，我使用的是 data.tables 而不是 data.frames。

library(data.table); library(plm)

### Generate dummy data. This way we know the "true" coefficients
set.seed(100)
n <- 500 # Run with more data if you want to get closer to the "true" coefficients
DT <- data.table(CJ(id = c("a","b","c","d","e"), time = c(1:(n / 5))))
DT[, x1 := rnorm(n)]
DT[, x2 := rnorm(n)]
DT[, y  := x1 + 2 * x2 + rnorm(n) / 10]

setkey(DT, id, time)
# # Make it an unbalanced panel & put in some NAs
DT <- DT[!(id == "a" & time == 4)]
DT[.("a", 3), x2 := as.numeric(NA)]
DT[.("d", 2), x2 := as.numeric(NA)]

str(DT)

### Run the model -- both individual and time effects; "within" model
summary(PLM <- plm(data = DT, id = c("id", "time"), formula = y ~ x1 + x2, model = "within", effect = "twoways", na.action = "na.omit"))

### Merge the fitted values back into the data.table DT
# Note that PLM$model$y is shorter than the data, i.e. the row(s) with NA have been dropped
cat("\nRows omitted (due to NA): ", nrow(DT) - length(PLM$model$y))

# Since the objects returned by plm() do not contain the index, need to generate it from the data
# The object returned by plm(), i.e. PLM$model$y, has names that point to the place in the index
# Note: The index can also be done as INDEX <- DT[, j = .(id, time)], but use the longer way with index() in case plm does not preserve the order
INDEX <- data.table(index(x = pdata.frame(x = DT, index = c("id", "time")), which = NULL)) # which = NULL extracts both the individual and time indexes
INDEX[, id := as.character(id)]
INDEX[, time := as.integer(time)] # it is returned as a factor, convert back to integer to match the variable type in DT

# Generate the fitted values as the difference between the y values and the residuals
if (all(names(PLM$residuals) == names(PLM$model$y))) { # this should not be needed, but just in case...
    FIT <- data.table(
        index   = as.integer(names(PLM$model$y)), # this index corresponds to the position in the INDEX, from where we get the "id" and "time" below
        fit.plm = as.numeric(PLM$model$y) - as.numeric(PLM$residuals)
    )
}

FIT[, id   := INDEX[index]$id]
FIT[, time := INDEX[index]$time]
# Now FIT has both the id and time variables, can match it back into the original dataset (i.e. we have the missing data accounted for)
DT <- merge(x = DT, y = FIT[, j = .(id, time, fit.plm)], by = c("id", "time"), all = TRUE) # Need all = TRUE, or some data from DT will be dropped!

关于r - 将 plm 拟合值合并到数据集，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/23143428/

36

4

0

文章推荐： highcharts - Highchart 图案填充

文章推荐： asp.net-core - 我可以调试到 asp.net 核心源代码吗？

文章推荐： git - git 缓存它的结果吗？

嵌套函数的 Gnuplot 拟合
gnuplot 中拟合函数的正确方法是什么 f(x)有下一个表格吗？ f(x) = A*exp(x - B*f(x)) 我尝试使用以下方法将其拟合为任何其他函数: fit f(x) "data.txt
pytorch动态神经网络(拟合)实现
（1）首先要建立数据集 ? 1
r - 如何使用非线性函数拟合数据并绘制数据并使用 ggplot() 拟合
测量显示一个信号，其形式类似于具有偏移量和因子的平方根函数。如何找到系数并在一个图中绘制原始数据和拟合曲线？ require(ggplot2) require(nlmrt) # may be thi
r - 对正弦模型进行非线性最小二乘 (nls) 拟合
我想将以下函数拟合到我的数据中: f(x) = Offset+Amplitudesin(FrequencyT+Phase)，或根据 Wikipedia : f(x) = C+alphasin(ome
c# - 拟合 Akima 样条曲线
我正在尝试使用与此工具相同的方法在 C# 中拟合 Akima 样条曲线:https://www.mycurvefit.com/share/4ab90a5f-af5e-435e-9ce4-652c95c
javascript - 添加功能之前的 OpenLayers 拟合
问题:开放层适合 map ，只有在添加特征之后(视觉)，我该如何避免这种情况？我在做这个第 1 步 - 创建特征 var feature = new ol.Feature({...}); 第 2
javascript - 拟合 D3 图表的数据以创建图例
我有一个数据变量，其中包含以下内容: [Object { score="2.8", word="Blue"}, Object { score="2.8", word="Red"}, Objec
python - 拟合 RandomForestClassifier 时内存使用量激增
我正在尝试用中等大小的 numpy float 组来填充森林 In [3]: data.shape Out[3]: (401125, 5) [...] forest = forest.fit(data
matlab - 使用 lsqcurvefit 拟合
我想用洛伦兹函数拟合一些数据，但我发现当我使用不同数量级的参数时拟合会出现问题。这是我的洛伦兹函数: function [ value ] = lorentz( x,x0,gamma,amp )
matlab - 区分居中和缩放的 Polyfit 拟合
我有一些数据，我希望对其进行建模，以便能够在与数据相同的范围内获得相对准确的值。为此，我使用 polyfit 来拟合 6 阶多项式，由于我的 x 轴值，它建议我将其居中并缩放以获得更准确的拟合。但
python - 拟合 beta 二项式
我一直在寻找一种方法来使数据符合 beta 二项分布并估计 alpha 和 beta，类似于 VGAM 库中的 vglm 包的方式。我一直无法找到如何在 python 中执行此操作。有一个 scipy
numpy - 拟合 scipy.optimize 参数的错误
我将 scipy.optimize.minimize ( https://docs.scipy.org/doc/scipy/reference/tutorial/optimize.html ) 函数与
python - 拟合 Von Mises 分布的圆形直方图
在过去的几天里，我一直在尝试使用 python 绘制圆形数据，方法是构建一个范围从 0 到 2pi 的圆形直方图并拟合 Von Mises 分布。我真正想要实现的是: 具有拟合 Von-Mises 分
python - Keras 拟合 LSTM 在循环中变得更慢
我有一个简单的循环，它在每次迭代中都会创建一个 LSTM(具有相同的参数)并将其拟合到相同的数据。问题是迭代过程中需要越来越多的时间。 batch_size = 10 optimizer = opti
Python/Scipy kde 拟合、缩放
我有一个 Python 系列，我想为其直方图拟合密度。问题:是否有一种巧妙的方法可以使用 np.histogram() 中的值来实现此结果？ (请参阅下面的更新) 我目前的问题是，我执行的 kde 拟
python - 拟合 Keras L1 模型
我有一个简单的 keras 模型(正常套索线性模型)，其中输入被移动到单个“神经元”Dense(1, kernel_regularizer=l1(fdr))(input_layer) 但是权重从这个模
python - 拟合 sklearn GridSearchCV 模型
我正在尝试解决 Boston Dataset 上的回归问题在random forest regressor的帮助下.我用的是GridSearchCV用于选择最佳超参数。问题一我是否应该将 Grid
python - Python 中的交互式 BSpline 拟合
使用以下函数，可以在输入点 P 上拟合三次样条: def plotCurve(P): pts = np.vstack([P, P[0]]) x, y = pts.T i = np.aran
python - 拟合 3D 点 python
我有 python 代码可以生成数字 x、y 和 z 的三元组列表。我想使用 scipy curve_fit 来拟合 z= f(x,y)。这是一些无效的代码 A = [(19,20,24), (10,
r - 使用 fitdistrplus 拟合 Gumbel 分布
我正在尝试从 this answer 中复制代码，但是我在这样做时遇到了问题。我正在使用包 VGAM 中的gumbel 发行版和 fitdistrplus . 做的时候出现问题: fit = fi

首页

博学

6Ren·AI

商城

r - 将 plm 拟合值合并到数据集