scikit-learn - class_weight 在 linearSVC 和 LogisticRegression 损失函数中的作用-6ren

scikit-learn - class_weight 在 linearSVC 和 LogisticRegression 损失函数中的作用

转载作者：行者123 更新时间：2023-12-04 18:15:41

27

4

我试图弄清楚损失函数公式到底是什么以及如何在 class_weight='auto' 时手动计算它在 svm.svc 的情况下, svm.linearSVC和 linear_model.LogisticRegression .

对于平衡数据，假设您有一个经过训练的分类器:clf_c .物流损失应该是(我说得对吗？):

def logistic_loss(x,y,w,b,b0):
    '''
    x: nxp data matrix where n is number of data points and p is number of features.
    y: nx1 vector of true labels (-1 or 1).
    w: nx1 vector of weights (vector of 1./n for balanced data).
    b: px1 vector of feature weights.
    b0: intercept.
    '''
    s = y
    if 0 in np.unique(y):
        print 'yes'
        s = 2. * y - 1
    l = np.dot(w, np.log(1 + np.exp(-s * (np.dot(x, np.squeeze(b)) + b0))))
    return l

我意识到逻辑回归有 predict_log_proba()当数据平衡时，它可以准确地为您提供:

b, b0 = clf_c.coef_, clf_c.intercept_
w = np.ones(len(y))/len(y)
-(clf_c.predict_log_proba(x[xrange(len(x)), np.floor((y+1)/2).astype(np.int8)]).mean() == logistic_loss(x,y,w,b,b0)

注意， np.floor((y+1)/2).astype(np.int8)简单地将 y=(-1,1) 映射到 y=(0,1)。

但是当数据不平衡时，这不起作用。

更重要的是，当数据处于平衡状态且 class_weight=None 时，您希望分类器(此处为逻辑回归)具有类似的性能(就损失函数值而言)。对比数据不平衡和 class_weight='auto' .我需要有一种方法来计算两种场景的损失函数(没有正则化项)并比较它们。

总之， class_weight = 'auto'有什么用？正好意思？是不是意思 class_weight = {-1 : (y==1).sum()/(y==-1).sum() , 1 : 1.}或者更确切地说 class_weight = {-1 : 1./(y==-1).sum() , 1 : 1./(y==1).sum()} ?

非常感谢任何帮助。我尝试浏览源代码，但我不是程序员，我被卡住了。
非常感谢。

最佳答案

class_weight启发式

我对您对 class_weight='auto' 的第一个提议感到有些困惑。启发式，如:

class_weight = {-1 : (y == 1).sum() / (y == -1).sum(), 
                1 : 1.}

如果我们对其进行标准化以使权重总和为 1，则与您的第二个命题相同。

反正明白什么 class_weight="auto"有，看这个问题:
what is the difference between class weight = none and auto in svm scikit learn .

我把它复制到这里供以后比较:

This means that each class you have (in classes) gets a weight equal to 1 divided by the number of times that class appears in your data (y), so classes that appear more often will get lower weights. This is then further divided by the mean of all the inverse class frequencies.

请注意这不是完全明显的;)。

此启发式已弃用，并将在 0.18 中删除。它将被另一个启发式替换， class_weight='balanced' .

“平衡”启发式方法根据类的频率倒数按比例加权。

从文档:

The "balanced" mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data: n_samples / (n_classes * np.bincount(y)).

np.bincount(y)是一个数组，其中元素 i 是 i 类样本的计数。

这里有一些代码来比较两者:

import numpy as np
from sklearn.datasets import make_classification
from sklearn.utils import compute_class_weight

n_classes = 3
n_samples = 1000

X, y = make_classification(n_samples=n_samples, n_features=20, n_informative=10, 
    n_classes=n_classes, weights=[0.05, 0.4, 0.55])

print("Count of samples per class: ", np.bincount(y))
balanced_weights = n_samples /(n_classes * np.bincount(y))
# Equivalent to the following, using version 0.17+:
# compute_class_weight("balanced", [0, 1, 2], y)

print("Balanced weights: ", balanced_weights)
print("'auto' weights: ", compute_class_weight("auto", [0, 1, 2], y))

输出:

Count of samples per class:  [ 57 396 547]
Balanced weights:  [ 5.84795322  0.84175084  0.60938452]
'auto' weights:  [ 2.40356854  0.3459682   0.25046327]

损失函数

现在真正的问题是:这些权重如何用于训练分类器？

不幸的是，我在这里没有一个彻底的答案。

对于 SVC和 linearSVC文档字符串很清楚

Set the parameter C of class i to class_weight[i]*C for SVC.

因此，高权重意味着该类的正则化较少，并且 svm 对其进行正确分类的激励更高。

我不知道他们如何处理逻辑回归。我会尝试研究它，但大部分代码都在 liblinear 或 libsvm 中，我对它们不太熟悉。

但是，请注意 class_weight 中的权重 不直接影响 predict_proba 等方法 .他们改变输出是因为分类器优化了不同的损失函数。
不确定这是否清楚，所以这里有一个片段来解释我的意思(您需要为导入和变量定义运行第一个):

lr = LogisticRegression(class_weight="auto")
lr.fit(X, y)
# We get some probabilities...
print(lr.predict_proba(X))

new_lr = LogisticRegression(class_weight={0: 100, 1: 1, 2: 1})
new_lr.fit(X, y)
# We get different probabilities...
print(new_lr.predict_proba(X))

# Let's cheat a bit and hand-modify our new classifier.
new_lr.intercept_ = lr.intercept_.copy()
new_lr.coef_ = lr.coef_.copy()

# Now we get the SAME probabilities.
np.testing.assert_array_equal(new_lr.predict_proba(X), lr.predict_proba(X))

希望这可以帮助。

关于scikit-learn - class_weight 在 linearSVC 和 LogisticRegression 损失函数中的作用，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/31657263/

27

4

0

文章推荐： functional-programming - 奇怪的 Haskell/GHCi 问题

使用小批量时累积的 pytorch 损失
我是pytorch的新手。请问添加'loss.item()'有什么区别？以下2部分代码: for epoch in range(epochs): trainingloss =0 for
MySQL 查询返回日期列表的总利润/损失
我有一个包含 4 列的 MySQL 表，如下所示。 TransactionID | Item | Amount | Date ------------------------------------
iphone - 有多个函数调用而不是一个大的函数调用有任何性能增益/损失？
我目前正在使用 cocos2d、Box2D 和 Objective-C 为 iPad 和 iPhone 制作游戏。每次更新都会发生很多事情，很多事情必须解决。我最近将我的很多代码重构为几个小方法，
tensorflow - 混合精度训练导致 NaN 损失
我一直在关注 Mixed Precision Guide .因此，我正在设置: keras.mixed_precision.set_global_policy(mixed_precision) 像这样
Java 大 double 损失
double lnumber = Math.pow(2, 1000); 打印 1.0715086071862673E301 我尝试过的事情我尝试使用 BigDecimal 类来扩展这个数字: St
python - 使用神经网络进行函数逼近 - 损失 0
我正在尝试创建一个神经网络来近似函数(正弦、余弦、自定义...)，但我在格式上遇到困难，我不想使用输入标签，而是使用输入输出。我该如何更改它？我正在关注this tutorial import te
python - 训练回归网络时的 NaN 损失
我有一个具有 260,000 行和 35 列的“单热编码”(全一和零)数据矩阵。我正在使用 Keras 训练一个简单的神经网络来预测一个连续变量。制作网络的代码如下: model = Sequenti
image-processing - 什么是像素级 softmax 损失？
什么是像素级 softmax 损失？在我的理解中，这只是一个交叉熵损失，但我没有找到公式。有人能帮我吗？最好有pytorch代码。最佳答案您可以阅读 here所有相关内容(那里还有一个指向源代码的
python - PyTorch 中多输出回归问题的 RMSE 损失
我正在训练一个 CNN 架构来使用 PyTorch 解决回归问题，其中我的输出是一个 20 个值的张量。我计划使用 RMSE 作为模型的损失函数，并尝试使用 PyTorch 的 nn.MSELoss(
tensorflow - 损失、准确性、验证损失、验证准确性之间有什么区别？
在每个时代结束时，我得到例如以下输出: Epoch 1/25 2018-08-06 14:54:12.555511: 2/2 [==============================] - 86
deep-learning - 计算网络两个输出之间的 cosine_proximity 损失
我正在使用 Keras 2.0.2 功能 API (Tensorflow 1.0.1) 来实现一个网络，该网络接受多个输入并产生两个输出 a 和 b。我需要使用 cosine_proximity 损失
python - 第一个纪元后的神经网络生成 NaN 值作为输出、损失
我正在尝试设置很少层的神经网络，这将解决简单的回归问题，这应该是f(x) = 0,1x 或 f(x) = 10x 所有代码如下所示(数据生成和神经网络) 4 个带有 ReLu 的全连接层损失函数 R
python - 越来越大的正 WGAN-GP 损失
我正在研究在 PyTorch 中使用带有梯度惩罚的 Wasserstein GAN，但始终得到大的、正的生成器损失，并且随着时间的推移而增加。我从 Caogang's implementation
python - TensorFlow 中的最大 margin 损失
我正在尝试在 TensorFlow 中实现最大利润损失。这个想法是我有一些积极的例子，我对一些消极的例子进行了采样，并想计算类似的东西其中 B 是我的批处理大小，N 是我要使用的负样本数。我是 t
python - 损失 : NaN in Keras while performing regression
我正在尝试预测一个连续值(第一次使用神经网络)。我已经标准化了输入数据。我不明白为什么我会收到 loss: nan从第一个纪元开始的输出。我阅读并尝试了以前对同一问题的回答中的许多建议，但没有一个对
python - 多层感知器异或，绘制误差(损失)图，收敛太快？
我目前正在学习神经网络，并尝试训练 MLP 以使用 Python 中的反向传播来学习 XOR。该网络有两个隐藏层(使用 Sigmoid 激活)和一个输出层(也是 Sigmoid)。网络(大约 20,
deep-learning - keras:平滑 L1 损失
尝试在 keras 中自定义损失函数(平滑 L1 损失)，如下所示 ValueError: Shape must be rank 0 but is rank 5 for 'cond/Switch' (
tensorflow - 使用卷积网络运行时，tensorflow 会产生 nan 损失
我试图在 tensorflow 中为门牌号图像创建一个卷积神经网络 http://ufldl.stanford.edu/housenumbers/ 当我运行我的代码时，我在第一步中得到了 nan 的成
neural-network - Keras - 变分自编码器 NaN 损失
我正在尝试使用我在 Keras 示例( https://github.com/keras-team/keras/blob/master/examples/variational_autoencoder
python - 了解 Keras 中语音识别的 CTC 损失
我试图了解 CTC 损失如何用于语音识别以及如何在 Keras 中实现它。我认为我理解的内容(如果我错了，请纠正我!)总体而言，CTC 损失被添加到经典网络之上，以便逐个元素(对于文本或语音而言逐个

首页

博学

6Ren·AI

商城

scikit-learn - class_weight 在 linearSVC 和 LogisticRegression 损失函数中的作用