python - 在标签不在训练集中的测试数据上使用 MultilabelBinarizer-6ren

python - 在标签不在训练集中的测试数据上使用 MultilabelBinarizer

转载作者：太空狗更新时间：2023-10-29 17:46:38

给定这个简单的多标签分类示例(取自这个问题，use scikit-learn to classify into multiple categories)

import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.svm import LinearSVC
from sklearn.feature_extraction.text import TfidfTransformer
from sklearn.multiclass import OneVsRestClassifier
from sklearn import preprocessing
from sklearn.metrics import accuracy_score

X_train = np.array(["new york is a hell of a town",
                "new york was originally dutch",
                "the big apple is great",
                "new york is also called the big apple",
                "nyc is nice",
                "people abbreviate new york city as nyc",
                "the capital of great britain is london",
                "london is in the uk",
                "london is in england",
                "london is in great britain",
                "it rains a lot in london",
                "london hosts the british museum",
                "new york is great and so is london",
                "i like london better than new york"])
y_train_text = [["new york"],["new york"],["new york"],["new york"],    ["new york"],
            ["new york"],["london"],["london"],["london"],["london"],
            ["london"],["london"],["new york","london"],["new york","london"]]

X_test = np.array(['nice day in nyc',
               'welcome to london',
               'london is rainy',
               'it is raining in britian',
               'it is raining in britian and the big apple',
               'it is raining in britian and nyc',
               'hello welcome to new york. enjoy it here and london too'])

y_test_text = [["new york"],["london"],["london"],["london"],["new york", "london"],["new york", "london"],["new york", "london"]]


lb = preprocessing.MultiLabelBinarizer()
Y = lb.fit_transform(y_train_text)
Y_test = lb.fit_transform(y_test_text)

classifier = Pipeline([
('vectorizer', CountVectorizer()),
('tfidf', TfidfTransformer()),
('clf', OneVsRestClassifier(LinearSVC()))])

classifier.fit(X_train, Y)
predicted = classifier.predict(X_test)


print "Accuracy Score: ",accuracy_score(Y_test, predicted)

代码运行良好，并打印准确度分数，但是如果我将 y_test_text 更改为

y_test_text = [["new york"],["london"],["england"],["london"],["new york", "london"],["new york", "london"],["new york", "london"]]

我明白了

Traceback (most recent call last):
  File "/Users/scottstewart/Documents/scikittest/example.py", line 52, in <module>
     print "Accuracy Score: ",accuracy_score(Y_test, predicted)
  File "/Library/Python/2.7/site-packages/sklearn/metrics/classification.py", line 181, in accuracy_score
differing_labels = count_nonzero(y_true - y_pred, axis=1)
File "/System/Library/Frameworks/Python.framework/Versions/2.7/Extras/lib/python/scipy/sparse/compressed.py", line 393, in __sub__
raise ValueError("inconsistent shapes")
ValueError: inconsistent shapes

请注意引入了不在训练集中的“england”标签。我如何使用多标签分类，以便在引入“测试”标签时，我仍然可以运行一些指标？或者这甚至可能吗？

编辑:感谢大家的回答，我想我的问题更多是关于 scikit 二值化器如何工作或应该如何工作。鉴于我的简短示例代码，如果我将 y_test_text 更改为

y_test_text = [["new york"],["new york"],["new york"],["new york"],["new york"],["new york"],["new york"]]

它会起作用——我的意思是我们已经适应了那个标签，但在这种情况下我明白了

ValueError: Can't handle mix of binary and multilabel-indicator

最佳答案

如果您也在训练 y 集中“引入”新标签，则可以，如下所示:

import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.svm import LinearSVC
from sklearn.feature_extraction.text import TfidfTransformer
from sklearn.multiclass import OneVsRestClassifier
from sklearn import preprocessing
from sklearn.metrics import accuracy_score

X_train = np.array(["new york is a hell of a town",
                "new york was originally dutch",
                "the big apple is great",
                "new york is also called the big apple",
                "nyc is nice",
                "people abbreviate new york city as nyc",
                "the capital of great britain is london",
                "london is in the uk",
                "london is in england",
                "london is in great britain",
                "it rains a lot in london",
                "london hosts the british museum",
                "new york is great and so is london",
                "i like london better than new york"])
y_train_text = [["new york"],["new york"],["new york"],["new york"],    
                ["new york"],["new york"],["london"],["london"],         
                ["london"],["london"],["london"],["london"],
                ["new york","England"],["new york","london"]]

X_test = np.array(['nice day in nyc',
               'welcome to london',
               'london is rainy',
               'it is raining in britian',
               'it is raining in britian and the big apple',
               'it is raining in britian and nyc',
               'hello welcome to new york. enjoy it here and london too'])

y_test_text = [["new york"],["new york"],["new york"],["new york"],["new york"],["new york"],["new york"]]


lb = preprocessing.MultiLabelBinarizer(classes=("new york","london","England"))
Y = lb.fit_transform(y_train_text)
Y_test = lb.fit_transform(y_test_text)

print Y_test

classifier = Pipeline([
('vectorizer', CountVectorizer()),
('tfidf', TfidfTransformer()),
('clf', OneVsRestClassifier(LinearSVC()))])

classifier.fit(X_train, Y)
predicted = classifier.predict(X_test)
print predicted

print "Accuracy Score: ",accuracy_score(Y_test, predicted)

输出:

Accuracy Score:  0.571428571429

关键部分是:

y_train_text = [["new york"],["new york"],["new york"],
                ["new york"],["new york"],["new york"],
                ["london"],["london"],["london"],["london"],
                ["london"],["london"],["new york","England"],
                ["new york","london"]]

我们也插入了“England”。这是有道理的，因为如果分类器以前没有看到它，其他方式如何预测分类器？所以我们以这种方式创建了一个三标签分类问题。

已编辑:

lb = preprocessing.MultiLabelBinarizer(classes=("new york","london","England"))

您必须将类作为 arg 传递给 MultiLabelBinarizer()，它可以与任何 y_test_text 一起使用。

关于python - 在标签不在训练集中的测试数据上使用 MultilabelBinarizer，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/31503874/

文章推荐： python - 编译并上传到pypicloud服务器后运行python包

文章推荐： python - 谁能给我解释一下 numpy.indices()？

文章推荐： c# - Silverlight 5 VS 2012 单元测试

Opencv 训练
real adaboost Logit boost discrete adaboost 和 gentle adaboost in train cascade parameter 有什么区别.. -bt
python - 训练/测试矩阵图书交叉推荐系统
我想为 book crossing 构建训练数据矩阵和测试数据矩阵数据集。但作为 ISBN 代码的图书 ID 可能包含字符。因此，我无法应用此代码(来自 tutorial ): #Create two
针对不同格式车牌的 JavaANPR 训练
我找到了 JavaANPR 库，我想对其进行自定义以读取我所在国家/地区的车牌。似乎包含的字母表与我们使用的字母表不同 ( http://en.wikipedia.org/wiki/FE-Schri
machine-learning - 训练/测试拆分之前或之后的欠采样
我有一个信用卡数据集，其中 98% 的交易是非欺诈交易，2% 是欺诈交易。我一直在尝试在训练和测试拆分之前对多数类别进行欠采样，并在测试集上获得非常好的召回率和精度。当我仅在训练集上进行欠采样并在
python - Keras NASNet 训练
我打算: 在数据集上从头开始训练 NASNet 只重新训练 NASNet 的最后一层(迁移学习) 并比较它们的相对性能。从文档中我看到: keras.applications.nasnet.NASNe
python - 训练 uNet 模型预测只有黑色
我正在训练用于分割的 uNet 模型。训练模型后，输出全为零，我不明白为什么。我看到建议我应该使用特定的损失函数，所以我使用了 dice 损失函数。这是因为黑色区域 (0) 比白色区域 (1) 大得
bash - Tesseract 训练 - 微调角色
我想为新角色训练我现有的 tesseract 模型。我已经尝试过上的教程 https://github.com/tesseract-ocr/tesseract/wiki/TrainingTesser
python - 如何执行多个 NN 训练？
我的机器中有两个 NVidia GPU，但我没有使用它们。我的机器上运行了三个神经网络训练。当我尝试运行第四个时，脚本出现以下错误: my_user@my_machine:~/my_project/
python - 具有稀疏数据的 tensorflow 训练
我想在python的tensorflow中使用稀疏张量进行训练。我找到了很多代码如何做到这一点，但没有一个有效。这里有一个示例代码来说明我的意思，它会抛出一个错误: import numpy as
python - 训练 MSE 损失大于理论最大值？
我正在训练一个 keras 模型，它的最后一层是单个 sigmoid单元: output = Dense(units=1, activation='sigmoid') 我正在用一些训练数据训练这个模型
python - 训练 Keras 模型会产生多个优化器错误
所以我需要使用我自己的数据集重新训练 Tiny YOLO。我正在使用的模型可以在这里找到:keras-yolo3 . 我开始训练并遇到多个优化器错误，添加了错误代码以防止混淆。我注意到即使它应该使用
nlp - 使用字符嵌入进行 BERT 训练
将 BERT 模型中的标记化范式更改为其他东西是否有意义？也许只是一个简单的单词标记化或字符级标记化？最佳答案这是论文“CharacterBERT: Reconciling ELMo and BE
neural-network - TensorFlow 训练
假设我有一个非常简单的神经网络，比如多层感知器。对于每一层，激活函数都是 sigmoid 并且网络是全连接的。在 TensorFlow 中，这可能是这样定义的: sess = tf.Inte
pybrain - 如何保存和恢复 PyBrain 训练？
有没有办法在 PyBrain 中保存和恢复经过训练的神经网络，这样我每次运行脚本时都不必重新训练它？最佳答案 PyBrain 的神经网络可以使用 python 内置的 pickle/cPickle
python - 训练 CNN 后准确率较低
我尝试使用 Keras 训练一个对手写数字进行分类的 CNN 模型，但训练的准确度很低(低于 10%)并且误差很大。我尝试了一个简单的神经网络，但没有效果。这是我的代码。 import tensor
ocr - 训练 tesseract 时的正确间距
我在 Windows 7 64 位上使用 tesseract 3.0.1。我用一种新语言训练图书馆。我的示例数据间隔非常好。当我为每个角色的盒子定义坐标时，盒子紧贴角色有多重要？我使用其中一个插件，
neural-network - dropout 训练
如何对由 dropout 产生的许多变薄层进行平均？在测试阶段要使用哪些权重？我真的很困惑这个。因为每个变薄的层都会学习一组不同的权重。那么反向传播是为每个细化网络单独完成的吗？这些细化网络之间的权重
java - 训练 Tesseract - 加载训练语言失败
我尝试训练超正方语言。我正在使用 Tess4J 进行 OCR 处理。我使用jTessBoxEditor和SerakTesseractTrainer进行训练操作。准备好训练数据后，我将其放在 Tesse
python - 训练 Keras 模型时使用稀疏数组表示标签
我正在构建一个 Keras 模型，将数据分类为 3000 个不同的类别，我的训练数据由大量样本组成，因此在用一种热编码对训练输出进行编码后，数据非常大(item_count * 3000 * 的大小)
python - 训练 pyBrain 需要多长时间？
关闭。这个问题需要多问focused 。目前不接受答案。想要改进此问题吗？更新问题，使其仅关注一个问题 editing this post . 已关闭 8 年前。 Improve this ques

太空狗

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 在标签不在训练集中的测试数据上使用 MultilabelBinarizer