gpt4 book ai didi

python - sklearn 管道适合 : AttributeError: lower not found

转载 作者:行者123 更新时间:2023-11-28 21:44:33 26 4
gpt4 key购买 nike

我想在 sklearn 中使用管道,如下所示:

corpus = load_files('corpus/train')

stop_words = [x for x in open('stopwords.txt', 'r').read().split('\n')] # Uppercase!

countvec = CountVectorizer(stop_words=stop_words, ngram_range=(1, 2))

X_train, X_test, y_train, y_test = train_test_split(corpus.data, corpus.target, test_size=0.9,
random_state=0)
x_train_counts = countvec.fit_transform(X_train)
x_test_counts = countvec.transform(X_test)

k_fold = KFold(n=len(corpus.data), n_folds=6)
confusion = np.array([[0, 0], [0, 0]])

pipeline = Pipeline([
('vectorizer', CountVectorizer(stop_words=stop_words, ngram_range=(1, 2))),
('classifier', MultinomialNB()) ])

for train_indices, test_indices in k_fold:

pipeline.fit(x_train_counts, y_train)
predictions = pipeline.predict(x_test_counts)

但是,我收到了这个错误:

AttributeError: lower not found

我看过这篇文章:

AttributeError: lower not found; using a Pipeline with a CountVectorizer in scikit-learn

但我将一个字节列表传递给矢量化器,所以这应该不是问题。

编辑

corpus = load_files('corpus')

stop_words = [x for x in open('stopwords.txt', 'r').read().split('\n')]

X_train, X_test, y_train, y_test = train_test_split(corpus.data, corpus.target, test_size=0.5,
random_state=0)

k_fold = KFold(n=len(corpus.data), n_folds=6)
confusion = np.array([[0, 0], [0, 0]])

pipeline = Pipeline([
('vectorizer', CountVectorizer(stop_words=stop_words, ngram_range=(1, 2))),
('classifier', MultinomialNB())])

for train_indices, test_indices in k_fold:
pipeline.fit(X_train[train_indices], y_train[train_indices])
predictions = pipeline.predict(X_test[test_indices])

现在我得到错误:

TypeError: only integer arrays with one element can be converted to an index

第二次编辑

corpus = load_files('corpus')

stop_words = [y for x in open('stopwords.txt', 'r').read().split('\n') for y in (x, x.title())]

k_fold = KFold(n=len(corpus.data), n_folds=6)
confusion = np.array([[0, 0], [0, 0]])

pipeline = Pipeline([
('vectorizer', CountVectorizer(stop_words=stop_words, ngram_range=(1, 2))),
('classifier', MultinomialNB())])

for train_indices, test_indices in k_fold:
pipeline.fit(corpus.data, corpus.target)

最佳答案

您没有正确使用管道。您不需要传递矢量化数据,想法是管道对数据进行矢量化。

# This is done by the pipeline
# x_train_counts = countvec.fit_transform(X_train)
# x_test_counts = countvec.transform(X_test)

k_fold = KFold(n=len(corpus.data), n_folds=6)
confusion = np.array([[0, 0], [0, 0]])

pipeline = Pipeline([
('vectorizer', CountVectorizer(stop_words=stop_words, ngram_range=(1, 2))),
('classifier', MultinomialNB()) ])

# also you are not using the indices...
for train_indices, test_indices in k_fold:

pipeline.fit(corpus.data[train_indices], corpus.target[train_indices])
predictions = pipeline.predict(corpus.data[test_indices])

关于python - sklearn 管道适合 : AttributeError: lower not found,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/40534082/

26 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com