python-3.x - 在 Keras 中使用单热编码创建模型-6ren

python-3.x - 在 Keras 中使用单热编码创建模型

转载作者：行者123 更新时间：2023-12-04 21:04:35

我正在研究句子分类问题并尝试使用 Keras 解决。
词汇表中的唯一单词总数为 36。

在这种情况下，总词汇是 [W1,W2,W3....W36]

所以，如果我有一个单词为 [W1 W2 W6 W7 W9] 的句子，如果我对其进行编码，我会得到一个 numpy 数组，如下所示

[[0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1]
 [0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0]
 [0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]
 [0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]]

形状是 (5,36)

我被困在这里。我已经生成了 20000 个形状各异的 numpy 数组，即 (N,36)
其中 N 是句子中的单词数。所以，我有 20,000 个用于训练的句子和 100 个用于测试的句子，所有句子都标有 (1,36) 单热编码

我有 x_train、x_test、y_train 和 y_test

x_test 和 y_test 是维度 (1,36)

任何人都可以请建议我该怎么做？

我做了一些下面的编码

model = Sequential()
model.add(Dense(512, input_shape=(??????))),
model.add(Activation('relu'))
model.add(Dropout(0.5))
model.add(Dense(num_classes))
model.add(Activation('softmax'))
model.compile(loss='categorical_crossentropy',
          optimizer='adam',
          metrics=['accuracy'])

任何帮助将非常感激。

更新和回应@putonspectacles

非常感谢您花费时间和精力进行详细回复。我对您的代码进行了一些小的修改，我认为需要完成这些修改才能使代码正常工作。请在下面找到它

num_classes = 5 
max_words = 20
sentences = ["The cat is in the house","The green boy","computer programs are not alive while the children are"]
labels = np.random.randint(0, num_classes, 3)
y = to_categorical(labels, num_classes=num_classes)
words = set(w for sent in sentences for w in sent.split())
word_map = {w : i+1 for (i, w) in enumerate(words)}
#-Changed the below line the inner for loop sent to sent.split()  
sent_ints = [[word_map[w] for w in sent.split()] for sent in sentences]
vocab_size = len(words)
print(vocab_size)
#-changed the below line - the outer for loop sentences to sent_ints
X = np.array([to_categorical(pad_sequences((sent,), max_words),vocab_size+1)  for sent in sent_ints])
print(X)
print(y)
model = Sequential()
model.add(Dense(512, input_shape=(max_words, vocab_size + 1)))
model.add(LSTM(128))
model.add(Dense(5, activation='softmax'))
model.compile(loss='categorical_crossentropy',
      optimizer='adam',
      metrics=['accuracy'])
model.fit(X,y)

如果没有这些更改，代码将无法工作。当我运行上面的代码时，我得到了如下所示的正确嵌入

[[[[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1.]
[0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0.]
[0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0.]
[0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0.]]]


[[[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
[0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0.]]]


 [[[1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0.]
 [0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0.]
 [0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]]]]



[[0. 0. 0. 0. 1.]
[1. 0. 0. 0. 0.]
[0. 1. 0. 0. 0.]]

但我得到的错误是“ 检查输入时出错:预期dense_44_input 有3 个维度，但得到了形状为(3, 1, 20, 16) 的数组”

当我将输入形状更改为
模型.添加(密集(512，input_shape =(无，max_words，vocab_size + 1)))

我收到错误“ 输入 0 与 lstm_27 层不兼容:预期 ndim=3，发现 ndim=4 ”

我正在努力解决这个问题。如果你能给我一个方向，那就太好了。

我接受了答案，因为它回答了嵌入单词的目标。再次感谢。

最佳答案

酷，你清理了问题。你想对一个句子进行分类。我假设你说我想要比词袋编码做得更好。您想重视序列。

我们将选择一个新模型然后- an RNN (the LSTM version) .该模型有效地总结了每个单词(按顺序)的重要性，因为它构建了最适合任务的句子表示。

但是我们将不得不以不同的方式处理预处理。为了效率(以便我们可以批量处理更多句子而不是一次处理单个句子)，我们希望所有句子都具有相同数量的单词。所以我们选择一个 max_words，比如 20，我们填充较短的句子以达到最大单词，然后我们将长度超过 20 个单词的句子删减。

Keras 将为此提供帮助。我们将用整数对每个单词进行编码。

from keras.preprocessing.sequence import pad_sequences
from keras.utils import to_categorical
from keras.models import Sequential
from keras.layers import Embedding, Dense, LSTM

num_classes = 5 
max_words = 20
sentences = ["The cat is in the house",
                           "The green boy",
            "computer programs are not alive while the children are"]
labels = np.random.randint(0, num_classes, 3)
y = to_categorical(labels, num_classes=num_classes)

words = set(w for sent in sentences for w in sent.split())
word_map = {w : i+1 for (i, w) in enumerate(words)}
sent_ints = [[word_map[w] for w in sent] for sent in sentences]
vocab_size = len(words)

所以“绿色男孩”现在可能是 [1, 3, 5]。
然后我们将填充和单热编码

# pad to max_words length and encode with len(words) + 1  
# + 1 because we'll reserve 0 add the padding sentinel.
X = np.array([to_categorical(pad_sequences((sent,), max_words),  
       vocab_size + 1)  for sent in sent_ints])
print(X.shape) # (3, 20, 16)

现在到模型:我们将添加一个 Dense层转换那些热
词到密集向量。然后我们使用 LSTM转换词向量
在句子中到一个密集的句子向量。最后，我们将使用 softmax 激活来生成类的概率分布。

model = Sequential()
model.add(Dense(512, input_shape=(max_words, vocab_size + 1)))
model.add(LSTM(128))
model.add(Dense(5, activation='softmax'))
model.compile(loss='categorical_crossentropy',
          optimizer='adam',
          metrics=['accuracy'])

那应该是完整的。然后你可以继续训练。

model.fit(X,y)

编辑:

这一行:

# we need to split the sentences in a words write now it reading every
# letter notice the sent.split() in the correct version below.
sent_ints = [[word_map[w] for w in sent] for sent in sentences]

应该:

sent_ints = [[word_map[w] for w in sent.split()] for sent in sentences]

关于python-3.x - 在 Keras 中使用单热编码创建模型，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/49604765/

文章推荐： sql - 行间相减

文章推荐： groovy - SoapUI LoadTest 执行失败

文章推荐： oop - 方法调用作为另一个方法调用的参数？

css - 网站的自定义 CSS 编码/编码 block
我对自定义 CSS 或在将图像作为 Logo 上传到页面时使用编码 block 有疑问。我正在为我的网站使用 squarespace，我需要帮助编码我的 Logo 以使其适合每个页面。一个选项是使用自
Golang 编码/json 编码(marshal)拆收器
如 encoding/json 包文档中所述， Marshal traverses the value v recursively. If an encountered value implement
Java 编码 - 相当于 Java 中的 sjisMS 编码
我必须做一些相当于Java中的iconv -f utf8 -t sjisMS $INPUT_FILE的事情。该命令在 Unix 中我在java中没有找到任何带有sjisMS的编码。 Java中有Sh
PHP 5.6 编码 latin1 MySQL 编码
从 PHP 5.3 迁移到 PHP 5.6 后，我遇到了编码问题。我的 MySQL 数据库是 latin1，我的 PHP 文件是 windows-1251。现在一切都显示为“ñëåäíèòå àäðå
r - 文件错误(文件名， "r"，编码=编码): cannot open the connection
我有一个 RScript文件(我们称之为 main.r )，它引用了另一个文件，使用以下代码: source("functions.R") 但是，当我运行 RScript 文件时，它提示以下错误:
java - 处理 RPC/编码 Web 服务中的 SOAP 编码
我无法设法从 WSDL 创建 RPC/编码风格的代码 - 有谁知道哪个框架可以做到这一点？带有 adb 和 xmlbeans 映射的 Axis2 无法正常工作(无法处理响应中的肥皂编码)直接使用 X
Node.Js Express-Generator 项目生成器错误(编码 && 编码.toLowerCase()
安装了最新版本的Node.Js()和npm包**(1.2.10)**当我运行 Express 命令来生成项目时，它向我抛出以下错误 buffer.js:240 switch (encoding &
javascript - JavaScript 中的 JSON 编码/解码 base64 编码/解码
JavaScript中有JSON编码/解码base64编码/解码函数吗？最佳答案是的，btoa() 和 atob() 在某些浏览器中可以工作: var enc = btoa("this is so
python - 为什么有些字符串采用 utf-16 编码，而另一些字符串仅采用 utf-8 编码？
>>> unicode('восстановление информации', 'utf-16') Traceback (most recent call last): File "", line
html - 是否有一个 JDK 类来进行 HTML 编码(但不是 URL 编码)？
我当然熟悉 java.net.URLEncoder 和 java.net.URLDecoder 类。但是，我只需要 HTML 样式的编码。 (我不想将 ' ' 替换为 '+' 等)。我不知道任何只做
utf-8 - SSIS - 平面文件始终采用 ANSI 编码，从不采用 UTF-8 编码
有一个非常简单的 SSIS 包: OLE DB Source 通过 View 获取数据(数据库表 nvarchar 或 nchar 中的所有字符串列)。派生列，用于格式化现有日期并将其添加到数据集(
node.js - golang base64 编码 vs nodejs 缓冲区 base64 编码
我正在使用一个在 Node 中进行base64编码的软件，如下所示: const enc = new Buffer('test', 'base64') console.log(enc) 显示: 我正
【编码】如何实现一套自定义网络协议？
前言下文介绍的自定义协议仅作为学习示例，纯粹是玩具项目，没有实际可用性。无需过度关注和讨论其合理性进行通信的双方是谁？常见的模型客户端-服务器，例如HTTP协议，浏览器<=>
hibernate 编码
我试图将带有日语字符的数据插入到 oracle 数据库中。事情是保存在数据库中的是一堆倒置的问号。我该如何解决这个问题最佳答案见 http://www.errcode.net/blogs/?p=6
Java解压奇怪的字符(编码？)
当我在 java 中解压 zip 文件时，我发现文件名中出现了带有重音字符的奇怪行为。西索: Add File user : L'equipe Technique -- Folder : spec
JavaScript 编码
在网上冲浪我找到了 ExtJS 的 Ext.Gantt 插件，该扩展有一个特殊的编码。任何人都知道如何编码那样或其他复杂的形式。 Encoded Gantt Chart 最佳答案它似乎被 Dean
编码，将一个整数写入文件中
我正在用C语言做一个编码任务，我进展顺利，直到读取符号并根据表格分配相应的代码的部分。我必须连接几个代码，直到它们的长度达到 32 位，为此我必须将它们写入一个文件中。这种写入文件的方法给我带来了很多
Javascript 编码
我有一个外部链接的 javascript 文件。在那个 javascript 里面，我有这个功能: function getMonthNumber(monthName){ monthName = mo
python 编码
使用mechanize，我检索到一个网页的源页面，其中包含一些非ASCII字符，比如汉字。代码如下: #using python2.6 from mechanize import Browser b
读取文件时的C#编码
我有一个包含字母 ø 的文件。当我用这段代码 File.ReadLines(filePath) 读取它时，我得到了一个问号而不是它。当我像这样添加编码时 File.ReadLines(filePat

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python-3.x - 在 Keras 中使用单热编码创建模型