python-3.x - 如何正确实现 MNIST 数据集机器学习的反向传播？-6ren

python-3.x - 如何正确实现 MNIST 数据集机器学习的反向传播？

转载作者：行者123 更新时间：2023-11-30 08:32:15

所以，我使用 Michael Nielson 的机器学习书籍作为我的代码的引用(它基本上是相同的):http://neuralnetworksanddeeplearning.com/chap1.html

有问题的代码:

    def backpropagate(self, image, image_value) :


        # declare two new numpy arrays for the updated weights & biases
        new_biases = [np.zeros(bias.shape) for bias in self.biases]
        new_weights = [np.zeros(weight_matrix.shape) for weight_matrix in self.weights]

        # -------- feed forward --------
        # store all the activations in a list
        activations = [image]

        # declare empty list that will contain all the z vectors
        zs = []
        for bias, weight in zip(self.biases, self.weights) :
            print(bias.shape)
            print(weight.shape)
            print(image.shape)
            z = np.dot(weight, image) + bias
            zs.append(z)
            activation = sigmoid(z)
            activations.append(activation)

        # -------- backward pass --------
        # transpose() returns the numpy array with the rows as columns and columns as rows
        delta = self.cost_derivative(activations[-1], image_value) * sigmoid_prime(zs[-1])
        new_biases[-1] = delta
        new_weights[-1] = np.dot(delta, activations[-2].transpose())

        # l = 1 means the last layer of neurons, l = 2 is the second-last, etc.
        # this takes advantage of Python's ability to use negative indices in lists
        for l in range(2, self.num_layers) :
            z = zs[-1]
            sp = sigmoid_prime(z)
            delta = np.dot(self.weights[-l+1].transpose(), delta) * sp
            new_biases[-l] = delta
            new_weights[-l] = np.dot(delta, activations[-l-1].transpose())
        return (new_biases, new_weights)

我的算法只能在出现此错误之前到达第一轮反向传播:

  File "D:/Programming/Python/DPUDS/DPUDS_Projects/Fall_2017/MNIST/network.py", line 97, in stochastic_gradient_descent
    self.update_mini_batch(mini_batch, learning_rate)
  File "D:/Programming/Python/DPUDS/DPUDS_Projects/Fall_2017/MNIST/network.py", line 117, in update_mini_batch
    delta_biases, delta_weights = self.backpropagate(image, image_value)
  File "D:/Programming/Python/DPUDS/DPUDS_Projects/Fall_2017/MNIST/network.py", line 160, in backpropagate
    z = np.dot(weight, activation) + bias
ValueError: shapes (30,50000) and (784,1) not aligned: 50000 (dim 1) != 784 (dim 0)

我明白为什么这是一个错误。权重中的列数与像素图像中的行数不匹配，因此我无法进行矩阵乘法。这就是我感到困惑的地方——反向传播中使用了 30 个神经元，每个神经元评估 50,000 张图像。我的理解是，50,000 个像素中的每一个都应该附加 784 个权重，每个像素一个。但是当我相应地修改代码时:

        count = 0
        for bias, weight in zip(self.biases, self.weights) :
            print(bias.shape)
            print(weight[count].shape)
            print(image.shape)
            z = np.dot(weight[count], image) + bias
            zs.append(z)
            activation = sigmoid(z)
            activations.append(activation)
            count += 1

我仍然遇到类似的错误:

ValueError: shapes (50000,) and (784,1) not aligned: 50000 (dim 0) != 784 (dim 0)

我对所有涉及的线性代数感到非常困惑，我想我只是错过了一些关于权重矩阵结构的东西。任何帮助将不胜感激。

最佳答案

看起来问题出在您对原始代码的更改中。

我从您提供的链接下载了示例，它工作正常，没有任何错误:

这是我使用的完整源代码:

import cPickle
import gzip
import numpy as np
import random

def load_data():
    """Return the MNIST data as a tuple containing the training data,
    the validation data, and the test data.
    The ``training_data`` is returned as a tuple with two entries.
    The first entry contains the actual training images.  This is a
    numpy ndarray with 50,000 entries.  Each entry is, in turn, a
    numpy ndarray with 784 values, representing the 28 * 28 = 784
    pixels in a single MNIST image.
    The second entry in the ``training_data`` tuple is a numpy ndarray
    containing 50,000 entries.  Those entries are just the digit
    values (0...9) for the corresponding images contained in the first
    entry of the tuple.
    The ``validation_data`` and ``test_data`` are similar, except
    each contains only 10,000 images.
    This is a nice data format, but for use in neural networks it's
    helpful to modify the format of the ``training_data`` a little.
    That's done in the wrapper function ``load_data_wrapper()``, see
    below.
    """
    f = gzip.open('../data/mnist.pkl.gz', 'rb')
    training_data, validation_data, test_data = cPickle.load(f)
    f.close()
    return (training_data, validation_data, test_data)

def load_data_wrapper():
    """Return a tuple containing ``(training_data, validation_data,
    test_data)``. Based on ``load_data``, but the format is more
    convenient for use in our implementation of neural networks.
    In particular, ``training_data`` is a list containing 50,000
    2-tuples ``(x, y)``.  ``x`` is a 784-dimensional numpy.ndarray
    containing the input image.  ``y`` is a 10-dimensional
    numpy.ndarray representing the unit vector corresponding to the
    correct digit for ``x``.
    ``validation_data`` and ``test_data`` are lists containing 10,000
    2-tuples ``(x, y)``.  In each case, ``x`` is a 784-dimensional
    numpy.ndarry containing the input image, and ``y`` is the
    corresponding classification, i.e., the digit values (integers)
    corresponding to ``x``.
    Obviously, this means we're using slightly different formats for
    the training data and the validation / test data.  These formats
    turn out to be the most convenient for use in our neural network
    code."""
    tr_d, va_d, te_d = load_data()
    training_inputs = [np.reshape(x, (784, 1)) for x in tr_d[0]]
    training_results = [vectorized_result(y) for y in tr_d[1]]
    training_data = zip(training_inputs, training_results)
    validation_inputs = [np.reshape(x, (784, 1)) for x in va_d[0]]
    validation_data = zip(validation_inputs, va_d[1])
    test_inputs = [np.reshape(x, (784, 1)) for x in te_d[0]]
    test_data = zip(test_inputs, te_d[1])
    return (training_data, validation_data, test_data)

def vectorized_result(j):
    """Return a 10-dimensional unit vector with a 1.0 in the jth
    position and zeroes elsewhere.  This is used to convert a digit
    (0...9) into a corresponding desired output from the neural
    network."""
    e = np.zeros((10, 1))
    e[j] = 1.0
    return e

class Network(object):

    def __init__(self, sizes):
        """The list ``sizes`` contains the number of neurons in the
        respective layers of the network.  For example, if the list
        was [2, 3, 1] then it would be a three-layer network, with the
        first layer containing 2 neurons, the second layer 3 neurons,
        and the third layer 1 neuron.  The biases and weights for the
        network are initialized randomly, using a Gaussian
        distribution with mean 0, and variance 1.  Note that the first
        layer is assumed to be an input layer, and by convention we
        won't set any biases for those neurons, since biases are only
        ever used in computing the outputs from later layers."""
        self.num_layers = len(sizes)
        self.sizes = sizes
        self.biases = [np.random.randn(y, 1) for y in sizes[1:]]
        self.weights = [np.random.randn(y, x)
                        for x, y in zip(sizes[:-1], sizes[1:])]

    def feedforward(self, a):
        """Return the output of the network if ``a`` is input."""
        for b, w in zip(self.biases, self.weights):
            a = sigmoid(np.dot(w, a)+b)
        return a

    def SGD(self, training_data, epochs, mini_batch_size, eta,
            test_data=None):
        """Train the neural network using mini-batch stochastic
        gradient descent.  The ``training_data`` is a list of tuples
        ``(x, y)`` representing the training inputs and the desired
        outputs.  The other non-optional parameters are
        self-explanatory.  If ``test_data`` is provided then the
        network will be evaluated against the test data after each
        epoch, and partial progress printed out.  This is useful for
        tracking progress, but slows things down substantially."""
        if test_data: n_test = len(test_data)
        n = len(training_data)
        for j in xrange(epochs):
            random.shuffle(training_data)
            mini_batches = [
                training_data[k:k+mini_batch_size]
                for k in xrange(0, n, mini_batch_size)]
            for mini_batch in mini_batches:
                self.update_mini_batch(mini_batch, eta)
            if test_data:
                print "Epoch {0}: {1} / {2}".format(
                    j, self.evaluate(test_data), n_test)
            else:
                print "Epoch {0} complete".format(j)

    def update_mini_batch(self, mini_batch, eta):
        """Update the network's weights and biases by applying
        gradient descent using backpropagation to a single mini batch.
        The ``mini_batch`` is a list of tuples ``(x, y)``, and ``eta``
        is the learning rate."""
        nabla_b = [np.zeros(b.shape) for b in self.biases]
        nabla_w = [np.zeros(w.shape) for w in self.weights]
        for x, y in mini_batch:
            delta_nabla_b, delta_nabla_w = self.backprop(x, y)
            nabla_b = [nb+dnb for nb, dnb in zip(nabla_b, delta_nabla_b)]
            nabla_w = [nw+dnw for nw, dnw in zip(nabla_w, delta_nabla_w)]
        self.weights = [w-(eta/len(mini_batch))*nw
                        for w, nw in zip(self.weights, nabla_w)]
        self.biases = [b-(eta/len(mini_batch))*nb
                       for b, nb in zip(self.biases, nabla_b)]

    def backprop(self, x, y):
        """Return a tuple ``(nabla_b, nabla_w)`` representing the
        gradient for the cost function C_x.  ``nabla_b`` and
        ``nabla_w`` are layer-by-layer lists of numpy arrays, similar
        to ``self.biases`` and ``self.weights``."""
        nabla_b = [np.zeros(b.shape) for b in self.biases]
        nabla_w = [np.zeros(w.shape) for w in self.weights]
        # feedforward
        activation = x
        activations = [x] # list to store all the activations, layer by layer
        zs = [] # list to store all the z vectors, layer by layer
        for b, w in zip(self.biases, self.weights):
            z = np.dot(w, activation)+b
            zs.append(z)
            activation = sigmoid(z)
            activations.append(activation)
        # backward pass
        delta = self.cost_derivative(activations[-1], y) * \
            sigmoid_prime(zs[-1])
        nabla_b[-1] = delta
        nabla_w[-1] = np.dot(delta, activations[-2].transpose())
        # Note that the variable l in the loop below is used a little
        # differently to the notation in Chapter 2 of the book.  Here,
        # l = 1 means the last layer of neurons, l = 2 is the
        # second-last layer, and so on.  It's a renumbering of the
        # scheme in the book, used here to take advantage of the fact
        # that Python can use negative indices in lists.
        for l in xrange(2, self.num_layers):
            z = zs[-l]
            sp = sigmoid_prime(z)
            delta = np.dot(self.weights[-l+1].transpose(), delta) * sp
            nabla_b[-l] = delta
            nabla_w[-l] = np.dot(delta, activations[-l-1].transpose())
        return (nabla_b, nabla_w)

    def evaluate(self, test_data):
        """Return the number of test inputs for which the neural
        network outputs the correct result. Note that the neural
        network's output is assumed to be the index of whichever
        neuron in the final layer has the highest activation."""
        test_results = [(np.argmax(self.feedforward(x)), y)
                        for (x, y) in test_data]
        return sum(int(x == y) for (x, y) in test_results)

    def cost_derivative(self, output_activations, y):
        """Return the vector of partial derivatives \partial C_x /
        \partial a for the output activations."""
        return (output_activations-y)

#### Miscellaneous functions
def sigmoid(z):
    """The sigmoid function."""
    return 1.0/(1.0+np.exp(-z))

def sigmoid_prime(z):
    """Derivative of the sigmoid function."""
    return sigmoid(z)*(1-sigmoid(z))

training_data, validation_data, test_data = load_data_wrapper()
net = Network([784, 30, 10])
net.SGD(training_data, 30, 10, 3.0, test_data=test_data)

其他信息:

但是，我建议使用现有框架之一，例如 Keras，以免重新发明轮子

此外，还用 python 3.6 进行了检查:

关于python-3.x - 如何正确实现 MNIST 数据集机器学习的反向传播？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/46855577/

文章推荐： WKWebView 中的 javascript 桥不起作用

文章推荐： javascript - 在 AngularJS 中将 base64 转换为图像文件

python - 学习 MNIST 后对非 MNIST 图像进行分类
我的机器学习算法已经学习了 MNIST 数据库中的 70000 张图像。我想在 MNIST 数据集中未包含的图像上对其进行测试。但是，我的预测函数无法读取我的测试图像的数组表示。如何在外部图像上测试
python - 制作自己的 MNIST 数据集(与 MNIST 格式相同)
我正在尝试创建我自己的 MNIST 数据版本。我已将训练和测试数据转换为以下文件； test-images-idx3-ubyte.gz test-labels-idx1-ubyte.gz train-
python - 无法在 Windows 上使用 python-mnist 包加载 MNIST 数据
我通过 pip 在我的 Windows 设备上安装了 python-mnist 包，正如 Github 文档中所述，方法是在我的 Anaconda 终端中输入以下命令: pip install pyt
一小时学会TensorFlow2之Fashion Mnist
描述 Fashion Mnist 是一个类似于 Mnist 的图像数据集. 涵盖 10 种类别的 7 万 (6 万训练集 + 1 万测试集) 个不同商品的图片. Tensor
tensorflow - MNIST 识别手写文字
该模型现在只能使用 tf. 识别单个字母。我怎样才能让它识别连续的字母单词？最佳答案手写数字识别。 ... MNIST 是一个广泛用于手写数字分类任务的数据集。它由 70,000 个标记为 28x
image - MNIST 图像是什么图像格式？
我已经从 MNIST 训练集中解压了第一张图像，并且可以访问 (28,28) 矩阵。 [[ 0 0 0 0 0 0 0 0 0 0 0 0 0 0
python - MNIST 数据反规范化不会给我返回相同的结果
这是我学习的一部分。我知道标准化确实有助于提高准确性，因此将 mnist 值除以 255。这会将所有像素除以 255，因此 28*28 的所有像素的值将在 0.0 到 1.0 范围内. 现在我厌倦了将
numpy - MNIST 中每个数字代表什么？
我已成功将 MNIST 数据下载到扩展名为 .npy 的文件中。当我打印第一张图像的几列时。我得到以下结果。这里每个数字代表什么？ a= np.load("training_set.npy") pri
TensorFlow - MNIST 数据中的训练准确性没有提高
我用tensorflow写了一个程序来处理Kaggle的数字识别问题。程序可以正常运行，但训练准确率总是很低，大约10%，如下: step 0, training accuracy 0.11 step
python - MNIST 数据集中的图像是如何转换的？
在 cnn_mnist.py例如，脚本首先加载训练和测试数据，如您从 120 行到 124 行中看到的那样。当我打印 print(train_data.shape) 时，我得到 (55000, 784
python - 神经网络 MNIST
我研究神经网络有一段时间了，用python和numpy做了一个实现。我用 XOR 做了一个非常简单的例子，它运行良好。所以我想我更进一步尝试 MNIST 数据库。这是我的问题。我正在使用具有 784
python - MNIST:试图获得高精度
我目前正在研究手写数字识别问题。首先，我针对 MNIST 数据集测试了示例手写数字。我的准确率为 53%，我需要 90% 以上的准确率。以下是我迄今为止为提高准确性所做的尝试。创建了我自己的数
python - 如何在我自己的数据集图像上测试 mnist
我正在尝试使用我自己的数字图像数据集测试 mnist。我为此写了一个 python 脚本，但它给出了一个错误。错误在代码的第 16 行。实际上我无法发送图像进行测试。给我一些建议。提前致谢。 imp
python - Mnist 数据图像和标签不匹配
我知道这可能是一个愚蠢的问题，但我真的不明白为什么。下面是我尝试从训练数据中打印单个图像和具有相同索引的标签的代码 import matplotlib.pyplot as plt from tenso
python - MNIST 手写数字
我尝试使用以下数据集在 python 中制作一个能够识别手写数字的脚本:http://deeplearning.net/data/mnist/mnist.pkl.gz . 关于这个问题和我试图实现的算
java - MNIST 的缩减图像
我正在尝试解决 Android 设备上的 MNIST 分类问题。我已经有一个经过训练的模型，现在我希望能够识别照片上的单个数字。拍完照片后，我会进行一些预处理，然后再将图像传递给模型。这是原始图像的
由浅入深学习TensorFlow MNIST 数据集
MNIST 数据集介绍 MNIST 包含 0~9 的手写数字, 共有 60000 个训练集和 10000 个测试集. 数据的格式为单通道 28*28 的灰度图. LeNet 模型
python - 为什么导入 mnist 数字数据集时总是漏掉一个子图？
我想导入 mnist digits 数字以在一个图中显示，并编写这样的代码， import keras from keras.datasets import mnist import matplotl
ocr - 去偏斜 MNIST 数据集
我目前正在研究数字手写识别问题。我发现很多state-of-art算法对mnist dateset采用了一些预处理方法，比如deskewing和jittering(我不知道'jittering'是什么
python - 对 MNIST 数据集进行标准化和缩放的正确方法
我到处找，但找不到我想要的。基本上，MNIST 数据集具有像素值在范围 [0, 255] 内的图像。 .人们说，一般来说，最好做到以下几点: 将数据缩放到 [0,1]范围。将数据标准化为具有零均值和

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python-3.x - 如何正确实现 MNIST 数据集机器学习的反向传播？