python-2.7 - Theano 在计算梯度方面的效率/智能如何？-6ren

python-2.7 - Theano 在计算梯度方面的效率/智能如何？

转载作者：行者123 更新时间：2023-12-04 19:51:00

24

4

假设我有一个带有 5 个隐藏层的人工神经网络。暂时忘记神经网络模型的细节，例如偏差、使用的激活函数、数据类型等......。当然，激活函数是可微的。

通过符号微分，以下计算目标函数相对于层权重的梯度:

w1_grad = T.grad(lost, [w1])
w2_grad = T.grad(lost, [w2])
w3_grad = T.grad(lost, [w3])
w4_grad = T.grad(lost, [w4])
w5_grad = T.grad(lost, [w5])
w_output_grad = T.grad(lost, [w_output])

这样，要计算梯度 w.r.t w1，必须首先计算梯度 w.r.t w2、w3、w4 和 w5。与计算梯度 w.r.t w2 类似，必须首先计算梯度 w.r.t w3、w4 和 w5。

但是，我可以使用以下代码计算每个权重矩阵的梯度 w.r.t:

w1_grad, w2_grad, w3_grad, w4_grad, w5_grad, w_output_grad = T.grad(lost, [w1, w2, w3, w4, w5, w_output])

我想知道，这两种方法在性能方面有什么区别吗？ Theano 是否足够智能，可以避免使用第二种方法重新计算梯度？智能我的意思是计算 w3_grad，Theano 应该 [最好] 使用 w_output_grad、w5_grad 和 w4_grad 的预先计算的梯度，而不是再次计算它们。

最佳答案

好吧，事实证明 Theano 并没有采用先前计算的梯度来计算计算图较低层的梯度。这是一个具有 3 个隐藏层和一个输出层的神经网络的虚拟示例。然而，它是不是这将是一个大问题，因为计算梯度是一生一次的操作，除非您必须在每次迭代时计算梯度。 Theano 将导数的符号表达式返回为 computational graph从那时起，您可以简单地将其用作函数。从那时起，我们只需使用 Theano 派生的函数来计算数字值并使用这些值更新权重。

import theano.tensor as T
import time
import numpy as np
class neuralNet(object):
    def __init__(self, examples, num_features, num_classes):
        self.w = shared(np.random.random((16384, 5000)).astype(T.config.floatX), borrow = True, name = 'w')
        self.w2 = shared(np.random.random((5000, 3000)).astype(T.config.floatX), borrow = True, name = 'w2')
        self.w3 = shared(np.random.random((3000, 512)).astype(T.config.floatX), borrow = True, name = 'w3')
        self.w4 = shared(np.random.random((512, 40)).astype(T.config.floatX), borrow = True, name = 'w4')
        self.b = shared(np.ones(5000, dtype=T.config.floatX), borrow = True, name = 'b')
        self.b2 = shared(np.ones(3000, dtype=T.config.floatX), borrow = True, name = 'b2')
        self.b3 = shared(np.ones(512, dtype=T.config.floatX), borrow = True, name = 'b3')
        self.b4 = shared(np.ones(40, dtype=T.config.floatX), borrow = True, name = 'b4')
        self.x = examples

        L1 = T.nnet.sigmoid(T.dot(self.x, self.w) + self.b)
        L2 = T.nnet.sigmoid(T.dot(L1, self.w2) + self.b2)
        L3 = T.nnet.sigmoid(T.dot(L2, self.w3) + self.b3)
        L4 = T.dot(L3, self.w4) + self.b4
        self.forwardProp = T.nnet.softmax(L4)
        self.predict = T.argmax(self.forwardProp, axis = 1)

    def loss(self, y):
        return -T.mean(T.log(self.forwardProp)[T.arange(y.shape[0]), y])

x = T.matrix('x')
y = T.ivector('y')

nnet = neuralNet(x)
loss = nnet.loss(y)

diffrentiationTime = []
for i in range(100):
    t1 = time.time()
    gw, gw2, gw3, gw4, gb, gb2, gb3, gb4 = T.grad(loss, [nnet.w, nnet.w2, logReg.w3, nnet.w4, nnet.b, nnet.b2, nnet.b3, nnet.b4])
    diffrentiationTime.append(time.time() - t1)
print 'Efficient Method: Took %f seconds with std %f' % (np.mean(diffrentiationTime), np.std(diffrentiationTime))

diffrentiationTime = []
for i in range(100):
    t1 = time.time()
    gw = T.grad(loss, [nnet.w])
    gw2 = T.grad(loss, [nnet.w2])
    gw3 = T.grad(loss, [nnet.w3])
    gw4 = T.grad(loss, [nnet.w4])
    gb = T.grad(loss, [nnet.b])
    gb2 = T.grad(loss, [nnet.b2])
    gb3 = T.grad(loss, [nnet.b3])
    gb4 = T.grad(loss, [nnet.b4])
    diffrentiationTime.append(time.time() - t1)
print 'Inefficient Method: Took %f seconds with std %f' % (np.mean(diffrentiationTime), np.std(diffrentiationTime))

这将打印出以下内容:

Efficient Method: Took 0.061056 seconds with std 0.013217
Inefficient Method: Took 0.305081 seconds with std 0.026024

这表明 Theano 使用动态规划方法来计算有效方法的梯度。

关于python-2.7 - Theano 在计算梯度方面的效率/智能如何？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/34425161/

24

4

0

文章推荐： google-bigquery - Bigquery Json 从 gs 导入，缺少字段 BUG 的行？

文章推荐： pandas HDFStore 按日期时间索引选择行

文章推荐： google-api - 谷歌 QPX API key

theano - Theano 中的索引
如何通过索引向量在 Theano 中索引矩阵？更准确地说: v 的类型为 theano.tensor.vector(例如 [0,2]) A 具有 theano.tensor.matrix 类型(例如
theano - theano 函数的错误输入参数
我是theano的新手。我正在尝试实现简单的线性回归，但我的程序抛出以下错误: TypeError: ('Bad input argument to theano function with name
theano - 如何在不重建图形的情况下重用具有不同共享变量的 Theano 函数？
我有一个被多次调用的 Theano 函数，每次都使用不同的共享变量。按照现在的实现方式，Theano 函数在每次运行时都会重新定义。我假设，这会使整个程序变慢，因为每次定义 Theano 函数时，都会
theano - Theano.function中 'givens'变量的用途
我正在阅读http://deeplearning.net/tutorial/logreg.html给出的逻辑函数代码。我对函数的inputs和givens变量之间的区别感到困惑。计算微型批次中的模型所
theano - 如何设置 theano 配置
我是 Theano 的新手。尝试设置配置文件。首先，我注意到我没有 .theanorc 文件: locate .theanorc - 不返回任何内容 echo $THEANORC - 不返回任何内
theano - 为什么我们需要 Theano reshape ？
我不明白为什么我们在 Theano 中需要 tensor.reshape() 函数。文档中说: Returns a view of this tensor that has been reshaped
theano - 如何在 Theano 中翻转张量？
给定一个张量 v = t.vector()，我该如何翻转它？例如，[1, 2, 3, 4, 5, 6] 翻转后是 [6, 5, 4, 3, 2, 1]。最佳答案您可以简单地执行 v[::-1].e
theano - 为什么 theano 运行这么慢？
我是 Theano 的新手，正在尝试一些示例。 import numpy import theano.tensor as T from theano import function import da
theano - 定期记录梯度而不需要 Theano 中的两个函数(或减速)
出于诊断目的，我定期获取网络的梯度。一种方法是将梯度作为 theano 函数的输出返回。然而，每次都将梯度从 GPU 复制到 CPU 内存可能代价高昂，所以我宁愿只定期进行。目前，我通过创建两个函数对
theano - 输入维度不匹配二元交叉熵 Lasagne 和 Theano
我阅读了网络上所有关于人们忘记将目标向量更改为矩阵的问题的帖子，由于更改后问题仍然存在，我决定在这里提出我的问题。下面提到了解决方法，但出现了新问题，我感谢您的建议! 使用卷积网络设置和带有 sigm
theano - 在 Theano 中从 scan 调用函数
我需要通过扫描多次执行 theano 函数，以便总结成本函数并将其用于梯度计算。我熟悉执行此操作的深度学习教程，但我的数据切片和其他一些复杂情况意味着我需要做一些不同的事情。下面是我正在尝试做的一个
theano - Caffe 与 Theano MNIST 示例
我正在尝试学习(和比较)不同的深度学习框架，到时候它们是 Caffe 和 Theano。 http://caffe.berkeleyvision.org/gathered/examples/mnist
theano - 当代码几乎相同时，为什么 theano scan 的工作方式不同？
下面的代码: import theano import numpy as np from theano import tensor as T h1=T.as_tensor_variable(np.ze
theano - 使用 Theano/Lasagne 在 ImageNet 等大规模数据集上进行训练的最佳实践？
我发现 Theano/Lasagne 的所有示例都处理像 mnist 和 cifar10 这样的小数据集，它们可以完全加载到内存中。我的问题是如何编写高效的代码来训练大规模数据集？具体来说，为了让
python - Theano:使用 CSV 文件中的数据训练 theano 神经网络
我正在做图像分类，我必须检测图像是否包含飞机。我完成了以下步骤: 1. 从图像数据集中提取特征作为描述符 2. 用 K 完成 - 表示聚类并生成描述符语料库 3.将语料数据在0-1范围内归一化并保存
python - PyMC3 & Theano - Theano 代码在导入 pymc3 后停止工作
一些简单的 theano 代码完美运行，当我导入 pymc3 时停止运行为了重现错误，这里有一些片段: #Initial Theano Code (this works) import the
theano - Theano 中的 1-of-k(one-hot)编码
我在做this对于 NumPy 。 seq 是一个带有索引的列表。 IE。这实现了 1-of-k 编码(也称为 one-hot)。 def 1_of_k(seq, num_classes): nu
theano - 如何将所有批量数据加载到 Keras(Theano 后端)的 GPU 内存中？
Keras 将数据批量加载到 GPU 上(作者注明here)。对于小型数据集，这是非常低效的。有没有办法修改 Keras 或直接调用 Theano 函数(在 Keras 中定义模型之后)以允许将所有
python - Theano 无法使用 theano 配置 cnmem = 1 导入
Theano导入失败，theano配置cnmem = 1 知道如何确保 GPU 完全分配给 theano python 脚本吗？ Note: Display is not used to avoid
Python/Theano : Is it possible to construct truly recursive theano functions?
例如，我可以定义一个递归 Python lambda 函数来计算斐波那契数列，如下所示: fn = lambda z: fn(z-1)+fn(z-2) if z > 1 else z 但是，如果我尝试

首页

博学

6Ren·AI

商城

python-2.7 - Theano 在计算梯度方面的效率/智能如何？