python - 为什么简单梯度下降会发散？-6ren

python - 为什么简单梯度下降会发散？

转载作者：太空宇宙更新时间：2023-11-03 10:55:41

25

4

这是我第二次尝试在一个变量中实现梯度下降，但它总是发散。有什么想法吗？

这是一个简单的线性回归，用于最小化一个变量中的残差平方和。

def gradient_descent_wtf(xvalues, yvalues):
    tolerance = 0.1

    #y=mx+b
    #some line to predict y values from x values
    m=1.
    b=1.

    #a predicted y-value has value mx + b

    for i in range(0,10):

        #calculate y-value predictions for all x-values
        predicted_yvalues = list()
        for x in xvalues:
            predicted_yvalues.append(m*x + b)

        # predicted_yvalues holds the predicted y-values

        #now calculate the residuals = y-value - predicted y-value for each point
        residuals = list()
        number_of_points = len(yvalues)
        for n in range(0,number_of_points):
            residuals.append(yvalues[n] - predicted_yvalues[n])

        ## calculate the residual sum of squares from the residuals, that is,
        ## square each residual and add them all up. we will try to minimize
        ## the residual sum of squares later.
        residual_sum_of_squares = 0.
        for r in residuals:
            residual_sum_of_squares += r**2
        print("RSS = %s" % residual_sum_of_squares)
        ##
        ##
        ##

        #now make a version of the residuals which is multiplied by the x-values
        residuals_times_xvalues = list()
        for n in range(0,number_of_points):
            residuals_times_xvalues.append(residuals[n] * xvalues[n])

        #now create the sums for the residuals and for the residuals times the x-values
        residuals_sum = sum(residuals)

        residuals_times_xvalues_sum = sum(residuals_times_xvalues)

        # now multiply the sums by a positive scalar and add each to m and b.

        residuals_sum *= 0.1
        residuals_times_xvalues_sum *= 0.1

        b += residuals_sum
        m += residuals_times_xvalues_sum

        #and repeat until convergence.
        #convergence occurs when ||sum vector|| < some tolerance.
        # ||sum vector|| = sqrt( residuals_sum**2 + residuals_times_xvalues_sum**2 )

        #check for convergence
        magnitude_of_sum_vector = (residuals_sum**2 + residuals_times_xvalues_sum**2)**0.5
        if magnitude_of_sum_vector < tolerance:
            break

    return (b, m)

结果:

gradient_descent_wtf([1,2,3,4,5,6,7,8,9,10],[6,23,8,56,3,24,234,76,59,567])
RSS = 370433.0
RSS = 300170125.7
RSS = 4.86943013045e+11
RSS = 7.90447409339e+14
RSS = 1.28312217794e+18
RSS = 2.08287421094e+21
RSS = 3.38110045417e+24
RSS = 5.48849288217e+27
RSS = 8.90939341376e+30
RSS = 1.44624932026e+34
Out[108]:
(-3.475524066284303e+16, -2.4195981188763203e+17)

最佳答案

梯度很大——因此您在长距离内跟随大矢量(大数的 0.1 倍是大的)。找到适当方向的单位向量。像这样的东西(用理解代替你的循环):

def gradient_descent_wtf(xvalues, yvalues):
    tolerance = 0.1

    m=1.
    b=1.

    for i in range(0,10):
        predicted_yvalues = [m*x+b for x in xvalues]

        residuals = [y-y_hat for y,y_hat in zip(yvalues,predicted_yvalues)]

        residual_sum_of_squares = sum(r**2 for r in residuals) #only needed for debugging purposes
        print("RSS = %s" % residual_sum_of_squares)

        residuals_times_xvalues = [r*x for r,x in zip(residuals,xvalues)]

        residuals_sum = sum(residuals)

        residuals_times_xvalues_sum = sum(residuals_times_xvalues)

        # (residuals_sum,residual_times_xvalues_sum) is a vector which points in the negative
        # gradient direction. *Find a unit vector which points in same direction*

        magnitude = (residuals_sum**2 + residuals_times_xvalues_sum**2)**0.5

        residuals_sum /= magnitude
        residuals_times_xvalues_sum /= magnitude

        b += residuals_sum * (0.1)
        m += residuals_times_xvalues_sum * (0.1)

        #check for convergence -- this needs work!
        magnitude_of_sum_vector = (residuals_sum**2 + residuals_times_xvalues_sum**2)**0.5
        if magnitude_of_sum_vector < tolerance:
            break

    return (b, m)

例如:

>>> gradient_descent_wtf([1,2,3,4,5,6,7,8,9,10],[6,23,8,56,3,24,234,76,59,567])
RSS = 370433.0
RSS = 368732.1655050716
RSS = 367039.18363896786
RSS = 365354.0543519137
RSS = 363676.7775934381
RSS = 362007.3533123621
RSS = 360345.7814567845
RSS = 358692.061974069
RSS = 357046.1948108295
RSS = 355408.17991291644
(1.1157111313023558, 1.9932828425473605)

这当然更合理。

制作数值稳定的梯度下降算法并非易事。您可能想查阅一本不错的数值分析教科书。

关于python - 为什么简单梯度下降会发散？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/41354177/

25

4

0

文章推荐： php - php验证中的解析错误

文章推荐： c# - 强制本地 DateTime 序列化

文章推荐： c# - 用指针调用非托管代码

文章推荐： c# - 来自 BitmapImage 的 CroppedImage

Java正则表达式，简单
我正在努力实现以下目标，假设我有字符串: ( z ) ( A ( z ) ( A ( z ) ( A ( z ) ( A ( z ) ( A ) ) ) ) ) 我想编写一个正则
CSS水平滚动(简单)
给定: 1 2 3 4 5 6
MySQL填充样例数据(简单)
很难说出这里要问什么。这个问题模棱两可、含糊不清、不完整、过于宽泛或夸夸其谈，无法以目前的形式得到合理的回答。如需帮助澄清此问题以便重新打开，visit the help center . 关闭 1
简单、好懂的Svelte实现原理
大家好，我卡颂。 Svelte问世很久了，一直想写一篇好懂的原理分析文章，拖了这么久终于写了。本文会围绕一张流程图和两个Demo讲解，正确的食用方式是用电脑打开本文，跟着流程图、Demo一
Javascript使用正则验证身份证号(简单)
身份证为15位或者18位，15位的全为数字，18位的前17位为数字，最后一位为数字或者大写字母”X“。与之匹配的正则表达式： ?
强大的jquery插件jqeuryUI做网页对话框效果！简单
我们先来最简单的，网页的登录窗口；不过开始之前，大家先下载jquery的插件本人习惯用了vs2008来做网页了，先添加一个空白页这是最简单的的做法。。。先在body里面插入 <
简单、易用的MySQL官方压力测试工具
1、MySQL自带的压力测试工具 Mysqlslap mysqlslap是mysql自带的基准测试工具,该工具查询数据,语法简单,灵活容易使用.该工具可以模拟多个客户端同时并发的向服务器发出
.NET开源、简单、实用的数据库文档生成工具
前言今天大姚给大家分享一款.NET开源（MIT License）、免费、简单、实用的数据库文档（字典）生成工具，该工具支持CHM、Word、Excel、PDF、Html、XML、Markdown等
【Go基础入门教程】Go语言代码风格清晰、简单
Go语言语法类似于C语言，因此熟悉C语言及其派生语言（ C++、 C#、Objective-C 等）的人都会迅速熟悉这门语言。 C语言的有些语法会让代码可读性降低甚至发生歧义。Go语言在C语言的
FFMpeg 简单/快速转换视频文件不适用于任何视频
我正在使用快速将 mkv 转换为 mp4 ffmpeg 命令 ffmpeg -i test.mkv -vcodec copy -acodec copy new.mp4 但不适用于任何 mkv 文件，当
VBA 计算具有特定名称的工作表数(简单)
我想计算我的工作簿中的工作表数量，然后从总数中减去特定的工作表。我错过了什么？这给了我一个对象错误: wsCount = ThisWorkbook.Sheets.Count - ThisWorkboo
Perl 配置::简单
我有一个 perl 文件，用于查看文件夹中是否存在 ini。如果是，它会从中读取，如果不是，它会根据我为它制作的模板创建一个。我在 ini 部分使用 Config::Simple。我的问题是，如果
ios - 如何访问iOS通知中传递的数据(简单)？
尝试让一个 ViewController 通过标准 Cocoa 通知与另一个 ViewController 进行通信。编写了一个简单的测试用例。在我最初的 VC 中，我将以下内容添加到 viewDi
optimization - (简单？)折线图的标签放置
我正在绘制高程剖面图，显示沿路径的高程增益/损失，类似于下面的: Sample Elevation Profile with hand-placed labels http://img38.image
javascript - 使用JS隐藏div(简单)
嗨，所以我需要做的是最终让 regStart 和 regPage 根据点击事件交替可见性，我不太担心编写 JavaScript 函数，但我根本无法让我的 regPage 首先隐藏。这是我的代码。请简单
c++ - 简单 for 循环中的大量时间损失
我有一个非常简单的程序来测量一个函数花费了多少时间。 #include #include #include struct Foo { void addSample(uint64_t s)
JavaScript 简单 BitConverter
我需要为 JavaScript 制作简单的 C# BitConverter。我做了一个简单的BitConverter class BitConverter{ constructor(){} GetBy
javascript - 简单 for 循环出现意外标记错误？
已关闭。这个问题是 not reproducible or was caused by typos 。目前不接受答案。这个问题是由拼写错误或无法再重现的问题引起的。虽然类似的问题可能是 on-top
group-by - 简单.数据分组依据
我是 Simple.Data 的新手。但我很难找到如何进行“分组依据”。我想要的是非常基本的。表格看起来像: +________+ | cards | +________+ | id |
Javascript 简单 UDF
我现在正在开发一个 JS UDF，它看起来遵循编码。通常情况下，由于循环计数为 2，Alert Msg 会出现两次。我想要的是即使循环计数为 3，Alert Msg 也只会出现一次。任何想法都

首页

博学

6Ren·AI

商城

python - 为什么简单梯度下降会发散？