scala - 超出物理限制运行的 Spark 容器-6ren

scala - 超出物理限制运行的 Spark 容器

转载作者：可可西里更新时间：2023-11-01 14:23:42

24

4

我一直在寻找以下问题的解决方案。我使用的是 Scala 2.11.8 和 Spark 2.1.0。

Application application_1489191400413_3294 failed 1 times due to AM Container for appattempt_1489191400413_3294_000001 exited with exitCode: -104
For more detailed output, check application tracking page:http://ip-172-31-17-35.us-west-2.compute.internal:8088/cluster/app/application_1489191400413_3294Then, click on links to logs of each attempt.
Diagnostics: Container [pid=23372,containerID=container_1489191400413_3294_01_000001] is running beyond physical memory limits. 
Current usage: 1.4 GB of 1.4 GB physical memory used; 3.5 GB of 6.9 GB virtual memory used. Killing container.

请注意，我分配的比此处错误中报告的 1.4 GB 多得多。因为我没有看到我的执行者失败，所以我从这个错误中读到这个驱动程序需要更多内存。但是，我的设置似乎没有传播。

我正在为 yarn 设置作业参数，如下所示:

val conf = new SparkConf()
  .setAppName(jobName)
  .set("spark.hadoop.mapred.output.committer.class", "com.company.path.DirectOutputCommitter")
additionalSparkConfSettings.foreach { case (key, value) => conf.set(key, value) }

// this is the implicit that we pass around
implicit val sparkSession = SparkSession
  .builder()
  .appName(jobName)
  .config(conf)
  .getOrCreate()

additionalSparkConfSettings 中的内存配置参数是使用以下代码段设置的:

HashMap[String, String](
  "spark.driver.memory" -> "8g",
  "spark.executor.memory" -> "8g",
  "spark.executor.cores" -> "5",
  "spark.driver.cores" -> "2",
  "spark.yarn.maxAppAttempts" -> "1",
  "spark.yarn.driver.memoryOverhead" -> "8192",
  "spark.yarn.executor.memoryOverhead" -> "2048"
)

我的设置真的没有传播吗？还是我误解了日志？

谢谢!

最佳答案

需要为执行程序和驱动程序设置开销内存，它应该是驱动程序和执行程序内存的一部分。

spark.yarn.executor.memoryOverhead = executorMemory * 0.10, with minimum of 384

The amount of off-heap memory (in megabytes) to be allocated per executor. This is memory that accounts for things like VM overheads, interned strings, other native overheads, etc. This tends to grow with the executor size (typically 6-10%).

spark.yarn.driver.memoryOverhead = driverMemory * 0.10, with minimum of 384.

The amount of off-heap memory (in megabytes) to be allocated per driver in cluster mode. This is memory that accounts for things like VM overheads, interned strings, other native overheads, etc. This tends to grow with the container size (typically 6-10%).

要了解有关内存优化的更多信息，请参阅 Memory Management Overview

另请参阅 SO Container is running beyond memory limits 上的以下主题

干杯!

关于scala - 超出物理限制运行的 Spark 容器，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/43055678/

24

4

0

文章推荐： hadoop - spark-submit 如何设置user.name

文章推荐： python - 如何在 Windows 机器上定义 .pdbrc？

文章推荐： Python写入hdfs文件

c# - 限制/限制 serviceBus 队列以触发 ServiceBusTrigger 形式的消息
我有一个 ServiceBusQueue(SBQ)，它获取大量消息负载。我有一个具有 accessRights(manage) 的 ServiceBusTrigger(SBT)，它不断轮询来自 SBQ
mysql - 对特定列应用 SQL 限制，而不是对完整结果集应用 SQL 限制
在下面给出的结果集中，有 2 个唯一用户 (id)，并且查询中可能会出现更多此类用户: 这是多连接查询: select id, name, col1Code, col2Code, col2Va
python - 限制/限制 GRequests 中 HTTP 请求的速率
我正在用 Python 2.7.3 编写一个带有 GRequests 的小脚本和 lxml 可以让我从各种网站收集一些收藏卡价格并进行比较。问题是其中一个网站限制了请求的数量，如果我超过它，就会发回
database - 跟进【删除(级联/限制)】和更新(级联/限制)
我想知道何时实际使用删除级联或删除限制以及更新级联或更新限制。我对使用它们或在我的数据库中应用感到很困惑。最佳答案在外键约束上使用级联运算符是一个热门话题。理论上，如果您知道删除父对象也将自动删
SQL where 限制
下面是我的输出，我只想显示那些重复的名字。每个名字都是飞行员，数字是飞行员驾驶的飞机类型。我想显示驾驶不止一架飞机的飞行员的姓名。我正在使用 sql*plus PIL_PILOTNAME
NativeScript 限制
我正在评估不同的移动框架，我认为 nativescript 是一个不错的选择。但我不知道开发过程是否存在限制。例如，我对样式有限制(这并不重要)，但我想知道将来我是否可以有限制并且不能使用某些 nat
GrailsDataBinder 限制？
我正在尝试使用 grails 数据绑定(bind)将一些表单参数映射到我的模型中，但我认为在映射嵌入式集合方面可能存在一些限制。例如，如果我提交一些这样的参数，那么映射工作正常: //this wo
Django模板timesince过滤器-限制
是否可以将 django 自过滤器起的时间限制为 7 天。如果日期超过 7 天，则不应用过滤器最佳答案 timesince 的源代码位于 django/django/utils/timesince.
Paypal 限制
我想在我的网站上嵌入一个 PayPal 捐赠按钮。但问题是我住在伊朗——这个国家受到制裁，人们不使用国际银行账户或主要信用卡。有什么想法吗？请帮忙! 问候沮丧最佳答案您可以在伊朗境内使用为伊朗
MySQL联合+限制
这是我的查询 select PhoneNumber as _data,PhoneType as _type from contact_phonenumbers where ContactID = 3
mongodb $in 限制
这个问题在这里已经有了答案: What is the maximum number of parameters passed to $in query in MongoDB? (4 个答案) 关闭
AndroidManifest 限制
我的一个项目的 AndroidManifest.xml 变得越来越大(> 1000 行)，因为我必须对某些文件类型使用react并且涵盖所有情况变得越来越复杂。我想知道 list 大小是否有任何限制。
MySQL 限制
在使用 Sybase、Infomix、DB2 等其他数据库产品多年后使用 MySQL 5.1 Enterprise 时；我遇到了 MySQL 不会做的事情。例如，它只能为 SELECT 查询生成 EX
mongodb $in 限制
这个问题在这里已经有了答案: What is the maximum number of parameters passed to $in query in MongoDB? (4 个回答) 关闭5年
限制 Apache日志文件大小的方法
通常我们是在{$apache}/conf/httpd.conf中设置Apache的参数，然而我们并没有发现可以设置日志文件大小的配置指令，通过参考http://httpd.apache.org/do
Android SharedPreferences 限制
我正在搜索最大的 Android SharedPreferences 键值对，但找不到任何好的答案。其次，我想问一下，如果我有一个键，它的字符串值限制是多少。多少字符可以放入其中。如果我需要频繁更改值
Soundcloud API 限制。
我目前正在试验 SoundCloud API，并注意到我对/tracks 资源的 GET 请求一次从不返回超过 200 个结果。关于这个的几个问题: 这个限制是故意的吗？有没有办法增加这个限制？如
Firebase TLS 限制
我正在与一家名为 Dwolla 的金融技术公司合作，该公司提供了一个 API，用于将银行信息附加到用户并收取/发送 ACH 付款。他们需要我将我的 TLS 最低版本升级到 1.2(禁用 TLS 1.
php - 根据重复元素的数量对PHP中的多维关联数组进行排序/限制
我在 PHP 中有一个多维数组，如下所示: $array = Array ( [0] => Array ( [bill] => 1 ) [1] => Array ( [
连接的 SQL 限制
我在获取下一个查询的第一行时遇到了问题: Select mar.Title MarketTitle, ololo.NUMBER, ololo.Title from Markets mar JOIN(

首页

博学

6Ren·AI

商城

scala - 超出物理限制运行的 Spark 容器