tensorflow - 如何在谷歌计算引擎上运行 tensorflow GPU 容器？-6ren

tensorflow - 如何在谷歌计算引擎上运行 tensorflow GPU 容器？

转载作者：行者123 更新时间：2023-12-02 00:49:52

25

4

我正在尝试使用 GPU 加速器在谷歌计算引擎上运行 tensorflow 容器。

试过命令

gcloud compute instances create-with-container job-name \
  --machine-type=n1-standard-4 \
  --accelerator=type=nvidia-tesla-k80 \
  --image-project=deeplearning-platform-release \
  --image-family=common-container \
  --container image gcr/io/my-container \
  --container-arg="--container-arguments=xxxx"

但得到警告

WARNING: This container deployment mechanism requires a Container-Optimized OS image in order to work. Select an image from a cos-cloud project (cost-stable, cos-beta, cos-dev image families).

我还尝试了来自 cos-cloud 的系统镜像项目，它似乎没有 CUDA 驱动程序，因为 tensorflow 记录警告 cuInit failed .

想知道在具有 GPU 支持的谷歌计算引擎上运行 tensorflow 容器的正确方法是什么？

最佳答案

您可以 docker run您的容器在 startup-script 内的 deeplearningvm .


gcloud beta compute instances create deeplearningvm-$(date +"%Y%m%d-%H%M%S") \
--zone=us-central1-c \
--machine-type=n1-standard-8 \
--subnet=default \
--service-account=<your google service account> \
--scopes='https://www.googleapis.com/auth/cloud-platform' \
--accelerator=type=nvidia-tesla-k80,count=1 \
--image-project=deeplearning-platform-release \
--image-family=tf-latest-gpu \
--maintenance-policy=TERMINATE \
--metadata=install-nvidia-driver=True,startup-script='#!/bin/bash

# Check the driver until installed
while ! [[ -x "$(command -v nvidia-smi)" ]];
do
  echo "sleep to check"
  sleep 5s
done
echo "nvidia-smi is installed"

gcloud auth configure-docker
echo "Docker run with GPUs"
docker run --gpus all --log-driver=gcplogs --rm gcr.io/<your container>

echo "Kill VM $(hostname)"
gcloud compute instances delete $(hostname) --zone \
$(curl -H Metadata-Flavor:Google http://metadata.google.internal/computeMetadata/v1/instance/zone -s | cut -d/ -f4) -q

'

由于安装 nvidia 驱动程序需要几分钟，因此您必须等到安装后才能启动容器。 https://cloud.google.com/ai-platform/deep-learning-vm/docs/tensorflow_start_instance#creating_a_tensorflow_instance_from_the_command_line

Compute Engine loads the latest stable driver on the first boot and performs the necessary steps (including a final reboot to activate the driver). It may take up to 5 minutes before your VM is fully provisioned. In this time, you will be unable to SSH into your machine. When the installation is complete, to guarantee that the driver installation was successful, you can SSH in and run nvidia-smi.

关于tensorflow - 如何在谷歌计算引擎上运行 tensorflow GPU 容器？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/58714973/

25

4

0

文章推荐： php - 当 websocket 客户端太慢时 websocket 将如何工作？

文章推荐： google-apps-script - 你能对 Google-Apps-Script 失败做些什么吗？

文章推荐： Electron 应用程序有时会尝试使用旧版 Oauth 路由登录用户

文章推荐： mysql - 防止两个外键同时为NULL

node.js - Passport 谷歌 oauth2 与 Passport 谷歌 oauth20 包
这两个包看起来非常相似: http://www.passportjs.org/packages/passport-google-oauth2/ http://www.passportjs.org/pa
javascript - 谷歌、推特认证
我想在我的网站上添加通过 Google 和 Twitter 登录的按钮。我需要只使用应用程序的客户端而不是服务器端来完成此操作。但我没有找到任何 API。对于我发现的所有内容，我需要使用带有 key
javascript - 谷歌+网址分享
我使用此链接通过 google plus 共享我的页面。 https://plus.google.com/share?url=http%3A%2F%2Fexample.com%2Fcompany%2
Python 谷歌 API
我正在尝试学习 google API，并且我的经验是使用 Python，因此我尝试使用 google api python 客户端来访问一些 google 服务，但在构建服务对象时遇到错误。从 ap
indexing - 谷歌，还没有索引
在其实际的实时托管平台上构建实时站点的努力中，有没有办法告诉谷歌不要索引该网站？我发现了以下内容: http://support.google.com/webmasters/bin/answer.py
ios - 谷歌+登录SDK不工作
我正在开发一个 iOS 应用程序。当我运行用于 google+ 登录的程序时，在我点击允许访问按钮后，会显示此消息。 You've reached this page because we have
javascript - 谷歌+1按钮不起作用
我有一个非常复杂的网站，每个页面包含 11 个 js 文件。我最近添加了 google +1 按钮，代码如下: 这会正确显示 +1 按钮，直到我单击它。当我单击它时，出现此错误:https://
javascript 谷歌 API
我正在尝试使用 google API 创建一个 html 文件，以便在 google MAPS 上显示 KML 文件。这是 HTML 代码: function initMap() {
c++ - 谷歌/基准测试结果不一致
我是使用 Google Benchmark 的新手，在本地运行代码与在 Quick-Bench.com 上运行代码时，我收到了运行相同基准测试(下方)的不同结果，该基准测试使用 C++ 检索本地时间.
Ajax 内容索引，谷歌
我已按照 Google 网站上的说明通过添加以下元标记在我的 AngularJS 网站上启用 Ajax 抓取: 呈现的内容有一些链接，如: User 1 User 2 User 3 还有一些呈现动态
java - 谷歌 AppInvite
通过 Google 手册实现 Google AppInvite - link . 启动 Invite Activity 并在 LogCat 中获取下一步: E/AppInviteAgent: Get
谷歌 Go 的表现如何？
那么有人用过 Google 的 Go 吗？我想知道数学性能(例如触发器)与其他具有垃圾收集器的语言(如 Java 或 .NET)相比如何？有人调查过吗？最佳答案理论性能:纯 Go 程序的理论性能
stackdriver - 谷歌 stackdriver 缓慢
Stackdriver 测试我的网站启动速度慢我们使用 cloudflare 作为我们的站点 CDN 提供商。我们使用 stackdriver 从外部测试站点可用性，我们将时间检查间隔设置为 1 分
python - 谷歌 JAX 一维卷积神经网络
我正在尝试使用 stax.GeneralConv() ( https://jax.readthedocs.io/en/latest/_modules/jax/experimental/stax.htm
api - 谷歌 API 更改了来自谷歌金融的数据
我有一个从谷歌金融中提取日内数据的软件。但是，由于昨天 Google 更新了 API，所以软件报错了 Conversion from string HTML HEAD meta http-equiv=
php - 谷歌 oAuth : redirect_uri_mismatch
我们在尝试从 Google 获取 oAuth token 时遇到“redirect_uri_mismatch”错误: [client 127.0.0.1:49892] {\n "error" : "
recaptcha - 谷歌 reCAPTCHA 在中国
我的网站正在使用 Google reCAPTCHA 控件，但我听说它被阻止了中国，反正我看到有人报告说将 API 更改为 https://www.recaptcha.net在中国工作？ Anyone
wordpress - 谷歌 anchor 广告高度过大
背景 WordPress Google Adsense 谷歌自动插入 anchor 定广告 https://pptmon.com 问题如下图所示，主播广告的容器高度太大了! 如何调整高度？这是谷歌
python - 谷歌 Colab 未加载
我在使用 Google Colab 时遇到问题。当我想制作一个新的 Python3 Notebook 时，由于我登录了我的 Google 帐户，因此无法加载刚刚打开的新页面。我该怎么办？感谢您的帮
express - 谷歌 Passport 回调后设置cookie
我正在使用 facebook和 google oauth2使用 passport js 登录, 有了这个流用户点击登录按钮重定向到 facebook/google auth 页面(取决于用户选择的

首页

博学

6Ren·AI

商城

tensorflow - 如何在谷歌计算引擎上运行 tensorflow GPU 容器？