mysql - 将大数据插入 Cloud Spanner 表-6ren

mysql - 将大数据插入 Cloud Spanner 表

转载作者：行者123 更新时间：2023-11-29 11:00:39

30

4

我想将大数据插入 Google 的 Cloud Spanner 表。

这就是我正在使用node.js应用程序所做的事情，但它停止了，因为txt文件太大(几乎2GB)。

1.load txt file

2.read line by line

3.split the line by "|"

4.build data object

5.insert data to Cloud Spanner table

Mysql支持使用.sql文件插入数据。 Cloud Spanner 也支持多种方式吗？

最佳答案

Cloud Spanner 目前不公开批量导入方法。听起来您打算单独插入每一行，这不是最佳方法。该文档提供了 efficient bulk loading 的最佳(和不良)实践。 :

To get optimal write throughput for bulk loads, partition your data by primary key with this pattern:

Each partition contains a range of consecutive rows. Each commit contains data for only a single partition. A good rule of thumb for your number of partitions is 10 times the number of nodes in your Cloud Spanner instance. So if you have N nodes, with a total of 10*N partitions, you can assign rows to partitions by:

Sorting your data by primary key. Dividing it into 10*N separate sections. Creating a set of worker tasks that upload the data. Each worker will write to a single partition. Within the partition, it is recommended that your worker write the rows sequentially. However, writing data randomly within a partition should also provide reasonably high throughput.

As more of your data is uploaded, Cloud Spanner automatically splits and rebalances your data to balance load on the nodes in your instance. During this process, you may experience temporary drops in throughput.

Following this pattern, you should see a maximum overall bulk write throughput of 10-20 MiB per second per node.

看起来您正在尝试在处理之前将整个大文件加载到内存中。对于大文件，您应该考虑加载和处理 block 而不是整个文件。我是一名 Node 专家，但您可能应该尝试将其作为流读取，而不是将所有内容都保留在内存中。

关于mysql - 将大数据插入 Cloud Spanner 表，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/42339544/

30

4

0

文章推荐： mysql - 如何在innodb中实现隐式行级锁定？

文章推荐： iphone - 核心蓝牙内部的可写特性

文章推荐： iphone - NSBlockOperation 在 NSOperation 中调用一个方法

google-cloud-platform - 从 Google Cloud 上的 Cloud Run 访问 Cloud SQL
我有一个 Cloud Run 服务，它通过 SQLAlchemy 访问 Cloud SQL 实例.但是，在 Cloud Run 的日志中，我看到 CloudSQL connection failed.
cloud - 为什么叫 "Cloud"？
关闭。这个问题是opinion-based .它目前不接受答案。想改善这个问题吗？更新问题，以便可以通过 editing this post 用事实和引文回答问题. 4年前关闭。 Improve t
google-cloud-platform - 如何为 Cloud Build 用于 Cloud Run 部署的 Cloud Storage 存储分区指定区域？
在将 docker 容器镜像部署到 Cloud Run 时，我可以选择一个区域，这很好。 Cloud Run 将构建委托(delegate)给 Cloud Build，后者显然会创建两个存储桶来实现这
google-cloud-platform - Cloud PubSub 重复消息触发的 Cloud Functions
我正在尝试将 Cloud Functions 用作由 PubSub 触发的异步后台工作程序，并进行更长时间的工作(以分钟为单位)。完整代码在这里https://github.com/zdenulo/c
user-data - cloud-init执行顺序不尊重/etc/cloud/cloud.cfg？
这是/etc/cloud/cloud.cfg的内容Ubuntu云16.04镜像: # The top level settings are used as module # and system co
google-cloud-platform - 从 Cloud Functions 启动 Cloud Dataflow
如何从 Google Cloud Function 启动 Cloud Dataflow 作业?我想使用 Google Cloud Functions 作为启用跨服务组合的机制。最佳答案我已经包含了
google-cloud-platform - 如何从 Cloud Shell 连接到 Cloud SQL？
我想使用 Cloud Shell 在我的第二代 Cloud Sql 实例上运行数据库迁移。我找到了一个 example in the docs关于如何使用 gcloud 进行连接.但是当我运行命令时
google-cloud-platform - Cloud Dataproc 和其他 Google Cloud 产品的身份验证错误
我正在尝试使用 Google Cloud PubSub和我的 Google Cloud Dataproc群集，我收到如下身份验证范围错误: { "code" : 403, "errors" :
google-cloud-platform - 使用用户帐户凭据访问私有(private) Cloud Run/Cloud Functions
这是我的用例。我已经有一个以私有(private)模式部署的 Cloud Run 服务。 (与云功能相同的问题) 我正在开发使用此 Cloud Run 的新服务。我在应用程序中使用默认凭据进行身份验
google-cloud-sql - 如何从 Cloud Run 安全地连接到 Cloud SQL？
如何连接到 Cloud SQL 上的数据库，而无需在容器中添加我的凭据文件？最佳答案使用 UNIX 域套接字 (Java) 从云运行(完全托管)连接到云 SQL At this time Clou
google-cloud-ml - 如何在google-cloud-ml作业或Google Cloud Storage中加载numpy npz文件？
我有一个google-cloud-ml作业，需要从gs存储桶加载numpy .npz文件。我遵循了this example上关于如何从gs加载.npy文件的操作，但是由于.npz文件已压缩，因此它对我
google-cloud-platform - Cloud build trigger 看不到另一个项目的 Cloud Source Repository
我想创建链接到另一个项目中的 Cloud Source Repository 的 Cloud Build 触发器。但是当我在应该选择存储库的步骤中时，列表是空的。我尝试了不同的许可，但没有运气。谁能告
google-cloud-functions - 从 Cloud Function 本身获取 Cloud Function 名称
向 Twilio 发送 SMS 时，Twilio 会向指定的 URL 发送多个请求，以通过 Webhook 提供该 SMS 传送的状态。我想让这个回调异步，所以我开发了一个 Cloud Functio
google-cloud-firestore - 将 Cloud Firestore 项目迁移到另一个 Cloud Firestore 项目
我需要更改我的项目 ID，因为要验证的 Firebase 身份验证链接在链接上显示了项目 ID，并且由于品牌 reshape ，项目名称已更改。根据我发现的信息，更改项目 ID 似乎不太可能。我正在考
google-cloud-platform - 如何在 Cloud Run 中自动部署来自 Cloud Build 的最新镜像
用于部署我的 Angular 应用程序的 CI/CD 管道已关闭，但我看到 Google Cloud Run 在容器镜像更新后没有部署新修订版。我已将 Cloud Build 设置为在 GitHub
google-cloud-platform - 将 Cloud Armor 与 Cloud Run 结合使用并避免绕过
报价https://cloud.google.com/load-balancing/docs/https/setting-up-https-serverless#enabling While Goog
google-cloud-platform - Cloud Spanner 读取与 Cloud Spanner SQL API
Cloud Spanner 提供了两种不同的 API。 Cloud Spanner 读取与 Cloud Spanner SQL API 之间有什么区别？最佳答案在幕后，它们都使用相同的执行机制，因
google-cloud-platform - Google Cloud Spanner 和 Cloud SQL 之间有什么区别？
我是 GCP 堆栈的新手，所以我对用于存储数据的 GCP 技术数量感到非常困惑: https://cloud.google.com/products/storage 虽然上面的文章中没有提到googl
google-cloud-platform - 如何避免从 Cloud Function 到 Cloud SQL 的网络出站费用？
我发现 Google Cloud Functions 的网络出站费用令人惊讶，我正在尝试了解发生这种情况的原因以及如何避免这种情况。 Stackdriver 监控表明有问题的函数是我的 ingest
google-cloud-sql - Prisma DATABASE_URL 错误(Cloud Run + Cloud SQL)
我使用 Prisma使用 Cloud Run 和 Cloud SQL。在向 prisma.schema 提供 DATABASE_URL 后，它会在运行时抛出一个错误。 Can't reach data

首页

博学

6Ren·AI

商城

mysql - 将大数据插入 Cloud Spanner 表