python-3.x - 从 Google Cloud Storage 中的 csv 读取 n 行以与 Python csv 模块一起使用-6ren

python-3.x - 从 Google Cloud Storage 中的 csv 读取 n 行以与 Python csv 模块一起使用

转载作者：行者123 更新时间：2023-12-01 21:56:05

28

4

我有各种包含不同格式的非常大(每个约 4GB)的 csv 文件。这些来自 10 多个不同制造商的数据记录器。我试图将所有这些整合到 BigQuery 中。为了每天加载这些文件，我想首先将这些文件加载到 Cloud Storage 中，确定架构，然后加载到 BigQuery 中。由于一些文件有额外的标题信息(从 2 到 ~30 行)，我已经生成了自己的函数来确定最可能的标题行和每个文件样本(~100 行)的模式，这然后我可以在将文件加载到 BQ 时在 job_config 中使用。

当我处理从本地存储直接到 BQ 的文件时，这很好用，因为我可以使用上下文管理器，然后使用 Python 的 csv 模块，特别是 Sniffer 和 reader 对象。但是，似乎没有直接从存储中使用上下文管理器的等效方法。我不想绕过云存储，以防这些文件在加载到 BQ 时被中断。

我可以做什么:

# initialise variables
with open(csv_file, newline  = '', encoding=encoding) as datafile:
    dialect = csv.Sniffer().sniff(datafile.read(chunk_size))
    reader = csv.reader(datafile, dialect)
    sample_rows = []
    row_num  = 0
    for row in reader:
         sample_rows.append(row)
         row_num+=1
         if (row_num >100):
             break
    sample_rows
# Carry out schema  and header investigation...

对于 Google Cloud Storage，我尝试使用 download_as_string 和 download_to_file，它们提供数据的二进制对象表示，但是我无法让 csv 模块处理任何数据。我试图使用 .decode('utf-8') 并返回一个带有\r\n 的 looong 字符串。然后我使用 splitlines() 来获取数据列表，但 csv 函数仍然提供方言和阅读器，将数据拆分为单个字符作为每个条目。

有没有人设法在不下载整个文件的情况下将 csv 模块与存储在 Cloud Storage 中的文件一起使用？

最佳答案

看了GitHub上的csv源码后，我设法使用Python中的io模块和csv模块解决了这个问题。 io.BytesIO 和 TextIOWrapper 是要使用的两个关键函数。可能不是一个常见的用例，但我想我会在此处发布答案，以便为需要它的任何人节省一些时间。

# Set up storage client and create a blob object from csv file that you are trying to read from GCS.
content = blob.download_as_string(start = 0, end = 10240) # Read a chunk of bytes that will include all header data and the recorded data itself.
bytes_buffer = io.BytesIO(content)
wrapped_text = io.TextIOWrapper(bytes_buffer, encoding = encoding, newline =  newline)
dialect = csv.Sniffer().sniff(wrapped_text.read()) 
wrapped_text.seek(0)
reader = csv.reader(wrapped_text, dialect)
# Do what you will with the reader object

关于python-3.x - 从 Google Cloud Storage 中的 csv 读取 n 行以与 Python csv 模块一起使用，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/56959627/

28

4

0

文章推荐： paypal - 尝试从智能按钮创建订阅时出错

文章推荐： Salesforce 批量 API 删除运算符

ios - Firebase swift 错误 Storage.storage() 在范围内找不到 'Storage'
嗨，当尝试将图像上传到 firebase 存储时，我正在使用 firebase 文档，但是出现此错误。在范围内找不到“存储” let storage = Storage.storage() le
firebase-storage - Firebase Storage 添加 firebase-storage@system.gserviceaccount.com 问题
我最近在使用 Firebase 存储时遇到了一些问题。当我们尝试访问刚刚上传的文件时，浏览器中出现此错误消息 { "error": { "code": 400,
azure-storage - 从 Microsoft.Azure.Storage.Blob 迁移到 Azure.Storage.Blobs - 缺少目录概念
这些是在不同版本的 NuGet 包之间迁移的重要指南: https://github.com/Azure/azure-sdk-for-net/blob/Azure.Storage.Blobs_12.6
angular - 警告 : Can't resolve all parameters for Storage in PATH/node_modules/@ionic/storage/es2015/storage. d.ts : (? )
警告: Warning: Can't resolve all parameters for Storage in /Users/zzm/Desktop/minan/node_modules/@ioni
storage - 圆形立方体问题 : connection to storage server failed
我在圆形立方体中收到此错误(“连接到存储服务器失败”)行。我已经检查了所有内容，配置和数据库用户名密码，服务器详细信息都是干净的。谁能告诉我可能是什么问题。这里我给出了整个配置文件。
docker - docker，-storage-opts和aufs storage-driver
我希望能够限制容器的大小，但是使用默认的存储驱动程序aufs(对于Ubuntu 14.04)，当我尝试使用--storage-opt参数时出现错误 $ docker create -it --name
google-cloud-storage - 为 Cloud Storage 使用不同的内容编码
我希望能够支持对使用 Google Cloud Storage 托管的静态 Assets 进行 Brotli 和 Gzip 编码。为此，我想在将文件上传为之前对其进行编码, .gz和 .br .问
google-cloud-storage - Google Cloud Storage 对象完成事件多次触发
场景我有几个由 Google Cloud Storage object.finalize 事件触发的 Google Cloud Functions。为此，我使用两个存储桶并使用“同步选项:覆盖目标位
google-cloud-storage - Google Cloud Storage - 使存储桶中的对象公开可见
我在 Google Cloud Storage 中有一个存储桶和一个网站。人们目前可以通过网站上传到存储桶(使用 Google 身份验证)。但是，我需要设置它以便任何人都可以查看上传的文件(并且不能
google-cloud-storage - Google Cloud Storage 是否已在搜索中编入索引？
如果文件被放入 Google Cloud 存储并公开，但该文件的网址在另一个网页上不存在，那么 Google 是否会在其搜索结果中将其编入索引？有人知道吗？最佳答案 Google 的搜索索引独立于其
google-cloud-storage - Google Cloud Storage 无法检索存储分区或存储分区的内容
截至今天早上，我无法访问我的存储桶。当我在导航上选择 Google Cloud Storage 选项卡时，一切都按预期加载，但不是显示我的两个存储桶，而是显示一个警告栏说: We were unab
google-cloud-storage - Google Cloud Storage 上传今天修改的文件
我想弄清楚是否可以在 Windows 平台上使用 gsutil 的 cp 命令将文件上传到 Google Cloud Storage。我的本地计算机上有 6 个文件夹，每天都会向其中添加新的 pdf
google-cloud-storage - Google Cloud Storage - 切换项目
我最近开始使用 Google Cloud Storage。最初我在安装 Cloud SDK 时创建了一个虚拟项目。现在我正在做另一个项目。 gsutil 仍然指向我以前的项目。我如何使它指向我的新项目
google-cloud-storage - 获取 Google Storage 存储桶大小的最快方法？
我目前正在这样做，但它非常慢，因为我的存储桶中有几 TB 的数据: gsutil du -sh gs://my-bucket-1/ 对于子文件夹也是如此: gsutil du -sh gs://my-
Azure - 'Blobs storage' 中的文件夹和 'File storage' 中的文件夹
这可能看起来很天真，我知道我们可以在 blob 中创建文件夹，并且这些文件夹仍然存储在容器中。我们仍然可以对这些“blob 中包含的文件夹”执行通常对文件存储中的文件夹执行的所有操作。我们仍然可以像
google-cloud-storage - Google Cloud Storage 中元数据值的长度有限制吗？
将文件上传到 Google Cloud Storage 时，有一个自定义数据字段元数据。 Google's example相当短: var metadata = { contentType: 'a
Azure - 'Blobs storage' 中的文件夹和 'File storage' 中的文件夹
这可能看起来很天真，我知道我们可以在 blob 中创建文件夹，并且这些文件夹仍然存储在容器中。我们仍然可以对这些“blob 中包含的文件夹”执行通常对文件存储中的文件夹执行的所有操作。我们仍然可以像
google-cloud-storage - 如何在短时间内列出 Google Storage 存储桶中的所有文件？
我有一个包含超过 2 万个文件名的 Google Storage 存储桶。有没有办法在短时间内列出存储桶中的所有文件名？最佳答案这取决于您所说的“短”是什么意思，但是: 您可以做的一件事来加快列出
google-cloud-storage - Google Cloud Storage 并为未找到的文件收费
有谁知道如果文件不存在，您是否需要为 Google Cloud Storage 中的文件请求付费？换句话说，有人访问您存储桶中不存在的文件是否计入您的请求？还是仅适用于存在的文件？最佳答案客户无需
google-cloud-storage - Google Cloud Storage 中的速率限制
在每一分钟结束时，我的代码总共会上传 20 到 40 个文件(从多台机器上并行上传大约 5 个文件，直到全部上传完毕)到 Google Cloud Storage。我经常收到 429 - Too Ma

首页

博学

6Ren·AI

商城

python-3.x - 从 Google Cloud Storage 中的 csv 读取 n 行以与 Python csv 模块一起使用