python - 使用Python xlsxwriter模块将srt数据写入excel-6ren

python - 使用Python xlsxwriter模块将srt数据写入excel

转载作者：行者123 更新时间：2023-11-30 23:07:17

24

4

这次我尝试使用Python的xlsxwriter模块将.srt中的数据写入excel。

字幕文件在 sublime text 中看起来像这样:

但我想将数据写入excel，所以它看起来像这样:

这是我第一次为此编写Python代码，所以我仍然处于尝试和错误的阶段......我尝试编写一些如下代码

但我认为这没有意义......

我会继续尝试，但如果您知道该怎么做，请告诉我。我将阅读您的代码并尝试理解它们!谢谢你! :)

最佳答案

以下将问题分解为几个部分:

解析输入文件。 parse_subtitles 是 generator获取行源并生成 {'index':'N', 'timestamp':'NN:NN:NN,NNN -> NN:NN:NN,NNN' 形式的记录序列, '副标题':'文本'}'。我采取的方法是跟踪我们处于三种不同状态中的哪一种:
1. 寻求下一个条目，当我们寻找下一个索引号时，它应该与正则表达式 ^\d*$ 匹配(只不过是一堆数字)
2. 查找时间戳，当找到索引时，我们期望时间戳出现在下一行，该时间戳应与正则表达式 ^\d{2}:\d{2 匹配}:\d{2},\d{3} -->\d{2}:\d{2}:\d{2},\d{3}$ (HH:MM:SS ,mmm -> HH:MM:SS,mmm) 和
3. 阅读字幕，同时使用实际的字幕文本，将空行和 EOF 解释为字幕终止点。
将上述记录写入工作表中的一行。 write_dict_to_worksheet 接受一行和工作表，以及一条记录和一个字典，为每个记录的键定义 Excel 0 索引的列号，然后适本地写入数据。
组织整个转换 convert 接受输入文件名(例如 'Wildlife.srt')，该文件名将被打开并传递给 parse_subtitles函数和输出文件名(例如将使用 xlsxwriter 创建的 'Subtitle.xlsx')。然后，它写入一个 header ，并且对于从输入文件解析的每条记录， writes that record to the XLSX file .

Logging statements出于 self 注释的目的而留下，并且因为在复制输入文件时，我在时间戳中将 : 插入到 ; 中，使其无法识别，并出现错误弹出窗口对于调试很方便!

我已将源文件的文本版本以及以下代码放在 this Gist 中

import xlsxwriter
import re
import logging

def parse_subtitles(lines):
    line_index = re.compile('^\d*$')
    line_timestamp = re.compile('^\d{2}:\d{2}:\d{2},\d{3} --> \d{2}:\d{2}:\d{2},\d{3}$')
    line_seperator = re.compile('^\s*$')

    current_record = {'index':None, 'timestamp':None, 'subtitles':[]}
    state = 'seeking to next entry'

    for line in lines:
        line = line.strip('\n')
        if state == 'seeking to next entry':
            if line_index.match(line):
                logging.debug('Found index: {i}'.format(i=line))
                current_record['index'] = line
                state = 'looking for timestamp'
            else:
                logging.error('HUH: Expected to find an index, but instead found: [{d}]'.format(d=line))

        elif state == 'looking for timestamp':
            if line_timestamp.match(line):
                logging.debug('Found timestamp: {t}'.format(t=line))
                current_record['timestamp'] = line
                state = 'reading subtitles'
            else:
                logging.error('HUH: Expected to find a timestamp, but instead found: [{d}]'.format(d=line))

        elif state == 'reading subtitles':
            if line_seperator.match(line):
                logging.info('Blank line reached, yielding record: {r}'.format(r=current_record))
                yield current_record
                state = 'seeking to next entry'
                current_record = {'index':None, 'timestamp':None, 'subtitles':[]}
            else:
                logging.debug('Appending to subtitle: {s}'.format(s=line))
                current_record['subtitles'].append(line)

        else:
            logging.error('HUH: Fell into an unknown state: `{s}`'.format(s=state))
    if state == 'reading subtitles':
        # We must have finished the file without encountering a blank line. Dump the last record
        yield current_record

def write_dict_to_worksheet(columns_for_keys, keyed_data, worksheet, row):
    """
    Write a subtitle-record to a worksheet. 
    Return the row number after those that were written (since this may write multiple rows)
    """
    current_row = row
    #First, horizontally write the entry and timecode
    for (colname, colindex) in columns_for_keys.items():
        if colname != 'subtitles': 
            worksheet.write(current_row, colindex, keyed_data[colname])

    #Next, vertically write the subtitle data
    subtitle_column = columns_for_keys['subtitles']
    for morelines in keyed_data['subtitles']:
        worksheet.write(current_row, subtitle_column, morelines)
        current_row+=1

    return current_row

def convert(input_filename, output_filename):
    workbook = xlsxwriter.Workbook(output_filename)
    worksheet = workbook.add_worksheet('subtitles')
    columns = {'index':0, 'timestamp':1, 'subtitles':2}

    next_available_row = 0
    records_processed = 0
    headings = {'index':"Entries", 'timestamp':"Timecodes", 'subtitles':["Subtitles"]}
    next_available_row=write_dict_to_worksheet(columns, headings, worksheet, next_available_row)

    with open(input_filename) as textfile:
        for record in parse_subtitles(textfile):
            next_available_row = write_dict_to_worksheet(columns, record, worksheet, next_available_row)
            records_processed += 1

    print('Done converting {inp} to {outp}. {n} subtitle entries found. {m} rows written'.format(inp=input_filename, outp=output_filename, n=records_processed, m=next_available_row))
    workbook.close()

convert(input_filename='Wildlife.srt', output_filename='Subtitle.xlsx')

编辑:更新为在输出中将多行字幕拆分为多行

关于python - 使用Python xlsxwriter模块将srt数据写入excel，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/32293013/

24

4

0

文章推荐： mysql - Multi-Tenancy : Hibernate with MySQL

文章推荐： javascript - 将 MYSQL 结果集从 JSP 转换为 Javascript 数组

文章推荐： java - 返回非持久数据作为对象的一部分

找不到 Python 模块 "cx_Oracle"模块
我最近在我的机器上安装了 cx_Oracle 模块，以便连接到远程 Oracle 数据库服务器。 (我身边没有 Oracle 客户端)。 Python:版本 2.7 x86 Oracle:版本 11.
python - Timeit 模块 - 将参数传递给 python timeit 模块
我想从 python timeit 模块检查打印以下内容需要多少时间，如何打印， import timeit x = [x for x in range(10000)] timeit.timeit("
javascript - 我该如何修复 --> 文件是 CommonJS 模块；它可以转换为 ES6 模块
我盯着 vs 代码编辑器上的 java 脚本编码，当我尝试将外部模块包含到我的项目中时，代码编辑器提出了这样的建议 -->(文件是 CommonJS 模块；它可能会转换为 ES6 模块。 )..有什么
javascript - 如何在 ES6 模块 Node 应用程序中包含 commonjs 模块？
我有一个 Node 应用程序，我想在标准 ES6 模块格式中使用(即 "type": "module" in the package.json ，并始终使用 import 和 export)而不转译为
css - BlueprintJS 未加载其图标 CSS 模块，但能够加载核心 CSS 模块
我正在学习将 BlueprintJS 合并到我的 React 网络应用程序中，并且在加载某些 CSS 模块时遇到了很多麻烦。我已经安装了 npm install @blueprintjs/core和
javascript - 将 RequireJS 模块 (AMD) 重构为 Webpack 模块 (CommonJS)
我需要重构一堆具有这样的调用的文件 define(['module1','module2','module3' etc...], function(a, b, c etc...) { //bun
javascript - 是否使用 : var app = angular. 模块...或简单地使用 : angular. 模块(
我是 Angular 的新手，正在学习各种教程(Codecademy、thinkster.io 等)，并且已经看到了声明应用程序容器的两种方法。首先: var app = angular.module
unit-testing - 在 OCaml 中使用 OUnit 模块 - 未绑定(bind)模块 OUnit 错误
我正在尝试将 OUnit 与 OCaml 一起使用。单元代码源码(unit.ml)如下: open OUnit let empty_list = [] let list_a = [1;2;3] le
javascript - "Argument ' 模块 ' is not a function, got Object"- 使用 webpack 导入 Angular 模块
我在 Angular 1.x 应用程序中使用 webpack 和 ES6 模块。在我设置的 webpack.config 中: resolve: { alias: { 'angular':
node-modules - 内部/模块/cjs/loader.js :750 return process. dlopen(模块，path.toNamespacedPath(文件名))；
internal/modules/cjs/loader.js:750 return process.dlopen(module, path.toNamespacedPath(filename));
JavaScript 模块
在本教程中，您将借助示例了解 JavaScript 中的模块。随着我们的程序变得越来越大，它可能包含许多行代码。您可以使用模块根据功能将代码分隔在单独的文件中，而不是将所有内容都放在一个文件
JavaScript 模块
我想知道是否可以将此代码更改为仅调用 MyModule.RED 而不是 MyModule.COLORS.RED。我尝试将 mod 设置为变量来存储颜色，但似乎不起作用。难道是我方法不对？ (funct
JavaScript 模块
我有以下代码。它是一个 JavaScript 模块。 (function() { // Object var Cahootsy; Cahootsy = { hello:
angular - 模块 : when and why?
关闭。这个问题是 opinion-based 。它目前不接受答案。想要改进这个问题？更新问题，以便 editing this post 可以用事实和引文来回答它。关闭 2 年前。 Improve
Lua极简入门指南（六）：模块
从用户的角度来看，一个模块能够通过 require 加载并返回一个 table，模块导出的接口都被定义在此 table 中（此 table 被作为一个 namespace）。所有的标准库都是模块。标
ruby 模块
Ruby的模块非常类似类,除了: 模块不可以有实体模块不可以有子类模块由module...end定义. 实际上...模块的'模块类'是'类的类'这个类的父类.搞懂了吗?不懂?让我们继续看
Perl GetOptions 模块
我有一个脚本，它从 CLI 获取 3 个输入变量并将其分别插入到 3 个变量: GetOptions("old_path=s" => \$old_path, "var=s" =
Python 模块，引用同一包中的其他模块？
我有一个简单的 python 包，其目录结构如下: wibble | |-----foo | |----ping.py | |-----bar | |----pong.py 简单的
ocaml - 无法解构仿函数(模块)
这种语法会非常有用——这不起作用有什么原因吗？谢谢! module Foo = { let bar: string = "bar" }; let bar = Foo.bar; /* works *
shell 模块 : < with ansible
我想运行一个命令: - name: install pip shell: "python {"changed": true, "cmd": "python <(curl https://boot

首页

博学

6Ren·AI

商城

python - 使用Python xlsxwriter模块将srt数据写入excel