python - 使用Python中的Azure语音服务读取音频文件并转换为文本，但只有第一句话转换为语音-6ren

python - 使用Python中的Azure语音服务读取音频文件并转换为文本，但只有第一句话转换为语音

转载作者：行者123 更新时间：2023-12-02 16:34:40

下面是代码，

import json
import os
from azure.storage.blob import BlobServiceClient, BlobClient, ContainerClient
import azure.cognitiveservices.speech as speechsdk

def main(filename):
    container_name="test-container"
            print(filename)
    blob_service_client = BlobServiceClient.from_connection_string("DefaultEndpoint")
    container_client=blob_service_client.get_container_client(container_name)
    blob_client = container_client.get_blob_client(filename)
    with open(filename, "wb") as f:
        data = blob_client.download_blob()
        data.readinto(f)

    speech_key, service_region = "1234567", "eastus"
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)

    audio_input = speechsdk.audio.AudioConfig(filename=filename)
    print("Audio Input:-",audio_input)
  
    speech_config.speech_recognition_language="en-US"
    speech_config.request_word_level_timestamps()
    speech_config.enable_dictation()
    speech_config.output_format = speechsdk.OutputFormat(1)

    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input)
    print("speech_recognizer:-",speech_recognizer)
    #result = speech_recognizer.recognize_once()
    all_results = []

    def handle_final_result(evt):
        all_results.append(evt.result.text)  
    done = False 

    def stop_cb(evt):
        #print('CLOSING on {}'.format(evt))
        speech_recognizer.stop_continuous_recognition()
        global done
        done= True

    #Appends the recognized text to the all_results variable. 
    speech_recognizer.recognized.connect(handle_final_result) 
    speech_recognizer.session_stopped.connect(stop_cb)
    speech_recognizer.canceled.connect(stop_cb)

    speech_recognizer.start_continuous_recognition()
    
    
    #while not done:
        #time.sleep(.5)
    
    print("Printing all results from speech to text:")
    print(all_results)


    
main(filename="test.wav")

从主函数调用时出错，

test.wav
Audio Input:- <azure.cognitiveservices.speech.audio.AudioConfig object at 0x00000204D72F4E88>
speech_recognizer:- <azure.cognitiveservices.speech.SpeechRecognizer object at 0x00000204D7065148>
[]

预期输出(不使用main函数的输出)

test.wav
Audio Input:- <azure.cognitiveservices.speech.audio.AudioConfig object at 0x00000204D72F4E88>
speech_recognizer:- <azure.cognitiveservices.speech.SpeechRecognizer object at 0x00000204D7065148>
Printing all results from speech to text:
['hi', '', '', 'Uh.', 'A good laugh.', '1487', "OK, OK, I think that's enough.", '']

如果我们不使用主函数，现有代码可以完美运行，但是当我使用主函数调用它时，我没有得到所需的输出。请指导我们弥补缺失的部分。

最佳答案

如文章 here 中所述,recognize_once_async()(您正在使用的方法) - 此方法只会检测从检测到的语音开始到下一次暂停的输入中已识别的话语。

根据我的理解，如果您使用start_continuous_recognition()，您的要求就会得到满足。启动函数将启动并继续处理所有话语，直到您调用停止函数。

此方法有很多与之相关的事件，当语音识别过程发生时，“识别”事件会触发。您需要有一个事件处理程序来处理识别和提取文本。您可以引用文章here了解更多信息。

分享一个使用 start_continuous_recognition() 将音频转换为文本的示例片段。

import azure.cognitiveservices.speech as speechsdk
import time
import datetime

# Creates an instance of a speech config with specified subscription key and service region.
# Replace with your own subscription key and region identifier from here: https://aka.ms/speech/sdkregion
speech_key, service_region = "YOURSUBSCRIPTIONKEY", "YOURREGION"
speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)

# Creates an audio configuration that points to an audio file.
# Replace with your own audio filename.
audio_filename = "sample.wav"
audio_input = speechsdk.audio.AudioConfig(filename=audio_filename)

# Creates a recognizer with the given settings
speech_config.speech_recognition_language="en-US"
speech_config.request_word_level_timestamps()
speech_config.enable_dictation()
speech_config.output_format = speechsdk.OutputFormat(1)

speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input)

#result = speech_recognizer.recognize_once()
all_results = []



#https://learn.microsoft.com/en-us/python/api/azure-cognitiveservices-speech/azure.cognitiveservices.speech.recognitionresult?view=azure-python
def handle_final_result(evt):
    all_results.append(evt.result.text) 
    
    
done = False

def stop_cb(evt):
    print('CLOSING on {}'.format(evt))
    speech_recognizer.stop_continuous_recognition()
    global done
    done= True

#Appends the recognized text to the all_results variable. 
speech_recognizer.recognized.connect(handle_final_result) 

#Connect callbacks to the events fired by the speech recognizer & displays the info/status
#Ref:https://learn.microsoft.com/en-us/python/api/azure-cognitiveservices-speech/azure.cognitiveservices.speech.eventsignal?view=azure-python   
speech_recognizer.recognizing.connect(lambda evt: print('RECOGNIZING: {}'.format(evt)))
speech_recognizer.recognized.connect(lambda evt: print('RECOGNIZED: {}'.format(evt)))
speech_recognizer.session_started.connect(lambda evt: print('SESSION STARTED: {}'.format(evt)))
speech_recognizer.session_stopped.connect(lambda evt: print('SESSION STOPPED {}'.format(evt)))
speech_recognizer.canceled.connect(lambda evt: print('CANCELED {}'.format(evt)))
# stop continuous recognition on either session stopped or canceled events
speech_recognizer.session_stopped.connect(stop_cb)
speech_recognizer.canceled.connect(stop_cb)

speech_recognizer.start_continuous_recognition()

while not done:
    time.sleep(.5)
    
print("Printing all results:")
print(all_results)

示例输出:

<小时/>

通过函数调用相同的内容

封装在一个函数中并尝试调用它。

只是做了一些调整并封装在一个函数中。确保变量“done”是非本地访问的。请检查并告诉我

import azure.cognitiveservices.speech as speechsdk
import time
import datetime

def speech_to_text():
    
    # Creates an instance of a speech config with specified subscription key and service region.
    # Replace with your own subscription key and region identifier from here: https://aka.ms/speech/sdkregion
    speech_key, service_region = "<>", "<>"
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)

    # Creates an audio configuration that points to an audio file.
    # Replace with your own audio filename.
    audio_filename = "whatstheweatherlike.wav"
    audio_input = speechsdk.audio.AudioConfig(filename=audio_filename)

    # Creates a recognizer with the given settings
    speech_config.speech_recognition_language="en-US"
    speech_config.request_word_level_timestamps()
    speech_config.enable_dictation()
    speech_config.output_format = speechsdk.OutputFormat(1)

    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input)

    #result = speech_recognizer.recognize_once()
    all_results = []



    #https://learn.microsoft.com/en-us/python/api/azure-cognitiveservices-speech/azure.cognitiveservices.speech.recognitionresult?view=azure-python
    def handle_final_result(evt):
        all_results.append(evt.result.text) 
    
    
    done = False

    def stop_cb(evt):
        print('CLOSING on {}'.format(evt))
        speech_recognizer.stop_continuous_recognition()
        nonlocal done
        done= True

    #Appends the recognized text to the all_results variable. 
    speech_recognizer.recognized.connect(handle_final_result) 

    #Connect callbacks to the events fired by the speech recognizer & displays the info/status
    #Ref:https://learn.microsoft.com/en-us/python/api/azure-cognitiveservices-speech/azure.cognitiveservices.speech.eventsignal?view=azure-python   
    speech_recognizer.recognizing.connect(lambda evt: print('RECOGNIZING: {}'.format(evt)))
    speech_recognizer.recognized.connect(lambda evt: print('RECOGNIZED: {}'.format(evt)))
    speech_recognizer.session_started.connect(lambda evt: print('SESSION STARTED: {}'.format(evt)))
    speech_recognizer.session_stopped.connect(lambda evt: print('SESSION STOPPED {}'.format(evt)))
    speech_recognizer.canceled.connect(lambda evt: print('CANCELED {}'.format(evt)))
    # stop continuous recognition on either session stopped or canceled events
    speech_recognizer.session_stopped.connect(stop_cb)
    speech_recognizer.canceled.connect(stop_cb)

    speech_recognizer.start_continuous_recognition()

    while not done:
        time.sleep(.5)
            
    print("Printing all results:")
    print(all_results)

#calling the conversion through a function    
speech_to_text()

关于python - 使用Python中的Azure语音服务读取音频文件并转换为文本，但只有第一句话转换为语音，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/62872929/

文章推荐： python-3.x - 使用Python(3.7)在OpenCV(4.2.0)中检测矩形，

文章推荐： c - 在 C 中赋值之前如何存储表达式值？

javascript - Web 音频/ radio 流客户端 : use Howler. js、 native 音频、其他库？
我一直在为实时流和静态文件(HTTP 上的 MP3)构建网络广播播放器。我选了Howler.js作为规范化 quirks 的后端的 HTML5 Audio (思考:自动播放、淡入/淡出、进度事件)。
vue实现移动端input上传视频、音频
vue移动端input上传视频、音频，供大家参考，具体内容如下 html部分 ?
PHP转换图像+音频=视频
关闭。这个问题需要更多 focused .它目前不接受答案。想改进这个问题？更新问题，使其仅关注一个问题 editing this post . 7年前关闭。 Improve this questi
iphone - 音频/视频编程
我想在我的程序中访问音频和视频。 MAC里面可以吗？我们的程序在 Windows 上运行，我使用 directshow 进行音频/视频编程。但我想在 MAC 中开发相同的东西。有没有像direct
iOS 音频/声音不会在后台模式处于事件状态时在后台播放
我的应用程序(使用 Flutter 制作，但这应该无关紧要)具有类似于计时器的功能，可以定期(10 秒到 3 分钟)发出滴答声。我在我的 Info.plist 中激活了背景模式 Audio、AirPl
javascript - 音频 JavaScript
我是 ionic 2 的初学者我使用了音频文件。 import { Component } from '@angular/core'; import {NavController, Alert
java - 插入声音/音频
我有一个包含ListView和图片的数据库，我想在每个语音数据中包含它们。我已经尝试过，但是有很多错误。以下是我的java和xml。数据库.java package com.example.data
php - 音频/音乐社交网站托管服务
我在zend framework 2上建立了一个音乐社交网络。您可以想象它与SoundCloud相同，用户上传歌曲，其他用户播放它们，这些是网站上的基本操作。我知道将要托管该页面的服务器将需要大量带
android - 音频-Android
我正在尝试在android应用中播放音频，但是在代码中AssetFileDescriptor asset1及其下一行存在错误。这是代码: MediaPlayer mp; @Override prote
wordpress - [音频] WordPress短代码中的网址错误
我对 WordPress Audio Shortcode有问题。我这样使用它: 但是在前面，在HTML代码中我得到了: document.createElement('audio');
matlab - 音频.wav文件的SNR和评估过滤技术的客观措施
我正在做一项关于降低噪音的滤波技术的实验。我在数据集中的样本是音频文件(.wav)，因此，我有:原始录制的音频文件，我将它们与噪声混合，因此变得混合(噪声信号)，我将这些噪声信号通过滤波算法传递，输出
audio - 音频/声音增强的神经网络
一个人会使用哪种类型的神经网络架构将声音映射到其他声音？神经网络擅长学习从序列到其他序列，因此声音增强/生成似乎是它们的一种非常流行的应用(但不幸的是，事实并非如此-我只能找到一个(相当古老的)洋红色
windows - 音频:如何设置默认麦克风的电平？
这个让我抓狂: 在专用于此声音播放/录制应用程序的 Vista+ 计算机上，我需要我的应用程序确保(默认)麦克风电平被推到最大。我该怎么做？我找到了 Core Audio lib ，找到了如何将 I
html - Chrome扩展程序和流式传输<音频>
{ "manifest_version": 2, "name": "Kitten Radio Extension", "description": "Listen while browsi
c# - 音频，FFT不起作用
class Main { WaveFileReader reader; short[] sample; Complex[] tmpComplexArray; publi
android - 音频，平衡2种来源的声音
我正在使用电话录音软件(android)，该软件可以记录2个人在电话中的通话。每个电话的输出是一个音频文件，其中包含来自 call 者和被 call 者的声音。但是，大多数情况下，运行此软件的电话发
javascript - 音频/语音比较和getUserMedia
我正在构建一个需要语音激活命令的Web应用程序。我正在使用getUserMedia作为音频输入。对于语音激活命令，该过程是用户将需要通过记录其语音来“校准”命令。例如，对于“停止”命令，用户将说出“
cordova - 在PouchDB中存储视频/音频
我正在开发一个Cordova应用程序，并将PouchDB用作数据库，当连接可用时，它将所有信息复制到CouchDB。我成功存储了简单的文本和图像。我一直在尝试存储视频和音频，但是没有运气。我存储
audio - 音频.MP3在Safari浏览器中不起作用
我正在开发web application，我必须在其中使用.MP3的地方使用播放声音，但是会发生问题。声音为play good in chrome, Firefox，但为safari its not
audio - 音频:软件中的位深度减少
如何减少音频文件的位深？是否忽略了MSB或LSB？两者混合吗？ (旁问:这叫什么？) 最佳答案 TL / DR:将音频曲线高度变量右移至较低位深度可以将音频视为幅度(Y轴)随时间(X轴)的模拟曲线。

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

python - 使用Python中的Azure语音服务读取音频文件并转换为文本，但只有第一句话转换为语音