numpy - Scipy、Numpy : Audio classifier, 语音/语音事件检测-6ren

numpy - Scipy、Numpy : Audio classifier, 语音/语音事件检测

转载作者：行者123 更新时间：2023-11-30 08:44:18

25

4

我正在编写一个程序来自动对录制的音频电话文件(wav 文件)进行分类，其中至少包含一些人声(仅 DTMF、拨号音、铃声、噪音)。

我的第一个方法是使用 ZCR(过零率)和计算能量来实现简单的 VAD(语音事件检测器)，但这两个参数都会混淆 DTMF、拨号音和语音。这项技术失败了，所以我实现了一种简单的方法来计算 200Hz 和 300Hz 之间的 FFT 方差。我的numpy代码如下

wavefft = np.abs(fft(frame))
n = len(frame)
fx = np.arange(0,fs,float(fs)/float(n))
stx = np.where(fx>=200)
stx = stx[0][0]
endx = np.where(fx>=300)
endx = endx[0][0]
return np.sqrt(np.var(wavefft[stx:endx]))/1000

这导致了 60% 的准确率。

接下来，我尝试使用 SVM(支持向量机)和 MFCC(梅尔频率倒谱系数)来实现基于机器学习的方法。结果完全不正确，几乎所有 sample 都被错误标记。应该如何使用 MFCC 特征向量训练 SVM？我使用scikit-learn的粗略代码如下

[samplerate, sample] = wavfile.read ('profiles/noise.wav')
noiseProfile = MFCC(samplerate, sample)
[samplerate, sample] = wavfile.read ('profiles/ring.wav')
ringProfile =  MFCC(samplerate, sample)
[samplerate, sample] = wavfile.read ('profiles/voice.wav')
voiceProfile = MFCC(samplerate, sample)

machineData = []
for noise in noiseProfile:
    machineData.append(noise)

for voice in voiceProfile:
    machineData.append(voice)

dataLabel = []
for i in range(0, len(noiseProfile)):
    dataLabel.append (0)
for i in range(0, len(voiceProfile)):
    dataLabel.append (1)

clf = svm.SVC()
clf.fit(machineData, dataLabel)

我想知道我可以实现什么替代方法？

最佳答案

如果您不必使用 scipy/numpy，您可以查看 webrtvad ，这是 Google 优秀的 Python 包装器 WebRTC语音事件检测代码。 WebRTC 使用高斯混合模型 (GMM)，效果很好，而且速度非常快。

以下是如何使用它的示例:

import webrtcvad

# audio must be 16 bit PCM, at 8 KHz, 16 KHz or 32 KHz.
def audio_contains_voice(audio, sample_rate, aggressiveness=0, threshold=0.5):
    # Frames must be 10, 20 or 30 ms.
    frame_duration_ms = 30

    # Assuming split_audio is a function that will split audio into
    # frames of the correct size.
    frames = split_audio(audio, sample_rate, frame_duration)

    # aggressiveness tells the VAD how aggressively to filter out non-speech.
    # 0 will have the most false-positives for speech, 3 the least.
    vad = webrtc.Vad(aggressiveness)

    num_voiced = len([f for f in frames if vad.is_voiced(f, sample_rate)])
    return float(num_voiced) / len(frames) > threshold

关于numpy - Scipy、Numpy : Audio classifier, 语音/语音事件检测，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/30409539/

25

4

0

文章推荐： javascript - jQuery $.getJSON 不返回数据

Scipy 和 CX_freeze - 导入 scipy : you cannot import scipy while being in scipy source directory 时出错
我在使用 cx_freeze 和 scipy 时无法编译 exe。特别是，我的脚本使用 from scipy.interpolate import griddata 构建过程似乎成功完成，但是当我尝试
scipy - SciPy 中由函数定义的稀疏矩阵
是否可以通过函数在 scipy 中定义一个稀疏矩阵，而不是列出所有可能的值？在文档中，我看到可以通过以下方式创建稀疏矩阵 There are seven available sparse matrix
scipy - SciPy:Minimumsq与Minimum_squares
SciPy为非线性最小二乘问题提供了两种功能： optimize.leastsq()仅使用Levenberg-Marquardt算法。 optimize.least_squares()允许我们选择Le
scipy - SciPy 中的复杂求解器
SciPy 中的求解器能否处理复数值(即 x=x'+i*x")？我对使用 Nelder-Mead 类型的最小化函数特别感兴趣。我通常是 Matlab 用户，我知道 Matlab 没有复杂的求解器。如果
scipy - 如何使用 scipy 计算三次样条插值的导数？
我有看起来像这样的数据集: position number_of_tag_at_this_position 3 4 8 6 13 25 23 12 我想对这个数据集应用三次样条插值来插值标签密度；为此
scipy - 如何使用 Scipy 处理巨大的稀疏矩阵构造？
所以，我正在处理维基百科转储，以计算大约 5,700,000 个页面的页面排名。这些文件经过预处理，因此不是 XML 格式。它们取自 http://haselgrove.id.au/wikipedi
scipy - 在 scipy 中获取非归一化特征向量
Scipy 和 Numpy 返回归一化的特征向量。我正在尝试将这些向量用于物理应用程序，我需要它们不被标准化。例如a = np.matrix('-3, 2; -1, 0') W,V = spl.ei
scipy - 有没有办法将 scipy.optimize.fsolve 与 jit_integrand_function 和 scipy.integrate.quad 一起使用？
基于此处提供的解释 1 ，我正在尝试使用相同的想法来加速以下积分: import scipy.integrate as si from scipy.optimize import root, fsol
scipy - 导入 scipy 或 scipy.signal 时 Pyinstaller --onefile 警告 pyconfig.h
这很容易重新创建。如果我的脚本 foo.py 是: import scipy 然后运行: python pyinstaller.py --onefile foo.py 当我启动 foo.exe 时，
python - 为什么 from scipy import spatial 有效，而 scipy.spatial 在 import scipy 后不起作用？
我想在我的代码中使用 scipy.spatial.distance.cosine。如果我执行类似 import scipy.spatial 或 from scipy import spatial 的操
scipy - 如何使用 scipy.integrate.quadpack(或 scipy 中的其他 c/fortran)直接作为来自 cython 的 c
Numpy 有一个基本的 pxd，声明它的 c 接口(interface)到 cython。是否有用于 scipy 组件(尤其是 scipy.integrate.quadpack)的 pxd？或者，
scipy - 理解 scipy.stats.chisquare
有人可以帮我处理 scipy.stats.chisquare 吗？我没有统计/数学背景，我正在使用来自 https://en.wikipedia.org/wiki/Chi-squared_test 的
scipy - 如何使用 scipy.odr 估计拟合优度？
我正在使用 scipy.odr 拟合数据与权重，但我不知道如何获得拟合优度或 R 平方的度量。有没有人对如何使用函数存储的输出获得此度量有建议？最佳答案 res_var Output 的属性是所谓的
scipy - pip 无法为 scipy 构建轮子
我刚刚下载了新的 python 3.8，我正在尝试使用以下方法安装 scipy 包: pip3.8 install scipy 但是构建失败并出现以下错误: **Failed to build sci
scipy - 如何使用带有自己的三角测量的 scipy.interpolate.LinearNDInterpolator
我有 my own triangulation algorithm它基于 Delaunay 条件和梯度创建三角剖分，使三角形与梯度对齐。这是一个示例输出: 以上描述与问题无关，但对于上下文是必要的。
scipy - scipy.stats.norm 上下文中的概率密度函数是什么？
这是一个非常基本的问题，但我似乎找不到好的答案。 scipy 到底计算什么内容 scipy.stats.norm(50,10).pdf(45) 据我了解，平均值为 50、标准差为 10 的高斯中像 4
scipy - 在 Scipy.signal 中拟合传递函数模型
我正在使用 curve_fit 来拟合一阶动态系统的阶跃响应，以估计增益和时间常数。我使用两种方法。第一种方法是在时域中拟合从函数生成的曲线。 # define the first order dyn
scipy - 使用 scipy.stats 计算条件期望
让我们假设 x ~ Poisson(2.5);我想计算类似 E(x | x > 2) 的东西。我认为这可以通过 .dist.expect 运算符来完成，即: D = stats.poisson(2.
scipy - 区分 OpenMDAO SciPy SLSQP 中的迭代和函数评估
我正在通过 OpenMDAO 使用 SLSQP 来解决优化问题。优化工作充分；最后的 SLSQP 输出如下: Optimization terminated successfully. (Exi
python - Scipy 最小化/Scipy 曲线拟合/lmfit
log( VA ) = gamma - (1/eta)log[alpha L ^(-eta) + 测试版 K ^(-eta)] 我试图用非线性最小二乘法估计上述函数。我为此使用了 3 个不同的包(Sc

首页

博学

6Ren·AI

商城

numpy - Scipy、Numpy : Audio classifier, 语音/语音事件检测