使用 NumPy 数据类型的 Python 字典查找速度-6ren

使用 NumPy 数据类型的 Python 字典查找速度

转载作者：太空狗更新时间：2023-10-29 22:13:16

46

4

背景

我在 NumPy 数组中有很多数字消息代码，我需要快速将它们转换为字符串。我在性能方面遇到了一些问题，想了解原因以及如何快速解决。

一些基准

I - 简单的方法

import numpy as np

# dictionary to use as the lookup dictionary
lookupdict = {
     1: "val1",
     2: "val2",
    27: "val3",
    35: "val4",
    59: "val5" }

# some test data
arr = np.random.choice(lookupdict.keys(), 1000000)

# create a list of words looked up
res = [ lookupdict[k] for k in arr ]

字典查找占用了我咖啡休息时间的大部分时间，758 毫秒。 (我也试过 res = map(lookupdict.get, arr) 但那更糟。)

II - 没有 NumPy

import random

# dictionary to use as the lookup dictionary
lookupdict = {
     1: "val1",
     2: "val2",
    27: "val3",
    35: "val4",
    59: "val5" }

# some test data
arr = [ random.choice(lookupdict.keys()) for _ in range(1000000) ]

# create a list of words looked up
res = [ lookupdict[k] for k in arr ]

计时结果变化相当大，为 76 毫秒!

应该注意的是，我对查找的时间很感兴趣。随机生成只是为了创建一些测试数据。花不花时间没意思。此处给出的所有基准测试结果仅针对一百万次查找。

III - 将 NumPy 数组转换为列表

我的第一个猜测是这与列表与数组问题有关。但是，通过修改 NumPy 版本以使用列表:

res = [ lookupdict[k] for k in list(arr) ]

给了我 778 毫秒，其中大约 110 毫秒用于转换列表，570 毫秒用于查找。所以，查找要快一点，但总时间是一样的。

IV - 从np.int32到int的类型转换

由于唯一的其他区别似乎是数据类型(np.int32 与 int)，我尝试即时转换类型。这有点愚蠢，因为 dict 可能也是这样做的:

res = [ lookupdict[int(k)] for k in arr ]

然而，这似乎做了一些有趣的事情，因为时间下降到 266 毫秒。似乎几乎但不完全相同的数据类型在字典查找中玩弄了讨厌的把戏，而且 dict 代码在转换方面不是很有效。

V - 字典键转换为 np.int32

为了对此进行测试，我修改了 NumPy 版本以在字典键和查找中使用完全相同的数据类型:

import numpy as np

# dictionary to use as the lookup dictionary
lookupdict = {
     np.int32(1): "val1",
     np.int32(2): "val2",
    np.int32(27): "val3",
    np.int32(35): "val4",
    np.int32(59): "val5" }

# some test data
arr = np.random.choice(lookupdict.keys(), 1000000)

# create a list of words looked up
res = [ lookupdict[k] for k in arr ]

这提高到 177 毫秒。这不是微不足道的改进，而是与 76 毫秒相去甚远。

VI - 使用 int 对象的数组转换

import numpy as np

# dictionary to use as the lookup dictionary
lookupdict = {
     1: "val1",
     2: "val2",
    27: "val3",
    35: "val4",
    59: "val5" }

# some test data
arr = np.array([ random.choice(lookupdict.keys()) for _ in range(1000000) ], 
               dtype='object')

# create a list of words looked up
res = [ lookupdict[k] for k in arr ]

这给出了 86 毫秒，这已经非常接近原生 Python 的 76 毫秒。

结果总结

dict 键 int，使用 int 进行索引(原生 Python):76 毫秒
字典键 int，使用 int 对象进行索引 (NumPy):86 毫秒
dict 键 np.int32，使用 np.int32 进行索引:177 毫秒
字典键 int，使用 np.int32 进行索引:758 毫秒

问题(S)

为什么？我该怎么做才能尽可能快地进行字典查找？我的输入数据是一个 NumPy 数组，所以到目前为止最好的(最快但丑陋的)是将字典键转换为 np.int32。 (不幸的是，dict 键可能分布在很大范围的数字上，因此逐个数组索引不是一个可行的替代方案。虽然很快，10 毫秒。)

最佳答案

正如您所怀疑的，它是 int32.__hash__ 的错，x11 和 int.__hash__ 一样慢:

%timeit hash(5)
10000000 loops, best of 3: 39.2 ns per loop
%timeit hash(np.int32(5))
1000000 loops, best of 3: 444 ns per loop

(int32 类型是用 C 语言实现的。如果你真的很好奇，你可以深入研究源代码，找出它在那里做了什么，花了这么长时间)。

编辑:

第二个减慢速度的部分是字典查找中的隐式==比较:

a = np.int32(5)
b = np.int32(5)
%timeit a == b  # comparing two int32's
10000000 loops, best of 3: 61.9 ns per loop
%timeit a == 5  # comparing int32 against int -- much slower
100000 loops, best of 3: 2.62 us per loop

这解释了为什么您的 V 比 I 和 IV 快得多。当然，坚持使用全 int 解决方案会更快。

在我看来，您有两个选择:

坚持使用纯 int 类型，或者在 dict-lookup 之前转换为 int
如果最大代码值不是太大，和/或内存不是问题，您可以将字典查找换成列表索引，这不需要散列ing。

例如:

lookuplist = [None] * (max(lookupdict.keys()) + 1)
for k,v in lookupdict.items():
    lookuplist[k] = v

res = [ lookuplist[k] for k in arr ] # using list indexing

(编辑:您可能还想在这里试验 np.choose)

关于使用 NumPy 数据类型的 Python 字典查找速度，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/24705644/

46

4

0

文章推荐： python - 在python中查找一组字符串的最小汉明距离

文章推荐： c# - 从另一个更新一个依赖属性

文章推荐： c# - 如何就地刷新组合框项目？

文章推荐： c# - MSIL 问题(基本)

haskell - 类型家庭类型黑客
我正在尝试编写一个相当多态的库。我遇到了一种更容易表现出来却很难说出来的情况。它看起来有点像这样: {-# LANGUAGE ScopedTypeVariables #-} {-# LANGUAGE
javascript - 这是如何运作的？类型 = 类型 || 'any' ;
谁能解释一下这个表达式是如何工作的？ type = type || 'any'; 这是否意味着如果类型未定义则使用“任意”？最佳答案如果 type 为“falsy”(即 false，或 undef
f# - 类型 'obj' 不是接口(interface)类型
我有一个界面，在IAnimal.fs中， namespace Kingdom type IAnimal = abstract member Eat : Food -> unit 以及另一个成功
c++ - 类型(变量)与(类型)变量
这个问题在这里已经有了答案: 关闭 10 年前。 Possible Duplicate: What is the difference between (type)value and type(va
c# - 默认(可空(类型))与默认(类型)
在 C# 中，default(Nullable) 之间有区别吗？ (或 default(long?) )和 default(long) ？ Long只是一个例子，它可以是任何其他struct类型。最
scala - 如何定义一个 HList 类型，但基于另一个 HList 类型
假设我有一个案例类: case class Foo(num: Int, str: String, bool: Boolean) 现在我还有一个简单的包装器: sealed trait Wrapper[
c# - 如何在运行时定义委托(delegate)类型(即动态委托(delegate)类型)
这个问题在这里已经有了答案: Create C# delegate type with ref parameter at runtime (1 个回答) 关闭 2 年前。为了即时创建委托(dele
python - dct 中的断言失败(类型 == CV_32FC1 || 类型 == CV_64FC1)
我正在尝试获取图像的 dct。一开始我遇到了错误 The function/feature is not implemented (Odd-size DCT's are not implemented
ios - PList 类型？应用程序/x-plist 类型
我正在尝试使用 AFNetworking 的 AFPropertyListRequestOperation，但是当我尝试下载它时，出现错误预期的内容类型{( “应用程序/x-plist” )}, 得
javascript - 元素隐式具有 'any' 类型，因为索引表达式不是 'number' 类型
我在下面收到错误。我知道这段代码的意思，但我不知道界面应该是什么样子: Element implicitly has an 'any' type because index expression is
swift2 - 类型 'Error' 约束为非协议(protocol)类型，即使类型是协议(protocol)
我尝试将 SignalType 从 ReactiveCocoa 扩展为自定义 ErrorType，代码如下所示 enum MyError: ErrorType { // .. cases }
scala - 如何使用 Scala 的 this 类型、抽象类型等来实现 Self 类型？
我无法在任何其他问题中找到答案。假设我有一个抽象父类(super class) Abstract0，它有两个子类 Concrete1 和 Concrete1。我希望能够在 Abstract0 中定义类
日期时间字段上的 MySQL 索引不是 RANGE 类型，而是使用 INDEX 类型
我想知道为什么这个索引没有用在 RANGE 类型中，而是用在 INDEX 中: 索引: CREATE INDEX myindex ON orders(order_date); 查询: EXPLAIN
java - IncompleteClassChangeError ...原本应该是 direct 类型，但结果却发现是 virtual 类型
我正在使用 RxJava，现在我尝试通过提供 lambda 来订阅可观察对象: observableProvider.stringForKey(CURRENT_DELETED_ID) .sub
javascript - MIME 类型 ('text/html' ) 不是受支持的样式表 MIME 类型
我已经尝试了几乎所有解决问题的方法，其中包括。为提供类型使用app.use(express.static('public'))还有更多，但我似乎无法为此找到解决方案。 index.js : imp
css - 哪个更快？输入[类型 ="submit"] 或 [类型 ="submit"]
以下哪个 CSS 选择器更快？ input[type="submit"] { /* styles */ } 或 [type="submit"] { /* styles */ } 只是好
java - 通过构造函数参数表达的不满足的依赖关系，索引为 0 类型 -> 没有合格的 bean 类型
我不知道这个设置有什么问题，我在 IDEA 中获得了所有注释(@Controller、@Repository、@Service)，它在行号左侧显示 bean，然后转到该 bean。这是错误: 14-
c - jni 回调适用于 java 类型，但不适用于 c 类型
我听从了建议 registering java function as a callback in C function并且可以使用“简单”类型(例如整数和字符串)进行回调，例如: jstring j
java - 将 java 类型 string[] 映射到 oracle 类型
有一些 java 类，加载到 Oracle 数据库(版本 11g)和 pl/sql 函数包装器: create or replace function getDataFromJava( in_uLis
javascript - 元素隐式具有 'any' 类型，因为索引表达式不是 'number' 类型 [7015]
我已经从 David Walsh 的 css 动画回调中获取代码并将其修改为 TypeScript。但是，我收到一个错误，我不知道为什么: interface IBrowserPrefix { [

首页

博学

6Ren·AI

商城

使用 NumPy 数据类型的 Python 字典查找速度