python - Pandas 哈希表 KeyError-6ren

python - Pandas 哈希表 KeyError

转载作者：太空宇宙更新时间：2023-11-03 15:12:58

25

4

我在 Kaggle 中找到了以下代码。

import re

from nltk.corpus import stopwords # Import the stop word list

def description_to_words(review_text):

    # 2. Remove non-letters        
    letters_only = re.sub("[^a-zA-Z]", " ", review_text)
    # 3. Convert to lower case, split into individual words
    words = letters_only.lower().split()
    # 4. In Python, searching a set is much faster than searching
    #   a list, so convert the stop words to a set
    stops = set(stopwords.words("english"))
    # 5. Remove stop words
    meaningful_words = [w for w in words if not w in stops]
    # 6. Join the words back into one string separated by space, 
    # and return the result.
    return( " ".join( meaningful_words ))

上面的代码与下面的函数调用配合良好

clean_review = description_to_words(df['MaterialDescription'][3] )
print(clean_review)

但是当我尝试上述相同的操作(例如将 DataFrame 分配给另一个变量时，如下所示)，

X = df['MaterialDescription']
clean_review = description_to_words(X[3] )
print(clean_review)

我收到以下错误，这非常荒谬。我确信我需要对 Pandas 进行一些澄清

---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\indexes\base.py in get_loc(self, key, method, tolerance)
   2133             try:
-> 2134                 return self._engine.get_loc(key)
   2135             except KeyError:

pandas\index.pyx in pandas.index.IndexEngine.get_loc (pandas\index.c:4433)()

pandas\index.pyx in pandas.index.IndexEngine.get_loc (pandas\index.c:4279)()

pandas\src\hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas\hashtable.c:13742)()

pandas\src\hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas\hashtable.c:13696)()

KeyError: 3

During handling of the above exception, another exception occurred:

KeyError                                  Traceback (most recent call last)
<ipython-input-15-5c63f93c009a> in <module>()
----> 1 clean_review = description_to_words(X[3] )
      2 print(clean_review)

C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\core\frame.py in __getitem__(self, key)
   2057             return self._getitem_multilevel(key)
   2058         else:
-> 2059             return self._getitem_column(key)
   2060 
   2061     def _getitem_column(self, key):

C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\core\frame.py in _getitem_column(self, key)
   2064         # get column
   2065         if self.columns.is_unique:
-> 2066             return self._get_item_cache(key)
   2067 
   2068         # duplicate columns & possible reduce dimensionality

C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\core\generic.py in _get_item_cache(self, item)
   1384         res = cache.get(item)
   1385         if res is None:
-> 1386             values = self._data.get(item)
   1387             res = self._box_item_values(item, values)
   1388             cache[item] = res

C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\core\internals.py in get(self, item, fastpath)
   3541 
   3542             if not isnull(item):
-> 3543                 loc = self.items.get_loc(item)
   3544             else:
   3545                 indexer = np.arange(len(self.items))[isnull(self.items)]

C:\Anaconda\envs\tensorflow\lib\site-packages\pandas\indexes\base.py in get_loc(self, key, method, tolerance)
   2134                 return self._engine.get_loc(key)
   2135             except KeyError:
-> 2136                 return self._engine.get_loc(self._maybe_cast_indexer(key))
   2137 
   2138         indexer = self.get_indexer([key], method=method, tolerance=tolerance)

pandas\index.pyx in pandas.index.IndexEngine.get_loc (pandas\index.c:4433)()

pandas\index.pyx in pandas.index.IndexEngine.get_loc (pandas\index.c:4279)()

pandas\src\hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas\hashtable.c:13742)()

pandas\src\hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas\hashtable.c:13696)()

KeyError: 3

我也尝试过在 pandas 对象中给出切片。这又产生了另一个错误

clean_review = description_to_words(X[:3] )
print(clean_review)

以下是上述两行代码的堆栈跟踪

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
<ipython-input-18-f8b01af18e4b> in <module>()
----> 1 clean_review = description_to_words(X[:3] )
      2 print(clean_review)

<ipython-input-6-70647cd7caba> in description_to_words(review_text)
      6 
      7     # 2. Remove non-letters
----> 8     letters_only = re.sub("[^a-zA-Z]", " ", review_text)
      9     # 3. Convert to lower case, split into individual words
     10     words = letters_only.lower().split()

C:\Anaconda\envs\tensorflow\lib\re.py in sub(pattern, repl, string, count, flags)
    180     a callable, it's passed the match object and must return
    181     a replacement string to be used."""
--> 182     return _compile(pattern, flags).sub(repl, string, count)
    183 
    184 def subn(pattern, repl, string, count=0, flags=0):

TypeError: expected string or bytes-like object

如果有人帮助我理解这里到底发生了什么，那将会有很大的帮助。

最佳答案

以下几行

X = df['MaterialDescription']
clean_review = description_to_words(X[3] )

用Python给出description_to_words(df['MaterialDescription'][3])

您必须通过以下方式找到您的索引:

clean_review = description_to_words(df.iloc[3]['MaterialDescription'] )

关于python - Pandas 哈希表 KeyError，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/44085633/

25

4

0

文章推荐： c# - 将每个对象的特定变量显示到组合框中

文章推荐： c# - 绘制和更新 PictureBox

文章推荐： c# - 单击按钮时出现 IndexOutOfRangeException

regex - Grep 所有不以#(哈希)或贪心空格和#(哈希)开头的行
我正在尝试 grep conf 文件中所有不以开头的有效行哈希(或) 任意数量的空格(0 个或多个)和一个散列下面的正则表达式似乎不起作用。 grep ^[^[[:blank:]]*#] /op
带斜线的 Laravel 哈希
我正在使用哈希通过 URL 发送 protected 电子邮件以激活帐户 Hash::make($data["email"]); 但是哈希结果是 %242y%2410%24xaiB/eO6knk8sL
来自文本文件的 Perl 哈希
我是 Perl 的新手，正在尝试从文本文件创建散列。我有一个代码外部的文本文件，旨在供其他人编辑。前提是他们应该熟悉 Perl 并且知道在哪里编辑。文本文件本质上包含几个散列的散列，具有正确的语法、缩
perl 哈希 - 比较键和值
我一直在阅读 perl 文档，但我不太了解哈希。我正在尝试查找哈希键是否存在，如果存在，则比较其值。让我感到困惑的是，我的搜索结果表明您可以通过 if (exists $files{$key}) 找到
当键和值都是数组引用时的 Perl 哈希
我遇到了数字对映射到其他数字对的问题。例如，(1,2)->(12,97)。有些对可能映射到多个其他对，所以我真正需要的是将一对映射到列表列表的能力，例如 (1,2)->((12,97),(4,1))。
Mustache:从模板中检索标签列表/哈希？
我见过的所有 Mustache 文档和示例都展示了如何使用散列来填充模板。我有兴趣去另一个方向。 EG，如果我有这个: Hello {{name}} mustache 能否生成这个(伪代码): tag
hash - ColdFusion 哈希
我正在尝试使用此公式创建密码摘要以获取以下变量，但我的代码不匹配。不确定我做错了什么，但当我需要帮助时我会承认。希望有人在那里可以提供帮助。文档中的公式:Base64(SHA1(NONCE + TI
arrays - 遍历数据数组/哈希
我希望遍历我传递给定路径的这些数据结构(基本上是目录结构)。目标是列出根/基本路径，然后列出所有子 path s 如果它们存在并且对于每个子 path存在，列出 file从那个子路径。我知道这可能
子函数的 Perl 哈希
我希望有一个包含对子函数的引用的散列，我可以在其中根据用户定义的变量调用这些函数，我将尝试给出我正在尝试做的事情的简化示例。 my %colors = ( vim => setup_vim()
vim - 为什么写入文件会更改内容(哈希)？
我注意到，在使用 vim 将它们复制粘贴到文件中后尝试生成一些散列时，散列不是它应该的样子。打开和写出文件时相同。与 nano 的行为相同，所以一定有我遗漏的地方。 $ echo -n "foo"
perl - 为什么我们不能在列表上下文中初始化状态数组/哈希？
数组和散列作为状态变量存在限制。从 Perl 5.10 开始，我们无法在列表上下文中初始化它们: 所以 state @array = qw(a b c); #Error! 为什么会这样？为什么这是不允
Varnish vcl_backend_response检测vcl_recv返回(哈希)
在端口 80 上使用 varnish 5.1 的多网站设置中，我不想缓存所有域。这在 vcl_recv 中很容易完成。 if ( req.http.Host == "cache.this.domai
Django 管道缓存破坏不更新缓存文件/哈希
基本上，缓存破坏文件上的哈希不会更新。 class S3PipelineStorage(PipelineMixin, CachedFilesMixin, S3BotoStorage): pa
eclipse - 调试Dart应用程序时变量的唯一ID(哈希？)
eclipse dart插件在“变量” View 中显示如下内容: 在“值”列中可见的“id”是什么意思？ “id”是唯一的吗？在调试期间，如何确定两个实例是否相同？我是否需要在所有类中重写toStr
arrays - 将相同类型的命令行参数读入Powershell中的数组/哈希
如何将Powershell中的命令行参数读入数组？就像是 myprogram -file file1 -file file2 -file file3 然后我有一个数组 [file1,file2,fil
用于安全支付网关的 coldfusion 哈希
我正尝试在 coldfusion 中为我们的安全支付网关创建哈希密码以接受交易。很遗憾，支付网关拒绝接受我生成的哈希值。表单发送交易的所有元素，并发送基于五个不同字段生成的哈希值。在 PHP 中
Ruby - 哈希 - 组合
例如，我有一个包含 5 个元素的哈希: my_hash = {a: 'qwe', b: 'zcx', c: 'dss', d: 'ccc', e: 'www' } 我的目标是每次循环哈希时都返回，但没
哈希问题的 Perl 哈希
我在这里看到了令人作呕的类似问题，但没有一个能具体回答我自己的问题。我正在尝试以编程方式创建哈希的哈希。我的问题代码如下: my %this_hash = (); if ($user_hash{$u
用于安全支付网关的 coldfusion 哈希
我正尝试在 coldfusion 中为我们的安全支付网关创建哈希密码以接受交易。很遗憾，支付网关拒绝接受我生成的哈希值。表单发送交易的所有元素，并发送基于五个不同字段生成的哈希值。在 PHP 中
Java 哈希(简单)
这个问题已经有答案了: Java - how to convert letters in a string to a number? (9 个回答) 已关闭 7 年前。我需要一种简短的方法将字符串转

首页

博学

6Ren·AI

商城

python - Pandas 哈希表 KeyError