python - 如何在 textacy 0.6.2 中初始化 `Doc`？-6ren

python - 如何在 textacy 0.6.2 中初始化 `Doc`？

转载作者：行者123 更新时间：2023-11-28 18:09:44

25

4

试图跟随 simple Doc initialization in the docs在 Python 2 中不起作用:

>>> import textacy
>>> content = '''
...     The apparent symmetry between the quark and lepton families of
...     the Standard Model (SM) are, at the very least, suggestive of
...     a more fundamental relationship between them. In some Beyond the
...     Standard Model theories, such interactions are mediated by
...     leptoquarks (LQs): hypothetical color-triplet bosons with both
...     lepton and baryon number and fractional electric charge.'''
>>> metadata = {
...     'title': 'A Search for 2nd-generation Leptoquarks at √s = 7 TeV',
...     'author': 'Burton DeWilde',
...     'pub_date': '2012-08-01'}
>>> doc = textacy.Doc(content, metadata=metadata)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/textacy/doc.py", line 120, in __init__
    {compat.unicode_, SpacyDoc}, type(content)))
ValueError: `Doc` must be initialized with set([<type 'unicode'>, <type 'spacy.tokens.doc.Doc'>]) content, not "<type 'str'>"

对于字符串或字符串序列，简单的初始化应该是什么样的？

更新:

将 unicode(content) 传递给 textacy.Doc() 吐出

ImportError: 'cld2-cffi' must be installed to use textacy's automatic language detection; you may do so via 'pip install cld2-cffi' or 'pip install textacy[lang]'.

从安装 textacy 的那一刻起，我会很高兴。

即使在安装 cld2-cffi 之后，尝试上面的代码也会失败

Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/textacy/doc.py", line 114, in __init__
    self._init_from_text(content, metadata, lang)
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/textacy/doc.py", line 136, in _init_from_text
    spacy_lang = cache.load_spacy(langstr)
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/cachetools/__init__.py", line 46, in wrapper
    v = func(*args, **kwargs)
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/textacy/cache.py", line 99, in load_spacy
    return spacy.load(name, disable=disable)
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/spacy/__init__.py", line 21, in load
    return util.load_model(name, **overrides)
  File "/Users/a/anaconda/envs/env1/lib/python2.7/site-packages/spacy/util.py", line 120, in load_model
    raise IOError("Can't find model '%s'" % name)
IOError: Can't find model 'en'

最佳答案

如回溯中所示，该问题位于 textacy/doc.py。在 _init_from_text() 函数中，该函数尝试检测语言并在第 136 行使用字符串 'en' 调用它。(spacy 存储库在 this issue comment. 中提到了这一点)

我通过提供有效的 lang (unicode) 字符串 u'en_core_web_sm' 并在 content 和 lang 参数字符串。

import textacy

content = u'''
    The apparent symmetry between the quark and lepton families of
    the Standard Model (SM) are, at the very least, suggestive of
    a more fundamental relationship between them. In some Beyond the
    Standard Model theories, such interactions are mediated by
    leptoquarks (LQs): hypothetical color-triplet bosons with both
    lepton and baryon number and fractional electric charge.'''

metadata = {
    'title': 'A Search for 2nd-generation Leptoquarks at √s = 7 TeV',
    'author': 'Burton DeWilde',
    'pub_date': '2012-08-01'}

doc = textacy.Doc(content, metadata=metadata, lang=u'en_core_web_sm')

字符串而不是 unicode 字符串(带有神秘的错误消息)改变了行为，缺少包的事实，以及使用 spacy 的可能过时/可能不全面的方式语言字符串对我来说都像是错误。 🤷‍♂️

关于python - 如何在 textacy 0.6.2 中初始化 `Doc`？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/51431112/

25

4

0

文章推荐： Python 跨多列返回第一个非零值

文章推荐： python - 再次出现UnicodeEncodeError(ascii codec无法编码)

文章推荐： javascript - 范围未定义，调用 Angularjs 模块中的函数

文章推荐： python - SqlAlchemy + Firebird + FDB 的 UnicodeError

c# - 为什么 "test user-doc.doc"==> TESTUS~1.DOC？
我编写了一个 c# 程序，并在未安装 MS-Office 的 PC 中将其与文件扩展名(如 DOC)相关联。然后，我双击名称中包含空白字符的任何文件，我的程序将启动以打开该文件。我使用了以下语句: s
google-docs - 如何使用 Google Docs API 编辑 Google Docs 标题？
我试过创建、批量更新、从 https://developers.google.com/docs/api/how-tos/overview 获取. 即使在 batchUpdate 中，我也看不到编辑 t
linux - 在 Linux 中运行 ls doc*.txt 和 ls doc?*.txt 和 ls doc*?.txt 有什么不同？
关闭。这个问题不符合Stack Overflow guidelines .它目前不接受答案。这个问题似乎不是关于 a specific programming problem, a softwar
google-docs - Google Docs API - 更新链接表格
我正在尝试使用新 API 更新 Google 文档中的表格。表格链接自 Google 表格。我尝试了谷歌云中的 API 资源管理器。我能够以 json 格式提取文档，然后过滤出表格。但是在表 jso
google-docs - Google Docs API - 模拟用户文件下载
将 Google Docs Java API 与 Google Apps 帐户一起使用，是否可以模拟用户并下载文件？当我运行下面的程序时，它显然是登录到域并冒充用户，因为它检索其中一个文件的详细信息
api-doc - 如何在 api-doc 中设置数组响应？
我试图通过 apidoc 生成 API 文档如果我的回应是一个数组 [ {"id" : 1, "name" : "John"}, {"id" : 2, "name" : "Mary"}
google-docs-api - 无需身份验证的 Google Docs API
是否可以在没有身份验证的情况下在 Google Docs 中查询公开共享的用户文档？我正在寻找的特定最终目标是能够提供用户 ID，然后列出所有公开共享的文档，并在集合中带有特定标记。谢谢。最佳答
elasticsearch - 在elasticsearch中，/doc/_mapping和/doc {“mappings”之间有什么区别……}
我对Elasticsearch映射感到困惑首先，我创建了一个带有映射请求的文档 PUT /person { "mappings":{ "properties":{ "firs
google-docs - Google Doc Query 在一张表中工作，但在另一张表中给出解析错误
我有一个可在一个电子表格中运行的 Google 文档查询。但是，当我复制电子表格时，查询不起作用，并且收到解析错误:无法解析函数 QUERY 参数 2 的查询字符串:NO_COLUMNCol2。我的
java - 如何使用现有 XML DOC 的属性创建新的 XML DOC？
我有一个如下所示的 XML 文档: _1 _2 TASK _3 TASK 我必须使用第一个文档中的节点属性创建另一
read-the-docs - 如何找到 Read-the-docs 项目的 PDF 版本
我没有看到什么？ RTD features页面说: PDF Generation When you build your project on RTD, we automatically build
google-docs - 嵌入式 Google Docs PDF 查看器显示登录页面而不是 PDF
我有一个网页，我在 iFrame 中嵌入了一个 Google 文档查看器 (其中 URL-encoded-URL 是实际编码的 URL)。对于我的许多/大多数用户，Google PDF 文档查看器
google-docs - 在 asp.net 应用程序中使用 google docs
我如何在我的项目中使用 GOOGLE DOCS，我正在使用 asp.net 和 C# 作为后面的代码。基本上我需要在浏览器中以只读形式显示一些 pdf、doc、dox、excel 文档。提前致谢
google-docs-api - 如何使用 Google Docs API 缩进项目符号列表
从看起来像的 Google Doc 开始: * Item 我希望进行一系列 API 调用以将文档转换为: * Item - Subitem 但是，我不知道如何使用 API 做到这一点。 Crea
google-docs - 使用 JavaScript 控制 Google Docs 嵌入式查看器
我需要控制我网站中嵌入的 Google 文档查看器。更具体地说，我需要能够启用/禁用 Google 幻灯片 View 的控件，并能够使用 JavaScript 启动/停止演示文稿。我无法为此找到任何
google-docs - 如何使用 Google Docs API 添加页眉/页脚
我想使用 Google Docs API 将页眉和页脚添加到现有的 Google 文档文件中. 看着documents.batchUpdate ( link ) 我们可以插入文本、替换文本、添加图像和
google-docs - 监控 Google Docs 上的 View 统计信息
已关闭。此问题不符合Stack Overflow guidelines 。目前不接受答案。这个问题似乎与 help center 中定义的范围内的编程无关。 . 已关闭 4 年前。 Improve
javascript - docs 文件夹中的 GitHub Pages 引用 docs 文件夹外部的文件
我已按照 GitHub 的文档进行操作，并使用 docs 成功发布了我的项目页面。我的项目存储库下的文件夹。但我想知道如何解决这个小问题: 我正在开发一个 JavaScript 库 wesa.js ，
java - 无法通过 Docs API 向新的 Google Doc 添加文本
我的程序正在创建文档，每个文档都有需要放入其中的文本。任何调用 InsertTextRequest 的尝试调用错误。 List requests = new ArrayList<>(); reques
如果 doc 的关键字发生变化，则 MySQL 会触发 doc 的更新时间戳
基于此: Set field to automatically insert time-stamp on UPDATE? 我正在尝试创建适合我需要的触发器，但我发现使用 OLD 和 NEW 关键字不方

首页

博学

6Ren·AI

商城

python - 如何在 textacy 0.6.2 中初始化 `Doc`？