python - 美丽汤错误 : '<class ' bs4. 元素。标签'>' object has no attribute ' 内容'？-6ren

python - 美丽汤错误 : '' object has no attribute ' 内容'？

转载作者：太空宇宙更新时间：2023-11-03 19:10:00

25

4

我正在编写一个脚本，从文章中提取内容并删除任何不必要的内容，例如。脚本和样式。 Beautiful Soup 不断引发以下异常:

'<class 'bs4.element.Tag'>' object has no attribute 'contents'

以下是trim函数的代码(element是包含网页内容的HTML元素):

def trim(element):
    elements_to_remove = ('script', 'style', 'link', 'form', 'object', 'iframe')
    for i in elements_to_remove:
        remove_all_elements(element, i)

    attributes_to_remove = ('class', 'id', 'style')
    for i in attributes_to_remove:
        remove_all_attributes(element, i)

    remove_all_comments(element)

    # Remove divs that have more non-p elements than p elements
    for div in element.find_all('div'):
        p = len(div.find_all('p'))
        img = len(div.find_all('img'))
        li = len(div.find_all('li'))
        a = len(div.find_all('a'))

        if p == 0 or img > p or li > p or a > p:
            div.decompose()

查看堆栈跟踪，问题似乎出在 for 语句之后的此方法:

    # Remove divs that have more non-p elements than p elements
    for div in element.find_all('div'):
        p = len(div.find_all('p')) # <-- div.find_all('p')

我不明白为什么 bs4.element.Tag 的这个实例没有“contents”属性？我在实际网页上尝试了一下，元素中充满了 p 和 img...

这是回溯(这是我正在开发的 Django 项目的一部分):

Environment:


Request Method: POST
Request URL: http://localhost:8000/read/add/

Django Version: 1.4.1
Python Version: 2.7.3
Installed Applications:
('django.contrib.auth',
 'django.contrib.contenttypes',
 'django.contrib.sessions',
 'django.contrib.sites',
 'django.contrib.messages',
 'django.contrib.staticfiles',
 'home',
 'account',
 'read',
 'review')
Installed Middleware:
('django.middleware.common.CommonMiddleware',
 'django.contrib.sessions.middleware.SessionMiddleware',
 'django.middleware.csrf.CsrfViewMiddleware',
 'django.contrib.auth.middleware.AuthenticationMiddleware',
 'django.contrib.messages.middleware.MessageMiddleware')


Traceback:
File "/home/marco/.virtualenvs/sandra/local/lib/python2.7/site-packages/django/core/handlers/base.py" in get_response
  111.                         response = callback(request, *callback_args, **callback_kwargs)
File "/home/marco/sandra/read/views.py" in add
  24.             Article.objects.create_article(request.user, url)
File "/home/marco/sandra/read/models.py" in create_article
  11.         title, content = logic.process_html(web_page.read())
File "/home/marco/sandra/read/logic.py" in process_html
  7.     soup = htmlbarber.give_haircut(BeautifulSoup(html_code, 'html5lib'))
File "/home/marco/sandra/read/htmlbarber/__init__.py" in give_haircut
  45.     scissor.trim(element)
File "/home/marco/sandra/read/htmlbarber/scissor.py" in trim
  35.         p = len(div.find_all('p'))
File "/home/marco/.virtualenvs/sandra/local/lib/python2.7/site-packages/bs4/element.py" in find_all
  1128.         return self._find_all(name, attrs, text, limit, generator, **kwargs)
File "/home/marco/.virtualenvs/sandra/local/lib/python2.7/site-packages/bs4/element.py" in _find_all
  413.                 return [element for element in generator
File "/home/marco/.virtualenvs/sandra/local/lib/python2.7/site-packages/bs4/element.py" in descendants
  1140.         if not len(self.contents):
File "/home/marco/.virtualenvs/sandra/local/lib/python2.7/site-packages/bs4/element.py" in __getattr__
  924.             "'%s' object has no attribute '%s'" % (self.__class__, tag))

Exception Type: AttributeError at /read/add/
Exception Value: '<class 'bs4.element.Tag'>' object has no attribute 'contents'

这是remove_all_*函数的源代码:

def remove_all_elements(element_to_clean, unwanted_element_name):
    for to_remove in element_to_clean.find_all(unwanted_element_name):
        to_remove.decompose()

def remove_all_attributes(element_to_clean, unwanted_attribute_name):
    for to_inspect in [element_to_clean] + element_to_clean.find_all():
        try:
            del to_inspect[unwanted_attribute_name]
        except KeyError:
            pass

def remove_all_comments(element_to_clean):
    for comment in element_to_clean.find_all(text=lambda text:isinstance(text, Comment)):
        comment.extract()

最佳答案

我认为问题在于，在 remove_all_elements 或代码中的其他位置，您正在删除某些标记的 contents 属性。

看起来当您调用to_remove.decompose()时就会发生这种情况。这是该方法的来源:

def decompose(self):
    """Recursively destroys the contents of this tree."""
    self.extract()
    i = self
    while i is not None:
        next = i.next_element
        i.__dict__.clear()
        i = next

如果您手动调用此函数，会发生以下情况:

>> soup = BeautifulSoup('<div><p>hi</p></div>')
>>> d0 = soup.find_all('div')[0]
>>> d0
<div><p>hi</p></div>
>>> d0.decompose()
>>> d0
Traceback (most recent call last):
...
Traceback (most recent call last):
AttributeError: '<class 'bs4.element.Tag'>' object has no attribute 'contents'

看来，一旦您在标签上调用了decompose，您就不能再尝试使用该标签。我不太确定这是在哪里发生的。

我要尝试检查的一件事是，在 trim() 函数中始终保持 len(element.__dict__) > 0 。

关于python - 美丽汤错误 : '<class ' bs4. 元素。标签'>' object has no attribute ' 内容'？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/13310936/

25

4

0

文章推荐： c# - XElement 及其属性

文章推荐： javascript - 覆盖标签？

c# - if((attributes and File Attributes.Hidden) == File Attributes.Hidden) { } 如何工作？
关于 this页面，我看到以下代码: if ((attributes & FileAttributes.Hidden) == FileAttributes.Hidden) 但我不明白为什么会变成这样。
attributes - pthread互斥锁的 “attribute”是什么？
函数pthread_mutex_init允许您指定指向属性的指针。但是我还没有找到关于pthread属性是什么的很好的解释。我一直只是提供NULL。这个论点有用吗？该文档，对于那些忘记它的人: PT
xml - 我怎样才能结合xsl :attribute and xsl:use-attribute-sets to conditionally use an attribute set?
我们有一个 xml 节点“item”，其属性为“style”，即“Header1”。但是，这种风格可以改变。我们有一个名为 Header1 的属性集，它定义了它在 PDF 中的外观，通过 xsl:fo
JavaScript: element.setAttribute(attribute,value) , element.attribute=value & element.[attribute]=value 不改变属性值
我的任务是在用户点击它时从输入框中删除占位符并使标签可见。如果用户未在其中再次填写任何内容，请放回占位符并使标签不可见。我可以隐藏它但不能重新分配它。我试过 element.setAttribute
attributes - ASP.NET 5 : Bind attribute with Include parameter - include is not a valid named attribute argument
我从文章中编写代码，并且有: public IActionResult Create([Bind(Include="Imie,Nazwisko,Stanowisko,Wiek")] Pracownik
attributes - 单点触控 : Understand Foundation Attributes
你能给我解释一下以下属性吗？ 1) [MonoTouch.Foundation.Register("SomeClass")] 这个属性是否只用于向IB注册类？以编程方式扩展 iOS 类时是否必须使用此
c++ - this.attribute 应该是 this->attribute 是什么意思
我正在编写一个 C++ 程序，在调试时我在执行以下函数: int CClass::do_something() { ... // I've put a breakpoint here } 我的 C
javascript - polymer 1.0 : Is there any way to use 'layout' as an attribute instead of as a CSS class or using Attribute serialization in the class attribute?
我已经在 polymer 0.5 中构建了我的应用程序。现在我已经将它更新到 polymer 1.0。对于响应式布局，我使用了一个布局属性，它使用 Polymer 0.5 中布局属性的自定义逻辑。
attributes - Jade : element attributes without value
我是使用 Jade 的新手——到目前为止它很棒。但是我需要发生的一件事是具有“itemscope”属性的元素: 我的 Jade 符是: header(itemscope, itemtype='ht
attributes - 为什么在 Chef 中使用普通属性(attribute.set[..])？
我正在研究一个厨师实现，有时在过去的地方使用了 attribute.set，attribute.default 会这样做。为了解决这个问题，我对 Chef 属性优先范式非常熟悉。我知道“正常”属性(使
HTML "data-attribute"与简单 "custom attribute"
我经常看到html data-attribute (s) 将特定值/参数添加到 html 元素，例如使用它们将按钮“链接”到要打开的模式对话框等的 Bootstrap。现在，我看到一个几乎著名的
ruby - self.attribute 与 @attribute 的优势？
假设如下: def create_new_salt self.salt = self.object_id.to_s + rand.to_s end 为什么使用“ self ”更好。而不是实例变量“
主干.js 访问模型中的模型属性 - this.attribute VS this.get ('attribute' )？
根据我的理解，Backbone.js 模型的属性应该通过以下方式声明为有点私有(private)的成员变量 this.set({ attributeName: attributeValue }) //
xml - 在Hive XML SerDe中使用 “Attribute to Attribute”映射
我有一个看起来像下面的XML文档: ... ... ... ...
JSF 复合 :attribute with f:attribute conversion error
我正在实现一个 JSF 组件，需要有条件地添加一些属性。这个问题类似于之前的 JSF: p:dataTable with f:attribute results in "argument type m
安卓市场发布: 'android:icon' attribute: attribute is not a string value
我正在尝试将应用程序发布到 Android 电子市场，但出现以下错误: W/ResourceType(16964): No known package when getting value for r
c++ - 玛雅编程 : Separating attributes into sections in the attribute editor
抱歉这么具体的应用程序，但我注意到另一篇关于 Maya 开发的回答很好的帖子。我刚刚为 Maya 编写了一个插件节点。它只是根据湍流函数杀死一堆粒子。湍流由许多可在属性编辑器中调整的属性驱动。在属
html - html元素中data-attribute=false与data-attribute ="false"有什么区别吗？
我在 html 元素中的数据属性为 Update .它具有数据属性的 bool 值。跟下面的元素Update有什么区别吗？因为数据属性用双引号引起来。 html是否支持 bool 值？最佳答案 b
c# - 错误 : "is not an attribute class" when using ConfigurationElementType attribute
我正在尝试为企业库 5.0 的异常处理 block 创建自定义异常处理程序。据我了解，我需要使用属性开始上课“[ConfigurationElementType(typeof(CustomHandle
css - [attribute~=value] 和 [attribute*=value] 的区别
我找不到这两个选择器之间的区别。两者似乎都做同样的事情，即根据包含给定字符串的特定属性值选择标签。对于 [attribute~=value] :http://www.w3schools.com/cs

首页

博学

6Ren·AI

商城

python - 美丽汤错误 : '' object has no attribute ' 内容'？