gpt4 book ai didi

python - 网站更改代码后,Webscraper 抛出错误

转载 作者:行者123 更新时间:2023-12-04 08:43:42 26 4
gpt4 key购买 nike

我为 realtor.com 构建了一个 webscraper,因为我正在寻找我所在地区的房屋和代理商,这对我来说很容易,但是他们只是更改了他们网站上的代码(可能是为了阻止人们这样做),现在我是得到一个属性错误。我收到的错误是这样的:
文件“webscraper.py”,第 22 行,在
name.getText().strip(),
AttributeError: 'NoneType' 对象没有属性 'getText'
下面的代码在他们更改代码之前完美地收集了名称和数字。看来他们所做的只是更改了类名。添加“jsx-1792441256”

import csv
import requests
from bs4 import BeautifulSoup
from time import sleep
from random import randint

sleep(randint(10,20))


realtor_data = []

for page in range(1, 10):
print(f"Scraping page {page}...")
url = f"https://www.realtor.com/realestateagents/san-diego_ca/pg-{page}"
soup = BeautifulSoup(requests.get(url).text, "html.parser")

for agent_card in soup.find_all("div", {"class": "jsx-1792441256 agent-list-card-title-text clearfix"}):
name = agent_card.find("div", {"class": "jsx-1792441256 agent-name text-bold"}).find("a")
number = agent_card.find("div", {"itemprop": "telephone"})
realtor_data.append(
[
name.getText().strip(),
number.getText().strip() if number is not None else "N/A"

],
)

with open("sandiego.csv", "w") as output:
w = csv.writer(output)
w.writerow(["NAME:", "PHONE NUMBER:", "CITY:"])
w.writerows(realtor_data)

import pandas as pd
a=pd.read_csv("sandiego.csv")
a2 = a.iloc[:,[0,1]]
a3 = a.iloc[:,[2]]
a3 = a3.fillna("San Diego")
b=pd.concat([a2,a3],axis=1)
b.to_csv("sandiego.csv")

最佳答案

固定代码:

import csv
import requests
from bs4 import BeautifulSoup
from time import sleep
from random import randint

# sleep(randint(10,20))


realtor_data = []

for page in range(1, 10):
print(f"Scraping page {page}...")
url = f"https://www.realtor.com/realestateagents/san-diego_ca/pg-{page}"
soup = BeautifulSoup(requests.get(url).text, "html.parser")

for agent_card in soup.select("div.agent-list-card-title.mobile-only"):
name = agent_card.find("div", {"class": "agent-name"})
number = agent_card.find("div", {"class": "agent-phone"})
realtor_data.append(
[
name.getText().strip(),
number.getText().strip() if number is not None else "N/A"
],
)

with open("data.csv", "w") as output:
w = csv.writer(output)
w.writerow(["NAME:", "PHONE NUMBER:", "CITY:"])
w.writerows(realtor_data)

import pandas as pd
a=pd.read_csv("data.csv")
a2 = a.iloc[:,[0,1]]
a3 = a.iloc[:,[2]]
a3 = a3.fillna("San Diego")
b=pd.concat([a2,a3],axis=1)
b.to_csv("data.csv")
创建 data.csv :
enter image description here

关于python - 网站更改代码后,Webscraper 抛出错误,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/64436337/

26 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com