gpt4 book ai didi

python - 尝试解析 JS 脚本中的特定值

转载 作者:行者123 更新时间:2023-12-01 08:57:17 25 4
gpt4 key购买 nike

我正在尝试抓取 POST 请求中所需的值。在 Chrome 上使用 Inspect Element 时可以多次找到该值,但由于 BS4 只查看源代码,我必须从该站点的 JS 脚本中抓取该值。

<script type ="text/javascript">        
var isSRFlow = true;
var isPpaOnSignIn =true;
var simplifyRegFlowSuccess = false;
var retUrl = "https&#x3a;&#x2f;&#x2f;www.ebay.com&#x2f;";
var isFB = false;
var isMobile = false;
var langCode = "en-US";


var emailAutoCompleteEnabled = true;

var dfpContext = '{"enableTMXTagging":"true","slURL":"ebay","flashTagUpgrade":"0","enableFlashTagging":"false","tmxDfpUrl":"https://signin.ebay.com/t_n.html?suppressFlash\u003dtrue\u0026org_id\u003dusllpic0\u0026session_id\u003d57be07a71660ad4e16f42acffffc95e8","swfURL":"ebay","enableSLTagging":"false","swfObjectJSLibURL":"ebay","mid":"AQAAAWZGrHELAAUxNjY1N2JlMDdhNy5hZDRlMTZmLjQyYWNmLmZmZmM5NWU5Jp0dBAKw4k3h8WAm/g97vwVzjcA*","tmxSessionId":"57be07a71660ad4e16f42acffffc95e8","enableHTML5Tagging":"true","flashTagVersion":"1","dfpjsURL":"https://secureir.ebaystatic.com/f/0vk0rkyoky1ltm32dhy0hthnxyx.js"}';

我设法通过使用获取整个脚本 r = requests.get('https://reg.ebay.com/reg/PartialReg')<br/>
soup = BeautifulSoup(r.text, 'lxml')
scripts = soup.find_all('script')
your_script = [script for script in scripts if 'tmxSessionId' in str(script)][0]

但是,我真正需要的是“57be07a71660ad4e16f42acffffc95e8”,它们是“tmxSessionId”之后的数字。如何才能做到这一点?

我也尝试过这些:

scripts = soup.find_all('script')
your_script = [script for script in scripts if 'tmxSessionId' in str(script)][0]
new = your_script.find("tmxSessionId")
print(new)

以及使用“find_all”而不仅仅是“find”。我的一位 friend 也建议拆分脚本,但我尝试了一下,发现效果不太好。有什么想法吗?

P.S:我不想使用基于浏览器的解决方案,例如 selenium 和 PhantomJS,因为我发现它缓慢且无效

编辑:我使用旧代码从源代码中获取脚本,然后使用 Selçuk 建议的内容

scripts = soup.find_all('script')
your_script = [script for script in scripts if 'tmxSessionId' in str(script)][0]
script_tag = your_script
soup = BeautifulSoup(script_tag, 'lxml')
script = soup.find_all('script')[0]
data = re.findall("{.*?}", script.text)[0]

print(json.loads(data)['tmxSessionId'])

最佳答案

我不知道你的脚本内容的其余部分,所以我不得不关闭标签。但它会起作用。

import requests
from bs4 import BeautifulSoup
import re
import json

script_tag = """
<script type ="text/javascript">
var isSRFlow = true;
var isPpaOnSignIn =true;
var simplifyRegFlowSuccess = false;
var retUrl = "https&#x3a;&#x2f;&#x2f;www.ebay.com&#x2f;";
var isFB = false;
var isMobile = false;
var langCode = "en-US";


var emailAutoCompleteEnabled = true;

var dfpContext = '{"enableTMXTagging":"true","slURL":"ebay","flashTagUpgrade":"0","enableFlashTagging":"false","tmxDfpUrl":"https://signin.ebay.com/t_n.html?suppressFlash\u003dtrue\u0026org_id\u003dusllpic0\u0026session_id\u003d57be07a71660ad4e16f42acffffc95e8","swfURL":"ebay","enableSLTagging":"false","swfObjectJSLibURL":"ebay","mid":"AQAAAWZGrHELAAUxNjY1N2JlMDdhNy5hZDRlMTZmLjQyYWNmLmZmZmM5NWU5Jp0dBAKw4k3h8WAm/g97vwVzjcA*","tmxSessionId":"57be07a71660ad4e16f42acffffc95e8","enableHTML5Tagging":"true","flashTagVersion":"1","dfpjsURL":"https://secureir.ebaystatic.com/f/0vk0rkyoky1ltm32dhy0hthnxyx.js"}';
</script>
"""

soup = BeautifulSoup(script_tag, 'lxml')
script = soup.find_all('script')[0]
data = re.findall("{.*?}", script.text)[0]

print(json.loads(data)['tmxSessionId'])

输出将是

57be07a71660ad4e16f42acffffc95e8

关于python - 尝试解析 JS 脚本中的特定值,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/52715847/

25 4 0
Copyright 2021 - 2024 cfsdn All Rights Reserved 蜀ICP备2022000587号
广告合作:1813099741@qq.com 6ren.com