redaiconglin 发表于 2022-4-15 17:22:53

如何使用BeautifulSoup提取说有href后面的网址

from bs4 import BeautifulSoup


html_doc = """
<html>
<head>
    <title>关联获取演示</title>
    <meta charset="utf-8"/>
</head>
<body>
<p class="p-1" value = "1"><a href="https://item.jd.com/12353915.html">零基础学Python</a></p>
第一个p节点下文本
<div class="div-1" value = "2"><a href="https://item.jd.com/12451724.html">Python从入门到项目实践</a></div>
<p class="p-3" value = "3"><a href="https://item.jd.com/12512461.html">Python项目开发案例集锦</a></p>
<div class="div-2" value = "4"><a href="https://item.jd.com/12550531.html">Python编程锦囊</a></div>
</body>
</html>
"""

soup = BeautifulSoup(html_doc, features="lxml")
print()                                                             如何使用Beautifulsoup才能提取所有href后面的网址谢谢老师够帮我解答谢谢了

isdkz 发表于 2022-4-15 17:33:59

from bs4 import BeautifulSoup


html_doc = """
<html>
<head>
    <title>关联获取演示</title>
    <meta charset="utf-8"/>
</head>
<body>
<p class="p-1" value = "1"><a href="https://item.jd.com/12353915.html">零基础学Python</a></p>
第一个p节点下文本
<div class="div-1" value = "2"><a href="https://item.jd.com/12451724.html">Python从入门到项目实践</a></div>
<p class="p-3" value = "3"><a href="https://item.jd.com/12512461.html">Python项目开发案例集锦</a></p>
<div class="div-2" value = "4"><a href="https://item.jd.com/12550531.html">Python编程锦囊</a></div>
</body>
</html>
"""

soup = BeautifulSoup(html_doc, features="lxml")
urls = for i in soup.find_all('a')]                         # 注意这行
print(urls)

redaiconglin 发表于 2022-4-15 18:08:27

isdkz 发表于 2022-4-15 17:33


谢谢老师
页: [1]
查看完整版本: 如何使用BeautifulSoup提取说有href后面的网址