|
马上注册,结交更多好友,享用更多功能^_^
您需要 登录 才可以下载或查看,没有账号?立即注册
x
from bs4 import BeautifulSoup
html_doc = """
<html>
<head>
<title>关联获取演示</title>
<meta charset="utf-8"/>
</head>
<body>
<p class="p-1" value = "1"><a href="https://item.jd.com/12353915.html">零基础学Python</a></p>
第一个p节点下文本
<div class="div-1" value = "2"><a href="https://item.jd.com/12451724.html">Python从入门到项目实践</a></div>
<p class="p-3" value = "3"><a href="https://item.jd.com/12512461.html">Python项目开发案例集锦</a></p>
<div class="div-2" value = "4"><a href="https://item.jd.com/12550531.html">Python编程锦囊</a></div>
</body>
</html>
"""
soup = BeautifulSoup(html_doc, features="lxml")
print() 如何使用Beautifulsoup才能提取所有href后面的网址谢谢老师够帮我解答谢谢了
- from bs4 import BeautifulSoup
- html_doc = """
- <html>
- <head>
- <title>关联获取演示</title>
- <meta charset="utf-8"/>
- </head>
- <body>
- <p class="p-1" value = "1"><a href="https://item.jd.com/12353915.html">零基础学Python</a></p>
- 第一个p节点下文本
- <div class="div-1" value = "2"><a href="https://item.jd.com/12451724.html">Python从入门到项目实践</a></div>
- <p class="p-3" value = "3"><a href="https://item.jd.com/12512461.html">Python项目开发案例集锦</a></p>
- <div class="div-2" value = "4"><a href="https://item.jd.com/12550531.html">Python编程锦囊</a></div>
- </body>
- </html>
- """
- soup = BeautifulSoup(html_doc, features="lxml")
- urls = [i['href'] for i in soup.find_all('a')] # 注意这行
- print(urls)
复制代码
|
|