使用 Python BeautifulSoup 查找网页超链接

要查找网页中的超链接,可以使用 BeautifulSoup 库中的 find_all() 方法,并指定标签名称为 'a',如下所示:

from bs4 import BeautifulSoup

html = """
<html>
<head></head>
<body>
<a href='https://www.google.com'>Google</a>
<a href='https://www.facebook.com'>Facebook</a>
<a href='https://www.twitter.com'>Twitter</a>
</body>
</html>
"""

soup = BeautifulSoup(html, 'html.parser')
links = soup.find_all('a')

for link in links:
    print(link.get('href'))

运行以上代码,将会输出超链接的地址:

https://www.google.com
https://www.facebook.com
https://www.twitter.com

更多示例:

  • 查找包含特定文本的超链接:soup.find_all('a', text='Google')
  • 查找具有特定属性的超链接:soup.find_all('a', href='https://www.google.com')

注意:

  • 确保已安装 BeautifulSoup 库。可以使用 pip install beautifulsoup4 命令进行安装。
  • 对于更复杂的网页结构,可能需要使用更精细的筛选条件来定位目标超链接。
  • 在使用 BeautifulSoup 之前,请确保你已经了解并遵守相关网站的 robots.txt 协议。
Python BeautifulSoup 查找网页超链接教程

原文地址: https://www.cveoy.top/t/topic/qw9W 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录