Implement a simple web crawler using Python

Document 对象参考手册Python3 Examples

We will use Python'srequestslibrary to send HTTP requests, and useBeautifulSouplibrary to parse HTML content. This simple web crawler will extract all links from a webpage.

Examples

import requests
from bs4 import BeautifulSoup

def simple_web_crawler(url):
    # Send HTTP request
    response = requests.get(url)

    # Check if the request was successful
    if response.status_code == 200:
        # Parse HTML content
        soup = BeautifulSoup(response.text, 'html.parser')

        # Find all links
        links = soup.find_all('a')

        # Extract and print links
        for link in links:
            href = link.get('href')
            if href:
                print(href)
    else:
        print(f"Failed to retrieve the webpage. Status code: {response.status_code}")

# Usage example
simple_web_crawler('https://www.example.com')

Code explanation:

  1. import requests: Importrequestslibrary, used to send HTTP requests.
  2. from bs4 import BeautifulSoup: ImportBeautifulSoupclass, used to parse HTML content.
  3. def simple_web_crawler(url):: Define a functionsimple_web_crawler, accepting a URL as a parameter.
  4. response = requests.get(url): Send a GET request to the specified URL, and store the response inresponsethe variable.
  5. if response.status_code == 200:: Check whether the request was successful (status code 200 indicates success).
  6. soup = BeautifulSoup(response.text, 'html.parser'): UseBeautifulSoupto parse HTML content.
  7. links = soup.find_all('a'): Find all<a>tags, which usually contain links.
  8. for link in links:: Iterate through all links.
  9. href = link.get('href'): Extract each link'shrefattribute.
  10. if href:: Checkhrefwhether it exists.
  11. print(href): Print the link.
  12. else:: If the request fails, print an error message.

Output result:https://www.example.comall links on the page. The specific output depends on the content of the target webpage. For example:

https://www.iana.org/domains/example

This is just an example; actual output may vary.

Document 对象参考手册Python3 Examples

Other Extensions