Implement a simple web crawler using Python
We will use Python'srequestslibrary to send HTTP requests, and useBeautifulSouplibrary to parse HTML content. This simple web crawler will extract all links from a webpage.
Examples
import requests
from bs4 import BeautifulSoup
def simple_web_crawler(url):
# Send HTTP request
response = requests.get(url)
# Check if the request was successful
if response.status_code == 200:
# Parse HTML content
soup = BeautifulSoup(response.text, 'html.parser')
# Find all links
links = soup.find_all('a')
# Extract and print links
for link in links:
href = link.get('href')
if href:
print(href)
else:
print(f"Failed to retrieve the webpage. Status code: {response.status_code}")
# Usage example
simple_web_crawler('https://www.example.com')
from bs4 import BeautifulSoup
def simple_web_crawler(url):
# Send HTTP request
response = requests.get(url)
# Check if the request was successful
if response.status_code == 200:
# Parse HTML content
soup = BeautifulSoup(response.text, 'html.parser')
# Find all links
links = soup.find_all('a')
# Extract and print links
for link in links:
href = link.get('href')
if href:
print(href)
else:
print(f"Failed to retrieve the webpage. Status code: {response.status_code}")
# Usage example
simple_web_crawler('https://www.example.com')
Code explanation:
import requests: Importrequestslibrary, used to send HTTP requests.from bs4 import BeautifulSoup: ImportBeautifulSoupclass, used to parse HTML content.def simple_web_crawler(url):: Define a functionsimple_web_crawler, accepting a URL as a parameter.response = requests.get(url): Send a GET request to the specified URL, and store the response inresponsethe variable.if response.status_code == 200:: Check whether the request was successful (status code 200 indicates success).soup = BeautifulSoup(response.text, 'html.parser'): UseBeautifulSoupto parse HTML content.links = soup.find_all('a'): Find all<a>tags, which usually contain links.for link in links:: Iterate through all links.href = link.get('href'): Extract each link'shrefattribute.if href:: Checkhrefwhether it exists.print(href): Print the link.else:: If the request fails, print an error message.
Output result:https://www.example.comall links on the page. The specific output depends on the content of the target webpage. For example:
https://www.iana.org/domains/example
This is just an example; actual output may vary.
Other Extensions
Python3 Examples