1. What is a Web Crawler

Web crawler: A program that automatically grabs information from the internet, capturing valuable information for us from the internet.

2. Python Web Crawler Architecture

The Python web crawler architecture mainly consists of five parts: the scheduler, URL manager, web page downloader, web page parser, and application (valuable data crawled).

  • Scheduler:Equivalent to a computer's CPU, mainly responsible for coordinating the work among the URL manager, downloader, and parser.
  • URL manager:Includes URLs to be crawled and already crawled URLs, preventing duplicate crawling and circular crawling of URLs. The URL manager is mainly implemented in three ways: through memory, database, and cache database.
  • Web page downloader:Downloads a web page by passing in a URL address, converting the web page into a string. Web page downloaders include urllib2 (Python's official basic module), including login, proxy, and cookie support, and requests (third-party package).
  • Web page parser:Parses a web page string, extracting useful information according to our requirements, and can also parse based on the DOM tree parsing method. Web page parsers include regular expressions (intuitive, converting web pages into strings and extracting valuable information through fuzzy matching; when the document is complex, this method becomes very difficult for extracting data), html.parser (built into Python), beautifulsoup (third-party plugin that can use Python's built-in html.parser or lxml for parsing; it is more powerful compared to the other methods), and lxml (third-party plugin that can parse XML and HTML). html.parser, beautifulsoup, and lxml all parse using the DOM tree method.
  • Application:It is an application composed of useful data extracted from web pages.

The following diagram explains how the scheduler coordinates work:

3. Three Ways urllib2 Implements Downloading Web Pages

#!/usr/bin/python # -*- coding: UTF-8 -*- import cookielib import urllib2 url = "http://www.baidu.com" response1 = urllib2.urlopen(url) print "Method 1" # Get the status code; 200 indicates success print response1.getcode() # Get the length of the web page content print len(response1.read()) print "Method 2" request = urllib2.Request(url) # Simulate a Mozilla browser for crawling request.add_header("user-agent","Mozilla/5.0") response2 = urllib2.urlopen(request) print response2.getcode() print len(response2.read()) print "Method 3" cookie = cookielib.CookieJar() # Add urllib2's ability to handle cookies opener = urllib2.build_opener(urllib2.HTTPCookieProcessor(cookie)) urllib2.install_opener(opener) response3 = urllib2.urlopen(url) print response3.getcode() print len(response3.read()) print cookie

4. Installation of the Third-Party Library Beautiful Soup

Beautiful Soup: A third-party Python plugin used to extract data from XML and HTML. Official website addresshttps://www.crummy.com/software/BeautifulSoup/

1. Install Beautiful Soup

Open cmd (Command Prompt), navigate to the scripts folder in the Python (Python 2.7 version) installation directory, type dir to check whether pip.exe exists. If it does, you can use Python's built-in pip command to install. Enter the following command to install:

pip install beautifulsoup4

2. Test Whether the Installation Was Successful

Write a Python file and enter:

import bs4
print bs4

Run the file. If it outputs normally, the installation was successful.

5. Using Beautiful Soup to Parse HTML Files

#!/usr/bin/python # -*- coding: UTF-8 -*- import re from bs4 import BeautifulSoup html_doc = """ <html><head><title>The Dormouse's story</title></head> <body> <p class="title"><b>The Dormouse's story</b></p> <p class="story">Once upon a time there were three little sisters; and their names were <a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>, <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>; and they lived at the bottom of a well.</p> <p class="story">...</p> """ # Create a BeautifulSoup parsing object soup = BeautifulSoup(html_doc,"html.parser",from_encoding="utf-8") # Get all links links = soup.find_all('a') print "All links" for link in links: print link.name,link['href'],link.get_text() print "Get a specific URL address" link_node = soup.find('a',href="http://example.com/elsie") print link_node.name,link_node['href'],link_node['class'],link_node.get_text() print "Regular expression matching" link_node = soup.find('a',href=re.compile(r"ti")) print link_node.name,link_node['href'],link_node['class'],link_node.get_text() print "Get the text of P paragraphs" p_node = soup.find('p',class_='story') print p_node.name,p_node['class'],p_node.get_text()

Original address: https://blog.csdn.net/sinat_29957455/article/details/70846427