Python Scrapy Library

Scrapy is a powerful Python web scraping framework specifically designed for crawling web pages and extracting information.

Scrapy is often used in applications such as data mining, information processing, or storing historical data.

Scrapy has many built-in useful features, such as handling requests, tracking state, handling errors, managing request rate limits, etc., making it very suitable for efficient, distributed web crawling.

Unlike simple web scraping libraries (such asrequestsandBeautifulSoup), Scrapy is a full-featured web scraping framework with a high degree of scalability and flexibility, suitable for complex and large-scale web scraping tasks.

Scrapy Official Website:https://scrapy.org/。

Scrapy Features and Introduction:https://www.example.com/w3cnote/scrapy-detail.html。

Scrapy Architecture Diagram (green line is data flow):

Scrapy's work is based on the following core components:

  • Spider: Spider class, used to define how to extract data from web pages and how to follow links on web pages.
  • Item: Used to define and store the scraped data. It is equivalent to a data model.
  • Pipeline: Used to process the scraped data, commonly used for cleaning, storing data, and other operations.
  • Middleware: Used to handle requests and responses, can be used to set proxies, handle cookies, user agents, etc.
  • Settings: Used to configure various settings of the Scrapy project, such as request delay, number of concurrent requests, etc.

Installing Scrapy

Before using Scrapy, you need to install it first. We use pip to install:

pip install scrapy

Scrapy Project Structure

A Scrapy project is a structured directory containing multiple folders and modules, designed to help you organize your spider code.

Scrapy uses command-line tools to create and manage spider projects. You can use the following command to create a new Scrapy project:

scrapy startproject myproject

This will create a project named myproject, with a project structure roughly as follows:

myproject/
    scrapy.cfg            # 项目的配置文件
    myproject/            # 项目源代码文件夹
        __init__.py
        items.py          # 定义抓取的数据结构
        middlewares.py    # 定义中间件
        pipelines.py      # 定义数据处理管道
        settings.py       # 项目的设置文件
        spiders/           # 存放爬虫代码的文件夹
            __init__.py
            myspider.py   # 自定义的爬虫代码

Writing a Simple Scrapy Spider

The following is a basic Scrapy spider example that shows how to scrape data from web pages.

We create a spider project:

scrapy startproject example_test_spiders

Execute the above command; if successful, it will output:

templates/project', created in:/Users/Example/example-test/example_test_spiders

You can start your first spider with:
    cd example_test_spiders
    scrapy genspider example example.com

The generated project structure is as follows:

Then enter that directory:

cd example_test_spiders

Next, use thescrapy genspidercommand to create a spider:

scrapy genspider douban_spider movie.douban.com

The directory structure is as follows:

A file named douban_spider.py is generated in the example_test_spiders directory, with the code as follows:

Example

import scrapy


class DoubanSpiderSpider(scrapy.Spider):
    name = "douban_spider"
    allowed_domains = ["movie.douban.com"]
    start_urls = ["https://movie.douban.com"]

    def parse(self, response):
        pass

Code explanation:

  • name: Defines the name of the spider; it must be unique.
  • allowed_domains: Restricts the spider's access domain, preventing the spider from crawling pages on other domains.
  • start_urls: Defines the spider's starting pages; the spider will start crawling from these pages.
  • parse:parseThe method is the core part of each spider, used to process responses and extract data. It receives aresponseobject, representing the page content returned by the server.

Write the Spider Code

Before writing spider code, we need to note a few points:

  • Websites such as Douban may detect spider behavior; it is recommended to set USER_AGENT and DOWNLOAD_DELAY to simulate normal user behavior.
  • When scraping data, please comply with the target website's robots.txt file rules to avoid putting too much pressure on the server.
  • If you crawl too frequently, you may trigger an IP ban.

Modify settings.py Configuration

Add the following configuration in settings.py to simulate browser requests and bypass anti-scraping mechanisms:

# 设置 User-Agent
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'

# 不遵守 robots.txt 规则
ROBOTSTXT_OBEY = False

# 设置下载延迟,避免过快请求
DOWNLOAD_DELAY = 2

# 启用自动限速扩展
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
AUTOTHROTTLE_MAX_DELAY = 5

In the spider code, add custom request headers (such as User-Agent and Referer) to further simulate browser behavior.

Open the douban_spider.py file and modify its contents as follows:

Example

import scrapy

class DoubanSpider(scrapy.Spider):
    name = "douban_spider"
    start_urls = [
        'https://movie.douban.com/top250',
    ]

    def start_requests(self):
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36',
            'Referer': 'https://movie.douban.com/',
        }
        for url in self.start_urls:
            yield scrapy.Request(url, headers=headers, callback=self.parse)

    def parse(self, response):
        for movie in response.css('div.item'):
            yield {
                'title': movie.css('span.title::text').get(),
                'rating': movie.css('span.rating_num::text').get(),
                'quote': movie.css('span.inq::text').get(),
            }

        # Handle pagination
        next_page = response.css('span.next a::attr(href)').get()
        if next_page is not None:
            yield response.follow(next_page, callback=self.parse)

Code analysis:

  1. name = "douban_spider": Defines the name of the spider.

  2. start_urls: Defines the initial URL the spider starts crawling from (the Douban Movie Top 250 page).

  3. parsemethod:

    • Uses CSS selectors to extract the title, rating, and introduction of each movie.

    • span.title::text: Extracts the movie title.

    • span.rating_num::text: Extracts the movie rating.

    • span.inq::text: Extracts the movie introduction.

  4. Pagination handling:

    • Usespan.next a::attr(href)to extract the next page link.

    • If the next page exists, useresponse.followto continue crawling.

Run the following command in the command line to start the spider:

scrapy crawl douban_spider -o douban_movies.csv

This will start the spider and save the extracted data to the douban_movies.csv file.

Note: The above content is for learning purposes only. When scraping data, please comply with the target website's robots.txt file rules.


Common Methods

1. Spider Methods

Method Name Description Example
start_requests() Generate the initial request; you can customize request headers, request methods, etc. yield scrapy.Request(url, callback=self.parse)
parse(response) Process the response and extract data; it is the core method of the spider. yield {'title': response.css('h1::text').get()}
follow(url, callback) Automatically handle relative URLs and generate new requests for pagination or link navigation. yield response.follow(next_page, callback=self.parse)
closed(reason) Called when the spider closes, used to clean up resources or record logs. def closed(self, reason): print('Spider closed:', reason)
log(message) Log information. self.log('This is a log message')

2. Data Extraction Methods

Method Name Description Example
response.css(selector) Use CSS selectors to extract data. title = response.css('h1::text').get()
response.xpath(selector) Use XPath selectors to extract data. title = response.xpath('//h1/text()').get()
get() fromSelectorListExtract the first matching result (string) from title = response.css('h1::text').get()
getall() fromSelectorListExtract all matching results (list) from titles = response.css('h1::text').getall()
attrib Extract attributes of the current node. link = response.css('a::attr(href)').get()

3. Request and Response Methods

Method Name Description Example
scrapy.Request(url, callback, method, headers, meta) Create a new request. yield scrapy.Request(url, callback=self.parse, headers=headers)
response.url Get the URL of the current response. current_url = response.url
response.status Get the status code of the response. if response.status == 200: print('Success')
response.meta Get the extra data passed in the request. value = response.meta.get('key')
response.headers Get the headers of the response. content_type = response.headers.get('Content-Type')

4. Middleware and Pipeline Methods

Method Name Description Example
process_request(request, spider) Process the request before it is sent (downloader middleware). request.headers['User-Agent'] = 'Mozilla/5.0'
process_response(request, response, spider) Process the response after it is returned (downloader middleware). if response.status == 403: return request.replace(dont_filter=True)
process_item(item, spider) Process the extracted data (pipeline). if item['price'] < 0: raise DropItem('Invalid price')
open_spider(spider) Called when the spider starts (pipeline). def open_spider(self, spider): self.file = open('items.json', 'w')
close_spider(spider) Called when the spider closes (pipeline). def close_spider(self, spider): self.file.close()

5. Tools and Extension Methods

Method Name Description Example
scrapy shell Start the interactive Shell for debugging and testing selectors. scrapy shell 'http://example.com'
scrapy crawl <spider_name> Run the specified spider. scrapy crawl myspider -o output.json
scrapy check Check the correctness of the spider code. scrapy check
scrapy fetch Download the content of the specified URL. scrapy fetch 'http://example.com'
scrapy view View the page downloaded by Scrapy in a browser. scrapy view 'http://example.com'

6. Common Settings (forsettings.py)

Setting Item Description Example
USER_AGENT Set the User-Agent in the request header. USER_AGENT = 'Mozilla/5.0'
ROBOTSTXT_OBEY Whether to obeyrobots.txtrules. ROBOTSTXT_OBEY = False
DOWNLOAD_DELAY Set download delay to avoid making requests too quickly. DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS Set the number of concurrent requests. CONCURRENT_REQUESTS = 16
ITEM_PIPELINES Enable pipelines. ITEM_PIPELINES = {'myproject.pipelines.MyPipeline': 300}
AUTOTHROTTLE_ENABLED Enable the auto-throttle extension. AUTOTHROTTLE_ENABLED = True

7. Other Common Methods

Method Name Description Example
response.follow_all(links, callback) Process links in batches and generate requests. yield from response.follow_all(links, callback=self.parse)
response.json() Parse the response content into JSON format. data = response.json()
response.text Get the text content of the response. html = response.text
response.selector Get the response content'sSelectorobject. title = response.selector.css('h1::text').get()

The table above lists the commonly used methods in Scrapy and their functions. These methods cover various aspects of crawler development, including request generation, data extraction, middleware processing, pipeline operations, etc. By mastering these methods, you can efficiently write and manage Scrapy crawlers. For more detailed features, please refer toScrapy official documentation。


Method Usage Examples

1. start_requests()

start_requests()The method is the entry point of a Scrapy crawler, used to generate initial requests. Usually, the start URL of the crawler is defined in this method.

Example

import scrapy

class MySpider(scrapy.Spider):
    name = 'myspider'
   
    def start_requests(self):
        urls = [
            'http://example.com/page1',
            'http://example.com/page2',
        ]
        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)

2. parse()

parse()The method is the default response handling method, used to parse responses and extract data or generate new requests.

Example

def parse(self, response):
    # Extract page title
    title = response.css('title::text').get()
    yield {
        'title': title
    }

3. parse_item()

parse_item()The method is used to parse the response of a single item (Item), usually to extract structured data.

Example

def parse_item(self, response):
    item = {}
    item['name'] = response.css('div.name::text').get()
    item['price'] = response.css('div.price::text').get()
    yield item

4. follow()

follow()The method is used to generate new requests and automatically handle responses, usually for link tracking.

Example

def parse(self, response):
    for link in response.css('a::attr(href)'):
        yield response.follow(link, self.parse_item)

5. yield

yieldThe keyword is used to generate requests or items (Item), and pass them to the Scrapy engine for processing.

Example

def parse(self, response):
    yield {
        'title': response.css('title::text').get()
    }

6. Item

ItemThe class is used to define data structures, usually for storing data extracted from web pages.

Example

import scrapy

class MyItem(scrapy.Item):
    name = scrapy.Field()
    price = scrapy.Field()

7. ItemLoader

ItemLoaderThe class is used to load and populate Item objects, simplifying the data extraction and processing workflow.

Example

from scrapy.loader import ItemLoader
from myproject.items import MyItem

def parse(self, response):
    loader = ItemLoader(item=MyItem(), response=response)
    loader.add_css('name', 'div.name::text')
    loader.add_css('price', 'div.price::text')
    yield loader.load_item()

8. Request

RequestThe class is used to generate HTTP request objects, usually for defining request URLs, callback methods, etc.

Example

import scrapy

def parse(self, response):
    yield scrapy.Request(url='http://example.com/page3', callback=self.parse_item)

9. Response

ResponseThe class represents an HTTP response object, containing information such as HTML content and status code returned from the server.

Example

def parse(self, response):
    print(response.status)  # Print response status code
    print(response.body)    # Print response content

10. Selector

SelectorThe class is used to extract data from HTML or XML documents, supporting XPath and CSS selectors.

Example

def parse(self, response):
    title = response.xpath('//title/text()').get()
    yield {
        'title': title
    }

11. CrawlSpider

CrawlSpiderIt is a special Spider class used to handle complex crawling rules and link tracking.

Example

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class MyCrawlSpider(CrawlSpider):
    name = 'mycrawlspider'
    allowed_domains = ['example.com']
    start_urls = ['http://example.com']

    rules = (
        Rule(LinkExtractor(allow=('page/\d+',)), callback='parse_item'),
    )

    def parse_item(self, response):
        yield {
            'title': response.css('title::text').get()
        }

12. LinkExtractor

LinkExtractorThe class is used to extract links from responses, usually for automatically tracking links in pages.

Example

from scrapy.linkextractors import LinkExtractor

def parse(self, response):
    extractor = LinkExtractor(allow=('page/\d+',))
    links = extractor.extract_links(response)
    for link in links:
        yield scrapy.Request(link.url, callback=self.parse_item)

13. Pipeline

PipelineThe class is used to process scraped data, usually for data cleaning, storage, and other operations.

Example

class MyPipeline:
    def process_item(self, item, spider):
        # Process item data
        return item

14. Middleware

MiddlewareThe class is used for middleware processing of requests and responses, usually for modifying request headers, handling exceptions, etc.

Example

class MyMiddleware:
    def process_request(self, request, spider):
        # Modify request headers
        request.headers['User-Agent'] = 'MyCustomUserAgent'
        return None
Other extensions