Python is a language that runs in an interpreter. According to research, Python has a global lock (GIL). When using multi-threading (Thread), it cannot take advantage of multi-core. However, using multi-process (Multiprocess) can leverage multi-core advantages to truly improve efficiency.

Comparison experiment

According to the information, if the multithreaded process isCPU-intensivethen multi-threading cannot improve efficiency much; on the contrary, it may even decrease efficiency due to frequent thread switching, so multi-process is recommended; if it isIO-intensivethen multi-threaded processes can use the idle time during IO blocking waits to execute other threads, improving efficiency. Therefore, we compare the efficiency of different scenarios through experiments.

Operating system CPU Memory Hard disk
Windows 10 Dual-core 8GB Mechanical hard disk

(1) Import the required modules

import requests
import time
from threading import Thread
from multiprocessing import Process

(2) Define a CPU-intensive computation function

def count(x, y):
    # 使程序完成50万计算
    c = 0
    while c < 500000:
        c += 1
        x += x
        y += y

(3) Define an IO-intensive file read/write function

def write():
    f = open("test.txt", "w")
    for x in range(5000000):
        f.write("testwrite\n")
    f.close()
def read():
    f = open("test.txt", "r")
    lines = f.readlines()
    f.close()

(4) Define a network request function

_head = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/48.0.2564.116 Safari/537.36'}
url = "http://www.tieba.com"
def http_request():
    try:
        webPage = requests.get(url, headers=_head)
        html = webPage.text
        return {"context": html}
    except Exception as e:
        return {"error": e}

(5) Test the time required for linear execution of IO-intensive operations, CPU-intensive operations, and network request-intensive operations

# CPU密集操作
t = time.time()
for x in range(10):
    count(1, 1)
print("Line cpu", time.time() - t)
# IO密集操作
t = time.time()
for x in range(10):
    write()
    read()
print("Line IO", time.time() - t)
# 网络请求密集型操作
t = time.time()
for x in range(10):
    http_request()
print("Line Http Request", time.time() - t)

Output

  • CPU-intensive: 95.6059999466, 91.57099986076355, 92.52800011634827, 99.96799993515015

  • IO-intensive: 24.25, 21.76699995994568, 21.769999980926514, 22.060999870300293

  • Network request-intensive: 4.519999980926514, 8.563999891281128, 4.371000051498413, 4.522000074386597, 14.671000003814697

(6) Test the time required for concurrent multi-threaded execution of CPU-intensive operations

counts = []
t = time.time()
for x in range(10):
    thread = Thread(target=count, args=(1,1))
    counts.append(thread)
    thread.start()
e = counts.__len__()
while True:
    for th in counts:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print(time.time() - t)

Output: 25.69700002670288、24.02400016784668

(7) Test the time required for concurrent multi-threaded execution of IO-intensive operations

def io():
    write()
    read()

t = time.time()
ios = []
t = time.time()
for x in range(10):
    thread = Thread(target=count, args=(1,1))
    ios.append(thread)
    thread.start()

e = ios.__len__()
while True:
    for th in ios:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print(time.time() - t)

Output: 99.9240000248 、101.26400017738342、102.32200002670288

(8) Test the time required for concurrent multi-threaded execution of network-intensive operations

t = time.time()
ios = []
t = time.time()
for x in range(10):
    thread = Thread(target=http_request)
    ios.append(thread)
    thread.start()

e = ios.__len__()
while True:
    for th in ios:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print("Thread Http Request", time.time() - t)

Output: 0.7419998645782471、0.3839998245239258、0.3900001049041748

(9) Test the time required for concurrent multi-process execution of CPU-intensive operations

counts = []
t = time.time()
for x in range(10):
    process = Process(target=count, args=(1,1))
    counts.append(process)
    process.start()
e = counts.__len__()
while True:
    for th in counts:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print("Multiprocess cpu", time.time() - t)

Output: 54.342000007629395、53.437999963760376

(10) Test concurrent multi-process execution of IO-intensive operations

t = time.time()
ios = []
t = time.time()
for x in range(10):
    process = Process(target=io)
    ios.append(process)
    process.start()

e = ios.__len__()
while True:
    for th in ios:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print("Multiprocess IO", time.time() - t)

Output: 12.509000062942505、13.059000015258789

(11) Test concurrent multi-process execution of HTTP request-intensive operations

t = time.time()
httprs = []
t = time.time()
for x in range(10):
    process = Process(target=http_request)
    ios.append(process)
    process.start()

e = httprs.__len__()
while True:
    for th in httprs:
        if not th.is_alive():
            e -= 1
    if e <= 0:
        break
print("Multiprocess Http Request", time.time() - t)

Output: 0.5329999923706055、0.4760000705718994

Experimental results

CPU-intensive operations IO-intensive operations Network request-intensive operations
Linear operations 94.91824996469 22.46199995279 7.3296000004
Multi-threaded operations 101.1700000762 24.8605000973 0.5053332647
Multi-process operations 53.8899999857 12.7840000391 0.5045000315

From the results above, we can see:

  • Multi-threading does not seem to have a significant advantage in IO-intensive operations either (perhaps if the IO tasks were heavier, the advantage would show). In CPU-intensive operations, it is clearly worse than single-threaded linear execution. However, for operations like network requests that cause busy-wait thread blocking, multi-threading has a very significant advantage.

  • Multi-process shows performance advantages in CPU-intensive, IO-intensive, and network request-intensive (operations where thread blocking frequently occurs) scenarios. However, for operations similar to network request-intensive ones, it is almost the same as multi-threading, but consumes more resources such as CPU. Therefore, in this case, we can choose multi-threading to execute.

Original address: http://blog.atomicer.cn/2016/09/30/Python%E4%B8%AD%E5%A4%9A%E7%BA%BF%E7%A8%8B%E5%92%8C%E5%A4%9A%E8%BF%9B%E7%A8%8B%E7%9A%84%E5%AF%B9%E6%AF%94/