Pandas pd.read_html() Function
read_html()It is a function in the pandas library used to parse HTML tables, which can read table data from web pages or HTML files and convert it into a DataFrame.
Table data in web pages is a very important data source. Many public data sets (such as stock information, statistical data, etc.) are presented in the form of HTML tables.read_html()UsinglxmlandBeautifulSoupthe library to parse HTML, it can automatically extract all tables or specified tables from the page.
Basic Syntax and Parameters
Syntax Format
pandas.read_html(io, match='.+', flavor=None, header=None, index_col=None,
skiprows=None, attrs=None, parse_dates=False, thousands=',',
decimal='.', converters=None, ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| io | str, path object, file-like object | HTML file path, URL, or string | Required |
| match | str, regex | Use regular expressions to match the text content of the table | '.+' |
| flavor | str | Parser: 'lxml', 'html5lib', 'bs4' | None |
| header | int, list of int | Row number used as column names | None |
| index_col | int, str | Column used as row index | None |
| skiprows | int, list, slice | Skip specified rows | None |
| attrs | dict | HTML tag attributes used to filter tables | None |
| parse_dates | bool, list | Whether to parse date columns | False |
Return Value
- Return Type:
listof DataFrames - Returns a list of DataFrames; the number of DataFrames returned corresponds to the number of tables on the page.
- If no matching table is found, returns an empty list.
Examples
Through the following examples, comprehensively masterread_html()the various usages.
Example 1: Read a table from a local HTML file
First create an HTML file containing a table, then useread_html()to read it.
Example
# Create an HTML file containing a table
html_content = '''
<title>Employee Information Table</title>
<h1>Pandas pd-read-html() Function</h1>
<table border="1" class="employee-table">
<tr>
<th>Name</th>
<th>Age</th>
<th>City</th>
<th>Salary</th>
</tr>
<tr>
<td>Tom</td>
<td>28</td>
<td>Beijing</td>
<td>8000</td>
</tr>
<tr>
<td>Jerry</td>
<td>35</td>
<td>Shanghai</td>
<td>12000</td>
</tr>
<tr>
<td>Mike</td>
<td>42</td>
<td>Guangzhou</td>
<td>15000</td>
</tr>
<tr>
<td>Lucy</td>
<td>26</td>
<td>Shenzhen</td>
<td>7000</td>
</tr>
</table>
<h2>Department List</h2>
<table border="1">
<tr>
<th>Department</th>
<th>Headcount</th>
</tr>
<tr>
<td>Technology Department</td>
<td>50</td>
</tr>
<tr>
<td>Sales Department</td>
<td>30</td>
</tr>
</table>
'''
# Write the HTML to a file
with open('tables.html', 'w', encoding='utf-8') as f:
f.write(html_content)
# Use read_html to read all tables
# io: HTML file path (required)
tables = pd.read_html('tables.html')
# View the read results
print(f"Found {len(tables)} tables")
print()
# Iterate over all tables
for i, df in enumerate(tables):
print(f"--- Table {i+1} ---")
print(df)
print()
Expected output:
共找到 2 个表格
--- 表格 1 ---
姓名 年龄 城市 薪资
0 Tom 28 Beijing 8000
1 Jerry 35 Shanghai 12000
2 Mike 42 Guangzhou 15000
3 Lucy 26 Shenzhen 7000
--- 表格 2 ---
部门 人数
0 技术部 50
1 销售部 30
Code explanation:
read_html()Returns a list of DataFrames, each table corresponding to one DataFrame.- By default, all tables are read.
- The first row is automatically recognized as the column names (because there are
thtags).
Example 2: Filter tables using attrs and match
When a page has multiple tables, you can use attribute or text matching to filter the desired tables.
Example
# Create an HTML file with attributes
html_with_attrs = '''
<table id="employees" class="data-table">
<tr><th>name</th><th>age</th></tr>
<tr><td>Tom</td><td>28</td></tr>
<tr><td>Jerry</td><td>35</td></tr>
</table>
<table id="products" class="data-table">
<tr><th>product</th><th>price</th></tr>
<tr><td>A</td><td>100</td></tr>
<tr><td>B</td><td>200</td></tr>
</table>
<table class="summary">
<tr><td>Total</td><td>2</td></tr>
</table>
'''
with open('tables_attrs.html', 'w', encoding='utf-8') as f:
f.write(html_with_attrs)
# Example 2a: Use attrs to filter by id attribute
# Read the table with id="employees"
tables_by_id = pd.read_html('tables_attrs.html', attrs={'id': 'employees'})
print("Using id to filter:")
print(tables_by_id[0])
print()
# Example 2b: Use attrs to filter by class attribute
# Read all tables with class="data-table"
tables_by_class = pd.read_html('tables_attrs.html', attrs={'class': 'data-table'})
print("Using class to filter (found {} tables):".format(len(tables_by_class)))
for i, df in enumerate(tables_by_class):
print(f"Table {i+1}:")
print(df)
print()
# Example 2c: Use match to filter tables containing specific text
# match uses regular expressions to match text in tables
tables_by_text = pd.read_html('tables_attrs.html', match='Tom')
print("Tables containing the text 'Tom':")
print(tables_by_text[0])
Expected output:
使用 id 筛选: name age 0 Tom 28 1 Jerry 35 使用 class 筛选 (找到 2 个表格): 表格 1: name age 0 Tom 28 1 Jerry 35 表格 2: product price 0 A 100 1 B 200 包含 'Tom' 文本的表格: name age 0 Tom 28 1 Jerry 35
Code explanation:
attrsThe parameter can specify HTML tag attributes (such as id, class, style, etc.) to filter tables.matchThe parameter uses regular expressions to match the text content of the table, returning tables that contain the matching text.- When only one table is returned, you can use
tables[0]to get the first DataFrame.
Example 3: Handle headers and indexes
HTML tables may have complex formats, so headers and indexes need to be handled flexibly.
# Create HTML with a complex structure
html_complex = '''
<!-- Table without a header -->
<table id="no_header">
<tr><td>Tom</td><td>28</td></tr>
<tr><td>Jerry</td><td>35</td></tr>
</table>
<!-- Multi-line header -->
<table id="multi_header">
<tr><th>Name</th><th colspan="2">Contact Information</th></tr>
<tr><th></th><th>Phone</th><th>Email</th></tr>
<tr><td>Tom</td><td>123456</td><td>tom@example.com</td></tr>
<tr><td>Jerry</td><td>789012</td><td>jerry@example.com</td></tr>
</table>
<!-- Table with attributes -->
<table id="with_index" data-type="employee">
<tr><th>Name</th><th>Age</th><th>City</th></tr>
<tr><th></th><th></th><th></th></tr>
<tr><td>Tom</td><td>28</td><td>Beijing</td></tr>
</table>
'''
with open('tables_complex.html', 'w', encoding='utf-8') as f:
f.write(html_complex)
# Example 3a: Case without a header
# header=None does not use the first row as column names
df_no_header = pd.read_html('tables_complex.html', attrs={'id': 'no_header'}, header=None)[0]
print("No header:")
print(df_no_header)
print()
# Example 3b: Multi-line header
# Use row 0 and row 1 together as the header
df_multi_header = pd.read_html('tables_complex.html', attrs={'id': 'multi_header'})[0]
print("Multi-line header:")
print(df_multi_header)
print()
# Example 3c: Set the index column
# index_col specifies column 0 as the index
df_with_index = pd.read_html('tables_complex.html', attrs={'id': 'with_index'}, index_col=0)[0]
print("Setting index:")
print(df_with_index)
Expected output:
没有表头:
0 1
0 Tom 28
1 Jerry 35
多行表头:
姓名 联系方式
NaN 电话 邮箱
0 Tom 123456 tom@example.com
1 Jerry 789012 jerry@example.com
设置索引:
年龄 城市
姓名
Tom 28 Beijing
</空行>
空字符串 NaN NaN
Tom 28 Beijing
Code explanation:
header=NoneYou can disable automatic header recognition and use the default integer index as column names.- Multi-line headers are processed as a MultiIndex.
index_colThe parameter can specify a column to be used as the row index.
Notes
- Using
read_html()requires installinglxmlthe library:pip install lxml。 - It returns a list of DataFrames; you need to select a specific table based on the index.
- When there are multiple tables on the page, use
attrsormatchthe parameter to filter. read_html()It will try to parse all tables, which may be slow. For large pages, you can usematchthe parameter to narrow down the scope.- Reading web pages requires network support, and there may be access restrictions or anti-crawling mechanisms.
Summary
read_html()It is a powerful tool in pandas for reading HTML table data. It can automatically parse tables in web pages and convert them into structured DataFrame format.
In practical work, if you need to obtain data from web pages,read_html()it is a very practical choice. Mastering the use of the attrs and match parameters allows you to efficiently extract the required data from complex pages. It is recommended that readers give priority to this function when they need to scrape web tables.
Pandas Common Functions