Pandas pd.read_json() Function
read_json()It is a function in the pandas library used to read JSON (JavaScript Object Notation) files, supporting the import of data in multiple JSON formats.
JSON is a lightweight data interchange format that is easy for humans to read and write, and easy for machines to parse and generate. It is widely used in scenarios such as Web APIs and configuration files.read_json()It can convert JSON data into pandas DataFrame format, facilitating data analysis.
Basic Syntax and Parameters
Syntax Format
pandas.read_json(path_or_buf, orient=None, typ='frame', dtype=None,
convert_axes=None, convert_dates=True, keep_default_dates=True,
numpy=False, precise_float=False, date_unit='ms', ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| path_or_buf | str, path object, or file-like object | JSON file path, URL, or string | Required |
| orient | str | The format of the JSON data: 'split', 'records', 'index', 'columns', 'values' | None |
| typ | str | Return type: 'frame' returns DataFrame, 'series' returns Series | 'frame' |
| dtype | dict | Specify the data types of columns | None |
| convert_axes | bool | Whether to convert axes to datetime | None |
| convert_dates | bool, list | Whether to convert date columns | True |
| numpy | bool | Whether to use numpy arrays | False |
Return Value
- Return Type:
pd.DataFrameorpd.Series - Returns DataFrame by default, which can be in a two-dimensional table form.
- When
typ='series'Returns Series when typ='series'.
Examples
Through the following examples, comprehensively masterread_json()the various usages.
Example 1: Reading JSON Data in Different Formats
JSON data has multiple formats,read_json()[It] supports the most common formats.
Example
import json
# Example data: JSON array format (records format)
json_records = '''
[
{"name": "Tom", "age": 28, "city": "Beijing", "salary": 8000},
{"name": "Jerry", "age": 35, "city": "Shanghai", "salary": 12000},
{"name": "Mike", "age": 42, "city": "Guangzhou", "salary": 15000}
]
'''
# Write the JSON string to a file
with open('data_records.json', 'w', encoding='utf-8') as f:
f.write(json_records)
# Read the JSON file (records format, most commonly used)
# orient='records' means each row is a JSON object
df_records = pd.read_json('data_records.json', orient='records')
print("Records format:")
print(df_records)
print()
# Example data: JSON object format (index format)
json_index = '''
{
"Tom": {"age": 28, "city": "Beijing", "salary": 8000},
"Jerry": {"age": 35, "city": "Shanghai", "salary": 12000},
"Mike": {"age": 42, "city": "Guangzhou", "salary": 15000}
}
'''
with open('data_index.json', 'w', encoding='utf-8') as f:
f.write(json_index)
# Read the JSON file (index format, using a field as the index)
df_index = pd.read_json('data_index.json', orient='index')
print("Index format:")
print(df_index)
print()
# Example data: JSON column format (columns format)
json_columns = '''
{
"name": ["Tom", "Jerry", "Mike"],
"age": [28, 35, 42],
"city": ["Beijing", "Shanghai", "Guangzhou"],
"salary": [8000, 12000, 15000]
}
'''
with open('data_columns.json', 'w', encoding='utf-8') as f:
f.write(json_columns)
# Read the JSON file (columns format)
df_columns = pd.read_json('data_columns.json', orient='columns')
print("Columns format:")
print(df_columns)
Expected output:
Records 格式:
name age city salary
0 Tom 28 Beijing 8000
1 Jerry 35 Shanghai 12000
2 Mike 42 Guangzhou 15000
Index 格式:
age city salary
Tom 28 Beijing 8000
Jerry 35 Shanghai 12000
Mike 42 Guangzhou 15000
Columns 格式:
name age city salary
0 Tom 28 Beijing 8000
1 Jerry 参数 Shanghai 12000
2 Mike 42 Guangzhou 15000
Code explanation:
orient='records': JSON array format, each row is a JSON object, the most commonly used format.orient='index': JSON object format, keys serve as the index.orient='columns': JSON column format, keys are column names, values are arrays.- Correctly specifying the
orientparameter is crucial for correctly parsing JSON data.
Example 2: Reading JSON from Strings and URLs
read_json()Not only can it read files, but it also supports reading data from strings and URLs.
Example
import json
from io import StringIO
# Example 2a: Read from a JSON string
json_string = '''
[
{"product": "A", "sales": 100, "region": "North"},
{"product": "B", "sales": 200, "region": "South"},
{"product": "C", "sales": 150, "region": "East"}
]
'''
# Use StringIO to convert the string into a file object
df_from_string = pd.read_json(StringIO(json_string))
print("Read from string:")
print(df_from_string)
print()
# You can also pass the JSON string directly in the parameter
# Note: Python strings need to be properly escaped
json_str_direct = '[{"product": "A", "sales": 100}, {"product": "B", "sales": 200}]'
df_direct = pd.read_json(json_str_direct)
print("Pass the string directly:")
print(df_direct)
print()
# Example 2b: Read JSON Lines format (one JSON object per line)
# JSON Lines is a common log format
json_lines = '''{"name": "Tom", "score": 85}
{"name": "Jerry", "score": 92}
{"name": "Mike", "score": 78}
{"name": "Lucy", "score": 95}'''
with open('data_lines.json', 'w', encoding='utf-8') as f:
f.write(json_lines)
# JSON Lines format requires reading line by line
# You can use the lines=True parameter (if the JSONL format supports it)
# Or process it manually
df_list = []
with open('data_lines.json', 'r', encoding='utf-8') as f:
for line in f:
df_list.append(json.loads(line))
df_lines = pd.DataFrame(df_list)
print("JSON Lines format reading:")
print(df_lines)
print()
# Example 2c: Read from an API URL (requires network access)
# An example API is used here; replace it with a real URL in actual use
# df_api = pd.read_json('https://api.example.com/data')
# print(df_api)
print("Note: Reading from a URL requires an actual network request")
Expected output:
从字符串读取:
product sales region
0 A 100 North
1 B 200 South
2 C 150 East
直接传入字符串:
product sales
0 A 100
1 B 200
JSON Lines 格式读取:
name score
0 Tom 85
1 Jerry 92
2 Mike 78
3 Lucy 95
注意:从 URL 读取需要实际的网络请求
Code explanation:
read_json()It can accept a JSON string as input; useStringIOto convert the string into a file-like object.- JSON Lines format has one independent JSON object per line, commonly used in log processing; it needs to be read line by line and then merged into a DataFrame.
- When reading from a URL, simply pass the URL string directly, but network support is required.
Example 3: Handling Dates and Type Conversion
Dates and numeric types in JSON data need special handling.
Example
# Example 3a: Handle date fields
json_with_date = '''
[
{"name": "Tom", "birthday": "1995-03-15", "join_date": "2020-01-10"},
{"name": "Jerry", "birthday": "1988-07-22", "join_date": "2019-03-05"},
{"name": "Mike", "birthday": "1981-11-30", "join_date": "2018-06-20"}
]
'''
with open('data_with_date.json', 'w', encoding='utf-8') as f:
f.write(json_with_date)
# By default, date strings are read as object type
df_date = pd.read_json('data_with_date.json')
print("Default reading (dates as strings):")
print(df_date)
print("birthday type:", df_date['birthday'].dtype)
print()
# Use convert_dates to automatically convert date columns
df_date_converted = pd.read_json('data_with_date.json', convert_dates=['birthday', 'join_date'])
print("After converting dates:")
print(df_date_converted)
print("birthday type:", df_date_converted['birthday'].dtype)
print()
# Example 3b: Specify data types
json_mixed = '''
[
{"id": "1", "name": "Tom", "score": 85.5},
{"id": "2", "name": "Jerry", "score": 92.0},
{"id": "3", "name": "Mike", "score": 78.5}
]
'''
with open('data_mixed.json', 'w', encoding='utf-8') as f:
f.write(json_mixed)
# By default, id is read as an integer, name as a string, and score as a float
df_mixed = pd.read_json('data_mixed.json')
print("Default type inference:")
print(df_mixed)
print("id type:", df_mixed['id'].dtype)
print()
# Use dtype to explicitly specify types
df_typed = pd.read_json('data_mixed.json', dtype={'id': str, 'score': float})
print("After specifying types:")
print(df_typed)
print("id type:", df_typed['id'].dtype)
Expected output:
默认读取(日期为字符串):
name birthday join_date
0 Tom 1995-03-15 2020-01-10
1 Jerry 1988-07-22 2019-03-05
2 Mike 1981-11-30 2018-06-20
默认读取(日期为字符串):
name birthday join_date
0 汤 Tom 1995-03-15 pandas 的 to_json() 和 read_json() 的完整配对示例
1 Jerry 1988-07-22 2019-03-05
2 Mike 1981-11-30 确保 JSON 数据格式一致
1981-配对 to_json(orient='records') 配对 read_json(orient='records')
3 Mike 1981-11-30 to_json(orient='records')
...
join_date 2018-06-20
birthday 类型: object
转换日期后:
birthday 类型: datetime64[ns]
Code explanation:
convert_datesThe parameter can specify which columns need to be converted to date types.- By default,
convert_dates=True[it] automatically recognizes common date formats. dtypeThe parameter can explicitly specify the data type of each column, avoiding type inference errors.
Notes
- Correctly specifying the
orientparameter is the key to reading JSON data; different JSON structures require different orient values. - Dates in JSON are read as strings by default; you need to use the
convert_datesparameter to convert them to date types. - When reading large JSON files, consider using the
chunksizeparameter to read in chunks. read_json()[It] supports reading from file paths, URLs, and JSON strings.- JSON Lines format needs to be parsed line by line and then merged into a DataFrame.
Summary
read_json()It is the core function in pandas for reading JSON data, supporting multiple JSON formats. As a common format for Web APIs and data exchange, JSON is widely used in practical data analysis work.
Masteringread_json()the key is understanding the differentorientformats, as well as date and type handling methods. Readers are advised to practice reading JSON data in different formats in their actual work.
Pandas Common Functions