MongoDB GridFS
GridFS is used to store and retrieve files that exceed 16M (the BSON document limit), such as images, audio, video, etc.
GridFS is also a way of file storage, but it stores files in MongoDB collections.
GridFS can store files larger than 16M more efficiently.
GridFS splits large file objects into multiple small chunks (file fragments), generally 256k each. Each chunk is stored in the chunks collection as a MongoDB document.
GridFS uses two collections to store a file: fs.files and fs.chunks.
The actual content of each file is stored in chunks (binary data), and file-related metadata (filename, content_type, and user-defined attributes) will be stored in the files collection.
The following is a simple fs.files collection document:
{
"filename": "test.txt",
"chunkSize": NumberInt(261120),
"uploadDate": ISODate("2014-04-13T11:32:33.557Z"),
"md5": "7b762939321e146569b07f72c62cca4f",
"length": NumberInt(646)
}
The following is a simple fs.chunks collection document:
{
"files_id": ObjectId("534a75d19f54bfec8a2fe44b"),
"n": NumberInt(0),
"data": "Mongo Binary Data"
}
GridFS Add File
Now we use the GridFS put command to store mp3 files. Invoke the mongofiles.exe tool in the bin directory under the MongoDB installation directory.
Open the command prompt, go to the bin directory of the MongoDB installation directory, find mongofiles.exe, and enter the following code:
>mongofiles.exe -d gridfs put song.mp3
-d gridfsSpecify the database name for storing the file. If the database does not exist, MongoDB will automatically create it. If the database does not exist, MongoDB will automatically create it. Song.mp3 is the audio file name.
Use the following command to view the file documents in the database:
>db.fs.files.find()
After executing the above command, the following document data is returned:
{
_id: ObjectId('534a811bf8b4aa4d33fdf94d'),
filename: "song.mp3",
chunkSize: 261120,
uploadDate: new Date(1397391643474), md5: "e4f53379c909f7bed2e9d631e15c1c41",
length: 10401959
}
We can see all the chunks in the fs.chunks collection. Below we get the _id value of the file. We can use this _id to obtain the chunk data:
>db.fs.chunks.find({files_id:ObjectId('534a811bf8b4aa4d33fdf94d')})
In the above example, the query returned data from 40 documents, meaning the mp3 file is stored in 40 chunks.
Other Extensions