MongoDB Map Reduce

Map-Reduce is a computational model. Simply put, it decomposes large batches of work (data) into smaller tasks (MAP) for execution, and then merges the results into the final result (REDUCE).

MongoDB's Map-Reduce is very flexible and quite practical for large-scale data analysis.

MapReduce Commands

The following is the basic syntax of MapReduce:

>db.collection.mapReduce(
   function() {emit(key,value);},  //map 函数
   function(key,values) {return reduceFunction},   //reduce 函数
   {
      out: collection,
      query: document,
      sort: document,
      limit: number
   }
)

To use MapReduce, you need to implement two functions: the Map function and the Reduce function. The Map function calls emit(key, value), iterates through all records in the collection, and passes the key and value to the Reduce function for processing.

The Map function must call emit(key, value) to return key-value pairs.

Parameter description:

  • map: mapping function (generates a sequence of key-value pairs, used as the reduce function parameter).
  • reduceReduce function. The task of the reduce function is to turn key-values into key-value, that is, to turn the values array into a single value.
  • outThe collection where the statistical results are stored (if not specified, a temporary collection is used, which is automatically deleted after the client disconnects).
  • queryA filter condition; only documents that meet the condition will call the map function. (query. limit, sort can be freely combined)
  • sortThe sort parameter combined with limit (also sorts documents before sending them to the map function), which can optimize the grouping mechanism
  • limitThe upper limit of the number of documents sent to the map function (without limit, using sort alone is not very useful)

The following example finds data with status:"A" in the orders collection, groups by cust_id, and calculates the sum of amount.


Using MapReduce

Consider the following document structure for storing user articles. The document stores the user's user_name and the article's status field:

>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "mark",
   "status":"active"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "mark",
   "status":"active"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "mark",
   "status":"active"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "mark",
   "status":"active"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "mark",
   "status":"disabled"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "example",
   "status":"disabled"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "example",
   "status":"disabled"
})
WriteResult({ "nInserted" : 1 })
>db.posts.insert({
   "post_text": "Example,最全的技术文档。",
   "user_name": "example",
   "status":"active"
})
WriteResult({ "nInserted" : 1 })

Now, we will use the mapReduce function on the posts collection to select published articles (status:"active"), group by user_name, and calculate the number of articles for each user:

>db.posts.mapReduce( 
   function() { emit(this.user_name,1); }, 
   function(key, values) {return Array.sum(values)}, 
      {  
         query:{status:"active"},  
         out:"post_total" 
      }
)

The output of the above mapReduce is:

{
        "result" : "post_total",
        "timeMillis" : 23,
        "counts" : {
                "input" : 5,
                "emit" : 5,
                "reduce" : 1,
                "output" : 2
        },
        "ok" : 1
}

The result shows that there are 5 documents matching the query condition (status:"active"), 5 key-value pair documents were generated in the map function, and finally the reduce function grouped the identical key values into 2 groups.

Specific parameter description:

  • result: the name of the collection storing the results. This is a temporary collection that is automatically deleted after the MapReduce connection is closed.
  • timeMillis: the time taken for execution, in milliseconds
  • input: the number of documents that meet the condition and are sent to the map function
  • emit: the number of times emit is called in the map function, i.e., the total amount of data in all collections
  • output: the number of documents in the result collection(count is very helpful for debugging)
  • ok: whether it succeeded, 1 for success
  • err: if it fails, the failure reason can be found here. However, from experience, the reason is relatively vague and not very useful.

Use the find operator to view the query results of mapReduce:

> var map=function() { emit(this.user_name,1); }
> var reduce=function(key, values) {return Array.sum(values)}
> var options={query:{status:"active"},out:"post_total"}
> db.posts.mapReduce(map,reduce,options)
{ "result" : "post_total", "ok" : 1 }
> db.post_total.find();

The above query displays the following results:

{ "_id" : "mark", "value" : 4 }
{ "_id" : "example", "value" : 1 }

In a similar way, MapReduce can be used to build large, complex aggregation queries.

The Map function and Reduce function can be implemented using JavaScript, making MapReduce very flexible and powerful.

Other Extensions