C++ OpenCV Machine Learning and Deep Learning

OpenCV not only supports traditional computer vision tasks, but also provides rich machine learning and deep learning functions, through which complex tasks such as image classification, object detection, and semantic segmentation can be realized.

  • Machine Learning: OpenCV provides a variety of traditional machine learning algorithms, such as KNN, SVM, decision trees, etc.

  • Deep Learning: OpenCV's DNN module supports loading and running pre-trained deep learning models (such as TensorFlow, PyTorch, Caffe, etc.)

Application Scenarios of Machine Learning and Deep Learning

Image Classification:

  • Use machine learning or deep learning models to classify images.

  • Application scenarios: medical image classification, industrial quality inspection, etc.

Object Detection:

  • Use deep learning models to detect objects in images.

  • Application scenarios: autonomous driving, security surveillance, etc.

Semantic Segmentation:

  • Use deep learning models to classify each pixel in an image.

  • Application scenarios: medical image analysis, remote sensing image analysis, etc.

Common Machine Learning Algorithms

OpenCV provides the following common machine learning algorithms:

AlgorithmDescription
KNNK-Nearest Neighbors algorithm, used for classification and regression.
SVMSupport Vector Machine, used for classification and regression.
Decision TreeA tree-based classification and regression algorithm.
Random ForestAn ensemble learning algorithm based on multiple decision trees.
K-MeansClustering algorithm, used to divide data into multiple clusters.

K-Means Clustering

K-Means clustering is an unsupervised learning algorithm mainly used to divide a dataset into K clusters. Each cluster consists of the samples closest to its center point.

The core idea of the K-Means algorithm is to iteratively optimize the cluster centers so that the distance from each sample point to its cluster center is minimized.

Implementation steps:

  1. Initialization: Randomly select K samples as the initial cluster centers.
  2. Assignment: Assign each sample to the nearest cluster center.
  3. Update: Recalculate the center point of each cluster.
  4. Iteration: Repeat steps 2 and 3 until the cluster centers no longer change or the maximum number of iterations is reached.

K-Means in OpenCV:

Example

#include <opencv2/opencv.hpp>
#include <iostream>

int main() {
    // Generate some random data
    cv::Mat data(100, 2, CV_32F);
    cv::randu(data, cv::Scalar(0, 0), cv::Scalar(100, 100));

    // Set K value and iteration conditions
    int K = 3;
    cv::Mat labels, centers;
    cv::kmeans(data, K, labels, cv::TermCriteria(cv::TermCriteria::EPS + cv::TermCriteria::MAX_ITER, 10, 1.0), 3, cv::KMEANS_PP_CENTERS, centers);

    // Output results
    std::cout << "Labels: " << labels << std::endl;
    std::cout << "Centers: " << centers << std::endl;

    return 0;
}

In the above code, we generated 100 two-dimensional data points and used the K-Means algorithm to divide them into 3 clusters.cv::kmeansThe function's parameters include the data, number of clusters, labels, termination conditions, etc.

Support Vector Machine (SVM)

Support Vector Machine (SVM) is a supervised learning algorithm mainly used for classification and regression tasks.

The core idea of SVM is to find a hyperplane that maximizes the margin between sample points of different classes.

Implementation steps:

  1. Training: Find an optimal hyperplane through the training data.
  2. Prediction: Use the trained model to classify new data.

SVM in OpenCV:

Example

#include <opencv2/opencv.hpp>
#include <iostream>

int main() {
    // Generate some training data
    cv::Mat trainData = (cv::Mat_<float>(4, 2) << 1, 1, 1, 2, 2, 1, 2, 2);
    cv::Mat labels = (cv::Mat_<int>(4, 1) << 1, 1, -1, -1);

    // Create an SVM model
    cv::Ptr<cv::ml::SVM> svm = cv::ml::SVM::create();
    svm->setKernel(cv::ml::SVM::LINEAR);
    svm->setType(cv::ml::SVM::C_SVC);
    svm->setC(1);

    // Train the model
    svm->train(trainData, cv::ml::ROW_SAMPLE, labels);

    // Predict new data
    cv::Mat testData = (cv::Mat_<float>(1, 2) << 1.5, 1.5);
    float response = svm->predict(testData);
    std::cout << "Predicted label: " << response << std::endl;

    return 0;
}

In the above code, we use SVM to classify two-dimensional data.

cv::ml::SVM::create()Used to create the SVM model,trainThe method is used to train the model,predictThe method is used to predict the class of new data.

Principal Component Analysis (PCA)

Principal Component Analysis (PCA) is a dimensionality reduction technique mainly used to reduce the dimensionality of a dataset while retaining the main features of the data.

PCA maps the original data to a new coordinate system through a linear transformation, so that the variance of the data in the new coordinate system is maximized.

Implementation steps:

  1. Compute the covariance matrix: Compute the covariance matrix of the dataset.
  2. Eigenvalue decomposition: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvectors and eigenvalues.
  3. Select principal components: Select the eigenvectors corresponding to the top K largest eigenvalues as the principal components.
  4. Data projection: Project the original data onto the principal components to obtain the reduced-dimensional data.

PCA in OpenCV:

Example

#include <opencv2/opencv.hpp>
#include <iostream>

int main() {
    // Generate some data
    cv::Mat data = (cv::Mat_<float>(4, 2) << 1, 2, 2, 3, 3, 4, 4, 5);

    // Create a PCA object
    cv::PCA pca(data, cv::Mat(), cv::PCA::DATA_AS_ROW, 1);

    // Project data
    cv::Mat projected = pca.project(data);
    std::cout << "Projected data: " << projected << std::endl;

    return 0;
}

In the above code, we use PCA to reduce the dimensionality of two-dimensional data.cv::PCAThe class constructor accepts input data, mean vector, data layout, and the number of retained principal components.projectThe method is used to project the data onto the principal components.


OpenCV and Deep Learning

With the wide application of deep learning in the field of computer vision, OpenCV also provides support for deep learning.

Through OpenCV's DNN module, developers can load pre-trained deep learning models and perform tasks such as image classification, object detection, and semantic segmentation.

Loading Pre-trained Deep Learning Models (DNN Module)

OpenCV's DNN module supports loading pre-trained models from multiple deep learning frameworks (such as TensorFlow, Caffe, PyTorch, etc.).

Through the DNN module, developers can easily integrate deep learning models into C++ applications.

Loading the model:

Example

#include <opencv2/opencv.hpp>
#include <opencv2/dnn.hpp>
#include <iostream>

int main() {
    // Load a pre-trained Caffe model
    cv::dnn::Net net = cv::dnn::readNetFromCaffe("deploy.prototxt", "model.caffemodel");

    // Check whether the model was loaded successfully
    if (net.empty()) {
        std::cerr << "Failed to load model!" << std::endl;
        return -1;
    }

    std::cout << "Model loaded successfully!" << std::endl;

    return 0;
}

In the above code, we usecv::dnn::readNetFromCaffethe function to load a Caffe model.deploy.prototxtis the model configuration file,model.caffemodelis the model weight file.

Image Classification, Object Detection, Semantic Segmentation

OpenCV's DNN module supports multiple deep learning tasks, including image classification, object detection, and semantic segmentation. The following is a simple image classification example:

Image Classification:

Example

#include <opencv2/opencv.hpp>
#include <opencv2/dnn.hpp>
#include <iostream>

int main() {
    // Load a pre-trained Caffe model
    cv::dnn::Net net = cv::dnn::readNetFromCaffe("deploy.prototxt", "model.caffemodel");

    // Load the image
    cv::Mat image = cv::imread("image.jpg");
    if (image.empty()) {
        std::cerr << "Failed to load image!" << std::endl;
        return -1;
    }

    // Preprocess the image
    cv::Mat blob = cv::dnn::blobFromImage(image, 1.0, cv::Size(224, 224), cv::Scalar(104, 117, 123));

    // Set the input
    net.setInput(blob);

    // Forward propagation
    cv::Mat prob = net.forward();

    // Get the classification result
    cv::Point classIdPoint;
    double confidence;
    cv::minMaxLoc(prob.reshape(1, 1), 0, &confidence, 0, &classIdPoint);
    int classId = classIdPoint.x;

    std::cout << "Class ID: " << classId << ", Confidence: " << confidence << std::endl;

    return 0;
}

In the above code, we loaded a Caffe model and classified the input image.cv::dnn::blobFromImageThe function is used to convert the image into the model input format,net.forward()is used to perform forward propagation,cv::minMaxLocis used to obtain the classification result.

Running Models like YOLO and SSD with OpenCV

OpenCV's DNN module also supports running object detection models such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector).

The following is an example of using the YOLO model for object detection:

YOLO Object Detection:

Example

#include <opencv2/opencv.hpp>
#include <opencv2/dnn.hpp>
#include <iostream>

int main() {
    // Load the YOLO model
    cv::dnn::Net net = cv::dnn::readNetFromDarknet("yolov3.cfg", "yolov3.weights");

    // Load the image
    cv::Mat image = cv::imread("image.jpg");
    if (image.empty()) {
        std::cerr << "Failed to load image!" << std::endl;
        return -1;
    }

    // Preprocess the image
    cv::Mat blob = cv::dnn::blobFromImage(image, 1/255.0, cv::Size(416, 416), cv::Scalar(0, 0, 0), true, false);

    // Set the input
    net.setInput(blob);

    // Forward propagation
    std::vector<cv::Mat> outs;
    net.forward(outs, net.getUnconnectedOutLayersNames());

    // Parse the detection results
    for (size_t i = 0; i < outs.size(); ++i) {
        float* data = (float*)outs[i].data;
        for (int j = 0; j < outs[i].rows; ++j, data += outs[i].cols) {
            cv::Mat scores = outs[i].row(j).colRange(5, outs[i].cols);
            cv::Point classIdPoint;
            double confidence;
            cv::minMaxLoc(scores, 0, &confidence, 0, &classIdPoint);
            if (confidence > 0.5) {
                int centerX = (int)(data[0] * image.cols);
                int centerY = (int)(data[1] * image.rows);
                int width = (int)(data[2] * image.cols);
                int height = (int)(data[3] * image.rows);
                int left = centerX - width / 2;
                int top = centerY - height / 2;

                cv::rectangle(image, cv::Point(left, top), cv::Point(left + width, top + height), cv::Scalar(0, 255, 0), 2);
            }
        }
    }

    // Display the results
    cv::imshow("Detection", image);
    cv::waitKey(0);

    return 0;
}

In the above code, we loaded a YOLO model and performed object detection on the input image.cv::dnn::readNetFromDarknetThe function is used to load the YOLO model,net.forwardis used to perform forward propagation,cv::rectangleis used to draw detection boxes.

Other Extensions