Sentiment Analysis

Sentiment Analysis is one of the most classic and widely applied tasks in the field of Natural Language Processing (NLP). It uses computational techniques to automatically identify, extract, and analyze subjective information in text, determining whether the author's attitude toward a specific topic, product, or service is positive, negative, or neutral.


Basic Types of Sentiment Analysis

Classification by Analysis Granularity

  1. Document-Level Sentiment Analysis: Determines the sentiment tendency of the entire document as a whole
  2. Sentence-Level Sentiment Analysis: Analyzes the sentiment polarity of a single sentence
  3. Aspect-Level Sentiment Analysis: Makes sentiment judgments on specific aspects mentioned in the text

Classification by Sentiment Dimension

  1. Binary Classification: Positive/Negative
  2. Three-Class Classification: Positive/Neutral/Negative
  3. Multi-Class Classification: More fine-grained sentiment classification (such as anger, happiness, sadness, etc.)
  4. Sentiment Intensity Analysis: Quantifies the intensity of sentiment

Dictionary-Based Sentiment Analysis Method

The dictionary-based method is the most traditional sentiment analysis technique, relying mainly on pre-built sentiment dictionaries.

Core Components

  1. Sentiment Dictionary: A collection of words with sentiment polarity and intensity

    • Common English dictionaries: SentiWordNet, AFINN, VADER
    • Common Chinese dictionaries: HowNet Sentiment Dictionary, Dalian University of Technology Sentiment Vocabulary Ontology Database
  2. Intensity Modifier: Handles the influence of degree adverbs and negation words

    • Degree adverbs: very (1.5), quite (1.3), a bit (0.8), etc.
    • Negation words: not, no, absolutely not, etc.

Basic Workflow

Example

# Pseudocode example: Dictionary-based sentiment analysis
def lexicon_based_sentiment(text):
    sentiment_score = 0
    words = tokenize(text)  # Word segmentation
    for word in words:
        if word in positive_lexicon:
            sentiment_score += positive_lexicon[word]
        elif word in negative_lexicon:
            sentiment_score -= negative_lexicon[word]
   
    # Handle negation and degree modification
    sentiment_score = apply_negation(words, sentiment_score)
    sentiment_score = apply_intensifier(words, sentiment_score)
   
    return normalize(sentiment_score)

Advantages and Disadvantages Analysis

Advantages:

  • No training data required
  • High computational efficiency
  • Strong interpretability

Disadvantages:

  • Difficult to handle complex linguistic phenomena (such as sarcasm, irony)
  • Relies on the coverage and quality of the dictionary
  • Cannot capture contextual semantics

Machine Learning-Based Sentiment Analysis Method

Machine learning methods perform sentiment analysis by learning patterns from annotated data.

Typical Feature Engineering

  1. Bag of Words Model (BOW): Represents text as a vector of word occurrence frequencies
  2. TF-IDF: Considers the importance of words in the document
  3. N-gram Features: Captures local word sequence patterns
  4. Sentiment Dictionary Features: Combines the advantages of dictionary methods

Common Algorithms

Code Example: Sentiment Classification with Scikit-learn

Example

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import LinearSVC
from sklearn.pipeline import Pipeline

# Build classification pipeline
sentiment_clf = Pipeline([
    ('tfidf', TfidfVectorizer(ngram_range=(1, 2))),
    ('clf', LinearSVC())
])

# Train the model
sentiment_clf.fit(train_texts, train_labels)

# Predict new text
prediction = sentiment_clf.predict(["This product is extremely easy to use, highly recommended!"])
print(prediction)  # Output: 'positive'

Fine-Grained Sentiment Analysis

Aspect-Based Sentiment Analysis (ABSA) is a more advanced sentiment analysis task that aims to identify specific aspects mentioned in text and their corresponding sentiments.

Core Subtasks of ABSA

  1. Aspect Extraction: Identifies entities or attributes discussed in the text

    • Explicit aspect: "The phone's battery life is very good" → "battery"
    • Implicit aspect: "The photos taken are very clear" → "camera"
  2. Sentiment Classification: Makes sentiment judgments for each identified aspect

Comparison of Implementation Methods

Method Type Representative Model Applicable Scenario Advantages Disadvantages
Pipeline Method First use CRF to extract aspects, then use a classifier to determine sentiment Scenarios with limited resources Clear modules, easy to debug Error propagation
End-to-End Method BERT-ABSA、AOA-LSTM High precision requirements Joint optimization, better performance Requires more data
Multi-Task Learning MT-DNN、Multi-Task BERT Related task assistance Knowledge sharing Difficulty in task balancing

Code Example: BERT-Based Aspect-Level Sentiment Analysis

Example

from transformers import BertTokenizer, BertForSequenceClassification
import torch

# Load pre-trained model
model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=3)
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

# Prepare input
text = "The restaurant's environment is great, but the service is too slow."
aspect = "service"
inputs = tokenizer(f"[CLS] {aspect} [SEP] {text} [SEP]", return_tensors="pt")

# Predict sentiment
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=1)
print(predictions)  # Possible output: 1 (negative)

Challenges and Development Directions of Sentiment Analysis

Current Main Challenges

  1. Context Dependence: The same word may have different sentiments in different contexts
  2. Domain Adaptability: Models trained in one domain perform worse in other domains
  3. Multilingual Processing: Sentiment expression varies greatly across different languages
  4. Sarcasm and Irony Detection: Cases where the literal text is opposite to the actual sentiment

Frontier Development Directions

  1. Multimodal Sentiment Analysis: Combines multiple types of information such as text, images, and speech
  2. Cross-Lingual Sentiment Analysis: Leverages commonalities between languages to improve performance on low-resource languages
  3. Sentiment Cause Extraction: Not only determines sentiment, but also analyzes the causes
  4. Personalized Sentiment Analysis: Considers the user's personal characteristics and historical behavior

Practical Exercises

Exercise 1: Building a Basic Sentiment Analyzer

  1. Implement a simple sentiment analyzer using NLTK's VADER dictionary
  2. Test its accuracy on a movie review dataset

Exercise 2: Comparing Different Machine Learning Methods

  1. Train sentiment classifiers using Naive Bayes, SVM, and Logistic Regression respectively
  2. Use cross-validation to compare their performance differences

Exercise 3: Aspect-Level Sentiment Analysis Practice

  1. Fine-tune a pre-trained BERT model on the SemEval 2014 restaurant review dataset
  2. Implement an end-to-end system that can simultaneously extract aspects and determine sentiment

Through this article, you should have mastered the basic concepts, main methods, and implementation techniques of sentiment analysis. As a fundamental NLP task, sentiment analysis technology continues to evolve and has broad value in practical applications, playing an important role from product review analysis to social media monitoring.

Other Extensions