Keras First Neural Network

Keras is a high-level neural network API written in Python that can run on top of TensorFlow, CNTK, or Theano. Its development focuses on supporting fast experimentation, enabling rapid transition from idea to result with minimal code.

Main Features of Keras

  • User-friendly: Keras has a simple and consistent interface
  • Modular: The various components of a neural network (layers, optimizers, initialization schemes, etc.) are composable modules
  • Easy extensibility: New modules can be easily added to express new research ideas
  • Multi-backend support: Can seamlessly switch between TensorFlow, Theano, and CNTK as the computational backend

Installing Keras

Before we begin, we need to install Keras and its backend engine (here we use TensorFlow):

pip install tensorflow keras

Note: Keras 2.4.0 and later versions have been integrated into TensorFlow and can be directly used throughtensorflow.kerasUse


Building the First Neural Network

Let's start with a simple fully connected neural network to solve the classic MNIST handwritten digit recognition problem.

1. Import the necessary libraries

Example

import numpy as np
from tensorflow import keras
from tensorflow.keras import layers

2. Prepare the data

The MNIST dataset contains 60,000 training images and 10,000 test images, each of which is a 28x28 pixel grayscale image of a handwritten digit.

Example

# Load data
(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()

# Preprocess data
x_train = x_train.reshape(60000, 784).astype("float32") / 255
x_test = x_test.reshape(10000, 784).astype("float32") / 255

# Convert labels to one-hot encoding
y_train = keras.utils.to_categorical(y_train, 10)
y_test = keras.utils.to_categorical(y_test, 10)

3. Build the model

We will build a simple fully connected network containing an input layer, a hidden layer, and an output layer.

Example

model = keras.Sequential([
    layers.Dense(512, activation="relu", input_shape=(784,)),
    layers.Dense(10, activation="softmax")
])

Model structure analysis

Example

graph TD
A[Input layer 784 neurons] --> B[Hidden layer 512 neurons, ReLU activation]
B --> C[Output layer 10 neurons, Softmax activation]

4. Compile the model

Before training the model, we need to configure the learning process:

Example

model.compile(
    optimizer="rmsprop",
    loss="categorical_crossentropy",
    metrics=["accuracy"]
)

Compile parameter description

Parameter Description Common values
optimizer Optimizer, used to update weights "rmsprop", "adam", "sgd"
loss Loss function, measures the difference between model predictions and true values "categorical_crossentropy" (classification), "mse" (regression)
metrics Evaluation metric, used to monitor training ["accuracy"]

5. Train the model

Now we can start training the model:

Example

history = model.fit(
    x_train, y_train,
    batch_size=128,
    epochs=10,
    validation_split=0.2
)

Training parameter description

Parameter Description Suggested values
batch_size Number of samples used per gradient update 32-256
epochs Number of training epochs Adjust based on data complexity
validation_split Proportion of training data used as validation set 0.1-0.3

6. Evaluate the model

After training is complete, we can evaluate model performance on the test set:

Example

test_loss, test_acc = model.evaluate(x_test, y_test)
print(f"Test accuracy: {test_acc:.4f}")

Complete code example

Example

import numpy as np
from tensorflow import keras
from tensorflow.keras import layers

# 1. Load data
(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()

# 2. Preprocess
x_train = x_train.reshape(60000, 784).astype("float32") / 255
x_test = x_test.reshape(10000, 784).astype("float32") / 255
y_train = keras.utils.to_categorical(y_train, 10)
y_test = keras.utils.to_categorical(y_test, 10)

# 3. Build the model
model = keras.Sequential([
    layers.Dense(512, activation="relu", input_shape=(784,)),
    layers.Dense(10, activation="softmax")
])

# 4. Compile the model
model.compile(
    optimizer="rmsprop",
    loss="categorical_crossentropy",
    metrics=["accuracy"]
)

# 5. Train the model
history = model.fit(
    x_train, y_train,
    batch_size=128,
    epochs=10,
    validation_split=0.2
)

# 6. Evaluate the model
test_loss, test_acc = model.evaluate(x_test, y_test)
print(f"Test accuracy: {test_acc:.4f}")

Model improvement suggestions

1. Add a Dropout layer: Prevent overfitting

Example

model.add(layers.Dropout(0.5))

2. Use a more advanced optimizer: Such as Adam

Example

model.compile(optimizer="adam", ...)

3. Add hidden layers: Build a deeper network

Example

model.add(layers.Dense(256, activation="relu"))

4. Use convolutional layers: More effective for image data

Example

model.add(layers.Conv2D(32, (3, 3), activation="relu"))

FAQ

Q1: Why is my model accuracy so low?

  • Check whether the data preprocessing is correct
  • Try adjusting the learning rate
  • Increase network capacity (more layers or more neurons)

Q2: What should I do if the loss does not decrease during training?

  • Check whether there are problems with the data
  • Try different optimizers
  • Adjust the learning rate (usually reduce it)

Q3: How to save and load a trained model?

Example

# Save the model
model.save("mnist_model.h5")

# Load the model
loaded_model = keras.models.load_model("mnist_model.h5")
Other extensions