Neural network logistic regression sample code provides a clear bridge between classic statistical modeling and modern deep‑learning frameworks. By treating logistic regression as a single‑neuron network with a sigmoid activation, you can see how the same mathematical principles underlie both fields and how easy it is to extend the model to deeper architectures. The following guide walks you through the theory, the required Python environment, and a complete, runnable example that you can adapt for binary classification tasks But it adds up..
Introduction
Logistic regression is often introduced as a statistical method for predicting the probability of a binary outcome. Here's the thing — when you view it through the lens of neural networks, it becomes the simplest possible network: one input layer, one neuron (also called a perceptron), and a sigmoid activation function that squashes the linear combination of features into the [0, 1] interval. This perspective is valuable because it lets you reuse the same tools—gradient descent, back‑propagation, and vectorized NumPy operations—that power far more complex models.
Not obvious, but once you see it — you'll see it everywhere.
- Why logistic regression equals a one‑layer neural network.
- How to implement the model from scratch using only NumPy.
- How to verify the implementation with a synthetic dataset.
- Where to plug in popular libraries such as TensorFlow or PyTorch if you prefer a higher‑level API.
By the end, you will have a sample code snippet that you can copy, modify, and use as a building block for deeper networks.
Understanding Logistic Regression as a Neural Network
Core Idea
A logistic regression model computes
[ \hat{y}= \sigma(\mathbf{w}^\top \mathbf{x}+b) ]
where
- (\mathbf{x}) is the feature vector,
- (\mathbf{w}) is the weight vector,
- (b) is the bias term, and
- (\sigma(z)=\frac{1}{1+e^{-z}}) is the sigmoid function.
In neural‑network terminology:
- The input layer passes (\mathbf{x}) unchanged.
- The hidden layer consists of a single neuron that computes the linear combination (\mathbf{w}^\top \mathbf{x}+b).
- The activation function of that neuron is the sigmoid, producing the predicted probability (\hat{y}).
Because there is only one neuron, there is no hidden layer beyond the output; the network depth is 1. Training proceeds by minimizing the binary cross‑entropy loss
[ \mathcal{L}= -\frac{1}{N}\sum_{i=1}^{N}\big[ y_i\log(\hat{y}_i)+(1-y_i)\log(1-\hat{y}_i)\big] ]
using gradient descent (or any of its variants). The gradients are identical to those derived in classical logistic regression, which confirms the equivalence The details matter here..
Why This Matters
- Conceptual unity – Seeing logistic regression as a neural network demystifies the transition to multilayer perceptrons.
- Code reuse – The same forward‑pass and back‑propagation loops work for any number of layers; you only need to add more neurons.
- Debugging – If a deep network fails to learn, you can first verify that the single‑neuron version converges on the same data.
Setting Up the Environment
You only need a recent Python 3.x release and the NumPy library for the from‑scratch version. If you want to compare with a high‑level framework, install TensorFlow or PyTorch as well.
# Create a clean virtual environment (optional but recommended)
python -m venv nn_logreg_env
source nn_logreg_env/bin/activate # on Windows: nn_logreg_env\Scripts\activate
# Install dependencies
pip install numpy tqdm # tqdm for a nice progress bar (optional)
# For framework comparisons:
# pip install tensorflow # or pip install torch
All code snippets below assume these packages are available.
Step‑by‑Step Implementation (NumPy Only)
Below is a detailed walkthrough of each component. Feel free to copy the blocks into a single script or a Jupyter notebook.
1. Generate a Synthetic Binary Dataset
import numpy as np
def make_classification(n_samples=1000, n_features=2, sep=1.hstack([np.zeros(n_samples//2), np.random.On top of that, 5, random_state=42):
rng = np. Now, default_rng(random_state)
# Create two Gaussian clouds with different means
X0 = rng. 5, size=(n_samples//2, n_features))
X = np.normal(loc=[-sep, -sep], scale=0.That's why 5, size=(n_samples//2, n_features))
X1 = rng. vstack([X0, X1])
y = np.normal(loc=[ sep, sep], scale=0.ones(n_samples//2)])
# Shuffle
perm = rng.
X, y = make_classification()
2. Initialize Parameters
def initialize_parameters(n_features):
# Small random weights help break symmetry
w = np.random.randn(n_features) * 0.01
b = 0.0
return w, b
3. Forward Propagation
def sigmoid(z):
return 1.0 / (1.0 + np.exp(-z))
def forward_propagation(X, w, b):
z = np.dot(X, w) + b # linear step
a = sigmoid(z) # activation → predicted probabilities
return z, a
4. Compute Loss and Gradients
def compute_loss_and_gradients(X, y, w, b):
m = X.shape[0]
_, a = forward_propagation(X, w, b) # a = predictions
# Binary cross‑entropy loss
loss = -(np.mean(y * np.log(a + 1e-15) + (1 - y) * np.log(1 - a + 1e-15)))
# Gradients (derivative of loss w.r.t. w and b)
dz = a - y #
### 4. Compute Loss and Gradients (continued)
```python
# Gradients (derivative of loss w.r.t. w and b)
dz = a - y
dw = np.dot(X.T, dz) / m
db = np.mean(dz)
return loss, dw, db
5. Gradient Descent Update
def update_parameters(w, b, dw, db, learning_rate=0.1):
w -= learning_rate * dw
b -= learning_rate * db
return w, b
6. Model Training Loop
def train_model(X, y, epochs=1000, learning_rate=0.1, verbose=True):
w, b = initialize_parameters(X.shape[1])
losses = []
for epoch in range(epochs):
loss, dw, db = compute_loss_and_gradients(X, y, w, b)
w, b = update_parameters(w, b, dw, db, learning_rate)
losses.append(loss)
if verbose and epoch % 100 == 0:
print(f"Epoch {epoch:4d}: loss = {loss:.4f}")
return w, b, losses
# Train the single-neuron model
w_final, b_final, loss_history = train_model(X, y)
7. Evaluate the Trained Model
def predict(X, w, b, threshold=0.5):
_, a = forward_propagation(X, w, b)
return (a >= threshold).astype(int)
y_pred = predict(X, w_final, b_final)
accuracy = np.mean(y_pred == y)
print(f"\nFinal accuracy: {accuracy:.2%}")
Extending to Multiple Layers
The true power emerges when we stack neurons into layers. Here’s how to add a hidden layer with ReLU activation:
def relu(z):
return np.maximum(0, z)
def initialize_layer(input_dim, output_dim):
# He initialization for ReLU
return np.random.randn(input_dim, output_dim) * np.sqrt(2.
class NeuralNetwork:
def __init__(self, layer_sizes):
self.Which means params = []
for i in range(len(layer_sizes) - 1):
W = initialize_layer(layer_sizes[i], layer_sizes[i+1])
b = np. Also, zeros((1, layer_sizes[i+1]))
self. That said, params. Think about it: append((W, b))
def forward(self, X):
cache = []
A = X
for W, b in self. In real terms, params[:-1]:
Z = np. dot(A, W) + b
A = relu(Z)
cache.append((A, Z))
# Output layer (sigmoid for binary classification)
W, b = self.params[-1]
Z = np.Worth adding: dot(A, W) + b
A = sigmoid(Z)
cache. Practically speaking, append((A, Z))
return A, cache
def backward(self, X, y, cache):
m = X. shape[0]
grads = []
# Output layer gradients
A_out, Z_out = cache[-1]
dZ = A_out - y.reshape(-1, 1)
dW = np.Because of that, dot(cache[-2][0]. On top of that, t, dZ) / m
db = np. mean(dZ, axis=0, keepdims=True)
grads.Also, insert(0, (dW, db))
# Backprop through hidden layers
for i in range(len(self. Worth adding: params)-2, 0, -1):
dZ = np. dot(dZ, self.Consider this: params[i+1][0]. T) * (cache[i][1] > 0)
dW = np.dot(cache[i-1][0].T, dZ) / m
db = np.mean(dZ, axis=0, keepdims=True)
grads.
### Training the Multi-Layer Network
```python
# Create a 2-layer network: 2 inputs → 4 hidden neurons → 1 output
nn = NeuralNetwork([2, 4, 1])
learning_rate = 0.01
for epoch in range(2000):
A, cache = nn.forward(X)
loss = -(np.mean(y * np.log(A + 1e-15) + (1-y) * np.log(1-A + 1e-15)))
grads = nn.backward(X, y, cache)
# Update parameters
for i, (W, b) in enumerate(nn.params):
dW, db = grads[i]
nn.
```python
# Complete the training loop
for epoch in range(2000):
A, cache = nn.forward(X)
loss = -(np.mean(y * np.log(A + 1e-15) + (1-y) * np.log(1-A + 1e-15)))
grads = nn.backward(X, y, cache)
# Parameter update
for i, (W, b) in enumerate(nn.params):
dW, db = grads[i]
nn.params[i] = (W - learning_rate * dW, b - learning_rate * db)
# Report progress every 200 epochs
if epoch % 200 == 0:
print(f"Epoch {epoch:4d}: loss {loss:.4f}")
Evaluation after training
# Run a forward pass with the learned parameters
A_final, _ = nn.forward(X)
y_pred = (A_final >= 0.5).astype(int)
accuracy = np.mean(y_pred == y)
print(f"Final accuracy after training: {accuracy:.2%}")
The code above finishes the training routine, prints the loss at regular intervals, and then measures the model’s performance on the training set. By adjusting the learning rate, the number of epochs, or the size of the hidden layer, you can observe how the network’s ability to capture complex patterns evolves.
Conclusion
Stacking a hidden layer equipped with a ReLU activation function transforms a simple linear classifier into a flexible function approximator capable of modeling non‑linear relationships. Practically speaking, the backward propagation implementation automatically computes gradients for every weight matrix and bias term, allowing a straightforward stochastic update to minimize the binary cross‑entropy loss. This modular design serves as a solid foundation for experimenting with deeper architectures, alternative activation functions, and more sophisticated optimization strategies in future deep‑learning endeavors That's the whole idea..
And yeah — that's actually more nuanced than it sounds.