Year 1 · group prototype, then individual extension · reworked in 2026
Facial-expression recognition: building a CNN, then checking it
My first substantial computer-vision project: a CNN that labels facial expressions in images and video, and a separate face-clustering experiment. The 2024 report gave 86.5%; a later review showed that figure came from a leaky split, with held-out results near 61%. In 2026 I rebuilt it as EmoSense, a modular app that compares four architectures.
- Context
- University projects, Year 1 · reworked as EmoSense in 2026
- My role
- Main developer (group prototype) · sole developer (individual extension)
- Team
- Group of 3 (prototype) · individual (extension)
- Dates
- Jan – May 2024 · EmoSense 2026
- Status
- January 2024 prototype and 2026 EmoSense on GitHub · May 2024 extension not on GitHub
Stack
- Python
- TensorFlow / Keras
- OpenCV
- face_recognition
- scikit-learn
- pandas
Scope of this pageTwo public repositories: Project-2 holds the January 2024 group prototype, and EmoSense is my 2026 rework. The May 2024 individual extension, which produced the report's headline figure, is not on GitHub. The saved-weight check and the clean retrain below are separate experiments, kept with this portfolio's source.
The short version
Each point is expanded, with diagrams, in the sections below.
- 01Problem and constraints
- Label seven facial expressions from small greyscale faces, then apply the model to video. The first group prototype struggled with a reduced dataset and notebook crashes.
- 02My responsibility
- I wrote the main code for the group prototype, then built the full-dataset model, video pipeline and face clustering alone.
- 03The design
- A four-layer CNN trained on FER-2013; OpenCV face detection feeds face crops to it; face clustering is a separate pipeline.
- 04What I implemented
- Data conversion, the CNN, frame-by-frame video annotation and K-means grouping of face encodings.
- 05The team’s part
- The January prototype was shared coursework with two teammates; the May extension was individual.
- 06What was delivered
- In 2024, an annotated-video prototype, cluster montages and the report. In 2026, EmoSense: named experiments, a webcam app and a small API.
- 07What testing established
- Reported 86.5% on a random split. Scored on held-out partitions, the saved weights reach 61.9%; a clean retrain reaches 60.7%.
- 08Learned, and unfinished
- A result is only as good as its data boundary, and a bigger network is not automatically better. Expression labels are not emotions.
On this page
From a failing prototype to a working model
The first version, in a January 2024 group project, trained a small network on a few thousand images because the full dataset kept crashing our notebook environment. Its accuracy was too low to be useful, and I was the one writing most of that code.
In an individual module that spring I rebuilt it on all 35,887 FER-2013 images with a deeper network, extended the video processing with OpenCV face detection, and added a separate experiment that groups similar faces. It was my first end-to-end computer-vision system, from pixel strings to an annotated video.
The network
Four convolution layers with pooling and dropout reduce each 48 × 48 face to a small stack of feature maps, which a dense layer turns into scores for seven labels: angry, disgust, fear, happy, neutral, sad and surprise. About 2.3 million of its 2.35 million parameters sit in the dense layer.
The CNN, layer by layer
- Input · greyscale48×48×1
- Conv 32 · 3×3, ReLU46×46×32
- Conv 64 · 3×3, ReLU44×44×64
- Max pool · 2×222×22×64
- Dropout 0.25
- Conv 128 · 3×3, ReLU20×20×128
- Max pool · 2×210×10×128
- Conv 128 · 3×3, ReLU8×8×128
- Max pool · 2×24×4×128
- Dropout 0.25
- Flatten2,048
- Dense 1,024 · ReLU1,024
- Dropout 0.5
- Dense 7 · softmax7
Text version of this diagram
- Input (greyscale) — output 48 × 48 × 1
- Conv 32 (3×3, ReLU) — output 46 × 46 × 32
- Conv 64 (3×3, ReLU) — output 44 × 44 × 64
- Max pool (2×2) — output 22 × 22 × 64
- Dropout 0.25
- Conv 128 (3×3, ReLU) — output 20 × 20 × 128
- Max pool (2×2) — output 10 × 10 × 128
- Conv 128 (3×3, ReLU) — output 8 × 8 × 128
- Max pool (2×2) — output 4 × 4 × 128
- Dropout 0.25
- Flatten — output 2048
- Dense 1,024 (ReLU) — output 1024
- Dropout 0.5
- Dense 7 (softmax) — output 7
Images, video and clustering
Expression labelling and face grouping answer different questions, so they are separate pipelines. The video path reads every frame, finds faces, crops each one to 48 × 48 and asks the CNN for a label. The clustering path turns faces in still images into 128-number encodings and groups them with K-means; it says nothing about who anyone is, or how they feel.
Expression classification and face clustering are separate pipelines
- Implemented
- Stored data
- Incomplete, or a gap found in testing
The only place clustering and expression labels meet is a printed sentence for a single image. I kept them separate here because they answer different questions, and neither identifies a person.
Text version of this diagram
Parts
- Train the expression classifier
- FER-2013 CSV — 35,887 greyscale faces, 7 labels
- Pixel strings → 48×48 images — saved into one folder per label
- Image generators — rescale ÷255, batches of 32
- Train the CNN — Adam, sparse categorical cross-entropy
- Saved weights
- Label faces in a video
- Video file — processed frame by frame
- Resize + greyscale — 1280×720 frames
- Haar-cascade face detection — OpenCV
- Crop face → 48×48 — not rescaled ÷255 as in training; incomplete or failed in testing
- Predict 1 of 7 labels
- Annotated video — box + label, written with VideoWriter
- Per-30-frame summary — a placeholder: scores averaged, never summarised; incomplete or failed in testing
- Separate extension: group similar faces
- 129 still images
- Locate faces + 128-d encodings — face_recognition
- K-means, k = 5 — no fixed seed
- One montage per cluster — not evaluated; not identity recognition
Connections
- FER-2013 CSV → Pixel strings → 48×48 images
- Pixel strings → 48×48 images → Image generators
- Image generators → Train the CNN
- Train the CNN → Saved weights
- Video file → Resize + greyscale
- Resize + greyscale → Haar-cascade face detection
- Haar-cascade face detection → Crop face → 48×48: each face
- Crop face → 48×48 → Predict 1 of 7 labels
- Saved weights → Predict 1 of 7 labels: loaded
- Predict 1 of 7 labels → Annotated video
- Predict 1 of 7 labels → Per-30-frame summary (incomplete or failed in testing)
- 129 still images → Locate faces + 128-d encodings
- Locate faces + 128-d encodings → K-means, k = 5
- K-means, k = 5 → One montage per cluster
Checking the result
The 2024 report gave 86.5% test accuracy. When the evaluation was re-checked in 2026, the problem was the split: the notebook drew its “test” images at random from the whole dataset, ignoring FER-2013's own partitions, and used the same images to monitor training.
Same model family, different data boundaries, different results
- used to fit the model
- used to choose when to stop
- used for the reported score
- scored “test” images that belong to the Training partition
The 2024 notebook drew its test set at random from all the data, and the same images also monitored training. Re-scoring the saved weights shows how much that matters: they score 92.8% on the official Training images but about 61% on each held-out partition.
The strong gap suggests the saved weights had seen most of the random test set, although the notebook alone cannot prove every image was used in training. The recorded retrain keeps fitting, model selection and the final score on separate partitions, and lands at a similar 60.7%.
Text version of this diagram
- FER-2013 has 28,709 Training, 3,589 PublicTest and 3,589 PrivateTest images.
- 2024 notebook: a random 80/20 split (28,709 / 7,178); 5,742 of the test images come from the official Training partition; reported accuracy 86.5%.
- Saved weights re-scored: Training 92.8%, PublicTest 60.5%, PrivateTest 61.9%.
- Recorded clean retrain: fit on Training, early stopping on PublicTest, PrivateTest scored once: 60.7% (macro F1 0.57).
- clean retrain, held-out PrivateTest
- 60.7%
- macro F1 across 7 labels
- 0.57
- always guessing the commonest label
- 24.5%
Recorded September 2026 retrain: same architecture, TensorFlow 2.17.0, seed 42; 15 epochs, best at epoch 10. Scripts and outputs are kept with the portfolio source.
The same network, scored under different protocols
Where the clean retrain gets confused
| True Predicted | Angry | Disgust | Fear | Happy | Neutral | Sad | Surprise |
|---|---|---|---|---|---|---|---|
| Angry491 | 55 | 1 | 12 | 8 | 15 | 6 | 3 |
| Disgust55 | 27 | 40 | 13 | 4 | 7 | 5 | 4 |
| Fear528 | 14 | 0 | 42 | 5 | 16 | 9 | 14 |
| Happy879 | 3 | 0 | 3 | 85 | 5 | 1 | 3 |
| Neutral626 | 7 | 0 | 9 | 8 | 66 | 6 | 3 |
| Sad594 | 15 | 1 | 17 | 11 | 25 | 29 | 3 |
| Surprise416 | 1 | 0 | 8 | 6 | 4 | 0 | 80 |
Values are the percentage of each true class. The small numbers are image counts.
More about the evidence
- 288 PrivateTest images have exact pixel duplicates in Training. Without them the retrained model scores 58.4%.
- The saved 2024 weights score almost the same on the notebook's random training and test rows, which fits weights trained on something close to the official Training partition. The full training history cannot be reconstructed from the notebook alone, so overlap is strongly suggested rather than proven.
- The video and image inference cells pass face crops to the model without the ÷255 rescaling used in training. Its effect was not measured.
Evidence
· 2026 rework
What came next: EmoSense
In 2026 I turned the notebooks into EmoSense, a structured application. Training, face detection, webcam inference and a small FastAPI service are separate modules. Every experiment is saved under its own name with its metrics, plots and checkpoint, and a manifest file decides which model the app loads.
EmoSense: six saved FER-2013 runs
| Run | Test accuracy and macro f1 |
|---|---|
| 62.4% | |
| 60.5% | |
| 54.6% | |
| 50.9% | |
| 50.2% | |
| 35.6% |
The two small greyscale CNNs beat both pretrained backbones. EfficientNet-B0 scored lowest, yet it is still the model the app's manifest selects: switching to the attention CNN is a single command that has not been run.
The eight-class AffectNet extension
A separate extension inside the same repository moved to EfficientNet backbones on a local copy of AffectNet, with eight labels (adding contempt), focal loss, class weighting and a two-model ensemble.
| Run | Accuracy | Top-2 |
|---|---|---|
| EfficientNet-B0 baseline | 59.6% | 78.8% |
| EfficientNet-B1, focal loss and class weighting | 64.3% | 82.0% |
| EfficientNet-B2, three-phase fine-tuning | 54.8% | 74.4% |
| Controlled B1 rerun (Experiment 6) | 58.4% | 78.4% |
| B0 + B1 ensemble, weights 25/75 * | 65.0% | 82.2% |
Transcribed from the saved reports in the EmoSense repository. * The ensemble weights were chosen on the same Test folder that reports its score, so 65.0% is a development result, not a held-out estimate. The controlled B1 rerun did not reproduce 64.3%; its settings differ and the gap is not yet explained. None of these numbers is comparable with the seven-label FER-2013 runs.
Where the work stands
- WorkJanuary 2024 group prototype (Project-2)
- StatusPublic: the reduced-data expression notebook, the group's trip-planner chatbot and planning documents.
- WorkMay 2024 individual extension
- StatusNot on GitHub. Its notebook produced the 86.5% random-split figure reviewed above.
- WorkSaved-weight check and clean retrain (2026)
- StatusRecorded with this portfolio's source; separate experiments, not a new model.
- WorkEmoSense (2026)
- StatusPublic. Results are recorded, not re-run. The app's manifest still selects the EfficientNet-B0 run rather than the strongest one.
What I took from it: a number means little until you know which data produced it. I also would not describe any of this as reading emotions. Benchmark labels describe posed or apparent expressions, not what a person feels.