Samik Hafeez
All work

Year 1 · group prototype, then individual extension · reworked in 2026

Facial-expression recognition: building a CNN, then checking it

My first substantial computer-vision project: a CNN that labels facial expressions in images and video, and a separate face-clustering experiment. The 2024 report gave 86.5%; a later review showed that figure came from a leaky split, with held-out results near 61%. In 2026 I rebuilt it as EmoSense, a modular app that compares four architectures.

Context
University projects, Year 1 · reworked as EmoSense in 2026
My role
Main developer (group prototype) · sole developer (individual extension)
Team
Group of 3 (prototype) · individual (extension)
Dates
Jan – May 2024 · EmoSense 2026
Status
January 2024 prototype and 2026 EmoSense on GitHub · May 2024 extension not on GitHub

Stack

  • Python
  • TensorFlow / Keras
  • OpenCV
  • face_recognition
  • scikit-learn
  • pandas

Scope of this pageTwo public repositories: Project-2 holds the January 2024 group prototype, and EmoSense is my 2026 rework. The May 2024 individual extension, which produced the report's headline figure, is not on GitHub. The saved-weight check and the clean retrain below are separate experiments, kept with this portfolio's source.

The short version

Each point is expanded, with diagrams, in the sections below.

01Problem and constraints
Label seven facial expressions from small greyscale faces, then apply the model to video. The first group prototype struggled with a reduced dataset and notebook crashes.
02My responsibility
I wrote the main code for the group prototype, then built the full-dataset model, video pipeline and face clustering alone.
03The design
A four-layer CNN trained on FER-2013; OpenCV face detection feeds face crops to it; face clustering is a separate pipeline.
04What I implemented
Data conversion, the CNN, frame-by-frame video annotation and K-means grouping of face encodings.
05The team’s part
The January prototype was shared coursework with two teammates; the May extension was individual.
06What was delivered
In 2024, an annotated-video prototype, cluster montages and the report. In 2026, EmoSense: named experiments, a webcam app and a small API.
07What testing established
Reported 86.5% on a random split. Scored on held-out partitions, the saved weights reach 61.9%; a clean retrain reaches 60.7%.
08Learned, and unfinished
A result is only as good as its data boundary, and a bigger network is not automatically better. Expression labels are not emotions.
On this page

From a failing prototype to a working model

The first version, in a January 2024 group project, trained a small network on a few thousand images because the full dataset kept crashing our notebook environment. Its accuracy was too low to be useful, and I was the one writing most of that code.

In an individual module that spring I rebuilt it on all 35,887 FER-2013 images with a deeper network, extended the video processing with OpenCV face detection, and added a separate experiment that groups similar faces. It was my first end-to-end computer-vision system, from pixel strings to an annotated video.

The network

Four convolution layers with pooling and dropout reduce each 48 × 48 face to a small stack of feature maps, which a dense layer turns into scores for seven labels: angry, disgust, fear, happy, neutral, sad and surprise. About 2.3 million of its 2.35 million parameters sit in the dense layer.

Explanatory diagramDerived from the model definition in the May 2024 notebook (the 2026 retrain uses the same architecture)

The CNN, layer by layer

The CNN, layer by layerInputgreyscale48×48×1Conv 323×3, ReLU46×46×32Conv 643×3, ReLU44×44×64DOMax pool2×222×22×64Conv 1283×3, ReLU20×20×128Max pool2×210×10×128Conv 1283×3, ReLU8×8×128DOMax pool2×24×4×128Flatten2,048DODense 1,024ReLU1,024Dense 7softmax7DO = dropout: 0.25 after the first and last pooling layers, 0.5 after the dense layer. Height shows spatial size; depth shows channels.
  1. Input · greyscale48×48×1
  2. Conv 32 · 3×3, ReLU46×46×32
  3. Conv 64 · 3×3, ReLU44×44×64
  4. Max pool · 2×222×22×64
  5. Dropout 0.25
  6. Conv 128 · 3×3, ReLU20×20×128
  7. Max pool · 2×210×10×128
  8. Conv 128 · 3×3, ReLU8×8×128
  9. Max pool · 2×24×4×128
  10. Dropout 0.25
  11. Flatten2,048
  12. Dense 1,024 · ReLU1,024
  13. Dropout 0.5
  14. Dense 7 · softmax7
Output shape after each layer for one 48 × 48 greyscale face. The final layer gives a probability for each of seven labels.

The CNN, layer by layer

100%
The CNN, layer by layerInputgreyscale48×48×1Conv 323×3, ReLU46×46×32Conv 643×3, ReLU44×44×64DOMax pool2×222×22×64Conv 1283×3, ReLU20×20×128Max pool2×210×10×128Conv 1283×3, ReLU8×8×128DOMax pool2×24×4×128Flatten2,048DODense 1,024ReLU1,024Dense 7softmax7DO = dropout: 0.25 after the first and last pooling layers, 0.5 after the dense layer. Height shows spatial size; depth shows channels.
Text version of this diagram
  1. Input (greyscale) — output 48 × 48 × 1
  2. Conv 32 (3×3, ReLU) — output 46 × 46 × 32
  3. Conv 64 (3×3, ReLU) — output 44 × 44 × 64
  4. Max pool (2×2) — output 22 × 22 × 64
  5. Dropout 0.25
  6. Conv 128 (3×3, ReLU) — output 20 × 20 × 128
  7. Max pool (2×2) — output 10 × 10 × 128
  8. Conv 128 (3×3, ReLU) — output 8 × 8 × 128
  9. Max pool (2×2) — output 4 × 4 × 128
  10. Dropout 0.25
  11. Flatten — output 2048
  12. Dense 1,024 (ReLU) — output 1024
  13. Dropout 0.5
  14. Dense 7 (softmax) — output 7

Images, video and clustering

Expression labelling and face grouping answer different questions, so they are separate pipelines. The video path reads every frame, finds faces, crops each one to 48 × 48 and asks the CNN for a label. The clustering path turns faces in still images into 128-number encodings and groups them with K-means; it says nothing about who anyone is, or how they feel.

Explanatory diagramDerived from Main_project_file.ipynb and face_clustering.ipynb (May 2024)

Expression classification and face clustering are separate pipelines

Expression classification and face clustering are separate pipelinesTraining, video labelling and face clustering are three separate pipelines. The dashed red steps are weaknesses found when the notebooks were re-read: a preprocessing mismatch and an unfinished summary.TRAIN THE EXPRESSION CLASSIFIERLABEL FACES IN A VIDEOSEPARATE EXTENSION: GROUP SIMILAR FACESFER-2013 CSV35,887 greyscale faces, 7labelsPixel strings → 48×48imagessaved into one folder per labelImage generatorsrescale ÷255, batches of 32Train the CNNAdam, sparse categoricalcross-entropySaved weightsVideo fileprocessed frame by frameResize + greyscale1280×720 framesHaar-cascade facedetectionOpenCVCrop face → 48×48not rescaled ÷255 as intrainingPredict 1 of 7 labelsAnnotated videobox + label, written withVideoWriterPer-30-frame summarya placeholder: scoresaveraged, never summarised129 still imagesLocate faces + 128-dencodingsface_recognitionK-means, k = 5no fixed seedOne montage per clusternot evaluated; not identityrecognitioneach faceloaded
Expression classification and face clustering are separate pipelinesTraining, video labelling and face clustering are three separate pipelines. The dashed red steps are weaknesses found when the notebooks were re-read: a preprocessing mismatch and an unfinished summary.TRAIN THE EXPRESSION CLASSIFIERLABEL FACES IN A VIDEOSEPARATE EXTENSION: GROUP SIMILAR FACESFER-2013 CSV35,887 greyscale faces, 7labelsPixel strings → 48×48imagessaved into one folder perlabelImage generatorsrescale ÷255, batches of32Train the CNNAdam, sparse categoricalcross-entropySaved weightsVideo fileprocessed frame by frameResize + greyscale1280×720 framesHaar-cascade facedetectionOpenCVCrop face → 48×48not rescaled ÷255 as intrainingPredict 1 of 7 labelsAnnotated videobox + label, written withVideoWriterPer-30-framesummarya placeholder: scoresaveraged, neversummarised129 still imagesLocate faces + 128-dencodingsface_recognitionK-means, k = 5no fixed seedOne montage perclusternot evaluated; not identityrecognition
  • Implemented
  • Stored data
  • Incomplete, or a gap found in testing
Training, video labelling and face clustering are three separate pipelines. The dashed red steps are weaknesses found when the notebooks were re-read: a preprocessing mismatch and an unfinished summary.

The only place clustering and expression labels meet is a printed sentence for a single image. I kept them separate here because they answer different questions, and neither identifies a person.

Expression classification and face clustering are separate pipelines

100%
Expression classification and face clustering are separate pipelinesTRAIN THE EXPRESSION CLASSIFIERLABEL FACES IN A VIDEOSEPARATE EXTENSION: GROUP SIMILAR FACESFER-2013 CSV35,887 greyscale faces, 7labelsPixel strings → 48×48imagessaved into one folder per labelImage generatorsrescale ÷255, batches of 32Train the CNNAdam, sparse categoricalcross-entropySaved weightsVideo fileprocessed frame by frameResize + greyscale1280×720 framesHaar-cascade facedetectionOpenCVCrop face → 48×48not rescaled ÷255 as intrainingPredict 1 of 7 labelsAnnotated videobox + label, written withVideoWriterPer-30-frame summarya placeholder: scoresaveraged, never summarised129 still imagesLocate faces + 128-dencodingsface_recognitionK-means, k = 5no fixed seedOne montage per clusternot evaluated; not identityrecognitioneach faceloaded
Text version of this diagram

Parts

  • Train the expression classifier
    • FER-2013 CSV — 35,887 greyscale faces, 7 labels
    • Pixel strings → 48×48 images — saved into one folder per label
    • Image generators — rescale ÷255, batches of 32
    • Train the CNN — Adam, sparse categorical cross-entropy
    • Saved weights
  • Label faces in a video
    • Video file — processed frame by frame
    • Resize + greyscale — 1280×720 frames
    • Haar-cascade face detection — OpenCV
    • Crop face → 48×48 — not rescaled ÷255 as in training; incomplete or failed in testing
    • Predict 1 of 7 labels
    • Annotated video — box + label, written with VideoWriter
    • Per-30-frame summary — a placeholder: scores averaged, never summarised; incomplete or failed in testing
  • Separate extension: group similar faces
    • 129 still images
    • Locate faces + 128-d encodings — face_recognition
    • K-means, k = 5 — no fixed seed
    • One montage per cluster — not evaluated; not identity recognition

Connections

  • FER-2013 CSV → Pixel strings → 48×48 images
  • Pixel strings → 48×48 images → Image generators
  • Image generators → Train the CNN
  • Train the CNN → Saved weights
  • Video file → Resize + greyscale
  • Resize + greyscale → Haar-cascade face detection
  • Haar-cascade face detection → Crop face → 48×48: each face
  • Crop face → 48×48 → Predict 1 of 7 labels
  • Saved weights → Predict 1 of 7 labels: loaded
  • Predict 1 of 7 labels → Annotated video
  • Predict 1 of 7 labels → Per-30-frame summary (incomplete or failed in testing)
  • 129 still images → Locate faces + 128-d encodings
  • Locate faces + 128-d encodings → K-means, k = 5
  • K-means, k = 5 → One montage per cluster

Checking the result

The 2024 report gave 86.5% test accuracy. When the evaluation was re-checked in 2026, the problem was the split: the notebook drew its “test” images at random from the whole dataset, ignoring FER-2013's own partitions, and used the same images to monitor training.

Explanatory diagramDerived from the 2024 notebook, the saved-weight verification (downloadable as JSON) and the recorded retrain outputs

Same model family, different data boundaries, different results

How the evaluation boundary changed the facial-expression resultFER-2013 official partitionshow the dataset is meant to be usedTraining 28,709PublicTestPrivateTest2024 notebook: random 80/20 splitUsage column ignored; test set also used as validation data while trainingTrain 28,7095,742 from Training86.5%reportedSaved 2024 weights, re-scored (Sept 2026)each official partition scored separately92.8% on Training60.5%61.9%61.9%PrivateTestClean retrain, same architecture (recorded Sept 2026)fit on Training, early stopping on PublicTest, PrivateTest scored oncefitselectscore60.7%PrivateTest
How the evaluation boundary changed the facial-expression resultFER-2013 official partitionshow the dataset is meant to be usedTraining 28,709PublicPrivate2024 notebook: random 80/20 splitUsage column ignored; test set also used as validationdata while trainingTrain 28,70986.5%reportedSaved 2024 weights, re-scored (Sept 2026)each official partition scored separately92.8% on Training61.9%PrivateTestClean retrain, same architecture (recordedSept 2026)fit on Training, early stopping on PublicTest, PrivateTestscored oncefit60.7%PrivateTest
  • used to fit the model
  • used to choose when to stop
  • used for the reported score
  • scored “test” images that belong to the Training partition
Bars are drawn to scale across all 35,887 FER-2013 images. Hatching marks the 5,742 random “test” images (80.0%) that belong to FER-2013's Training partition.

The 2024 notebook drew its test set at random from all the data, and the same images also monitored training. Re-scoring the saved weights shows how much that matters: they score 92.8% on the official Training images but about 61% on each held-out partition.

The strong gap suggests the saved weights had seen most of the random test set, although the notebook alone cannot prove every image was used in training. The recorded retrain keeps fitting, model selection and the final score on separate partitions, and lands at a similar 60.7%.

Same model family, different data boundaries, different results

100%
How the evaluation boundary changed the facial-expression resultFER-2013 official partitionshow the dataset is meant to be usedTraining 28,709PublicTestPrivateTest2024 notebook: random 80/20 splitUsage column ignored; test set also used as validation data while trainingTrain 28,7095,742 from Training86.5%reportedSaved 2024 weights, re-scored (Sept 2026)each official partition scored separately92.8% on Training60.5%61.9%61.9%PrivateTestClean retrain, same architecture (recorded Sept 2026)fit on Training, early stopping on PublicTest, PrivateTest scored oncefitselectscore60.7%PrivateTest
Text version of this diagram
  • FER-2013 has 28,709 Training, 3,589 PublicTest and 3,589 PrivateTest images.
  • 2024 notebook: a random 80/20 split (28,709 / 7,178); 5,742 of the test images come from the official Training partition; reported accuracy 86.5%.
  • Saved weights re-scored: Training 92.8%, PublicTest 60.5%, PrivateTest 61.9%.
  • Recorded clean retrain: fit on Training, early stopping on PublicTest, PrivateTest scored once: 60.7% (macro F1 0.57).
clean retrain, held-out PrivateTest
60.7%
macro F1 across 7 labels
0.57
always guessing the commonest label
24.5%

Recorded September 2026 retrain: same architecture, TensorFlow 2.17.0, seed 42; 15 epochs, best at epoch 10. Scripts and outputs are kept with the portfolio source.

Chart · recorded resultsFrom the saved-weight verification and the recorded retrain (September 2026)

The same network, scored under different protocols

Accuracy by evaluation setup
Evaluation setupAccuracy
Reported in 2024 (random 20% split)86.5%
2024 model, random-split rows from Training92.6%
2024 model, random-split rows from test partitions62.2%
2024 model, official PrivateTest partition61.9%
Retrained with official splits, PrivateTest60.7%
Majority-class baseline24.5%
The top two bars come from the leaky random split. On data the saved weights had not seen, they score 61.9%–62.2%, close to the clean retrain.
Chart · recorded resultsRecorded retrain, held-out PrivateTest partition (3,589 images)

Where the clean retrain gets confused

Confusion matrix: rows are the true class, columns the predicted class, values are the share of each true class.
True PredictedAngryDisgustFearHappyNeutralSadSurprise
Angry4915511281563
Disgust552740134754
Fear52814042516914
Happy87930385513
Neutral62670986663
Sad594151171125293
Surprise41610864080

Values are the percentage of each true class. The small numbers are image counts.

Each row shows how the images of one true label were predicted; the diagonal is recall. Happy (85%) and surprise (80%) are recognised most reliably; sad is more often predicted as neutral or fear than as sad, and disgust is often read as angry.
More about the evidence
  • 288 PrivateTest images have exact pixel duplicates in Training. Without them the retrained model scores 58.4%.
  • The saved 2024 weights score almost the same on the notebook's random training and test rows, which fits weights trained on something close to the official Training partition. The full training history cannot be reconstructed from the notebook alone, so overlap is strongly suggested rather than proven.
  • The video and image inference cells pass face crops to the model without the ÷255 rescaling used in training. Its effect was not measured.

· 2026 rework

What came next: EmoSense

In 2026 I turned the notebooks into EmoSense, a structured application. Training, face detection, webcam inference and a small FastAPI service are separate modules. Every experiment is saved under its own name with its metrics, plots and checkpoint, and a manifest file decides which model the app loads.

Chart · recorded resultsTranscribed from the saved metrics.json reports in the EmoSense repository; not re-run

EmoSense: six saved FER-2013 runs

EmoSense: six saved FER-2013 runs
RunTest accuracy and macro f1
Attention CNNBestExperiment 02 · light augmentation62.4%F1 0.59
Custom CNNExperiment 01 · light augmentation60.5%F1 0.57
MobileNetV2Experiment 03 · light augmentation54.6%F1 0.52
MobileNetV2Baseline · strong augmentation50.9%F1 0.47
MobileNetV2Default run · strong augmentation50.2%F1 0.46
EfficientNet-B0Active in appExperiment 04 · light augmentation35.6%F1 0.27
  • Test accuracy
  • Macro F1
Seven expression labels, scored on EmoSense's own stratified random test split (15% of FER-2013). Bars show test accuracy; the dark tick marks macro F1 on the same 0–100 scale.

The two small greyscale CNNs beat both pretrained backbones. EfficientNet-B0 scored lowest, yet it is still the model the app's manifest selects: switching to the attention CNN is a single command that has not been run.

The eight-class AffectNet extension

A separate extension inside the same repository moved to EfficientNet backbones on a local copy of AffectNet, with eight labels (adding contempt), focal loss, class weighting and a two-model ensemble.

Recorded AffectNet results (eight labels, local dataset copy)
RunAccuracyTop-2
EfficientNet-B0 baseline59.6%78.8%
EfficientNet-B1, focal loss and class weighting64.3%82.0%
EfficientNet-B2, three-phase fine-tuning54.8%74.4%
Controlled B1 rerun (Experiment 6)58.4%78.4%
B0 + B1 ensemble, weights 25/75 *65.0%82.2%

Transcribed from the saved reports in the EmoSense repository. * The ensemble weights were chosen on the same Test folder that reports its score, so 65.0% is a development result, not a held-out estimate. The controlled B1 rerun did not reproduce 64.3%; its settings differ and the gap is not yet explained. None of these numbers is comparable with the seven-label FER-2013 runs.

Where the work stands

WorkJanuary 2024 group prototype (Project-2)
StatusPublic: the reduced-data expression notebook, the group's trip-planner chatbot and planning documents.
WorkMay 2024 individual extension
StatusNot on GitHub. Its notebook produced the 86.5% random-split figure reviewed above.
WorkSaved-weight check and clean retrain (2026)
StatusRecorded with this portfolio's source; separate experiments, not a new model.
WorkEmoSense (2026)
StatusPublic. Results are recorded, not re-run. The app's manifest still selects the EfficientNet-B0 run rather than the strongest one.

What I took from it: a number means little until you know which data produced it. I also would not describe any of this as reading emotions. Benchmark labels describe posed or apparent expressions, not what a person feels.

Contact

I’m looking for graduate and early-career roles in AI/ML and software engineering.