Samik Hafeez
All work

Year 3 individual project · Advanced AI Methods and Tools

Does prompt wording change what CLIP finds?

An experiment on a frozen vision-language model: I worded the same ten queries four ways, including deliberately misleading ones, and measured what changed. Overall accuracy barely moved (86.2–87.5%), individual classes did, and my retrieval metric turned out too easy to show anything.

Context
Individual project, Year 3 · Advanced AI Methods and Tools, University of Bradford
My role
Sole author: experiment design, code, evaluation, demo and report
Team
Individual
Dates
Semester 1, 2025–26 · submitted December 2025
Status
Completed academic study · notebook on GitHub

Stack

  • Python
  • PyTorch
  • Hugging Face Transformers
  • CLIP ViT-B/32
  • scikit-learn
  • Gradio
  • Google Colab (CPU)

The short version

Each point is expanded, with diagrams, in the sections below.

01Problem and constraints
Does the wording of a text prompt change which images CLIP matches it with? And how does zero-shot CLIP compare with a small classifier trained on its features?
02My responsibility
Individual work: I chose the prompt conditions, wrote the evaluation, built a demo and wrote the report.
03The design
One frozen CLIP model on CPU, a fixed 1,000-image CIFAR-10 gallery (seed 0), four prompt conditions, and a linear probe on the same image embeddings.
04What I implemented
Embedding and ranking code, four prompt sets, retrieval and classification metrics, confusion matrices, a linear probe and a four-tool Gradio demo.
05The team’s part
None: individual coursework.
06What was delivered
A notebook with its stored outputs (now public on GitHub), a report, a recorded walkthrough and a demo run from Colab.
07What testing established
Classifying the 1,000 gallery images, the four wordings scored between 86.2% and 87.5%. On the same 200 held-out images, a linear probe reached 91.5% against zero-shot's 87.0%.
08Learned, and unfinished
Pick a metric that can move: my retrieval check scored a perfect 1.0 in every condition. And look below the total, where the trade-offs between classes were.
On this page

The question

CLIP matches images and text by embedding both in the same space, so a text query can find images it was never trained to label. That makes the wording of the query part of the model. I wanted to measure how much it mattered.

I kept everything else fixed: the model, the images and the random seed. Only the prompts changed. Alongside the single canonical prompt, I wrote three sets of five templates: natural rewordings, deliberately misleading descriptions and descriptions of poor image quality. A set's five prompts are averaged into one query per class.

The four prompt conditions ({label} is the class name)
ConditionTemplatesWhat it tests
Canonicala photo of a {label}The standard single prompt: the baseline
Ensemble A: naturala photo of / a picture of / an image of / a close-up photo of / a bright photo of a {label}Everyday rewordings
Ensemble B: misleadinga random object unrelated to a {label} · an incorrect depiction of a {label} · a distorted sculpture not resembling a {label} · a photo of something different instead of a {label} · an abstract wrong representation of a {label}Wording that contradicts the image
Ensemble C: degradeda blurry photo of a {label} · a low resolution image of a {label} · a {label} partially hidden behind objects · an overexposed photo of a {label} · a small {label} in the distanceDescriptions of poor image quality

The set-up

The images are encoded once and reused for every condition, so any difference comes from the text side. Each image is compared with each class query by cosine similarity. The same scores support two measurements: retrieving images for a class name, and classifying each image by its most similar class.

Explanatory diagramDerived from the submitted notebook (sampling, encoding, evaluation and probe cells) and the final report

One frozen model, one fixed gallery, four ways of wording the query

One frozen model, one fixed gallery, four ways of wording the queryHow every condition was measured. Images are encoded once; only the text side changes between conditions. The dashed red box is the retrieval check, which turned out unable to tell the conditions apart. Individual work throughout.IMAGES: ENCODED ONCETEXT: ONCE PER CONDITIONEVALUATIONCIFAR-10 test split10,000 images, 32 × 32 pixelsFixed gallery100 images per class, seed 0: 1,000imagesCLIP image encoderViT-B/32, frozen, on CPU; upscaled to224 × 224Image embeddings1,000 × 512, unit lengthTen class namesairplane, automobile … truckFour prompt conditionscanonical: 1 template; A, B and C:5 templates eachCLIP text encoderan ensemble's templates are averagedper classClass embeddings10 × 512 for each conditionCosine similaritydot product of unit vectors, then rankZero-shot classificationeach image takes its most similar class:accuracy, F1, confusionRetrieval check (Recall@1, @5)10 class-name queries: is a correctimage in the top K?Linear probelogistic regression on the sameembeddings: 800 train, 200 test
One frozen model, one fixed gallery, four ways of wording the queryHow every condition was measured. Images are encoded once; only the text side changes between conditions. The dashed red box is the retrieval check, which turned out unable to tell the conditions apart. Individual work throughout.IMAGESTEXTEVALUATIONCIFAR-10 test split10,000 images, 32 × 32pixelsFixed gallery100 images per class,seed 0: 1,000 imagesCLIP image encoderViT-B/32, frozen, on CPU;upscaled to 224 × 224Image embeddings1,000 × 512, unit lengthTen class namesairplane, automobile …truckFour promptconditionscanonical: 1template; A, B andC: 5 templates eachCLIP text encoderan ensemble's templatesare averaged per classClass embeddings10 × 512 for eachconditionCosine similaritydot product of unitvectors, then rankZero-shotclassificationeach image takes its mostsimilar class: accuracy,F1, confusionRetrieval check(Recall@1, @5)10 class-name queries: isa correct image in the topK?Linear probelogistic regression on thesame embeddings: 800train, 200 test
  • Implemented
  • External service or data
  • Decision
  • Stored data
  • Incomplete, or a gap found in testing
How every condition was measured. Images are encoded once; only the text side changes between conditions. The dashed red box is the retrieval check, which turned out unable to tell the conditions apart. Individual work throughout.

One frozen model, one fixed gallery, four ways of wording the query

100%
One frozen model, one fixed gallery, four ways of wording the queryIMAGES: ENCODED ONCETEXT: ONCE PER CONDITIONEVALUATIONCIFAR-10 test split10,000 images, 32 × 32 pixelsFixed gallery100 images per class, seed 0: 1,000imagesCLIP image encoderViT-B/32, frozen, on CPU; upscaled to224 × 224Image embeddings1,000 × 512, unit lengthTen class namesairplane, automobile … truckFour prompt conditionscanonical: 1 template; A, B and C:5 templates eachCLIP text encoderan ensemble's templates are averagedper classClass embeddings10 × 512 for each conditionCosine similaritydot product of unit vectors, then rankZero-shot classificationeach image takes its most similar class:accuracy, F1, confusionRetrieval check (Recall@1, @5)10 class-name queries: is a correctimage in the top K?Linear probelogistic regression on the sameembeddings: 800 train, 200 test
Text version of this diagram

Parts

  • Images: encoded once
    • CIFAR-10 test split — 10,000 images, 32 × 32 pixels
    • Fixed gallery — 100 images per class, seed 0: 1,000 images
    • CLIP image encoder — ViT-B/32, frozen, on CPU; upscaled to 224 × 224
    • Image embeddings — 1,000 × 512, unit length
  • Text: once per condition
    • Ten class names — airplane, automobile … truck
    • Four prompt conditions — canonical: 1 template; A, B and C: 5 templates each
    • CLIP text encoder — an ensemble's templates are averaged per class
    • Class embeddings — 10 × 512 for each condition
  • Evaluation
    • Cosine similarity — dot product of unit vectors, then rank
    • Zero-shot classification — each image takes its most similar class: accuracy, F1, confusion
    • Retrieval check (Recall@1, @5) — 10 class-name queries: is a correct image in the top K?; incomplete or failed in testing
    • Linear probe — logistic regression on the same embeddings: 800 train, 200 test

Connections

  • CIFAR-10 test split → Fixed gallery
  • Fixed gallery → CLIP image encoder
  • CLIP image encoder → Image embeddings
  • Ten class names → Four prompt conditions
  • Four prompt conditions → CLIP text encoder
  • CLIP text encoder → Class embeddings
  • Image embeddings → Cosine similarity
  • Class embeddings → Cosine similarity
  • Cosine similarity → Zero-shot classification
  • Cosine similarity → Retrieval check (Recall@1, @5)
  • Image embeddings → Linear probe

What the numbers showed

My main retrieval metric, Recall@1 and Recall@5, was 1.0 for every condition. It asked whether each of the ten class-name queries put a correct image first (Recall@1) or anywhere in the top five (Recall@5), and every query did, however it was worded. A metric that cannot fall cannot compare anything, so the useful evidence came from classifying all 1,000 gallery images instead.

Chart · recorded resultsStored output of the notebook's full evaluation cell; not re-run

Four wordings, almost the same accuracy

Four wordings, almost the same accuracy
Prompt conditionAccuracy
Canonical“a photo of a {label}”87.2% · 872
Ensemble Anatural rewordings87.1% · 871
Ensemble BHighestmisleading descriptions87.5% · 875
Ensemble Cdegraded-image descriptions86.2% · 862
Zero-shot accuracy when each of the 1,000 gallery images is given its most similar class, under each prompt condition. The whole spread is 13 images.

The misleading prompts came out highest, by three images out of 1,000, and the degraded ones lowest, ten images below the baseline. Differences that small are too close to rank. What did change was the pattern underneath.

Chart · recorded resultsRead from the four stored confusion matrices in the submitted notebook (the diagonal of each)

Totals barely move; individual classes do

Correct predictions per class, out of 100, for the canonical prompt and ensembles A, B and C
ClassCanonicalEnsemble A: naturalEnsemble B: misleadingEnsemble C: degraded
frog6866 −2 fewer than canonical73 +5 more than canonical69 +1 more than canonical
deer7981 +2 more than canonical84 +5 more than canonical79 ±0
cat8383 ±088 +5 more than canonical82 −1 fewer than canonical
bird8886 −2 fewer than canonical82 −6 fewer than canonical85 −3 fewer than canonical
airplane9091 +1 more than canonical86 −4 fewer than canonical84 −6 fewer than canonical
dog9089 −1 fewer than canonical88 −2 fewer than canonical88 −2 fewer than canonical
truck9090 ±093 +3 more than canonical88 −2 fewer than canonical
ship9394 +1 more than canonical95 +2 more than canonical94 +1 more than canonical
horse9594 −1 fewer than canonical94 −1 fewer than canonical96 +1 more than canonical
automobile9697 +1 more than canonical92 −4 fewer than canonical97 +1 more than canonical
All 1,000872871 −1 fewer than canonical875 +3 more than canonical862 −10 fewer than canonical
  • more correct than canonical
  • fewer correct than canonical
Correct predictions out of 100 gallery images per class, sorted by the canonical result. Each ensemble cell shows its change from the canonical prompt; darker shading means a bigger change.

The misleading prompts (B) gained on frog, deer and cat but lost on airplane, bird and automobile, so their total ended up close to the others. The images are the same in every column, so the wording caused these changes; but they are counts from one sample of 100 images per class, and a change of a few images may not hold on a different sample.

What I expectedAveraging several prompts beats one
What the stored results showNot supported: the three ensembles landed 3 images above to 10 below the single prompt.
What I expectedNatural wording (A) does best
What the stored results showNot supported: A finished level with the baseline, one image behind.
What I expectedMisleading wording (B) does worst
What the stored results showNot supported: B finished highest, though only by three images. Every B template still contains the class name, which may be why; that was not tested.
What I expectedSome classes are harder than others
What the stored results showSupported: frog was the weakest class under every wording, and cat–dog and deer–horse confusions appeared under all four.
What I expectedZero-shot comes close to a trained probe
What the stored results showPartly: 87.0% against 91.5% on the same 200 images, without any labelled training data.

Looking at what came back

The notebook stores the top five images for every class under each ensemble: 150 images in all. 147 are the right class. The three misses are cats returned for “dog” under the misleading prompts, and a dog returned for “cat” under the degraded ones.

Five small 32-pixel CIFAR-10 images ranked 1 to 5 for the query ‘dog’ under Ensemble B. Ranks 1, 3 and 4 are dogs; ranks 2 and 5 are cats, one a dark kitten on a pale background and one a brown cat lying down.
Notebook figureStored notebook output: the top five gallery images for “dog” under the misleading prompts. Ranks 2 and 5 are cats. Every image is 32 × 32 pixels, upscaled for CLIP.Open full size

The same cat–dog and deer–horse confusions appear under every wording, which points at the image side rather than the prompts: at 32 × 32 pixels, small animals are hard to tell apart.

Zero-shot against a trained probe

Zero-shot classification needs no labelled examples, only the class names. To see what it gives up, I trained a logistic-regression probe on the same frozen image embeddings: 80 images per class for training and 20 held out, and I scored both methods on those same 200 images.

Chart · recorded resultsStored output of the notebook's side-by-side comparison cell; not re-run

Zero-shot CLIP and a linear probe on the same 200 images

Zero-shot CLIP and a linear probe on the same 200 images
MethodAccuracy
Linear probetrained on 800 labelled images91.5% · 183 of 200
Zero-shot CLIPno training: class names only87.0% · 174 of 200
Both score the same 20 held-out images per class. The probe is logistic regression trained on the other 800 gallery images' CLIP embeddings; zero-shot uses only the canonical class prompts.

The probe got nine more of the 200 right, gaining most on frog and deer, while zero-shot did slightly better on dog, horse and ship. Zero-shot's mistakes were also more scattered: its deer errors went to five different classes, the probe's to two.

The stored confusion matrices
Four confusion matrices, one per prompt condition, each ten by ten with true labels in rows and predicted labels in columns. The diagonals are dark; the largest off-diagonal cells are frog predicted as cat or bird, deer predicted as horse, and cat predicted as dog.
Notebook figureStored notebook output: all 1,000 gallery images classified under each prompt condition.Open full size
Two ten-by-ten confusion matrices side by side for the same 200 held-out images: zero-shot CLIP in blue on the left and the linear probe in green on the right. The probe's diagonal is darker and has fewer off-diagonal cells.
Notebook figureStored notebook output: zero-shot CLIP and the linear probe on the same 200 held-out images.Open full size

A demo to try it by hand

I built a small Gradio app in the notebook so that anyone could test the model directly, using the same encoding functions as the experiment. It ran from Colab with a temporary public link and was never deployed.

  • Image to text: score an uploaded image against your own comma-separated prompts.
  • Text to image: return the closest image in the CIFAR-10 gallery.
  • Zero-shot top five: the five most likely CIFAR-10 classes for an uploaded image.
  • Ensemble comparison: the class each prompt set (A, B and C) picks for the same image.

Responsible use

CIFAR-10 contains no people, so the study itself raises few direct risks. The point of the report's governance section was what the behaviour means for real systems built on the same model.

  • If wording shifts which classes succeed, prompts are part of the system and should be fixed, tested and versioned, not left to each user.
  • Some classes are weaker whatever the wording, so a system needs to know where it is unreliable before it is trusted there.
  • CLIP was trained on internet-scale data with undocumented biases; anything touching people needs its own evaluation and human review of uncertain results.
  • I did not build any safeguards, such as uncertainty estimates or input checks. The report lists them as recommendations.

What I would change

  • Replace hit-at-K with a metric that can move, such as precision at K or mean average precision over all 100 relevant images, and use many more queries than ten.
  • Report differences as counts with uncertainty. A gap of 13 images in 1,000 needs repeated samples or a significance test before it means anything.
  • Re-normalise each averaged prompt embedding. Without it, classes whose five prompts agree more closely get a slight edge when classifying, which could be as large as the differences I was measuring.
  • Add prompts without the class name, to test whether the class word alone carries the result.
  • Repeat on higher-resolution images, where the image side limits results less.
  • Keep comparisons like for like: one notebook cell sets zero-shot accuracy on all 1,000 images beside the probe's score on 200. The fair comparison, on the same 200, is the one used here.
More about the evidence
  • Every number here comes from outputs stored in the submitted notebook; nothing was re-run for this website. The copy on GitHub has the same cells and outputs; only Colab's widget metadata was removed so that GitHub can display it.
  • Gallery: 100 images per class drawn from the CIFAR-10 test split with NumPy seed 0; one further image per class was set aside as a query image but not used by the retrieval metric.
  • Top-5 classification accuracy was 99.3–99.5% under every condition, so the correct class was almost always among the five most similar.
  • The report's AI-use statement records that ChatGPT helped with parts of the code, debugging and the report's structure; the experimental choices and interpretation were mine.

Contact

I’m looking for graduate and early-career roles in AI/ML and software engineering.