CLIP · Attention follow-up 1

CLIP: From Fixed Labels to Language-Defined Vision

Following our Vision Transformer lecture

How Images and Text Learn a Shared Comparison Space

Last time: an image became a useful ViT representation.

Today: let language say what we want to find.

Cat
a photo of a cat
Dog
a photo of a dog

An image goes in. We choose the descriptions to compare it with.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

An image goes in. We choose the descriptions to compare it with.

Teaching note

An image goes in. We choose the descriptions to compare it with.

Section 01

From fixed labels to language

Key idea

Why must ViT’s answer vocabulary be fixed?

01 · From fixed labels to language

Last lecture: one image, 1,000 stored answers

SCHEMATIC · our earlier ViT-Tiny classifier

Newfoundland
ViT · width 192
hCLSh_{\mathrm{CLS}}
Linear(192, 1000)

ImageNet scores

The image representation is useful. The classifier chooses from its training labels.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

The image representation is useful. The classifier chooses from its training labels.

Teaching note

The image representation is useful. The classifier chooses from its training labels.

01 · From fixed labels to language

Can we choose the labels ourselves?

APPLICATION 01 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Newfoundland

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply

a photo of a dog?
a photo of a cat?
a photo of a bird?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Which supplied description will match the dog photograph best?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which supplied description will match the dog photograph best? One matching score for each supplied description, sorted from highest to lowest. The description supplies a comparison vector. Changing the candidates does not train or replace the image encoder. Replace the menu with Newfoundland, pug, and Persian. Run live again; the model weights stay fixed.

Teaching note

Which supplied description will match the dog photograph best? One matching score for each supplied description, sorted from highest to lowest. The description supplies a comparison vector. Changing the candidates does not train or replace the image encoder. Replace the menu with Newfoundland, pug, and Persian. Run live again; the model weights stay fixed.

01 · From fixed labels to language

What happened? · New labels

APPLICATION 01 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Newfoundland

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply · cosine

a photo of a dog+0.2723
a photo of a cat+0.2152
a photo of a bird+0.2184

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Dog has the largest cosine. Every score comes from the same image vector.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which supplied description will match the dog photograph best? One matching score for each supplied description, sorted from highest to lowest. The description supplies a comparison vector. Changing the candidates does not train or replace the image encoder. Replace the menu with Newfoundland, pug, and Persian. Run live again; the model weights stay fixed.

Teaching note

Which supplied description will match the dog photograph best? One matching score for each supplied description, sorted from highest to lowest. The description supplies a comparison vector. Changing the candidates does not train or replace the image encoder. Replace the menu with Newfoundland, pug, and Persian. Run live again; the model weights stay fixed.

01 · From fixed labels to language

Coffee, tea, juice or soup?

APPLICATION 02 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Coffee

Rachel Michetti · CC0

Descriptions we supply

a cup of coffee?
a cup of tea?
a glass of orange juice?
a bowl of soup?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Can we ask what is inside the cup, rather than just name the cup?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can we ask what is inside the cup, rather than just name the cup? A ranking of the four supplied descriptions. Natural language lets us propose attributes, materials and contexts. These are hypotheses to test, not guaranteed abilities. Change only the descriptions to ask about a different property. Each new menu needs a live run.

Teaching note

Can we ask what is inside the cup, rather than just name the cup? A ranking of the four supplied descriptions. Natural language lets us propose attributes, materials and contexts. These are hypotheses to test, not guaranteed abilities. Change only the descriptions to ask about a different property. Each new menu needs a live run.

01 · From fixed labels to language

What happened? · Attributes

APPLICATION 02 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Coffee

Rachel Michetti · CC0

Descriptions we supply · cosine

a cup of coffee+0.2820
a cup of tea+0.2522
a glass of orange juice+0.2398
a bowl of soup+0.2258

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Coffee scores above tea. The output is a ranking of the descriptions we supplied.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can we ask what is inside the cup, rather than just name the cup? A ranking of the four supplied descriptions. Natural language lets us propose attributes, materials and contexts. These are hypotheses to test, not guaranteed abilities. Change only the descriptions to ask about a different property. Each new menu needs a live run.

Teaching note

Can we ask what is inside the cup, rather than just name the cup? A ranking of the four supplied descriptions. Natural language lets us propose attributes, materials and contexts. These are hypotheses to test, not guaranteed abilities. Change only the descriptions to ask about a different property. Each new menu needs a live run.

01 · From fixed labels to language

What happened? · Text → image

APPLICATION 03 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Our text querysomething to drink

Compare it with gallery images.

Four pictured candidates from the 22-image gallery; the saved file contains every score.

Text vector → compare with image vectors → rank the gallery.

The generated mugs score above the coffee photograph for this query.

Radford et al., 2021 · OpenAI implementation · Image credits · Supabase · image search, 6:54 · Try this example · Saved vectors and scores

What changed from the previous slide?

What will a search for ‘something to drink’ return from this gallery? All 22 gallery images ranked against one text query. Encode the gallery once. A new query needs only a text embedding and dot products. Search always has a nearest item even when nothing matches. Try ‘a laptop’, which is absent from this gallery. A nearest result still exists; decide whether it is relevant.

Teaching note

What will a search for ‘something to drink’ return from this gallery? All 22 gallery images ranked against one text query. Encode the gallery once. A new query needs only a text embedding and dot products. Search always has a nearest item even when nothing matches. Try ‘a laptop’, which is absent from this gallery. A nearest result still exists; decide whether it is relevant.

01 · From fixed labels to language

Which supplied sentence describes this scene?

APPLICATION 04 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Rocket launch

SpaceX · CC0

Descriptions we supply

a rocket launching into the night sky?
a tall building beside a river?
fireworks above a city?
an airplane on a runway?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Which of these four complete sentences best matches the rocket photograph?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which of these four complete sentences best matches the rocket photograph? A score for each candidate sentence. CLIP ranks supplied text. It does not generate captions. A caption generator requires additional machinery. Remove the rocket sentence and rerun. CLIP can only rank the sentences still on the menu.

Teaching note

Which of these four complete sentences best matches the rocket photograph? A score for each candidate sentence. CLIP ranks supplied text. It does not generate captions. A caption generator requires additional machinery. Remove the rocket sentence and rerun. CLIP can only rank the sentences still on the menu.

01 · From fixed labels to language

What happened? · Image → text

APPLICATION 04 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Rocket launch

SpaceX · CC0

Descriptions we supply · cosine

a rocket launching into the night sky+0.2635
a tall building beside a river+0.2041
fireworks above a city+0.2085
an airplane on a runway+0.1966

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

The rocket sentence wins. We supplied all four sentences; CLIP generated none of them.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which of these four complete sentences best matches the rocket photograph? A score for each candidate sentence. CLIP ranks supplied text. It does not generate captions. A caption generator requires additional machinery. Remove the rocket sentence and rerun. CLIP can only rank the sentences still on the menu.

Teaching note

Which of these four complete sentences best matches the rocket photograph? A score for each candidate sentence. CLIP ranks supplied text. It does not generate captions. A caption generator requires additional machinery. Remove the rocket sentence and rerun. CLIP can only rank the sentences still on the menu.

01 · From fixed labels to language

Find another image like this one

APPLICATION 05 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Persian cat

Query image · exclude the image itself

These are the top four retrieved images, shown in result order. Inspect their similarities.

Image vector → compare with other image vectors. The text encoder is not used.

Which other gallery image will be nearest to the Persian cat photograph?

Radford et al., 2021 · OpenAI implementation · Image credits · Roboflow · embedding analysis, 10:40 and 17:22 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which other gallery image will be nearest to the Persian cat photograph? The other 21 gallery images ranked against the query image. The image encoder is useful on its own. Similarity can suggest groups or possible duplicates; it cannot certify that two files are duplicates. Switch the query to the pug or a portrait. No text candidates are needed.

Teaching note

Which other gallery image will be nearest to the Persian cat photograph? The other 21 gallery images ranked against the query image. The image encoder is useful on its own. Similarity can suggest groups or possible duplicates; it cannot certify that two files are duplicates. Switch the query to the pug or a portrait. No text candidates are needed.

01 · From fixed labels to language

What happened? · Image → image

APPLICATION 05 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Persian cat

Query image · exclude the image itself

These are the top four retrieved images, shown in result order. Inspect their similarities.

Image vector → compare with other image vectors. The text encoder is not used.

Chelsea is the nearest other image. This comparison uses only the image encoder.

Radford et al., 2021 · OpenAI implementation · Image credits · Roboflow · embedding analysis, 10:40 and 17:22 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which other gallery image will be nearest to the Persian cat photograph? The other 21 gallery images ranked against the query image. The image encoder is useful on its own. Similarity can suggest groups or possible duplicates; it cannot certify that two files are duplicates. Switch the query to the pug or a portrait. No text candidates are needed.

Teaching note

Which other gallery image will be nearest to the Persian cat photograph? The other 21 gallery images ranked against the query image. The image encoder is useful on its own. Similarity can suggest groups or possible duplicates; it cannot certify that two files are duplicates. Switch the query to the pug or a portrait. No text candidates are needed.

01 · From fixed labels to language

What if we add a hat?

APPLICATION 06 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · predict before revealing

Before
After

A straw hat is added to the portrait.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply

hat?
cup?
cat?
boat?

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Which candidate word best describes the change between the two pictures?

Radford et al., 2021 · OpenAI implementation · Image credits · CodeEmporium · hat experiment, 7:25 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which candidate word best describes the change between the two pictures? One signed cosine between the after − before direction and each candidate word. The person, pose, clothing, framing, and background are intended to stay the same. These are generated pictures. Hair, lighting, and fine details can also change; the pair does not perfectly isolate a hat. Inspired by CodeEmporium’s hat experiment, using our own generated pair. This is an exploratory direction, not a trained change detector. Subtle image differences remain. Cosine is not a probability of the change. Swap before / after and run again. With fixed candidates, every change score should reverse sign.

Teaching note

Which candidate word best describes the change between the two pictures? One signed cosine between the after − before direction and each candidate word. The person, pose, clothing, framing, and background are intended to stay the same. These are generated pictures. Hair, lighting, and fine details can also change; the pair does not perfectly isolate a hat. Inspired by CodeEmporium’s hat experiment, using our own generated pair. This is an exploratory direction, not a trained change detector. Subtle image differences remain. Cosine is not a probability of the change. Swap before / after and run again. With fixed candidates, every change score should reverse sign.

01 · From fixed labels to language

What happened? · Hat change

APPLICATION 06 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · raw cosines, not probabilities

Before
After

A straw hat is added to the portrait.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply · cosine

hat+0.0808
cup+0.0132
cat+0.0026
boat+0.0254

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Hat has the largest change cosine for this pair and these four words.

Radford et al., 2021 · OpenAI implementation · Image credits · CodeEmporium · hat experiment, 7:25 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which candidate word best describes the change between the two pictures? One signed cosine between the after − before direction and each candidate word. The person, pose, clothing, framing, and background are intended to stay the same. These are generated pictures. Hair, lighting, and fine details can also change; the pair does not perfectly isolate a hat. Inspired by CodeEmporium’s hat experiment, using our own generated pair. This is an exploratory direction, not a trained change detector. Subtle image differences remain. Cosine is not a probability of the change. Swap before / after and run again. With fixed candidates, every change score should reverse sign.

Teaching note

Which candidate word best describes the change between the two pictures? One signed cosine between the after − before direction and each candidate word. The person, pose, clothing, framing, and background are intended to stay the same. These are generated pictures. Hair, lighting, and fine details can also change; the pair does not perfectly isolate a hat. Inspired by CodeEmporium’s hat experiment, using our own generated pair. This is an exploratory direction, not a trained change detector. Subtle image differences remain. Cosine is not a probability of the change. Swap before / after and run again. With fixed candidates, every change score should reverse sign.

01 · From fixed labels to language

What if the mug changes colour?

APPLICATION 07 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · predict before revealing

Before
After

The mug’s surface color changes from red to blue.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply

blue?
red?
cup?
hat?

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Which word matches the change from the red mug to the blue mug?

Radford et al., 2021 · OpenAI implementation · Image credits · Extension of the CodeEmporium subtraction idea · Try this example · Saved vectors and scores

What changed from the previous slide?

Which word matches the change from the red mug to the blue mug? Signed matches to the change direction, not probabilities that an image contains a blue or red object. The mug shape, handle, position, viewpoint, tabletop, and background are intended to stay the same. Generated reflections, edges, and background details can differ too. The measured direction contains every encoded difference, not only color. Difference directions depend on images, normalization and wording. They are exploratory measurements, not guaranteed semantic algebra. Swap before / after and run again. The scores reverse sign; the highest-ranked word may change.

Teaching note

Which word matches the change from the red mug to the blue mug? Signed matches to the change direction, not probabilities that an image contains a blue or red object. The mug shape, handle, position, viewpoint, tabletop, and background are intended to stay the same. Generated reflections, edges, and background details can differ too. The measured direction contains every encoded difference, not only color. Difference directions depend on images, normalization and wording. They are exploratory measurements, not guaranteed semantic algebra. Swap before / after and run again. The scores reverse sign; the highest-ranked word may change.

01 · From fixed labels to language

What happened? · Color change

APPLICATION 07 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · raw cosines, not probabilities

Before
After

The mug’s surface color changes from red to blue.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply · cosine

blue+0.1121
red-0.1778
cup-0.0120
hat-0.0117

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Blue is positive, red is negative, and cup is near zero: the cup is shared by both images.

Radford et al., 2021 · OpenAI implementation · Image credits · Extension of the CodeEmporium subtraction idea · Try this example · Saved vectors and scores

What changed from the previous slide?

Which word matches the change from the red mug to the blue mug? Signed matches to the change direction, not probabilities that an image contains a blue or red object. The mug shape, handle, position, viewpoint, tabletop, and background are intended to stay the same. Generated reflections, edges, and background details can differ too. The measured direction contains every encoded difference, not only color. Difference directions depend on images, normalization and wording. They are exploratory measurements, not guaranteed semantic algebra. Swap before / after and run again. The scores reverse sign; the highest-ranked word may change.

Teaching note

Which word matches the change from the red mug to the blue mug? Signed matches to the change direction, not probabilities that an image contains a blue or red object. The mug shape, handle, position, viewpoint, tabletop, and background are intended to stay the same. Generated reflections, edges, and background details can differ too. The measured direction contains every encoded difference, not only color. Difference directions depend on images, normalization and wording. They are exploratory measurements, not guaranteed semantic algebra. Swap before / after and run again. The scores reverse sign; the highest-ranked word may change.

01 · From fixed labels to language

Can words help us find a brick kiln?

APPLICATION 08 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Brick kiln · generated

AI-generated teaching image · OpenAI image tool · 29 Sep 2026

Descriptions we supply

an aerial image of a brick kiln?
an aerial image of a solar farm?
an aerial image of agricultural fields?
an aerial image of industrial warehouses?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Which supplied scene description matches the generated kiln tile?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which supplied scene description matches the generated kiln tile? Four scene-description scores for the selected synthetic aerial tile. The search idea transfers to satellite tiles, but synthetic images are not evidence of field performance. Real deployment needs georeferenced imagery, independent labels, and evaluation across regions and resolutions. Compare ‘brick kiln’ with a longer description in a separate live run. Record the images and wording used.

Teaching note

Which supplied scene description matches the generated kiln tile? Four scene-description scores for the selected synthetic aerial tile. The search idea transfers to satellite tiles, but synthetic images are not evidence of field performance. Real deployment needs georeferenced imagery, independent labels, and evaluation across regions and resolutions. Compare ‘brick kiln’ with a longer description in a separate live run. Record the images and wording used.

01 · From fixed labels to language

What happened? · Sustainability

APPLICATION 08 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Brick kiln · generated

AI-generated teaching image · OpenAI image tool · 29 Sep 2026

Descriptions we supply · cosine

an aerial image of a brick kiln+0.3135
an aerial image of a solar farm+0.2882
an aerial image of agricultural fields+0.2903
an aerial image of industrial warehouses+0.2902

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Kiln wins this generated scene. This is a demonstration, not a satellite benchmark.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Which supplied scene description matches the generated kiln tile? Four scene-description scores for the selected synthetic aerial tile. The search idea transfers to satellite tiles, but synthetic images are not evidence of field performance. Real deployment needs georeferenced imagery, independent labels, and evaluation across regions and resolutions. Compare ‘brick kiln’ with a longer description in a separate live run. Record the images and wording used.

Teaching note

Which supplied scene description matches the generated kiln tile? Four scene-description scores for the selected synthetic aerial tile. The search idea transfers to satellite tiles, but synthetic images are not evidence of field performance. Real deployment needs georeferenced imagery, independent labels, and evaluation across regions and resolutions. Compare ‘brick kiln’ with a longer description in a separate live run. Record the images and wording used.

01 · From fixed labels to language

What if we add a chimney?

APPLICATION 09 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · predict before revealing

Before
After

A tall brick chimney is added to a kiln scene.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply

chimney?
tree?
truck?
solar panel?

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Which word matches the change between the two kiln scenes?

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Which word matches the change between the two kiln scenes? A ranked match between the image-change direction and four words. The kiln, brick stacks, camera position, and surrounding landscape are intended to stay the same. The chimney brings a shadow; generated fine details also vary. This pair is not a verified physical intervention or a trained kiln detector. Synthetic paired scenes test a hypothesis. Shadows and fine details also change; this is not a causal proof or a validated kiln detector. Swap the order and rerun. Inspect the pictures for changed shadows and background details.

Teaching note

Which word matches the change between the two kiln scenes? A ranked match between the image-change direction and four words. The kiln, brick stacks, camera position, and surrounding landscape are intended to stay the same. The chimney brings a shadow; generated fine details also vary. This pair is not a verified physical intervention or a trained kiln detector. Synthetic paired scenes test a hypothesis. Shadows and fine details also change; this is not a causal proof or a validated kiln detector. Swap the order and rerun. Inspect the pictures for changed shadows and background details.

01 · From fixed labels to language

What happened? · Chimney change

APPLICATION 09 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · raw cosines, not probabilities

Before
After

A tall brick chimney is added to a kiln scene.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply · cosine

chimney+0.0336
tree+0.0160
truck+0.0057
solar panel+0.0048

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Chimney wins this synthetic pair. Shadows and scene details can also affect the difference.

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Which word matches the change between the two kiln scenes? A ranked match between the image-change direction and four words. The kiln, brick stacks, camera position, and surrounding landscape are intended to stay the same. The chimney brings a shadow; generated fine details also vary. This pair is not a verified physical intervention or a trained kiln detector. Synthetic paired scenes test a hypothesis. Shadows and fine details also change; this is not a causal proof or a validated kiln detector. Swap the order and rerun. Inspect the pictures for changed shadows and background details.

Teaching note

Which word matches the change between the two kiln scenes? A ranked match between the image-change direction and four words. The kiln, brick stacks, camera position, and surrounding landscape are intended to stay the same. The chimney brings a shadow; generated fine details also vary. This pair is not a verified physical intervention or a trained kiln detector. Synthetic paired scenes test a hypothesis. Shadows and fine details also change; this is not a causal proof or a validated kiln detector. Swap the order and rerun. Inspect the pictures for changed shadows and background details.

01 · From fixed labels to language

Does removing the stripes change the match?

MEASURED ORIGINAL CLIP · ViT-B/32 · CPU float32 · generated before/after images · raw cosines

Before: stripes
After: generated edit
Supplied descriptionBeforeAfter
a photo of a zebra0.3358?
a photo of a horse0.2689?
a photo of a donkey0.2638?
a photo of an okapi0.2646?

Classify each image separately. Same model, same descriptions; no vector subtraction here.

The first image matches “zebra”. Which description will win after the edit?

Radford et al., 2021 · OpenAI implementation · Image credits · Both measured score sets · Generated image provenance

What changed from the previous slide?

This is an input-image intervention, not an edit to an internal concept bottleneck. The generated edit can change shape, texture and fine detail as well as stripes. It suggests sensitivity to the edit; one pair does not establish that stripes alone caused the prediction change. Softmax shares are introduced only after the loss walkthrough.

Teaching note

This is an input-image intervention, not an edit to an internal concept bottleneck. The generated edit can change shape, texture and fine detail as well as stripes. It suggests sensitivity to the edit; one pair does not establish that stripes alone caused the prediction change. Softmax shares are introduced only after the loss walkthrough.

01 · From fixed labels to language

Does removing the stripes change the match?

MEASURED ORIGINAL CLIP · ViT-B/32 · CPU float32 · generated before/after images · raw cosines

Before: stripes
After: generated edit
Supplied descriptionBeforeAfter
a photo of a zebra0.33580.2233
a photo of a horse0.26890.2980
a photo of a donkey0.26380.2718
a photo of an okapi0.26460.1824

Classify each image separately. Same model, same descriptions; no vector subtraction here.

Zebra drops from 0.3358 to 0.2233; horse becomes the best match. The edit changes the prediction.

Radford et al., 2021 · OpenAI implementation · Image credits · Both measured score sets · Generated image provenance

What changed from the previous slide?

This is an input-image intervention, not an edit to an internal concept bottleneck. The generated edit can change shape, texture and fine detail as well as stripes. It suggests sensitivity to the edit; one pair does not establish that stripes alone caused the prediction change. Softmax shares are introduced only after the loss walkthrough.

Teaching note

This is an input-image intervention, not an edit to an internal concept bottleneck. The generated edit can change shape, texture and fine detail as well as stripes. It suggests sensitivity to the edit; one pair does not establish that stripes alone caused the prediction change. Softmax shares are introduced only after the loss walkthrough.

01 · From fixed labels to language

What kind of medical image is this?

APPLICATION 10 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Chest radiograph

Stillwaterising · Wikimedia Commons · CC0

Descriptions we supply

a chest X-ray?
an X-ray of a hand?
a brain MRI scan?
a photograph of a person?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Can CLIP match this public radiograph to an imaging modality?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can CLIP match this public radiograph to an imaging modality? A ranking of image-type descriptions, not a diagnosis. This public radiograph supports a modality-matching lesson, not diagnosis. General CLIP scores are not clinical evidence; medical tasks need domain-specific validation. Keep the distinction clear: identifying an image type is a different task from diagnosing disease.

Teaching note

Can CLIP match this public radiograph to an imaging modality? A ranking of image-type descriptions, not a diagnosis. This public radiograph supports a modality-matching lesson, not diagnosis. General CLIP scores are not clinical evidence; medical tasks need domain-specific validation. Keep the distinction clear: identifying an image type is a different task from diagnosing disease.

01 · From fixed labels to language

What happened? · Healthcare

APPLICATION 10 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Chest radiograph

Stillwaterising · Wikimedia Commons · CC0

Descriptions we supply · cosine

a chest X-ray+0.2673
an X-ray of a hand+0.2295
a brain MRI scan+0.2047
a photograph of a person+0.2304

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Chest X-ray wins among these four descriptions. We have not tested a diagnosis.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can CLIP match this public radiograph to an imaging modality? A ranking of image-type descriptions, not a diagnosis. This public radiograph supports a modality-matching lesson, not diagnosis. General CLIP scores are not clinical evidence; medical tasks need domain-specific validation. Keep the distinction clear: identifying an image type is a different task from diagnosing disease.

Teaching note

Can CLIP match this public radiograph to an imaging modality? A ranking of image-type descriptions, not a diagnosis. This public radiograph supports a modality-matching lesson, not diagnosis. General CLIP scores are not clinical evidence; medical tasks need domain-specific validation. Keep the distinction clear: identifying an image type is a different task from diagnosing disease.

01 · From fixed labels to language

Will an added device give a useful change vector?

APPLICATION 11 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · predict before revealing

Before
After

A pacemaker-like device and leads are added to a synthetic radiograph-style image.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply

pacemaker?
rib?
stethoscope?
hat?

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Will the added device make ‘pacemaker’ the highest-scoring word?

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Will the added device make ‘pacemaker’ the highest-scoring word? Four exploratory word scores for a synthetic change; no clinical finding is inferred. The chest outline, viewpoint, and radiograph appearance are intended to stay the same. Both images are generated; anatomy and device placement are approximate, and other pixels differ. This is not a patient before/after study. Both radiographs are generated illustrations. Anatomy and device placement are approximate. This measures concept matching on synthetic pictures, not diagnostic or clinical performance. Try fuller descriptions in a separate run. Record the original failure alongside any improvement.

Teaching note

Will the added device make ‘pacemaker’ the highest-scoring word? Four exploratory word scores for a synthetic change; no clinical finding is inferred. The chest outline, viewpoint, and radiograph appearance are intended to stay the same. Both images are generated; anatomy and device placement are approximate, and other pixels differ. This is not a patient before/after study. Both radiographs are generated illustrations. Anatomy and device placement are approximate. This measures concept matching on synthetic pictures, not diagnostic or clinical performance. Try fuller descriptions in a separate run. Record the original failure alongside any improvement.

01 · From fixed labels to language

What happened? · Device change

APPLICATION 11 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · raw cosines, not probabilities

Before
After

A pacemaker-like device and leads are added to a synthetic radiograph-style image.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply · cosine

pacemaker+0.0127
rib-0.0133
stethoscope+0.0318
hat+0.0372

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Hat wins; pacemaker does not. The intended edit does not guarantee a useful direction.

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Will the added device make ‘pacemaker’ the highest-scoring word? Four exploratory word scores for a synthetic change; no clinical finding is inferred. The chest outline, viewpoint, and radiograph appearance are intended to stay the same. Both images are generated; anatomy and device placement are approximate, and other pixels differ. This is not a patient before/after study. Both radiographs are generated illustrations. Anatomy and device placement are approximate. This measures concept matching on synthetic pictures, not diagnostic or clinical performance. Try fuller descriptions in a separate run. Record the original failure alongside any improvement.

Teaching note

Will the added device make ‘pacemaker’ the highest-scoring word? Four exploratory word scores for a synthetic change; no clinical finding is inferred. The chest outline, viewpoint, and radiograph appearance are intended to stay the same. Both images are generated; anatomy and device placement are approximate, and other pixels differ. This is not a patient before/after study. Both radiographs are generated illustrations. Anatomy and device placement are approximate. This measures concept matching on synthetic pictures, not diagnostic or clinical performance. Try fuller descriptions in a separate run. Record the original failure alongside any improvement.

01 · From fixed labels to language

Can a sketch still match “dog”?

APPLICATION 12 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Dog sketch · generated

AI-generated teaching image · OpenAI image tool · 29 Sep 2026

Descriptions we supply

a drawing of a dog?
a drawing of a cat?
a drawing of a horse?
a drawing of a fox?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Can the dog concept still match when the image is a pencil sketch?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can the dog concept still match when the image is a pencil sketch? One similarity per description for a single image; there is no subtraction in this example. This is a small domain-shift experiment. A success here does not establish robustness to every drawing style. Open Style change to ask a different question: what changed between the two images?

Teaching note

Can the dog concept still match when the image is a pencil sketch? One similarity per description for a single image; there is no subtraction in this example. This is a small domain-shift experiment. A success here does not establish robustness to every drawing style. Open Style change to ask a different question: what changed between the two images?

01 · From fixed labels to language

What happened? · Photo → sketch

APPLICATION 12 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Dog sketch · generated

AI-generated teaching image · OpenAI image tool · 29 Sep 2026

Descriptions we supply · cosine

a drawing of a dog+0.3176
a drawing of a cat+0.2601
a drawing of a horse+0.2757
a drawing of a fox+0.2639

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Dog wins for the sketch too. The subject can remain recognizable across styles.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Can the dog concept still match when the image is a pencil sketch? One similarity per description for a single image; there is no subtraction in this example. This is a small domain-shift experiment. A success here does not establish robustness to every drawing style. Open Style change to ask a different question: what changed between the two images?

Teaching note

Can the dog concept still match when the image is a pencil sketch? One similarity per description for a single image; there is no subtraction in this example. This is a small domain-shift experiment. A success here does not establish robustness to every drawing style. Open Style change to ask a different question: what changed between the two images?

01 · From fixed labels to language

What changes when a photo becomes a sketch?

APPLICATION 13 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · predict before revealing

Before
After

A photograph-style dog scene is rendered as a pencil sketch.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply

pencil sketch?
photograph?
dog?
cat?

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Does the change direction emphasize pencil sketch, photograph, dog, or cat?

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Does the change direction emphasize pencil sketch, photograph, dog, or cat? A signed match to the style-change direction, distinct from recognizing the dog in either image. The dog identity, pose, and composition are intended to remain similar. Color, texture, background, and small shapes all change. This direction does not isolate a single style coordinate. This subtraction includes background and rendering changes. Compare original image-to-text matching with the difference direction; a successful guess does not isolate a causal concept. Swap the order and rerun. Then use Photo → sketch to classify each image separately.

Teaching note

Does the change direction emphasize pencil sketch, photograph, dog, or cat? A signed match to the style-change direction, distinct from recognizing the dog in either image. The dog identity, pose, and composition are intended to remain similar. Color, texture, background, and small shapes all change. This direction does not isolate a single style coordinate. This subtraction includes background and rendering changes. Compare original image-to-text matching with the difference direction; a successful guess does not isolate a causal concept. Swap the order and rerun. Then use Photo → sketch to classify each image separately.

01 · From fixed labels to language

What happened? · Style change

APPLICATION 13 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · generated image pair · raw cosines, not probabilities

Before
After

A photograph-style dog scene is rendered as a pencil sketch.

Two generated inputs. CLIP does not edit either picture.

Descriptions we supply · cosine

pencil sketch+0.1468
photograph-0.0105
dog-0.0078
cat-0.0022

Unit image vectors → normalize(after − before) → compare with unit text vectors.

Pencil sketch wins; dog is near zero because the subject is shared by both pictures.

Radford et al., 2021 · OpenAI implementation · Image credits · Our extension of CodeEmporium’s subtraction experiment · Try this example · Saved vectors and scores

What changed from the previous slide?

Does the change direction emphasize pencil sketch, photograph, dog, or cat? A signed match to the style-change direction, distinct from recognizing the dog in either image. The dog identity, pose, and composition are intended to remain similar. Color, texture, background, and small shapes all change. This direction does not isolate a single style coordinate. This subtraction includes background and rendering changes. Compare original image-to-text matching with the difference direction; a successful guess does not isolate a causal concept. Swap the order and rerun. Then use Photo → sketch to classify each image separately.

Teaching note

Does the change direction emphasize pencil sketch, photograph, dog, or cat? A signed match to the style-change direction, distinct from recognizing the dog in either image. The dog identity, pose, and composition are intended to remain similar. Color, texture, background, and small shapes all change. This direction does not isolate a single style coordinate. This subtraction includes background and rendering changes. Compare original image-to-text matching with the difference direction; a successful guess does not isolate a causal concept. Swap the order and rerun. Then use Photo → sketch to classify each image separately.

01 · From fixed labels to language

What if the right answer is missing?

APPLICATION 14 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Newfoundland

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply

a photo of a cat?
a photo of a bird?
a photo of a car?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

What will win when none of the descriptions says ‘dog’?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

What will win when none of the descriptions says ‘dog’? Scores and relative shares across the supplied menu, even if every option is wrong. Softmax distributes 100% across the choices you provide. It is not calibrated confidence that the best label is true. Add ‘a photo of a dog’ and rerun. The old pairwise cosines stay fixed, while their relative shares change.

Teaching note

What will win when none of the descriptions says ‘dog’? Scores and relative shares across the supplied menu, even if every option is wrong. Softmax distributes 100% across the choices you provide. It is not calibrated confidence that the best label is true. Add ‘a photo of a dog’ and rerun. The old pairwise cosines stay fixed, while their relative shares change.

01 · From fixed labels to language

What happened? · Candidate trap

APPLICATION 14 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Newfoundland

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply · cosine

a photo of a cat+0.2152
a photo of a bird+0.2184
a photo of a car+0.2179

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Bird wins a menu with no dog option. The highest score can still select a wrong description.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

What will win when none of the descriptions says ‘dog’? Scores and relative shares across the supplied menu, even if every option is wrong. Softmax distributes 100% across the choices you provide. It is not calibrated confidence that the best label is true. Add ‘a photo of a dog’ and rerun. The old pairwise cosines stay fixed, while their relative shares change.

Teaching note

What will win when none of the descriptions says ‘dog’? Scores and relative shares across the supplied menu, even if every option is wrong. Softmax distributes 100% across the choices you provide. It is not calibrated confidence that the best label is true. Add ‘a photo of a dog’ and rerun. The old pairwise cosines stay fixed, while their relative shares change.

01 · From fixed labels to language

Does “without” change what the model sees?

APPLICATION 15 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · predict before revealing

Persian cat

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply

a photo of a cat?
a photo without a cat?
a photo of a dog?

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

Will ‘a photo without a cat’ score much lower than ‘a photo of a cat’?

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Will ‘a photo without a cat’ score much lower than ‘a photo of a cat’? A similarity for each complete phrase; no rule explicitly implements logical negation. A flexible vocabulary does not guarantee precise compositional reasoning. Report what this checkpoint does on these examples, not universal claims. Try another phrasing while keeping the photograph fixed. Save failures as well as successes.

Teaching note

Will ‘a photo without a cat’ score much lower than ‘a photo of a cat’? A similarity for each complete phrase; no rule explicitly implements logical negation. A flexible vocabulary does not guarantee precise compositional reasoning. Report what this checkpoint does on these examples, not universal claims. Try another phrasing while keeping the photograph fixed. Save failures as well as successes.

01 · From fixed labels to language

What happened? · Language limits

APPLICATION 15 / 15 · frozen pretrained CLIP ViT-B/32 · saved Node CPU q8 measurements · raw cosines, not probabilities

Persian cat

Oxford-IIIT Pet · CC BY-SA 4.0

Descriptions we supply · cosine

a photo of a cat+0.2737
a photo without a cat+0.2699
a photo of a dog+0.2236

Image vector + one vector per supplied description → cosine scores. CLIP writes no answer text.

“A cat” and “without a cat” differ by only 0.0038 here. Wording alone does not ensure logic.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP · Radford et al., ICML 2021 · Try this example · Saved vectors and scores

What changed from the previous slide?

Will ‘a photo without a cat’ score much lower than ‘a photo of a cat’? A similarity for each complete phrase; no rule explicitly implements logical negation. A flexible vocabulary does not guarantee precise compositional reasoning. Report what this checkpoint does on these examples, not universal claims. Try another phrasing while keeping the photograph fixed. Save failures as well as successes.

Teaching note

Will ‘a photo without a cat’ score much lower than ‘a photo of a cat’? A similarity for each complete phrase; no rule explicitly implements logical negation. A flexible vocabulary does not guarantee precise compositional reasoning. Report what this checkpoint does on these examples, not universal claims. Try another phrasing while keeping the photograph fixed. Save failures as well as successes.

01 · From fixed labels to language

What made all these questions possible?

APPLICATION TOUR · from observed behaviour to the model that makes it possible

Newfoundland
Rocket launch
Brick kiln · generated
Chest radiograph
Dog sketch · generated

The model stayed fixed. We changed the pictures, the words or the comparison.

Now let’s follow those inputs through CLIP, then work out how it learns to match them.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

All fifteen original lab experiments are back in the opening. Keep the medical-device failure and the wrong-menu example visible. The next section returns to the two encoders; the complete loss is worked through later.

Teaching note

All fifteen original lab experiments are back in the opening. Keep the medical-device failure and the wrong-menu example visible. The next section returns to the two encoders; the complete loss is worked through later.

Section 02

Two encoders, one comparison

Key idea

How can a photograph and a sentence become comparable?

02 · Two encoders, one comparison

The classifier has 1,000 stored class vectors

EARLIER VIT-TINY · 192 image features · actual stored checkpoint weights, rounded for display

Newfoundland
↓
Earlier ViT
↓ final CLS
h∈R192\mathbf h\in\mathbb R^{192}

192 image features

Stored output layer W1,000 rows × 192 weights
Fixed categoryIts stored weights · first 3 of 192Bias
Newfoundlanddog class
wdog=[+0.0068,  +0.0195,  −0.0066,  …]\mathbf w_{\mathrm{dog}}=[+0.0068,\;+0.0195,\;-0.0066,\;\ldots]
+0.0017+0.0017
Persian catcat class
wcat=[+0.0218,  +0.0067,  −0.0400,  …]\mathbf w_{\mathrm{cat}}=[+0.0218,\;+0.0067,\;-0.0400,\;\ldots]
−0.0024-0.0024
macawbird class
wbird=[−0.0683,  −0.0459,  −0.0311,  …]\mathbf w_{\mathrm{bird}}=[-0.0683,\;-0.0459,\;-0.0311,\;\ldots]
+0.0026+0.0026
⋮ 997 more class rows, each with its own 192 weights and bias

Each row: dot with the same h, then add its bias.

sdog=wdog⊤h+bdogs_{\mathrm{dog}}=\mathbf w_{\mathrm{dog}}^\top\mathbf h+b_{\mathrm{dog}}
s=Wh+b⟶1,000  class scores\mathbf s=\mathbf W\mathbf h+\mathbf b\quad\longrightarrow\quad 1{,}000\;\text{class scores}

A new photo changes h. The 1,000 class vectors stay fixed when we predict.

Radford et al., 2021 · OpenAI implementation · Image credits · Actual class weights and checkpoint provenance · ImageNet class names · Linear layer

Where would the weights for a new category come from?

The earlier ImageNet classifier has 1,000 output categories. Each row of its PyTorch head.weight tensor stores 192 learned weights; head.bias stores one bias for each row. We show actual checkpoint rows for Newfoundland (dog), Persian cat and macaw (bird), reordered for this illustration. The dog/cat/bird subscripts are shorthand for these three specific categories, not broad ImageNet labels. Only the first three coordinates are printed; all 192 enter the dot product. The names label the output rows and are not sent through a text encoder. Labelled training learned both the ViT and these class parameters. During prediction, only the computed image features change. These are affine scores, not cosine similarities. Typing a new class name does not supply another row of learned weights; the following slide asks how language could provide the comparison vector instead.

Teaching note

The earlier ImageNet classifier has 1,000 output categories. Each row of its PyTorch head.weight tensor stores 192 learned weights; head.bias stores one bias for each row. We show actual checkpoint rows for Newfoundland (dog), Persian cat and macaw (bird), reordered for this illustration. The dog/cat/bird subscripts are shorthand for these three specific categories, not broad ImageNet labels. Only the first three coordinates are printed; all 192 enter the dot product. The names label the output rows and are not sent through a text encoder. Labelled training learned both the ViT and these class parameters. During prediction, only the computed image features change. These are affine scores, not cosine similarities. Typing a new class name does not supply another row of learned weights; the following slide asks how language could provide the comparison vector instead.

02 · Two encoders, one comparison

Let the words supply the class vector

BRIDGE TO CLIP · a stored class row → a vector computed from our description

Previous slide: 1,000 fixed categories

Each category has a stored vector, such as wdog\mathbf w_{\mathrm{dog}}.

h⊤wdog+bdog\mathbf h^\top\mathbf w_{\mathrm{dog}}+b_{\mathrm{dog}}one class score
Dog
CLIP image encoder
u\mathbf uimage vector
a photo of a dog
CLIP text encoder
vdog\mathbf v_{\mathrm{dog}}vector for this description
u⊤vdog\mathbf u^\top\mathbf v_{\mathrm{dog}}
match score

Compare the image with this description.

Change the description, and CLIP computes a new class vector. The model weights stay fixed.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP: descriptions as classifier vectors

Why can we try a description that was not one of the earlier classifier’s 1,000 labels?

The previous classifier compared image features h with a stored weight vector for each fixed category. That vector is what we mean by the class-side comparison vector: it determines how the image is scored against that category. CLIP supplies this side of the comparison by encoding a description into v, so our candidate list is not limited to the earlier 1,000 stored classifier rows. The image is encoded by CLIP’s own image encoder to u; we do not attach text vectors to the earlier 192-dimensional ViT features. Both displayed CLIP encoders include their learned projection and unit normalization, expanded in the following slides. Their final vectors have 512 coordinates and were trained together to be comparable. The diagram shows raw cosine u dot v, while the earlier classifier uses an affine score with a bias. This is a bridge between the roles of the vectors, not equality of the two models or their spaces. A new description changes the computed text vector, not the pretrained parameters; it also does not guarantee a reliable match.

Teaching note

The previous classifier compared image features h with a stored weight vector for each fixed category. That vector is what we mean by the class-side comparison vector: it determines how the image is scored against that category. CLIP supplies this side of the comparison by encoding a description into v, so our candidate list is not limited to the earlier 1,000 stored classifier rows. The image is encoded by CLIP’s own image encoder to u; we do not attach text vectors to the earlier 192-dimensional ViT features. Both displayed CLIP encoders include their learned projection and unit normalization, expanded in the following slides. Their final vectors have 512 coordinates and were trained together to be comparable. The diagram shows raw cosine u dot v, while the earlier classifier uses an affine score with a bias. This is a bridge between the roles of the vectors, not equality of the two models or their spaces. A new description changes the computed text vector, not the pretrained parameters; it also does not guarantee a reliable match.

02 · Two encoders, one comparison

First turn the image into one vector

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

Use the final CLS state after the ViT blocks and LayerNorm.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Use the final CLS state after the ViT blocks and LayerNorm.

Teaching note

Use the final CLS state after the ViT blocks and LayerNorm.

02 · Two encoders, one comparison

Project the image summary

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

A learned 768 × 512 matrix maps the image summary into the comparison space.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

A learned 768 × 512 matrix maps the image summary into the comparison space.

Teaching note

A learned 768 × 512 matrix maps the image summary into the comparison space.

02 · Two encoders, one comparison

Normalize it

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

Divide by the vector’s length. The resulting image vector has length 1.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Divide by the vector’s length. The resulting image vector has length 1.

Teaching note

Divide by the vector’s length. The resulting image vector has length 1.

02 · Two encoders, one comparison

The complete image–text comparison

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

Next, zoom into the highlighted text Transformer and its EOT output.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Show the complete route first: each input passes through its own encoder, learned projection and unit normalization, then the two vectors meet at a cosine score. The purple highlight marks the text Transformer and its final EOT readout. The next slide enlarges this part; the following slides walk through text projection and normalization.

Teaching note

Show the complete route first: each input passes through its own encoder, learned projection and unit normalization, then the two vectors meet at a cosine score. The purple highlight marks the text Transformer and its final EOT readout. The next slide enlarges this part; the following slides walk through text projection and normalization.

02 · Two encoders, one comparison

Zoom in: EOT is our sentence readout

ORIGINAL CLIP · causal attention in the text branch; full attention in the image ViT

Recall Beyond AttentionWe selected the updated CLS state to represent a sentence.

SOTStart of text
a
photo
of
a
dog
EOTEnd of text
Causal text Transformer

Each token reads itself and earlier tokens.
EOT can read the whole description.

Readout: select the final EOT state
hT=LN⁡(hEOT)∈R512\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})\in\mathbb R^{512}
LayerNorm (LN) → 512 numbers for this description. Projection comes next.

The same readout idea: use one updated token state as the representation of the whole description.

Radford et al., 2021 · OpenAI implementation · Image credits · Causal mask and EOT readout · SOT and EOT token names

What changed from the previous slide?

Connect to the sentence-readout diagram in the earlier From Attention to Applications lecture, referred to here as Beyond Attention. There we selected the updated CLS state after full self-attention. Here we use the same readout idea, but select the final EOT state after causal self-attention; these are analogous roles, not identical tokens or vectors. SOT means start of text (<|startoftext|>); EOT means end of text (<|endoftext|>). Both are special input tokens added around the description, not start or end of an individual word. EOT follows all description tokens; padding, omitted here, can follow EOT. For example, “photo” can attend to SOT, “a” and itself, but not the later “dog” token. EOT can attend to all preceding tokens and itself. SOT cannot collect the later words under this causal mask. The input EOT embedding is not the final contextual state: the Transformer computes that state. LN is the final LayerNorm; h_T has 512 features in ViT-B/32, before the learned text projection and unit normalization on the next slides. This is an encoder-like use of the Transformer: produce one representation of the supplied text. It does not change the causal attention mask. We train with CLIP’s contrastive loss; there is no next-token prediction head or next-token training loss. The paper says masked attention preserved the option of adding language modelling as an auxiliary objective, but leaves that experiment to future work. The image ViT uses full self-attention across its class token and patches.

Teaching note

Connect to the sentence-readout diagram in the earlier From Attention to Applications lecture, referred to here as Beyond Attention. There we selected the updated CLS state after full self-attention. Here we use the same readout idea, but select the final EOT state after causal self-attention; these are analogous roles, not identical tokens or vectors. SOT means start of text (<|startoftext|>); EOT means end of text (<|endoftext|>). Both are special input tokens added around the description, not start or end of an individual word. EOT follows all description tokens; padding, omitted here, can follow EOT. For example, “photo” can attend to SOT, “a” and itself, but not the later “dog” token. EOT can attend to all preceding tokens and itself. SOT cannot collect the later words under this causal mask. The input EOT embedding is not the final contextual state: the Transformer computes that state. LN is the final LayerNorm; h_T has 512 features in ViT-B/32, before the learned text projection and unit normalization on the next slides. This is an encoder-like use of the Transformer: produce one representation of the supplied text. It does not change the causal attention mask. We train with CLIP’s contrastive loss; there is no next-token prediction head or next-token training loss. The paper says masked attention preserved the option of adding language modelling as an auxiliary objective, but leaves that experiment to future work. The image ViT uses full self-attention across its class token and patches.

02 · Two encoders, one comparison

Project the text summary

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

A learned 512 × 512 projection can change direction without changing width.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

A learned 512 × 512 projection can change direction without changing width.

Teaching note

A learned 512 × 512 projection can change direction without changing width.

02 · Two encoders, one comparison

Normalize the text vector too

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

We now have one image vector and one text vector, each with 512 coordinates and unit length.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

We now have one image vector and one text vector, each with 512 coordinates and unit length.

Teaching note

We now have one image vector and one text vector, each with 512 coordinates and unit length.

02 · Two encoders, one comparison

Only now compare the directions

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

uvBoth lengths are 145°
Read the diagram labels
  • u
  • v
  • Both lengths are 1
  • 45°
u⊤v=∥u∥ ∥v∥cos⁡θ\mathbf u^\top\mathbf v=\|\mathbf u\|\,\|\mathbf v\|\cos\theta

The dot product depends on both lengths and the angle between the vectors.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

The dot product depends on both lengths and the angle between the vectors.

Teaching note

The dot product depends on both lengths and the angle between the vectors.

02 · Two encoders, one comparison

Unit length leaves only the angle

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

uvBoth lengths are 145°
Read the diagram labels
  • u
  • v
  • Both lengths are 1
  • 45°
u⊤v=1⋅1⋅cos⁡θ=cos⁡θ\mathbf u^\top\mathbf v=1\cdot1\cdot\cos\theta=\cos\theta

Normalization makes the dot product equal to cosine similarity.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Normalization makes the dot product equal to cosine similarity.

Teaching note

Normalization makes the dot product equal to cosine similarity.

02 · Two encoders, one comparison

Same direction: cosine 1

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

uvBoth lengths are 10°
Read the diagram labels
  • u
  • v
  • Both lengths are 1
  • 0°
cos⁡(0∘)=1\cos(0^\circ)=1

This is geometry. Training must make this geometry useful.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

This is geometry. Training must make this geometry useful.

Teaching note

This is geometry. Training must make this geometry useful.

02 · Two encoders, one comparison

A right angle: cosine 0

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

uvBoth lengths are 190°
Read the diagram labels
  • u
  • v
  • Both lengths are 1
  • 90°
cos⁡(90∘)=0\cos(90^\circ)=0

This is geometry. Training must make this geometry useful.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

This is geometry. Training must make this geometry useful.

Teaching note

This is geometry. Training must make this geometry useful.

02 · Two encoders, one comparison

Opposite directions: cosine −1

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

uvBoth lengths are 1180°
Read the diagram labels
  • u
  • v
  • Both lengths are 1
  • 180°
cos⁡(180∘)=−1\cos(180^\circ)=-1

This is geometry. Training must make this geometry useful.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

This is geometry. Training must make this geometry useful.

Teaching note

This is geometry. Training must make this geometry useful.

02 · Two encoders, one comparison

Two independent branches meet at the score

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
ViT
hI∈R768\mathbf h_I\in\mathbb R^{768}
WI\mathbf W_I
zI∈R512\mathbf z_I\in\mathbb R^{512}
÷∥zI∥\div\|\mathbf z_I\|
u\mathbf u
a photo of a dog
Causal text Transformer
hT=LN⁡(hEOT)\mathbf h_T=\operatorname{LN}(h_{\mathrm{EOT}})
WT\mathbf W_T
zT∈R512\mathbf z_T\in\mathbb R^{512}
÷∥zT∥\div\|\mathbf z_T\|
v\mathbf v
u⊤v\mathbf u^\top\mathbf v
cosine similarity

The branches do not exchange tokens. They meet after encoding, at the dot product.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

The branches do not exchange tokens. They meet after encoding, at the dot product.

Teaching note

The branches do not exchange tokens. They meet after encoding, at the dot product.

Section 03

Why should a cosine mean “match”?

Key idea

Equal dimensions let us calculate a score. What makes the score meaningful?

03 · Why should a cosine mean “match”?

Equal width is only the starting point

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

u,v∈R512\mathbf u,\mathbf v\in\mathbb R^{512}

Would two random encoders understand each other?

Same width permits comparison. Training gives the comparison meaning.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Same width permits comparison. Training gives the comparison meaning.

Teaching note

Same width permits comparison. Training gives the comparison meaning.

03 · Why should a cosine mean “match”?

Before training: the closest caption can be wrong

SCHEMATIC · illustrative directions, not measured checkpoints

Before alignment: the closest text directions belong to different pairs. cat, dog, car images and their full captions label six unit arrows. Directions are illustrative, not measured.u (cat image)v (cat text)“a photo of a cat”u (dog image)v (dog text)“a photo of a dog”u (car image)v (car text)“a photo of a car”same originimage vectortext vectorSmaller angle↓larger cosine
Read the diagram labels
  • u (cat image)
  • v (cat text)
  • “a photo of a cat”
  • u (dog image)
  • v (dog text)
  • “a photo of a dog”
  • u (car image)
  • v (car text)
  • “a photo of a car”
  • same origin
  • image vector
  • text vector
  • Smaller angle
  • ↓
  • larger cosine

The model has no useful pairing information merely because both outputs have width 512.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

The model has no useful pairing information merely because both outputs have width 512.

Teaching note

The model has no useful pairing information merely because both outputs have width 512.

03 · Why should a cosine mean “match”?

After training: paired directions move closer

SCHEMATIC · illustrative directions, not measured checkpoints

After alignment: each image vector points close to its paired text vector. cat, dog, car images and their full captions label six unit arrows. Directions are illustrative, not measured.u (cat image)v (cat text)“a photo of a cat”u (dog image)v (dog text)“a photo of a dog”u (car image)v (car text)“a photo of a car”same originimage vectortext vectorSmaller angle↓larger cosine
Read the diagram labels
  • u (cat image)
  • v (cat text)
  • “a photo of a cat”
  • u (dog image)
  • v (dog text)
  • “a photo of a dog”
  • u (car image)
  • v (car text)
  • “a photo of a car”
  • same origin
  • image vector
  • text vector
  • Smaller angle
  • ↓
  • larger cosine

Learn a geometry in which observed pairs beat competing pairs.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Learn a geometry in which observed pairs beat competing pairs.

Teaching note

Learn a geometry in which observed pairs beat competing pairs.

03 · Why should a cosine mean “match”?

An individual coordinate has no assigned concept

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

u=[u1,u2,…,u512]\mathbf u=[u_1,u_2,\ldots,u_{512}]

We did not label coordinates “dogness”, “redness” or “furriness”.

The matching information is carried by the learned geometry.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

The matching information is carried by the learned geometry.

Teaching note

The matching information is carried by the learned geometry.

03 · Why should a cosine mean “match”?

If every caption scores 1, how do we choose?

HYPOTHETICAL FAILURE · these are illustrative scores, not measured CLIP predictions

Imagine the model maps every image and every caption to the same unit vector.

Dog
Compare this image
with each caption
Supplied captionCosine with this image
a photo of a dogPaired description for this image1.00
a photo of a cat1.00
a photo of a car1.00

Three equal scores. The paired description has no advantage.

A high score helps only if the paired caption scores higher than the competing captions.

Radford et al., 2021 · OpenAI implementation · Image credits

Why does a cosine of 1 fail to identify the paired description in this example?

Return to the toy dog image and the descriptions used in the alignment examples. Suppose both branches map every input to the same unit vector. The dot product of that vector with itself is 1, so every image–caption comparison ties at cosine 1. In the table, each row names one supplied caption; the number on its right is the cosine between that caption vector and this one image’s vector. The highlighted row identifies the supplied training pair, not a model-selected winner. These are hypothetical values, not measured CLIP outputs or probabilities. We cannot identify the observed partner from tied scores. This failure is called representation collapse. The next slide motivates comparison with competitors; we introduce batches and the full image-by-caption matrix afterward.

Teaching note

Return to the toy dog image and the descriptions used in the alignment examples. Suppose both branches map every input to the same unit vector. The dot product of that vector with itself is 1, so every image–caption comparison ties at cosine 1. In the table, each row names one supplied caption; the number on its right is the cosine between that caption vector and this one image’s vector. The highlighted row identifies the supplied training pair, not a model-selected winner. These are hypothetical values, not measured CLIP outputs or probabilities. We cannot identify the observed partner from tied scores. This failure is called representation collapse. The next slide motivates comparison with competitors; we introduce batches and the full image-by-caption matrix afterward.

03 · Why should a cosine mean “match”?

A paired description must beat its competitors

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Dog
a photo of a dog

observed partner

a photo of a cat

competitor in this batch

Contrastive training compares the observed pairing with alternative pairings.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Contrastive training compares the observed pairing with alternative pairings.

Teaching note

Contrastive training compares the observed pairing with alternative pairings.

Section 04

Follow one batch through CLIP

Key idea

Keep the examples in view as we build the computation, one step at a time.

04 · Follow one batch through CLIP

Start with images and the text found with them

SCHEMATIC · toy illustration and supplied caption; not an original CLIP web record

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cata photo of a cat
One image–text record

Where does the text come from?

A webpage may have a caption, title or description associated with an image.

Collect those pairs. Filter the collection, then sample complete records for training.

The pairing supplies the training target; the model does not write the caption.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The original CLIP paper describes a query-filtered collection of 400M image–text pairs. Our 12 records are authored teaching examples. The illustrations, supplied captions and chosen vectors are the same as in the tutorial. They are teaching examples, not measured CLIP embeddings. Associated web text can be noisy, incomplete, duplicated or only loosely related.

Teaching note

The original CLIP paper describes a query-filtered collection of 400M image–text pairs. Our 12 records are authored teaching examples. The illustrations, supplied captions and chosen vectors are the same as in the tutorial. They are teaching examples, not measured CLIP embeddings. Associated web text can be noisy, incomplete, duplicated or only loosely related.

04 · Follow one batch through CLIP

Here are the 12 pairs in our small dataset

TOY DATASET · 12 authored image–caption pairs from the interactive

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cata photo of a cat
Doga photo of a dog
Cara photo of a car
Bicyclea bicycle
Airplanean airplane
Treea tree
Shoea shoe
Muga mug
Bananaa banana
Chaira chair
Buildinga building
Birda bird

For this walkthrough, take cat, dog and car as the first batch.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

For this walkthrough, take cat, dog and car as the first batch.

Teaching note

For this walkthrough, take cat, dog and car as the first batch.

04 · Follow one batch through CLIP

Sample three complete pairs

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cata photo of a cat
Doga photo of a dog
Cara photo of a car
Batch 1: cat · dog · car9 other records wait for later batches.

Keep image 1 with text 1, image 2 with text 2, and image 3 with text 3.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Keep image 1 with text 1, image 2 with text 2, and image 3 with text 3.

Teaching note

Keep image 1 with text 1, image 2 with text 2, and image 3 with text 3.

04 · Follow one batch through CLIP

The batch gives us partners and competitors

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Images ↓
Text →
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
Image I₁Cat
I₁ ↔ T₁Observed pairI₁ ↔ T₂CompetitorI₁ ↔ T₃Competitor
Image I₂Dog
I₂ ↔ T₁CompetitorI₂ ↔ T₂Observed pairI₂ ↔ T₃Competitor
Image I₃Car
I₃ ↔ T₁CompetitorI₃ ↔ T₂CompetitorI₃ ↔ T₃Observed pair

3 observed pairs · 6 competitors · 9 comparisons

I₁ pairs with T₁, I₂ with T₂, and I₃ with T₃. The other cells are competitors in this batch.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

I denotes an image and T its supplied text. The matching indices identify the sampled pair records. An off-diagonal caption is treated as a competitor here; it need not be a false description in general.

Teaching note

I denotes an image and T its supplied text. The matching indices identify the sampled pair records. An off-diagonal caption is treated as a competitor here; it need not be a false description in general.

04 · Follow one batch through CLIP

One image, one caption, two paths

TEACHING EXAMPLE · toy illustration, supplied caption, chosen vectors · 3 coordinates to show every calculation

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁
Image encoder
features
WIProjection
zᴵ₁[2.0, 0.5, 0.2]Projected vector
÷ 2.0712
u₁[0.9656, 0.2414, 0.0966]Unit vector · length 1
a photo of a catText T₁
Text encoder
features
WTProjection
zᵀ₁[1.6, 0.8, 0.4]Projected vector
÷ 1.8330
v₁[0.8729, 0.4364, 0.2182]Unit vector · length 1

We will compare the two final unit vectors with u₁ · v₁. Next, work through each step.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors. The overview previews both paths; the following slides expand features, projection and normalization. Normalization changes vector length, not the number of coordinates. This restores the old worksheet-notation link.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors. The overview previews both paths; the following slides expand features, projection and normalization. Normalization changes vector length, not the number of coordinates. This restores the old worksheet-notation link.

04 · Follow one batch through CLIP

Follow the cat image all the way to its unit vector

CHOSEN ARITHMETIC · 4 image features → 3 coordinates · same cat vector as the interactive

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁
Image encoder
hᴵ₁
[1]1[2]2[3]1[4]1
4 toy features
×
WI
4 × 3 weights
zᴵ₁
[1]2[2]0.5[3]0.2
3 coordinates
÷ ‖z‖₂
u₁
[1]0.9656[2]0.2414[3]0.0966
Unit length

The encoder makes features. The projection changes coordinates. Normalization sets the length to 1.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

First, the encoder turns the image into features

CHOSEN ARITHMETIC · 4 image features → 3 coordinates · same cat vector as the interactive

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · cat
Image encoder
hᴵ₁
[1]1[2]2[3]1[4]1
1 row × 4 features

In the real ViT, we read the final CLS state after LayerNorm: 768 image features.

For this arithmetic example, take hᴵ₁ = [1, 2, 1, 1]. These four values are chosen.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

Project the image features into three coordinates

CHOSEN ARITHMETIC · 4 image features → 3 coordinates · same cat vector as the interactive

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cat image I₁Image encoderFeatures hᴵ₁Projection WIVector zᴵ₁Unit vector u₁
WI · 4 input features × 3 output coordinates
z[1]z[2]z[3]
h[1] = 1100.1
h[2] = 20.50.20
h[3] = 100.10
h[4] = 1000.1

Multiply h by W

h1I⏟1×4  WI⏟4×3=z1I⏟1×3\underbrace{\mathbf h^I_1}_{1\times4}\;\underbrace{\mathbf W_I}_{4\times3}=\underbrace{\mathbf z^I_1}_{1\times3}

Each column of W supplies the weights for one output coordinate.

Use all four features for each column.

The projection learns how to combine features. Its output will be compared with text.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

Calculate one output coordinate at a time

CHOSEN ARITHMETIC · 4 image features → 3 coordinates · same cat vector as the interactive

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cat image I₁Image encoderFeatures hᴵ₁Projection WIVector zᴵ₁Unit vector u₁
z[1](1 × 1) + (2 × 0.5) + (1 × 0) + (1 × 0)= 2
z[2](1 × 0) + (2 × 0.2) + (1 × 0.1) + (1 × 0)= 0.5
z[3](1 × 0.1) + (2 × 0) + (1 × 0) + (1 × 0.1)= 0.2
z1I=[2,  0.5,  0.2]∈R3\mathbf z^I_1=[2,\;0.5,\;0.2]\in\mathbb R^3

This is the interactive’s raw cat vector. Next, divide it by its length.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

Normalize the cat image vector

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Cat image I₁Image encoderFeatures hᴵ₁Projection WIVector zᴵ₁Unit vector u₁
Image I₁

Projected vector
[2.0, 0.5, 0.2]

u1=z/∥z∥2\mathbf u_1=\mathbf z/\|\mathbf z\|_2

Divide each coordinate by the same length

∥z∥2=22+0.52+0.22=2.071232\|\mathbf z\|_2=\sqrt{2^2+0.5^2+0.2^2}=2.071232
2.0 ÷ 2.071232 = 0.965609
0.5 ÷ 2.071232 = 0.241402
0.2 ÷ 2.071232 = 0.096561
u₁[0.9656, 0.2414, 0.0966]

The vector still has three coordinates. Normalization changes its length, not its direction.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

The caption follows its own encoder and projection

CHOSEN 3D EXAMPLE · text vector from the interactive · real ViT-B/32 text projection: 512 → 512

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
a photo
of a cat
Text T₁
Causal text
Transformer
hᵀ₁

Final EOT
readout

Before projection
×
WT
Text projection
zᵀ₁
[1]1.6[2]0.8[3]0.4
3 coordinates
÷ ‖z‖₂
v₁
[1]0.8729[2]0.4364[3]0.2182
Unit length

The two branches use different weights. Their projected vectors have the same number of coordinates.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The text branch reads the final contextual EOT state after LayerNorm, applies its own learned projection, and normalizes the resulting vector. The displayed raw text vector and unit vector are the original tutorial’s chosen values, not measured embeddings. The projection uses W_T for text and W_I for images. The stacked coordinate lists are not matrix-orientation notation; the multiplication uses row vectors.

Teaching note

The text branch reads the final contextual EOT state after LayerNorm, applies its own learned projection, and normalizes the resulting vector. The displayed raw text vector and unit vector are the original tutorial’s chosen values, not measured embeddings. The projection uses W_T for text and W_I for images. The stacked coordinate lists are not matrix-orientation notation; the multiplication uses row vectors.

04 · Follow one batch through CLIP

Normalize the cat text vector

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
a photo of a catText T₁

Projected vector
[1.6, 0.8, 0.4]

v1=z/∥z∥2\mathbf v_1=\mathbf z/\|\mathbf z\|_2

Divide each coordinate by the same length

∥z∥2=1.62+0.82+0.42=1.833030\|\mathbf z\|_2=\sqrt{1.6^2+0.8^2+0.4^2}=1.833030
1.6 ÷ 1.833030 = 0.872872
0.8 ÷ 1.833030 = 0.436436
0.4 ÷ 1.833030 = 0.218218
v₁[0.8729, 0.4364, 0.2182]

The vector still has three coordinates. Normalization changes its length, not its direction.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

These are the source interactive’s chosen text outputs after projection. Divide by their Euclidean norm; no extra model is run.

Teaching note

These are the source interactive’s chosen text outputs after projection. Divide by their Euclidean norm; no extra model is run.

04 · Follow one batch through CLIP

Now compare the image with its caption

CHOSEN 3D VECTORS · unit-vector dot product · shown coordinates are rounded; calculation uses full precision

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ → u₁[0.9656, 0.2414, 0.0966]
a photo of a catText T₁ → v₁[0.8729, 0.4364, 0.2182]

(0.9656 × 0.8729) + (0.2414 × 0.4364) + (0.0966 × 0.2182)

Cosine for I₁ and T₁0.969281

One image and one caption give one score. Next, repeat this calculation for the other pairs.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

Teaching note

The toy cat illustration identifies the same record throughout. Its supplied text is “a photo of a cat”. We choose h and W to explain the projection; no trained image model produced these four feature values. The multiplication reproduces the interactive’s existing raw cat vector exactly. Use row vectors here: h is 1×4, W is 4×3 and z is 1×3. Real CLIP ViT-B/32 uses 768 image features projected to 512 coordinates. Projection is linear and learned; unit normalization is a separate operation. The source interactive begins at the projected vectors.

04 · Follow one batch through CLIP

Repeat both branches for every pair

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
RecordRaw image output zᴵRaw text output zᵀ
Cata photo of a cat
[2.0, 0.5, 0.2][1.6, 0.8, 0.4]
Doga photo of a dog
[0.4, 2.0, 0.1][0.7, 1.4, 0.3]
Cara photo of a car
[0.2, 0.3, 2.0][0.4, 0.5, 1.5]

Use the same image weights for all images and the same text weights for all captions. Normalize every output.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

These are the original interactive’s six chosen post-projection outputs. Only the cat-image projection was expanded as a worked hW calculation; the other raw vectors remain the provided toy outputs.

Teaching note

These are the original interactive’s six chosen post-projection outputs. Only the cat-image projection was expanded as a worked hW calculation; the other raw vectors remain the provided toy outputs.

04 · Follow one batch through CLIP

Six unit vectors, ready to compare

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

U · image vectors

u₁ [0.9656, 0.2414, 0.0966]

u₂ [0.1959, 0.9794, 0.0490]

u₃ [0.0984, 0.1476, 0.9841]

V · text vectors

v₁ [0.8729, 0.4364, 0.2182]

v₂ [0.4392, 0.8784, 0.1882]

v₃ [0.2453, 0.3066, 0.9197]

xyzu₁v₁v₂u₂v₃u₃

● solid = image   □ dashed = text   ·   colour = pair

Colour identifies a pair. Each vector has length 1; the directions can still differ.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The geometry is the source tutorial’s orthographic view of the actual chosen 3D unit vectors. It is not a t-SNE/PCA projection of measured CLIP features. The source geometry labels I/T are changed to the lecture’s u/v; positions and vectors are unchanged.

Teaching note

The geometry is the source tutorial’s orthographic view of the actual chosen 3D unit vectors. It is not a t-SNE/PCA projection of measured CLIP features. The source geometry labels I/T are changed to the lecture’s u/v; positions and vectors are unchanged.

04 · Follow one batch through CLIP

Stack the image vectors as rows

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

U: image rows

Rows ↓ / columns →Coordinate 1Coordinate 2Coordinate 3
Image 1 · cat0.96560.24140.0966
Image 2 · dog0.19590.97940.0490
Image 3 · car0.09840.14760.9841

3 images × 3 coordinates

Row 1 is the cat image. Its three columns are coordinates, not classes.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Row 1 is the cat image. Its three columns are coordinates, not classes.

Teaching note

Row 1 is the cat image. Its three columns are coordinates, not classes.

04 · Follow one batch through CLIP

Stack the text vectors in the same pair order

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

V: text rows

Rows ↓ / columns →Coordinate 1Coordinate 2Coordinate 3
Text 1 · cat0.87290.43640.2182
Text 2 · dog0.43920.87840.1882
Text 3 · car0.24530.30660.9197

3 texts × 3 coordinates

Row 1 is “a photo of a cat”. Keep it aligned with image row 1.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Row 1 is “a photo of a cat”. Keep it aligned with image row 1.

Teaching note

Row 1 is “a photo of a cat”. Keep it aligned with image row 1.

04 · Follow one batch through CLIP

Transpose V: each caption now has a column

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Vᵀ: text columns

Rows ↓ / columns →Text 1catText 2dogText 3car
Coordinate 10.87290.43920.2453
Coordinate 20.43640.87840.3066
Coordinate 30.21820.18820.9197

3 coordinates × 3 texts

Transpose moves the numbers. Text 1’s coordinates now run down column 1.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Transpose moves the numbers. Text 1’s coordinates now run down column 1.

Teaching note

Transpose moves the numbers. Text 1’s coordinates now run down column 1.

04 · Follow one batch through CLIP

An image row meets a text column

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁catI₁↔T₁positiveI₁↔T₂negativeI₁↔T₃negativeImage I₂dogI₂↔T₁negativeI₂↔T₂positiveI₂↔T₃negativeImage I₃carI₃↔T₁negativeI₃↔T₂negativeI₃↔T₃positiveDIAGONAL = PAIRED PARTNERSObserved pairs

One dot product per cell

U⏟3×3  V⊤⏟3×3=C⏟3×3\underbrace{\mathbf U}_{3\times3}\;\underbrace{\mathbf V^\top}_{3\times3}=\underbrace{\mathbf C}_{3\times3}

U: image rows × coordinates

Vᵀ: coordinates × text columns

C: image rows × text columns

In real ViT-B/32: (3 × 512) times (512 × 3) still gives nine scores.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The tutorial labels normalized matrices EI and ET and cosine scores S. The lecture keeps U, V and C introduced earlier. These names refer to exactly the same numeric arrays; no values change.

Teaching note

The tutorial labels normalized matrices EI and ET and cosine scores S. The lecture keeps U, V and C introduced earlier. These names refer to exactly the same numeric arrays; no values change.

04 · Follow one batch through CLIP

Follow the cat image and the cat caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat???Image I₂dog???Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Build C₁₁

C11=u1⊤v1C_{11}=\mathbf u_1^\top\mathbf v_1

u₁ = [0.9656, 0.2414, 0.0966]

v₁ = [0.8729, 0.4364, 0.2182]

Multiply matching coordinates, then add the three products.

The same coordinate positions meet: first with first, second with second, third with third.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The same coordinate positions meet: first with first, second with second, third with third.

Teaching note

The same coordinate positions meet: first with first, second with second, third with third.

04 · Follow one batch through CLIP

Cat image × cat caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair??Image I₂dog???Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₁₁

(0.9656 × 0.8729)
+ (0.2414 × 0.4364)
+ (0.0966 × 0.2182)

0.842853 + 0.105357 + 0.021071

Cosine0.969281

This is the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Cat image × dog caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂?Image I₂dog???Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₁₂

(0.9656 × 0.4392)
+ (0.2414 × 0.8784)
+ (0.0966 × 0.1882)

0.424114 + 0.212057 + 0.018176

Cosine0.654347

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Cat image × car caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog???Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₁₃

(0.9656 × 0.2453)
+ (0.2414 × 0.3066)
+ (0.0966 × 0.9197)

0.236821 + 0.074007 + 0.088808

Cosine0.399636

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Dog image × cat caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁??Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₂₁

(0.1959 × 0.8729)
+ (0.9794 × 0.4364)
+ (0.0490 × 0.2182)

0.170979 + 0.427447 + 0.010686

Cosine0.609112

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Dog image × dog caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair?Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₂₂

(0.1959 × 0.4392)
+ (0.9794 × 0.8784)
+ (0.0490 × 0.1882)

0.086035 + 0.860346 + 0.009218

Cosine0.955599

This is the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Dog image × car caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair0.393I₂ · T₃Image I₃car???DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₂₃

(0.1959 × 0.2453)
+ (0.9794 × 0.3066)
+ (0.0490 × 0.9197)

0.048041 + 0.300256 + 0.045038

Cosine0.393335

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Car image × cat caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair0.393I₂ · T₃Image I₃car0.365I₃ · T₁??DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₃₁

(0.0984 × 0.8729)
+ (0.1476 × 0.4364)
+ (0.9841 × 0.2182)

0.085902 + 0.064427 + 0.214756

Cosine0.365085

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Car image × dog caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair0.393I₂ · T₃Image I₃car0.365I₃ · T₁0.358I₃ · T₂?DIAGONAL = PAIRED PARTNERSC · cosine

Cell C₃₂

(0.0984 × 0.4392)
+ (0.1476 × 0.8784)
+ (0.9841 × 0.1882)

0.043225 + 0.129675 + 0.185250

Cosine0.358151

This is an in-batch competitor. Use the same dot-product calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Car image × car caption

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair0.393I₂ · T₃Image I₃car0.365I₃ · T₁0.358I₃ · T₂0.975correct pairDIAGONAL = PAIRED PARTNERSC · cosine

Cell C₃₃

(0.0984 × 0.2453)
+ (0.1476 × 0.3066)
+ (0.9841 × 0.9197)

0.024136 + 0.045256 + 0.905118

Cosine0.974511

This is the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

Teaching note

Rounded coordinates are printed; the displayed sum is computed with full-precision normalized vectors. The figures use the interactive’s own SVG matrix renderer, keeping every cell in the same position.

04 · Follow one batch through CLIP

Now every image has three caption scores

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.969correct pair0.654I₁ · T₂0.400I₁ · T₃Image I₂dog0.609I₂ · T₁0.956correct pair0.393I₂ · T₃Image I₃car0.365I₃ · T₁0.358I₃ · T₂0.975correct pairDIAGONAL = PAIRED PARTNERSC · cosine

Read the cat row

Cat caption: 0.9693

Dog caption: 0.6543

Car caption: 0.3996

C=UV⊤\mathbf C=\mathbf U\mathbf V^\top

Cosines describe directions. We have not calculated probabilities yet.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Cosines describe directions. We have not calculated probabilities yet.

Teaching note

Cosines describe directions. We have not calculated probabilities yet.

Section 05

From competition to a training loss

Key idea

Which caption belongs to each image—and which image belongs to each caption?

05 · From competition to a training loss

Scale the cosines before making a choice

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.939correct pair1.309I₁ · T₂0.799I₁ · T₃Image I₂dog1.218I₂ · T₁1.911correct pair0.787I₂ · T₃Image I₃car0.730I₃ · T₁0.716I₃ · T₂1.949correct pairDIAGONAL = PAIRED PARTNERSS · logits

Temperature τ = 0.5

Sij=Cij/τS_{ij}=C_{ij}/\tau
S11=0.969281/0.5S_{11}=0.969281/0.5
Cat–cat logit1.938561

Every entry is multiplied by 2.

Scaling preserves the ranking. It changes how sharply softmax will separate the choices.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Notation bridge: this lecture uses C for cosine, S for logits and calligraphic L for loss. The source interactive uses S for cosine and L for logits. Temperature is fixed at 0.5 in this toy run; real CLIP learns a positive logit scale.

Teaching note

Notation bridge: this lecture uses C for cosine, S for logits and calligraphic L for loss. The source interactive uses S for cosine and L for logits. Temperature is fixed at 0.5 in this toy run; real CLIP learns a positive logit scale.

05 · From competition to a training loss

Temperature changes the strength of competition

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car

Keep the cat-image cosines: [0.9693, 0.6543, 0.3996].

Smaller τ gives sharper shares; larger τ spreads them more evenly. The top caption stays the same.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Smaller τ gives sharper shares; larger τ spreads them more evenly. The top caption stays the same.

Teaching note

Smaller τ gives sharper shares; larger τ spreads them more evenly. The top caption stays the same.

05 · From competition to a training loss

Take the cat image’s row of logits

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993

These are the three scores for Image I₁. Softmax turns them into shares of one total.

Values shown are rounded; calculations use full precision.

Read across: T₁ is the cat caption, T₂ the dog caption, and T₃ the car caption.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Read across: T₁ is the cat caption, T₂ the dog caption, and T₃ the car caption.

Teaching note

Read across: T₁ is the cat caption, T₂ the dog caption, and T₃ the car caption.

05 · From competition to a training loss

Exponentiate each score in this row

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239

Each score becomes a positive weight. A larger score gives a larger weight.

Values shown are rounded; calculations use full precision.

The three captions still occupy the same three columns.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The three captions still occupy the same three columns.

Teaching note

The three captions still occupy the same three columns.

05 · From competition to a training loss

Add the three weights in this row

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239
Add across this row:
D1=e1.9386+e1.3087+e0.7993D_1=e^{1.9386}+e^{1.3087}+e^{0.7993}
≈6.9487+3.7013+2.2239≈12.8740\approx 6.9487+3.7013+2.2239\approx 12.8740
Values shown are rounded; calculations use full precision.

This one sum is the denominator for all three captions competing for Image I₁.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This one sum is the denominator for all three captions competing for Image I₁.

Teaching note

This one sum is the denominator for all three captions competing for Image I₁.

05 · From competition to a training loss

Divide the cat caption’s weight by the row sum

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239

Every caption uses the same row sum: D₁ ≈ 12.8740.

Values shown are rounded; calculations use full precision.

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Teaching note

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

05 · From competition to a training loss

Divide the dog caption’s weight by the row sum

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239

Every caption uses the same row sum: D₁ ≈ 12.8740.

Values shown are rounded; calculations use full precision.

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Teaching note

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

05 · From competition to a training loss

Divide the car caption’s weight by the row sum

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239

Every caption uses the same row sum: D₁ ≈ 12.8740.

Values shown are rounded; calculations use full precision.

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

Teaching note

Softmax share = this caption’s exponential weight ÷ the sum of all three weights in the row.

05 · From competition to a training loss

Now the cat image’s three shares sum to 1

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.5398I₁ · T₁0.2875I₁ · T₂0.1727I₁ · T₃Image I₂dog…later…later…laterImage I₃car…later…later…laterONE IMAGE PER ROW · THREE TEXT CANDIDATESP · row shares

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS1jS_{1j}1.93861.30870.7993
ExponentiateeS1je^{S_{1j}}e1.9386e^{1.9386}≈ 6.9487e1.3087e^{1.3087}≈ 3.7013e0.7993e^{0.7993}≈ 2.2239
0.5398+0.2875+0.1727=1.00000.5398+0.2875+0.1727=1.0000
Values shown are rounded; calculations use full precision.

P stores the softmax shares. We have completed only the first image row so far.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

P stores the softmax shares. We have completed only the first image row so far.

Teaching note

P stores the softmax shares. We have completed only the first image row so far.

05 · From competition to a training loss

Repeat the same calculation for the dog image

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₂ · dogKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.5398I₁ · T₁0.2875I₁ · T₂0.1727I₁ · T₃Image I₂dog0.2740I₂ · T₁0.5480I₂ · T₂0.1780I₂ · T₃Image I₃car…later…later…laterONE IMAGE PER ROW · THREE TEXT CANDIDATESP · row shares

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS2jS_{2j}1.21821.91120.7867
ExponentiateeS2je^{S_{2j}}e1.2182e^{1.2182}≈ 3.3812e1.9112e^{1.9112}≈ 6.7612e0.7867e^{0.7867}≈ 2.1961

Every caption uses the same row sum: D₂ ≈ 12.3384.

Values shown are rounded; calculations use full precision.

This image gets its own denominator: add the exponential weights across its row.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This image gets its own denominator: add the exponential weights across its row.

Teaching note

This image gets its own denominator: add the exponential weights across its row.

05 · From competition to a training loss

Repeat the same calculation for the car image

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₃ · carKeep this image fixed.
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.5398I₁ · T₁0.2875I₁ · T₂0.1727I₁ · T₃Image I₂dog0.2740I₂ · T₁0.5480I₂ · T₂0.1780I₂ · T₃Image I₃car0.1862I₃ · T₁0.1837I₃ · T₂0.6301I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESP · row shares

Compare the three captions across this row →

Same image
in every cell
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
LogitsS3jS_{3j}0.73020.71631.9490
ExponentiateeS3je^{S_{3j}}e0.7302e^{0.7302}≈ 2.0754e0.7163e^{0.7163}≈ 2.0468e1.9490e^{1.9490}≈ 7.0218

Every caption uses the same row sum: D₃ ≈ 11.1441.

Values shown are rounded; calculations use full precision.

This image gets its own denominator: add the exponential weights across its row.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This image gets its own denominator: add the exponential weights across its row.

Teaching note

This image gets its own denominator: add the exponential weights across its row.

05 · From competition to a training loss

Every image row now sums to 1

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.5398I₁ · T₁0.2875I₁ · T₂0.1727I₁ · T₃Image I₂dog0.2740I₂ · T₁0.5480I₂ · T₂0.1780I₂ · T₃Image I₃car0.1862I₃ · T₁0.1837I₃ · T₂0.6301I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESP · row shares

One softmax per image row

Image I₁
0.5398+0.2875+0.1727≈10.5398+0.2875+0.1727\approx1
Image I₂
0.2740+0.5480+0.1780≈10.2740+0.5480+0.1780\approx1
Image I₃
0.1862+0.1837+0.6301≈10.1862+0.1837+0.6301\approx1
Pij=eSij∑k=13eSikP_{ij}=\frac{e^{S_{ij}}}{\sum_{k=1}^3 e^{S_{ik}}}

Keep this completed matrix P. We will use its paired entries for the image → text loss.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The shares sum to one across each row. They are conditional on the supplied candidates and temperature, not calibrated probabilities that a caption is true. Keep P unchanged for the image-to-text loss. The next slide returns to the original temperature-scaled cosine scores S to start a separate column softmax.

Teaching note

The shares sum to one across each row. They are conditional on the supplied candidates and temperature, not calibrated probabilities that a caption is true. Keep P unchanged for the image-to-text loss. The next slide returns to the original temperature-scaled cosine scores S to start a separate column softmax.

05 · From competition to a training loss

Keep P. Return to the original scores S.

CHOSEN 3D EXAMPLE · C = cosine similarities · S = C / τ · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Keep P for the image → text loss

TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.5398I₁ · T₁0.2875I₁ · T₂0.1727I₁ · T₃Image I₂dog0.2740I₂ · T₁0.5480I₂ · T₂0.1780I₂ · T₃Image I₃car0.1862I₃ · T₁0.1837I₃ · T₂0.6301I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESP · row shares
S→ row softmax →P✓ complete

Return to the same scores: S = C / τ

TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.9386I₁ · T₁1.3087I₁ · T₂0.7993I₁ · T₃Image I₂dog1.2182I₂ · T₁1.9112I₂ · T₂0.7867I₂ · T₃Image I₃car0.7302I₃ · T₁0.7163I₃ · T₂1.9490I₃ · T₃ONE IMAGE PER ROW · THREE TEXT CANDIDATESS · logits
S→ column softmax →Qnext

Both softmax calculations start from S = C / τ. Q will supply the text → image loss.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

C contains the original unit-vector dot products, and S = C / τ contains their temperature-scaled logits. Row softmax creates P without overwriting S. Column softmax starts again from S, not from P. It creates a separate matrix Q. Later, the image-to-text loss reads the paired entries of P and the text-to-image loss reads those of Q; their mean is the CLIP loss. No encoder rerun or new dot products are needed here.

Teaching note

C contains the original unit-vector dot products, and S = C / τ contains their temperature-scaled logits. Row softmax creates P without overwriting S. Column softmax starts again from S, not from P. It creates a separate matrix Q. Later, the image-to-text loss reads the paired entries of P and the text-to-image loss reads those of Q; their mean is the CLIP loss. No encoder rerun or new dot products are needed here.

05 · From competition to a training loss

Back at S: which image belongs to the cat caption?

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.939correct pair1.309I₁ · T₂0.799I₁ · T₃Image I₂dog1.218I₂ · T₁1.911correct pair0.787I₂ · T₃Image I₃car0.730I₃ · T₁0.716I₃ · T₂1.949correct pairDIAGONAL = PAIRED PARTNERSS · logits

Read column 1 of S ↓

I₁catI₂dogI₃car
Cosine0.96930.60910.3651
÷ 0.5 → logit1.93861.21820.7302
Qi1=eSi1∑k=13eSk1Q_{i1}=\frac{e^{S_{i1}}}{\sum_{k=1}^3 e^{S_{k1}}}

P stays unchanged. Start from S again, fix caption T₁, and compare the three images down its column.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

P stays unchanged. Start from S again, fix caption T₁, and compare the three images down its column.

Teaching note

P stays unchanged. Start from S again, fix caption T₁, and compare the three images down its column.

05 · From competition to a training loss

Add the weights down this column

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat1.939correct pair1.309I₁ · T₂0.799I₁ · T₃Image I₂dog1.218I₂ · T₁1.911correct pair0.787I₂ · T₃Image I₃car0.730I₃ · T₁0.716I₃ · T₂1.949correct pairDIAGONAL = PAIRED PARTNERSS · logits

A different set of competitors

I₁catI₂dogI₃car
Cosine0.96930.60910.3651
÷ 0.5 → logit1.93861.21820.7302
exp(logit)6.94873.38122.0754
Sum this column ↓
6.9487 + 3.3812 + 2.0754
= 12.405358

The cat–cat numerator is unchanged. The column denominator includes different cells.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The cat–cat numerator is unchanged. The column denominator includes different cells.

Teaching note

The cat–cat numerator is unchanged. The column denominator includes different cells.

05 · From competition to a training loss

The cat caption assigns the cat image a share

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Follow I₁ × T₁

Q11=eS11∑keSk1Q_{11}=\frac{e^{S_{11}}}{\sum_k e^{S_{k1}}}
=6.94874612.405358=\frac{6.948746}{12.405358}
Share0.560141

Divide by the column sum because the candidate images compete for one caption.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Divide by the column sum because the candidate images compete for one caption.

Teaching note

Divide by the column sum because the candidate images compete for one caption.

05 · From competition to a training loss

The cat caption assigns the dog image a share

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Follow I₂ × T₁

Q21=eS21∑keSk1Q_{21}=\frac{e^{S_{21}}}{\sum_k e^{S_{k1}}}
=3.38117812.405358=\frac{3.381178}{12.405358}
Share0.272558

Divide by the column sum because the candidate images compete for one caption.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Divide by the column sum because the candidate images compete for one caption.

Teaching note

Divide by the column sum because the candidate images compete for one caption.

05 · From competition to a training loss

The cat caption assigns the car image a share

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Follow I₃ × T₁

Q31=eS31∑keSk1Q_{31}=\frac{e^{S_{31}}}{\sum_k e^{S_{k1}}}
=2.07543412.405358=\frac{2.075434}{12.405358}
Share0.167301

Divide by the column sum because the candidate images compete for one caption.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Divide by the column sum because the candidate images compete for one caption.

Teaching note

Divide by the column sum because the candidate images compete for one caption.

05 · From competition to a training loss

Now let the dog caption choose its image

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Column 2 ↓

I₁catI₂dogI₃car
÷ 0.5 → logit1.30871.91120.7163
exp(logit)3.70136.76122.0468
÷ sum → share0.29590.54050.1636
Sum this column ↓
3.7013 + 6.7612 + 2.0468
= 12.509366

Every caption gets its own denominator, computed down its column.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Every caption gets its own denominator, computed down its column.

Teaching note

Every caption gets its own denominator, computed down its column.

05 · From competition to a training loss

Now let the car caption choose its image

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Column 3 ↓

I₁catI₂dogI₃car
÷ 0.5 → logit0.79930.78671.9490
exp(logit)2.22392.19617.0218
÷ sum → share0.19440.19190.6137
Sum this column ↓
2.2239 + 2.1961 + 7.0218
= 11.441806

Every caption gets its own denominator, computed down its column.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Every caption gets its own denominator, computed down its column.

Teaching note

Every caption gets its own denominator, computed down its column.

05 · From competition to a training loss

Each column now sums to 1

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

Text → image

Qij=eSij∑keSkjQ_{ij}=\frac{e^{S_{ij}}}{\sum_k e^{S_{kj}}}

P is kept for the image → text loss.

Q is ready for the text → image loss.

We now have two probability matrices, both computed from the same scores S. We will average their two losses.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

We now have two probability matrices, both computed from the same scores S. We will average their two losses.

Teaching note

We now have two probability matrices, both computed from the same scores S. We will average their two losses.

05 · From competition to a training loss

Same pair and numerator. Different competition.

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Cat image → caption

P11=6.94874612.874005=0.539750P_{11}=\frac{6.948746}{12.874005}=0.539750

Competes with the dog and car captions.

Cat caption → image

Q11=6.94874612.405358=0.560141Q_{11}=\frac{6.948746}{12.405358}=0.560141

Competes with the dog and car images.

We train both retrieval directions.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

We train both retrieval directions.

Teaching note

We train both retrieval directions.

05 · From competition to a training loss

The recorded pair tells us which answer is correct

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.540correct pair0.288I₁ · T₂0.173I₁ · T₃Image I₂dog0.274I₂ · T₁0.548correct pair0.178I₂ · T₃Image I₃car0.186I₃ · T₁0.184I₃ · T₂0.630correct pairDIAGONAL = PAIRED PARTNERSP · row shares

Target for the cat image

[1, 0, 0]

Caption 1 is the observed partner.

Its share0.539750

A loss should be small when this share is close to 1.

The target comes from the data record, not from whichever cell currently has the largest score.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The target comes from the data record, not from whichever cell currently has the largest score.

Teaching note

The target comes from the data record, not from whichever cell currently has the largest score.

05 · From competition to a training loss

A small share for the partner should cost more

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
p=0.9p=0.9
−ln(p)0.105
p=0.333333p=0.333333
−ln(p)1.099
p=0.01p=0.01
−ln(p)4.605
ℓ=−ln⁡p\ell=-\ln p

As the paired share approaches 1, its penalty approaches 0. All logs here are natural.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

As the paired share approaches 1, its penalty approaches 0. All logs here are natural.

Teaching note

As the paired share approaches 1, its penalty approaches 0. All logs here are natural.

05 · From competition to a training loss

Cross-entropy reads the cat image’s row of P

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catBring back row 1 of P: this image’s three caption probabilities.
Same caption
in each column ↓
Text T₁a photo of a catText T₂a photo of a dogText T₃a photo of a car
Row 1 of Ppj=P1jp_j=P_{1j}0.5397500.2875050.172745
Target y1Recorded partner0Other caption0Other caption
Each loss term−yjln⁡pj-y_j\ln p_j−1ln⁡(0.539750)-1\ln(0.539750)−0ln⁡(0.287505)-0\ln(0.287505)−0ln⁡(0.172745)-0\ln(0.172745)
CE⁡(y,p)=−∑j=13yjln⁡pj\operatorname{CE}(\mathbf y,\mathbf p)=-\sum_{j=1}^3 y_j\ln p_j
Add the three terms:
−ln⁡(0.539750)+0+0≈0.616649-\ln(0.539750)+0+0\approx 0.616649

The target [1, 0, 0] selects −ln(P₁₁). All three captions influenced P₁₁ through softmax.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Read the first row of the saved row-softmax matrix P: [0.539750, 0.287505, 0.172745]. Each caption keeps the same column across the probabilities, target labels and negative-log terms. Image I₁ was paired with text T₁, giving target [1,0,0]. These values are rounded for display; the penalty is computed from the full-precision probability. Mathematical cross-entropy is shown on probabilities here. PyTorch F.cross_entropy takes logits and performs log-softmax plus the target negative log internally. Do not softmax first when calling that API.

Teaching note

Read the first row of the saved row-softmax matrix P: [0.539750, 0.287505, 0.172745]. Each caption keeps the same column across the probabilities, target labels and negative-log terms. Image I₁ was paired with text T₁, giving target [1,0,0]. These values are rounded for display; the penalty is computed from the full-precision probability. Mathematical cross-entropy is shown on probabilities here. PyTorch F.cross_entropy takes logits and performs log-softmax plus the target negative log internally. Do not softmax first when calling that API.

05 · From competition to a training loss

Cat image: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.540correct pair0.288I₁ · T₂0.173I₁ · T₃Image I₂dog0.274I₂ · T₁0.548correct pair0.178I₂ · T₃Image I₃car0.186I₃ · T₁0.184I₃ · T₂0.630correct pairDIAGONAL = PAIRED PARTNERSP · row shares

One correct-match penalty

ℓ1=−ln⁡P11\ell_1=-\ln P_{11}
=−ln⁡(0.539750)=-\ln(0.539750)
Penalty0.616649

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Dog image: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.540correct pair0.288I₁ · T₂0.173I₁ · T₃Image I₂dog0.274I₂ · T₁0.548correct pair0.178I₂ · T₃Image I₃car0.186I₃ · T₁0.184I₃ · T₂0.630correct pairDIAGONAL = PAIRED PARTNERSP · row shares

One correct-match penalty

ℓ2=−ln⁡P22\ell_2=-\ln P_{22}
=−ln⁡(0.547977)=-\ln(0.547977)
Penalty0.601521

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Car image: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.540correct pair0.288I₁ · T₂0.173I₁ · T₃Image I₂dog0.274I₂ · T₁0.548correct pair0.178I₂ · T₃Image I₃car0.186I₃ · T₁0.184I₃ · T₂0.630correct pairDIAGONAL = PAIRED PARTNERSP · row shares

One correct-match penalty

ℓ3=−ln⁡P33\ell_3=-\ln P_{33}
=−ln⁡(0.630093)=-\ln(0.630093)
Penalty0.461888

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Average the three image → text penalties

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
cat0.616649
dog0.601521
car0.461888
LI→T=0.616649+0.601521+0.4618883=0.560020\mathcal L_{I\to T}=\frac{0.616649+0.601521+0.461888}3=0.560020

Three prediction tasks contribute to this mean. We do not average nine cells.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Three prediction tasks contribute to this mean. We do not average nine cells.

Teaching note

Three prediction tasks contribute to this mean. We do not average nine cells.

05 · From competition to a training loss

Cat caption: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

One correct-match penalty

ℓ1=−ln⁡Q11\ell_1=-\ln Q_{11}
=−ln⁡(0.560141)=-\ln(0.560141)
Penalty0.579567

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Dog caption: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

One correct-match penalty

ℓ2=−ln⁡Q22\ell_2=-\ln Q_{22}
=−ln⁡(0.540489)=-\ln(0.540489)
Penalty0.615280

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Car caption: read its paired cell

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.560correct pair0.296I₁ · T₂0.194I₁ · T₃Image I₂dog0.273I₂ · T₁0.540correct pair0.192I₂ · T₃Image I₃car0.167I₃ · T₁0.164I₃ · T₂0.614correct pairDIAGONAL = PAIRED PARTNERSQ · column shares

One correct-match penalty

ℓ3=−ln⁡Q33\ell_3=-\ln Q_{33}
=−ln⁡(0.613698)=-\ln(0.613698)
Penalty0.488252

A smaller penalty means more share went to the observed partner.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

A smaller penalty means more share went to the observed partner.

Teaching note

A smaller penalty means more share went to the observed partner.

05 · From competition to a training loss

Average the three text → image penalties

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
cat0.579567
dog0.615280
car0.488252
LT→I=0.579567+0.615280+0.4882523=0.561033\mathcal L_{T\to I}=\frac{0.579567+0.615280+0.488252}3=0.561033

Three prediction tasks contribute to this mean. We do not average nine cells.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Three prediction tasks contribute to this mean. We do not average nine cells.

Teaching note

Three prediction tasks contribute to this mean. We do not average nine cells.

05 · From competition to a training loss

Give both retrieval directions equal weight

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image → text0.560020
+
Text → image0.561033
L=LI→T+LT→I2\mathcal L=\frac{\mathcal L_{I\to T}+\mathcal L_{T\to I}}2
This batch’s CLIP loss0.560526

The final value averages all six correct-match penalties.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The final value averages all six correct-match penalties.

Teaching note

The final value averages all six correct-match penalties.

05 · From competition to a training loss

What does “no preference” look like for one image?

HYPOTHETICAL BASELINE · three candidates · equal logits chosen as 0 · natural logs

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁ · catSuppose the model gives all three captions the same score.
Same image
in every cell
Text T₁a photo of a catRecorded partnerText T₂a photo of a dogOther captionText T₃a photo of a carOther caption
Equal logits000
Equal weightse0=1e^0=1e0=1e^0=1e0=1e^0=1
Divide by 1 + 1 + 113\frac1313\frac1313\frac13
Loss for the recorded partner T₁
ℓ1=−ln⁡P11=−ln⁡(1/3)≈1.098612\ell_1=-\ln P_{11}=-\ln(1/3)\approx1.098612

The correct caption gets only one third of the probability, just like either competitor.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This is a comparison case, not the cat row computed earlier. All equal finite logits give the same softmax, whatever their shared value: exp(c)/(3 exp(c)) = 1/3. The recorded partner is still T₁, so its negative-log penalty is ln(3). This uniform baseline does not claim that every randomly initialized model produces exactly equal scores.

Teaching note

This is a comparison case, not the cat row computed earlier. All equal finite logits give the same softmax, whatever their shared value: exp(c)/(3 exp(c)) = 1/3. The recorded partner is still T₁, so its negative-log penalty is ln(3). This uniform baseline does not claim that every randomly initialized model produces exactly equal scores.

05 · From competition to a training loss

If every choice is a tie, averaging keeps the same loss

HYPOTHETICAL BASELINE · three paired records · both retrieval directions tie

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Now suppose every row and every column gives equal shares.

Correct-match penalties
Cat
Dog
Car
Image → textEach paired share = ⅓I₁ → T₁1.098612I₂ → T₂1.098612I₃ → T₃1.098612
Text → imageEach paired share = ⅓T₁ → I₁1.098612T₂ → I₂1.098612T₃ → I₃1.098612
3 image penalties + 3 text penalties; average all six.
Ltie=3ln⁡3+3ln⁡36=ln⁡3≈1.098612\mathcal L_{\mathrm{tie}}=\frac{3\ln3+3\ln3}{6}=\ln3\approx1.098612

Six equal penalties, averaged: the batch baseline is 1.098612.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Imagine a score matrix whose nine entries are equal. Row softmax P and column softmax Q both contain 1/3 everywhere. Each of the three image penalties and each of the three text penalties is ln(3). Averaging within each direction, then across the two directions, is equivalent to averaging all six. It does not multiply the baseline by six.

Teaching note

Imagine a score matrix whose nine entries are equal. Row softmax P and column softmax Q both contain 1/3 everywhere. Each of the three image penalties and each of the three text penalties is ln(3). Averaging within each direction, then across the two directions, is equivalent to averaging all six. It does not multiply the baseline by six.

05 · From competition to a training loss

Now 0.5605 has something to compare with

READING THE LOSS · one recorded partner per image and caption · natural logs

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Ideal limit

→ 0

Every paired answer gets a share approaching 1.

Our cat / dog / car batch

0.560526

The paired answers do better than an equal split.

Equal-choice baseline

1.098612

Every candidate gets ⅓. The model has no preference.

Lower is better. A confidently wrong prediction can cost more than this baseline.

N pairs:ppaired=1/N⇒Ltie=ln⁡NN\text{ pairs:}\quad p_{\mathrm{paired}}=1/N\quad\Rightarrow\quad\mathcal L_{\mathrm{tie}}=\ln N

For the same three-pair batch: 0.560526 < 1.098612. The loss rewards putting more share on the paired answers.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The equal-choice baseline is ln(N) for a batch of N pairs with one positive each. It is not an upper bound: a paired probability below 1/N incurs a penalty above ln(N). Do not compare raw losses across different batch sizes without accounting for that changing reference. Zero is the ideal limit as all target probabilities approach one; at a fixed finite temperature and with bounded cosine scores it need not be attainable. A low training loss alone is not evidence of generalization.

Teaching note

The equal-choice baseline is ln(N) for a batch of N pairs with one positive each. It is not an upper bound: a paired probability below 1/N incurs a penalty above ln(N). Do not compare raw losses across different batch sizes without accounting for that changing reference. Zero is the ideal limit as all target probabilities approach one; at a fixed finite temperature and with bounded cosine scores it need not be attainable. A low training loss alone is not evidence of generalization.

05 · From competition to a training loss

This is the whole calculation we just built

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
C=UV⊤,S=C/τ\mathbf C=\mathbf U\mathbf V^\top,\qquad\mathbf S=\mathbf C/\tau
Pij=eSij∑keSik,Qij=eSij∑keSkjP_{ij}=\frac{e^{S_{ij}}}{\sum_k e^{S_{ik}}},\qquad Q_{ij}=\frac{e^{S_{ij}}}{\sum_k e^{S_{kj}}}
L=−12N(∑iln⁡Pii+∑jln⁡Qjj)\mathcal L=-\frac1{2N}\left(\sum_i\ln P_{ii}+\sum_j\ln Q_{jj}\right)

Normalize vectors, compare all pairs, normalize each direction, then penalize the recorded partners.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Normalize vectors, compare all pairs, normalize each direction, then penalize the recorded partners.

Teaching note

Normalize vectors, compare all pairs, normalize each direction, then penalize the recorded partners.

05 · From competition to a training loss

Trace another cell in the interactive

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a photo of a cat
Image I₂Text T₂a photo of a dog
Image I₃Text T₃a photo of a car
TEXT COLUMNS →Text T₁catText T₂dogText T₃carImage I₁cat0.540correct pair0.288I₁ · T₂0.173I₁ · T₃Image I₂dog0.274I₂ · T₁0.548correct pair0.178I₂ · T₃Image I₃car0.186I₃ · T₁0.184I₃ · T₂0.630correct pairDIAGONAL = PAIRED PARTNERSP · row shares

The dog caption competes for the cat image

T₁catT₂dogT₃car
÷ 0.5 → logit1.93861.30870.7993
exp(logit)6.94873.70132.2239
÷ sum → share0.53980.28750.1727

Select a cell, switch row/column, or change temperature. The examples and numbers are the same.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Select a cell, switch row/column, or change temperature. The examples and numbers are the same.

Teaching note

Select a cell, switch row/column, or change temperature. The examples and numbers are the same.

05 · From competition to a training loss

The loss tells us how to change the two branches

REAL CLIP · gradients pass through both branches; the logit scale is also learned

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Loss
Logits
U,V\mathbf U,\mathbf V
Both projections
Both encoders

← forward computation   ·   backpropagation follows the arrows above →

θ←θ−η ∇θL\theta\leftarrow\theta-\eta\,\nabla_\theta\mathcal L

The optimizer updates the shared parameters that produced the vectors.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Embeddings and scores are intermediate outputs, not independent stored parameters in real CLIP. The loss backpropagates through normalization, projections and encoders. The interactive uses directly trainable raw vectors so the geometric effect is visible; the next slide labels that simplification.

Teaching note

Embeddings and scores are intermediate outputs, not independent stored parameters in real CLIP. The loss backpropagates through normalization, projections and encoders. The interactive uses directly trainable raw vectors so the geometric effect is visible; the next slide labels that simplification.

05 · From competition to a training loss

Inspect one gradient update in the toy

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Update the cat’s raw image vector

Before: [2.000000, 0.500000, 0.200000]

Gradient: [-0.017257, 0.050730, 0.045746]

znew=z−0.25∇zL\mathbf z_{\mathrm{new}}=\mathbf z-0.25\nabla_{\mathbf z}\mathcal L

After: [2.004314, 0.487317, 0.188564]

The toy updates all six raw vectors through normalization. Real CLIP instead updates encoder and projection weights.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

These are exact analytical gradients from the source interactive, checked by finite differences. The gradient also includes the contributions from both retrieval directions.

Teaching note

These are exact analytical gradients from the source interactive, checked by finite differences. The gradient also includes the contributions from both retrieval directions.

05 · From competition to a training loss

One step: compare the same batch before and after

COMPUTED TOY UPDATE · six trainable raw vectors; no neural encoder is trained

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Before

xyzu₁v₁v₂u₂v₃u₃

After one step

xyzu₁v₁v₂u₂v₃u₃

● solid = image   □ dashed = text   ·   colour = pair

Same batch · τ = 0.5 · η = 0.250.560526 → 0.539743

The first step makes the paired choices easier. The movement is small; more updates accumulate.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The first step makes the paired choices easier. The movement is small; more updates accumulate.

Teaching note

The first step makes the paired choices easier. The movement is small; more updates accumulate.

05 · From competition to a training loss

Repeat 100 updates on this batch

COMPUTED TOY EXPERIMENT · fixed cat / dog / car batch · η = 0.25

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate

Starting directions

xyzu₁v₁v₂u₂v₃u₃

After 100 updates

xyzv₁u₁v₂u₂v₃u₃

● solid = image   □ dashed = text   ·   colour = pair

Same three pairs · fixed temperature0.560526 → 0.124937

Paired directions come closer and become easier to distinguish from competing pairs.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This fixed-batch run isolates the effect of repeated updates. It does not demonstrate generalization. The source interactive’s default training run changes batches; we show that next.

Teaching note

This fixed-batch run isolates the effect of repeated updates. It does not demonstrate generalization. The source interactive’s default training run changes batches; we show that next.

05 · From competition to a training loss

Now sample the next three records

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Bicyclea bicycle
Airplanean airplane
Treea tree
Batch 2: bicycle · airplane · tree

The row numbers now refer to these records. Build a fresh 3 × 3 matrix and repeat the calculation.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

This is the source interactive’s deterministic first-epoch demonstration order; later epochs shuffle. Real training commonly shuffles the dataset before sampling. In real CLIP shared network parameters persist between batches. In the toy, the 12 pairs have separate raw vectors and only sampled vectors update.

Teaching note

This is the source interactive’s deterministic first-epoch demonstration order; later epochs shuffle. Real training commonly shuffles the dataset before sampling. In real CLIP shared network parameters persist between batches. In the toy, the 12 pairs have separate raw vectors and only sampled vectors update.

05 · From competition to a training loss

New pairs, new scores, the same loss calculation

CHOSEN 3D EXAMPLE · batch 2: bicycle / airplane / tree · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Image I₁Text T₁a bicycle
Image I₂Text T₂an airplane
Image I₃Text T₃a tree
TEXT CANDIDATEST₁T₂T₃I₁0.953correct pair-0.065I₁ · T₂-0.406I₁ · T₃I₂0.006I₂ · T₁0.942correct pair0.928I₂ · T₃I₃-0.384I₃ · T₁0.501I₃ · T₂0.921correct pairDIAGONAL = PAIRED PARTNERSC · cosine

Batch 2: before an update

Image → text0.447912
Text → image0.456861
L=0.447912+0.4568612\mathcal L=\frac{0.447912+0.456861}2
Batch loss0.452386

Image 1 is now bicycle, and text 1 is “a bicycle”. The target is still the diagonal.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Image 1 is now bicycle, and text 1 is “a bicycle”. The target is still the diagonal.

Teaching note

Image 1 is now bicycle, and text 1 is “a bicycle”. The target is still the diagonal.

05 · From competition to a training loss

Compare batch 2 with itself after one step

COMPUTED TOY UPDATE · batch 2 · τ = 0.5 · η = 0.25

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Bicyclea bicycle
Airplanean airplane
Treea tree
Batch 2 · same pairs before and after0.452386 → 0.442814

Do not compare a new batch’s loss with the old batch’s loss and call that progress.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Do not compare a new batch’s loss with the old batch’s loss and call that progress.

Teaching note

Do not compare a new batch’s loss with the old batch’s loss and call that progress.

05 · From competition to a training loss

One pass through the dataset is an epoch

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Batch 1

cat · dog · car

Batch 2

bicycle · airplane · tree

Batch 3

shoe · mug · banana

Batch 4

chair · building · bird

After all 12 records have appeared, shuffle the pairs and begin the next epoch.

Changing the batch changes the competitors. Keep the learned parameters; do not restart the model.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

Changing the batch changes the competitors. Keep the learned parameters; do not restart the model.

Teaching note

Changing the batch changes the competitors. Keep the learned parameters; do not restart the model.

05 · From competition to a training loss

Watch the next batch learn

CHOSEN 3D EXAMPLE · same cat / dog / car batch as the interactive · τ = 0.5

PairsEncodeNormalizeDot productsTemperatureRowsColumnsLossUpdate
Bicyclea bicycle
Airplanean airplane
Treea tree

One training step

  1. Sample complete pairs.
  2. Encode and normalize.
  3. Compare every image and caption.
  4. Average both cross-entropies.
  5. Backpropagate, update, repeat.

Compare before and after on the same batch. A different batch can have a different loss.

Radford et al., 2021 · OpenAI implementation · Image credits · CLIP loss interactive · Exact numbers and source figures

What changed from the previous slide?

The tutorial evaluates its progress curve on all 12 fixed candidates, so it does not confuse a changing batch with a changing evaluation set. A minibatch update need not lower that full-dataset evaluation loss. The toy stores per-example vectors; real CLIP shares encoder weights across every record.

Teaching note

The tutorial evaluates its progress curve on all 12 fixed candidates, so it does not confuse a changing batch with a changing evaluation set. A minibatch update need not lower that full-dataset evaluation loss. The toy stores per-example vectors; real CLIP shares encoder weights across every record.

Section 06

How web-scale training works

Key idea

How does the same recipe scale from 3 pairs to hundreds of millions?

06 · How web-scale training works

The dataset stores pairs, not a giant score matrix

SCHEMATIC · hypothetical teaching dataset: 4 million complete pair records

record 1: image_001 ↔ associated text_001

record 2: image_002 ↔ associated text_002

…

record 4,000,000: image_M ↔ associated text_M

A caption may be a title, description or other associated web text. It need not be a clean class label.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

A caption may be a title, description or other associated web text. It need not be a clean class label.

Teaching note

A caption may be a title, description or other associated web text. It need not be a clean class label.

06 · How web-scale training works

Sample four complete records for this step

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

record 1
a photo of a Newfoundland
record 2
a photo of a Persian cat
record 3
a photo of a pug
record 4
a cup of coffee

Apply image preprocessing and tokenize each paired text. Keep their pair order together.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

This is a B=4 teaching batch. The coffee caption is written for the lesson. CLIP uses image preprocessing and byte-pair tokenization; the pairing itself supplies supervision.

Teaching note

This is a B=4 teaching batch. The coffee caption is written for the lesson. CLIP uses image preprocessing and byte-pair tokenization; the pairing itself supplies supervision.

06 · How web-scale training works

Run the same recipe on the four sampled records

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

Preprocessed images
+ tokenized text
Two encoders
+ projections
+ normalization
S∈R4×4\mathbf S\in\mathbb R^{4\times4}
Two directional losses
→ update

Then sample the next mini-batch. Each step creates new in-batch competition.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Then sample the next mini-batch. Each step creates new in-batch competition.

Teaching note

Then sample the next mini-batch. Each step creates new in-batch competition.

06 · How web-scale training works

Dataset size and batch size are different quantities

SCHEMATIC · CLIP ViT-B/32 unless stated otherwise

400M pair records
Sample B pairs
B×B scoresB\times B\text{ scores}
42=16,10242=1,048,5764^2=16,\qquad 1024^2=1{,}048{,}576

400 million records does not mean a 400M × 400M score matrix.

Radford et al., 2021 · OpenAI implementation · Image credits · Original training description

What changed from the previous slide?

400 million records does not mean a 400M × 400M score matrix.

Teaching note

400 million records does not mean a 400M × 400M score matrix.

06 · How web-scale training works

Original CLIP: this objective at web scale

HISTORICAL · original CLIP, 2021

Radford et al. · ICML 2021

~400 million image–text pairs

32,768 candidate texts per image in the reported contrastive batch

ResNet or ViT image encoders; a text Transformer

The published recipe learns both branches from naturally associated image–text data.

Radford et al., 2021 · OpenAI implementation · Image credits · OpenAI: verified 32,768 batch description

What changed from the previous slide?

The published recipe learns both branches from naturally associated image–text data.

Teaching note

The published recipe learns both branches from naturally associated image–text data.

06 · How web-scale training works

Two captions can describe the same image

TEACHING EXAMPLE · supplied captions for the same toy dog

Image I₁ · the same dog
T1a dogFits this image
T2a dog with floppy earsFits this image
T3a catDoes not fit

A positive is an image–text pair we know should match. This dog has two here.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Read the two dog captions and point to the same image. Both are valid descriptions. The cat caption is an unrelated comparison. These are authored teaching captions, not measured predictions or original CLIP web records. Positive describes a known relationship between an image and a text; it is not a positive-valued cosine.

Teaching note

Read the two dog captions and point to the same image. Both are valid descriptions. The cat caption is an unrelated comparison. These are authored teaching captions, not measured predictions or original CLIP web records. Positive describes a known relationship between an image and a text; it is not a positive-valued cosine.

06 · How web-scale training works

The one-pair target can miss a valid caption

TEACHING EXAMPLE · supplied captions for the same toy dog

Image I₁ · the same dog
For image I₁, compare these three candidate texts
Supplied textFits I₁?Target weight
T1a dogYes1
T2a dog with floppy earsYes0
T3a catNo0

T₂ describes this dog too, but the target [1, 0, 0] makes it a competitor.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Suppose I₁ was recorded with T₁, while T₂ entered the batch with another dog image. The ordinary diagonal target for I₁ is [1,0,0]. Point to the Yes beside T₂ and its target weight of zero. This is a false negative: an unpaired text is treated as a competitor even though it describes the image. Zero is a training target, not a measured probability or a factual claim that the caption is false. Show this problem before discussing the two options on the next slides.

Teaching note

Suppose I₁ was recorded with T₁, while T₂ entered the batch with another dog image. The ordinary diagonal target for I₁ is [1,0,0]. Point to the Yes beside T₂ and its target weight of zero. This is a false negative: an unpaired text is treated as a competitor even though it describes the image. Zero is a training target, not a measured probability or a factual claim that the caption is false. Show this problem before discussing the two options on the next slides.

06 · How web-scale training works

Option 1: choose one valid caption each time

TRAINING OPTION · extends the one-pair-per-image setup we just used

We have saved both dog captions with this image.

This visit to the image

Same image I₁
a dog

One image + one selected caption

A later visit to the image

Same image I₁
a dog with floppy ears

One image + one selected caption

In this option, the other caption for I₁ stays out of that batch.

Keep the same one-pair loss. Across visits, the model can learn from both captions.

Radford et al., 2021 · OpenAI implementation · Image credits

What changed from the previous slide?

Sample each image once in a batch. When it is selected, choose one of its known valid captions; the drawing shows two possible selections on different visits, not a required alternation. Keep the other captions for that same image out of the current batch. Other images contribute their own selected captions as usual, so the original one-to-one batch objective still applies. This avoids competition between the two saved captions of this image in this example; it does not eliminate semantic overlap with captions belonging to other images. This is a possible data-sampling choice, not a claim about the original CLIP pipeline.

Teaching note

Sample each image once in a batch. When it is selected, choose one of its known valid captions; the drawing shows two possible selections on different visits, not a required alternation. Keep the other captions for that same image out of the current batch. Other images contribute their own selected captions as usual, so the original one-to-one batch objective still applies. This avoids competition between the two saved captions of this image in this example; it does not eliminate semantic overlap with captions belonging to other images. This is a possible data-sampling choice, not a claim about the original CLIP pipeline.

06 · How web-scale training works

Option 2: include both and give both credit

ONE MULTIPLE-POSITIVE LOSS · image-to-text direction shown

Image I₁ · the same dog
For image I₁, compare these three candidate texts
Supplied textFits I₁?Target weight
T1a dogYes½
T2a dog with floppy earsYes½
T3a catNo0

Share the target equally: y = [½, ½, 0]. Average the two caption penalties.

ℓI1=−12ln⁡p(T1∣I1)−12ln⁡p(T2∣I1)\ell_{I_1}=-\tfrac12\ln p(T_1\mid I_1)-\tfrac12\ln p(T_2\mid I_1)

The ½ values are target weights we set, not probabilities predicted by the model.

Radford et al., 2021 · OpenAI implementation · Image credits · Multiple positives: Khosla et al., Eq. 2 (adapted here to image–text pairs)

What changed from the previous slide?

Use the same dog and candidate texts as before. We now explicitly label T₁ and T₂ as valid for I₁. One possible multiple-positive loss averages their negative-log penalties, equivalent to cross-entropy against [0.5,0.5,0]. p(T_j|I₁) is the row softmax share over all three candidates, so T₃ still participates in the denominator. This is one choice, not the only multi-positive objective. Its ideal row shares are [0.5,0.5,0], with limiting loss ln(2), not zero; the earlier zero baseline applied to a one-hot target. The slide shows one image row, not an entire symmetric loss. For a full two-direction batch loss, normalize known-match indicators separately within each image row and each text column, then average the directional losses. The encoders and cosine calculation are unchanged. This extends the original one-positive CLIP objective and requires known pair relationships.

Teaching note

Use the same dog and candidate texts as before. We now explicitly label T₁ and T₂ as valid for I₁. One possible multiple-positive loss averages their negative-log penalties, equivalent to cross-entropy against [0.5,0.5,0]. p(T_j|I₁) is the row softmax share over all three candidates, so T₃ still participates in the denominator. This is one choice, not the only multi-positive objective. Its ideal row shares are [0.5,0.5,0], with limiting loss ln(2), not zero; the earlier zero baseline applied to a one-hot target. The slide shows one image row, not an entire symmetric loss. For a full two-direction batch loss, normalize known-match indicators separately within each image row and each text column, then average the directional losses. The encoders and cosine calculation are unchanged. This extends the original one-positive CLIP objective and requires known pair relationships.

Section 07

Return to the opening question

Key idea

We can now explain the hat experiment, one operation at a time.

07 · Return to the opening question

How did we calculate what changed?

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

Before: no hat
After: hat added
hatboatcupcat

Which vectors would you compute, and what would you compare?

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

Should we classify the second image, subtract pixels, or compare the two image representations?

Return to the same two generated images and exact four candidate words from the opening. Ask the class to reconstruct the calculation before advancing. The supplied edit happened before inference; CLIP is used only to encode and compare.

Teaching note

Return to the same two generated images and exact four candidate words from the opening. Ask the class to reconstruct the calculation before advancing. The supplied edit happened before inference; CLIP is used only to encode and compare.

07 · Return to the opening question

Encode both images with the same model

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

Before
→
Same frozen image encoder
→
b[+0.051608, -0.018763, -0.043882, …]First 3 of 512 coordinates

Projection included · divide by length

After
→
Same frozen image encoder
→
a[+0.066042, -0.016296, -0.049582, …]First 3 of 512 coordinates

Projection included · divide by length

a and b are unit image vectors. Their 512 coordinates use the same learned space.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

The saved evidence contains the complete normalized image vectors. encode_image in the OpenAI API includes the learned projection, but unit normalization is done afterward. Here a is the after image and b is the before image. Values shown are rounded; all subsequent calculations use full saved precision.

Teaching note

The saved evidence contains the complete normalized image vectors. encode_image in the OpenAI API includes the learned projection, but unit normalization is done afterward. Here a is the after image and b is the before image. Values shown are rounded; all subsequent calculations use full saved precision.

07 · Return to the opening question

Subtract before from after, coordinate by coordinate

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

δ=a−b\boldsymbol\delta=\mathbf a-\mathbf b
CoordinateAfter: aBefore: bδ = a − b
1+0.066042+0.051608+0.014434
2-0.016296-0.018763+0.002468
3-0.049582-0.043882-0.005700

Continue the same subtraction through coordinate 512.

The subtraction happens between image vectors. The pixels and model weights stay as they are.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

Read the first coordinate aloud: 0.066042 minus 0.051608 gives about 0.014434. Apply the same operation at every coordinate. Negative coordinates are allowed. These are measured embeddings, not the earlier chosen 3D vectors.

Teaching note

Read the first coordinate aloud: 0.066042 minus 0.051608 gives about 0.014434. Apply the same operation at every coordinate. Negative coordinates are allowed. These are measured embeddings, not the earlier chosen 3D vectors.

07 · Return to the opening question

The difference needs its own normalization

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

∥δ∥2=δ12+δ22+⋯+δ5122≈0.483430\|\boldsymbol\delta\|_2=\sqrt{\delta_1^2+\delta_2^2+\cdots+\delta_{512}^2}\approx0.483430
d=δ∥δ∥2\mathbf d=\frac{\boldsymbol\delta}{\|\boldsymbol\delta\|_2}
+0.014434 ÷ 0.483430+0.029858
+0.002468 ÷ 0.483430+0.005104
-0.005700 ÷ 0.483430-0.011791

a and b each have length 1. Their difference usually does not; d does.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

All 512 squared differences contribute to the norm, not just these three preview coordinates. Divide every coordinate by 0.48343039473281413. If the difference is zero, it has no direction; tiny differences can be numerically unstable. The code checks for that before normalizing.

Teaching note

All 512 squared differences contribute to the norm, not just these three preview coordinates. Divide every coordinate by 0.48343039473281413. If the difference is zero, it has no direction; tiny differences can be numerically unstable. The code checks for that before normalizing.

07 · Return to the opening question

Encode the exact words we offered at the start

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

Same frozen text encoder

Tokenize each word → encode + project → divide by length

hat→[+0.013264, +0.001263, -0.016369, …]vhat
boat→[+0.012867, -0.004047, +0.009648, …]vboat
cup→[+0.005676, -0.001876, -0.014071, …]vcup
cat→[+0.020644, -0.024952, -0.021981, …]vcat

Each candidate becomes a unit text vector with 512 coordinates.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

Use the bare strings hat, boat, cup and cat, exactly as in the recorded run. Replacing them with prompts such as a photo of a hat would change the text embeddings and the scores. Projection is inside encode_text; normalize its output. These are supplied candidate words, not words generated by CLIP.

Teaching note

Use the bare strings hat, boat, cup and cat, exactly as in the recorded run. Replacing them with prompts such as a photo of a hat would change the text embeddings and the scores. Projection is inside encode_text; normalize its output. These are supplied candidate words, not words generated by CLIP.

07 · Return to the opening question

Work out the score for “hat”

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

shat=d⊤vhat=∑k=1512dkvhat,ks_{\mathrm{hat}}=\mathbf d^\top\mathbf v_{\mathrm{hat}}=\sum_{k=1}^{512}d_kv_{\mathrm{hat},k}
(+0.029858) × (+0.013264)+0.00039603
(+0.005104) × (+0.001263)+0.00000645
(-0.011791) × (-0.016369)+0.00019300
Sum of the remaining 509 products+0.08016226
shat≈0.080758s_{\mathrm{hat}}\approx0.080758

Both vectors have length 1, so this dot product is a cosine.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

Every product and the full sum are in the linked calculation file. The displayed six-decimal operands are rounded, while the products and final cosine are computed with full saved precision. We use raw cosine here: no temperature scaling and no softmax.

Teaching note

Every product and the full sum are in the linked calculation file. The displayed six-decimal operands are rounded, while the products and final cosine are computed with full saved precision. We use raw cosine here: no temperature scaling and no softmax.

07 · Return to the opening question

Now repeat the same calculation for all four words

SAVED CLIP MEASUREMENTS · original hat pair · ViT-B/32 · 512D · rounded previews · Node CPU q8

Candidate wordBefore · wordAfter · wordChange cosine
hat0.2116200.250661+0.080758
boat0.2025440.214823+0.025401
cup0.2030430.209409+0.013169
cat0.1924270.193662+0.002556
d⊤v=a⊤v−b⊤v∥a−b∥2\mathbf d^\top\mathbf v=\frac{\mathbf a^\top\mathbf v-\mathbf b^\top\mathbf v}{\|\mathbf a-\mathbf b\|_2}

For “hat”: (0.250661 − 0.211620) ÷ 0.483430 ≈ 0.080758

“Hat” wins this candidate list. The score measures alignment with the change direction.

Radford et al., 2021 · OpenAI implementation · Image credits · Every coordinate and product · Try the live CLIP playground

What changed from the previous slide?

This reproduces the opening result from the same saved vectors. All four cosines are positive; only their relative ranking identifies hat within this supplied menu. A modest winning cosine is not a calibrated probability or proof of isolated hat causality. The generated edit can change other details. Reversing before and after negates every score. CLIP did not generate the edit or the candidate words.

Teaching note

This reproduces the opening result from the same saved vectors. All four cosines are positive; only their relative ranking identifies hat within this supplied menu. A modest winning cosine is not a calibrated probability or proof of isolated hat causality. The generated edit can change other details. Reversing before and after negates every score. CLIP did not generate the edit or the candidate words.

Section 08

Linear probing

Key idea

Can we reuse the learned image features for a labelled task?

08 · Linear probing

Linear probing: train only the new classifier

SCHEMATIC · after CLIP pretraining · labelled task examples · no measured probe result

label: dog
label: cat
FROZEN
CLIP image encoder

Projection + normalization

Image vector

u∈R512\mathbf u\in\mathbb R^{512}

One fixed vector
per image

TRAIN THIS PART

Linear classifier

sc=wc⊤u+bcs_c=\mathbf w_c^\top\mathbf u+b_c

One score per class
dog · cat

Known label → cross-entropy → update only W and b“Linear”: each score is a weighted sum of u’s coordinates, plus a bias.

The classifier learns from labels. CLIP’s image encoder and projection stay fixed.

Radford et al., 2021 · OpenAI implementation · Image credits · OpenAI: linear-probe evaluation

What changed from the previous slide?

Collect labelled examples for the task. Compute and optionally cache one fixed vector per image. Here we reuse the normalized image representation u; normalization is a feature choice, not part of the definition of a linear probe. Fit a linear classifier such as logistic regression using these vectors and known class labels. Its score is a weighted sum of the coordinates plus a bias. The image encoder can be very nonlinear; linear refers only to the new classifier. There is no text encoder in this probe, and the loss is ordinary class cross-entropy, not a batch image-text contrastive loss. Fine-tuning would also update some or all pretrained image-model weights.

Teaching note

Collect labelled examples for the task. Compute and optionally cache one fixed vector per image. Here we reuse the normalized image representation u; normalization is a feature choice, not part of the definition of a linear probe. Fit a linear classifier such as logistic regression using these vectors and known class labels. Its score is a weighted sum of the coordinates plus a bias. The image encoder can be very nonlinear; linear refers only to the new classifier. There is no text encoder in this probe, and the loss is ordinary class cross-entropy, not a batch image-text contrastive loss. Fine-tuning would also update some or all pretrained image-model weights.

08 · Linear probing

Two ways to classify with frozen CLIP

SCHEMATIC · same frozen image representation · different sources of class vectors

Zero-shot

a photo of a doga photo of a cat
→
Frozen text encoder
→
vdog,  vcat\mathbf v_{\mathrm{dog}},\;\mathbf v_{\mathrm{cat}}

Class vectors from words

No labelled task examples

Linear probe

label: dog
label: cat
→
Frozen image vectors
+ known labels

Fit a linear classifier

→
wdog,  wcat,  b\mathbf w_{\mathrm{dog}},\;\mathbf w_{\mathrm{cat}},\;\mathbf b

Class weights learned from examples

Test on held-out images. Their labels must not be used to fit the classifier.

A linear probe asks: how useful are the image features for this labelled task?

Radford et al., 2021 · OpenAI implementation · Image credits · OpenAI: linear-probe evaluation

What changed from the previous slide?

Both routes start with a frozen image branch. Zero-shot CLIP uses normalized text embeddings as comparison vectors, with a positive score scale. A linear probe learns weights and biases from labelled task examples; these weights need not be unit vectors. Evaluate on held-out images, not the photographs used to fit the classifier. This is a conceptual detour; no probe accuracy is claimed. Fine-tuning, unlike a linear probe, changes the pretrained representation too.

Teaching note

Both routes start with a frozen image branch. Zero-shot CLIP uses normalized text embeddings as comparison vectors, with a positive score scale. A linear probe learns weights and biases from labelled task examples; these weights need not be unit vectors. Evaluate on held-out images, not the photographs used to fit the classifier. This is a conceptual detour; no probe accuracy is claimed. Fine-tuning, unlike a linear probe, changes the pretrained representation too.

08 · Linear probing

The hat comparison in a few lines of PyTorch

OPENAI CLIP API · same method; saved lecture numbers use the Node q8 checkpoint

before and after are the two supplied PIL RGB images. Load the model once, then compare.

import torch, clipimport torch.nn.functional as F model, preprocess = clip.load("ViT-B/32", device="cpu")model.eval()images = torch.stack([preprocess(before), preprocess(after)])words = ["hat", "boat", "cup", "cat"]with torch.no_grad():    u = F.normalize(model.encode_image(images).float(), dim=-1)    v = F.normalize(model.encode_text(clip.tokenize(words)).float(), dim=-1)    delta = u[1] - u[0]                 # after minus before    if delta.norm() < 1e-6:        raise ValueError("No stable change direction")    d = F.normalize(delta, dim=-1)      # normalize again    scores = v @ d                     # one cosine per wordprint(sorted(zip(words, scores.tolist()), key=lambda x: -x[1]))

Radford et al., 2021 · OpenAI implementation · Image credits · OpenAI CLIP API · Copy the code

Which lines change if we compare one image with descriptions instead of describing an edit?

This is inference, not a training step. model and preprocess are loaded once on CPU; model.eval() fixes inference behavior and torch.no_grad() disables gradient recording. before and after are supplied PIL RGB images. encode_image and encode_text include CLIP’s projection; F.normalize supplies unit length. The second normalization is for the difference. Raw cosines are ranked directly. The API code uses original OpenAI CLIP rather than the saved Node q8 implementation, so exact numerical equality is not claimed. The code is syntax checked; no additional model run is claimed. The training step and later topics remain available in the preserved extended deck. This code returns to the hat calculation just worked through; it is not linear-probe training code. Three architecture summaries close the main lecture next.

Teaching note

This is inference, not a training step. model and preprocess are loaded once on CPU; model.eval() fixes inference behavior and torch.no_grad() disables gradient recording. before and after are supplied PIL RGB images. encode_image and encode_text include CLIP’s projection; F.normalize supplies unit length. The second normalization is for the difference. Raw cosines are ranked directly. The API code uses original OpenAI CLIP rather than the saved Node q8 implementation, so exact numerical equality is not claimed. The code is syntax checked; no additional model run is claimed. The training step and later topics remain available in the preserved extended deck. This code returns to the hat calculation just worked through; it is not linear-probe training code. Three architecture summaries close the main lecture next.

08 · Linear probing

Two encoders build a shared comparison space

SUMMARY · architecture and geometry are schematic

Image and text independently become 512D unit vectors. The image ViT reads CLS; the causal text Transformer reads EOT. Their dot product is the match score.“a photo of a dog”Image ViTFull self-attentionFinal CLS state → LayerNormLearned projection → unit normalizationu · 512 coordinatesText TransformerCausal self-attentionFinal EOT state → LayerNormLearned projection → unit normalizationv · 512 coordinatesu · vcosine
Read the diagram labels
  • “a photo of a dog”
  • Image ViT
  • Full self-attention
  • Final CLS state → LayerNorm
  • Learned projection → unit normalization
  • u · 512 coordinates
  • Text Transformer
  • Causal self-attention
  • Final EOT state → LayerNorm
  • Learned projection → unit normalization
  • v · 512 coordinates
  • u · v
  • cosine

Read one vector from each branch, project, normalize, then compare.

Radford et al., 2021 · OpenAI implementation · Image credits · Beyond Attention: visual summary reference

What changed from the previous slide?

Summary 1 of 3. Follow the blue and purple branches separately. The image ViT uses full self-attention and the text branch uses causal self-attention. The output readouts are final CLS and EOT states after LayerNorm. The shared space is learned through the contrastive objective; equal width alone is insufficient.

Teaching note

Summary 1 of 3. Follow the blue and purple branches separately. The image ViT uses full self-attention and the text branch uses causal self-attention. The output readouts are final CLS and EOT states after LayerNorm. The shared space is learned through the contrastive objective; equal width alone is insufficient.

08 · Linear probing

The loss teaches the two branches what should match

SUMMARY · architecture and geometry are schematic

A batch supplies image rows and text columns. Scaled dot products form a score matrix. Paired cells supply row and column targets; their average loss updates both shared encoder branches, both projections and the score scale.Observed pairsShared encodersUnit vectors“a cat”“a dog”“a car”Image encoder+ projectionText encoder+ projectionUVS = α U Vᵀcolumns: textT1T2T3I1●··I2·●·I3··●Paired cells are the targetsRows+ columns→ lossAverage → backpropagateUpdate both encoders, both projections and the score scale.
Read the diagram labels
  • Observed pairs
  • Shared encoders
  • Unit vectors
  • “a cat”
  • “a dog”
  • “a car”
  • Image encoder
  • + projection
  • Text encoder
  • + projection
  • U
  • V
  • S = α U Vᵀ
  • columns: text
  • T1
  • T2
  • T3
  • I1
  • ●
  • ·
  • ·
  • I2
  • ·
  • ●
  • ·
  • I3
  • ·
  • ·
  • ●
  • Paired cells are the targets
  • Rows
  • + columns
  • → loss
  • Average → backpropagate
  • Update both encoders, both projections and the score scale.

Pairing supplies the targets. Gradients change the shared model parameters.

Radford et al., 2021 · OpenAI implementation · Image credits · Beyond Attention: visual summary reference

What changed from the previous slide?

Summary 2 of 3. The miniature matrix marks paired cells, not measured values. Each image row competes over texts and each text column over images. Cross-entropy is applied in both directions and averaged. Backpropagation goes through normalization, projections and both encoders, and learns the score scale. The whole batch is encoded by shared weights.

Teaching note

Summary 2 of 3. The miniature matrix marks paired cells, not measured values. Each image row competes over texts and each text column over images. Cross-entropy is applied in both directions and averaged. Backpropagation goes through normalization, projections and both encoders, and learns the score scale. The whole batch is encoded by shared weights.

08 · Linear probing

Once trained, the same vectors answer different questions

SUMMARY · architecture and geometry are schematic

Three applications of frozen CLIP: match an image to supplied descriptions, match supplied text to gallery images, or compare a normalized after-minus-before image-vector difference with candidate words.Name an imageImage → unit vector uu · vdogu · vcatRank supplied descriptionsSearch with words“a dog”Text → unit vector vRank gallery images by uᵢ · vDescribe a changebefore → bafter → aUnit image vectorsd = unit(a − b)d · vhatd · vboatRank candidate changesReuse the frozen encoders. Choose which vectors to compare.
Read the diagram labels
  • Name an image
  • Image → unit vector u
  • u · vdogu · vcat
  • Rank supplied descriptions
  • Search with words
  • “a dog”
  • Text → unit vector v
  • Rank gallery images by uᵢ · v
  • Describe a change
  • before → b
  • after → a
  • Unit image vectors
  • d = unit(a − b)
  • d · vhatd · vboat
  • Rank candidate changes
  • Reuse the frozen encoders. Choose which vectors to compare.

Choose the comparison that matches the question you want to ask.

Radford et al., 2021 · OpenAI implementation · Image credits · Beyond Attention: visual summary reference

What changed from the previous slide?

Summary 3 of 3, visually inspired by the three-family comparison in Beyond Attention. Image-to-text matching, text-to-image retrieval and image-difference matching all reuse frozen encoders. The subtraction route operates on normalized image vectors, normalizes the difference again, and compares it with normalized candidate text vectors. These routes rank supplied candidates; they do not generate captions or edited images.

Teaching note

Summary 3 of 3, visually inspired by the three-family comparison in Beyond Attention. Image-to-text matching, text-to-image retrieval and image-difference matching all reuse frozen encoders. The subtraction route operates on normalized image vectors, normalizes the difference again, and compares it with normalized candidate text vectors. These routes rank supplied candidates; they do not generate captions or edited images.

Lecture overview

The complete model

Forward: input → encoder → projection → unit normalization → dot-product scores.

Training: compare observed pairs in both directions, then send gradients through both branches.

Read or present

P: present/read. → or Space: reveal the next part, then move to the next slide. ←: undo a reveal, then move back. The build counter shows your place. Reveals wait for you; nothing advances automatically.

O: overview. S: teaching notes. Q: show or hide the question. A: answer. M: full model map. R: show or hide amber review marks. A focused slider keeps its native arrow keys. N advances even while a slider has focus. Esc closes a dialog first, then exits presentation. Home/End: first/last slide.

Reading and PDF views show completed diagrams. Amber marks identify possible repetition for the teacher to review. They do not remove or skip a slide. The overview links every section and flags the same overlaps. Phone reading unfolds the slides with explanations. Diagram labels are expandable. All media are silent. The linked CLIP Lab runs real inference in your browser and has an executed notebook.