01 / THE DATAA tiny, fully computed example
How CLIP learns.
THE CONTRASTIVE LOSS, STEP BY STEP
12 paired examples / a miniature datasetSame index = a positive pair
The image and caption at the same dataset index are a positive pair.
ALL PAIRWISE COMPARISONS
0.05 · sharper1.0 · softer
FOLLOW ONE NUMBER
IMAGE → TEXT · ROW PROBABILITIES
TEXT → IMAGE · COLUMN PROBABILITIES
Same scores.
Two directions.
Both probability matrices update after every gradient step. The diagonal entries are the probabilities of the correct partners.
Each row on the left sums to 1. Each column on the right sums to 1.
The whole algorithm, in six lines.
- Take N matching image–text pairs.
- Encode both modalities into vectors and normalize each vector.
- Compare every image with every text: N² cosine similarities.
- Scale by temperature, then apply softmax across rows and down columns.
- Take the negative log of each correct match’s probability; average both directions.
- Backpropagate through both encoders. Repeat with the next batch.
Training makes matched image–text representations easier to distinguish from mismatched pairs.