01 / THE DATAEvery number comes from the vectors
How CLIP learns.
THE CONTRASTIVE LOSS, STEP BY STEP
Start here
12 matched pairs
Which pairs are in this batch?
Selected for this calculation
12 paired examples / a miniature datasetSame index = a positive pair
The image and caption at the same dataset index are a positive pair.
ALL PAIRWISE COMPARISONS
0.05 · sharper1.0 · softer
FOLLOW ONE NUMBER
IMAGE → TEXT · ROW PROBABILITIES
TEXT → IMAGE · COLUMN PROBABILITIES
Compare both retrieval directions.
Both probability matrices update after every gradient step. The diagonal entries are the probabilities of the correct partners.
Each row on the left sums to 1. Each column on the right sums to 1.
Put the steps together.
- Take N matching image and text pairs.
- Encode both modalities into vectors and normalize each vector.
- Compare every image with every text: N² cosine similarities.
- Scale by temperature, then apply softmax across rows and down columns.
- Take the negative log of each correct match’s probability; average both directions.
- Backpropagate through both encoders. Repeat with the next batch.
Training makes matched image and text representations easier to distinguish from mismatched pairs.