A separate, tiny training experiment
What changes when the two encoders learn?
270 paired images. Two randomly initialized encoders. The same 90 held-out images before and after training.
Each point is a recorded training loss. No simulated training animation. Source: train_tiny.py · complete run data.
What exactly trained?
Both a small CNN image encoder and a learned two-token text encoder, plus a logit scale. Neither uses pretrained CLIP features. Each description is a colour token and a shape token. We normalize their 24-dimensional outputs and optimize the bidirectional contrastive loss.
Each batch contains one image per concept, so two identical descriptions are not false negatives. Training and test have different seeds and no identical image bytes. Test images contain new positions, sizes, shades and backgrounds within the same nine concepts. This does not test unseen-concept zero-shot learning.
The fixed 500-step run uses seed 7. Held-out accuracy is computed before and after training, with no early stopping based on test accuracy. The notebook reruns this training from scratch.
Objective: Radford et al., 2021. Dataset and small-model implementation newly authored for this teaching lab.