Sources & credits
This lab teaches existing ideas with newly written exercises and code. Source ideas, public photographs, generated pictures and measured results are distinguished below.
Ideas and teaching references
Review method: public captions and targeted timestamped sections, plus the primary paper and code. We do not claim uninterrupted audiovisual viewing. Full transcripts and third-party notebooks are not redistributed.
- Radford et al., ICML 2021
Primary source for the encoders, bidirectional contrastive loss, prompt-based transfer and reported benchmarks. Original training data: 400 million pairs. The lecture’s small experiments are not replications of its benchmarks. - OpenAI CLIP implementation
Original weights and Python implementation used in the notebook; MIT. The browser uses the Xenova ONNX conversion of ViT-B/32, revision d15189d7028b43f1d3e65039190477f6af591c2a, with q8 weights. - CodeEmporium — CLIP Explained!
At 7:25: compare a portrait with and without a hat by subtracting normalized image embeddings and comparing the difference with word embeddings. This is the source of our hat experiment idea. No creator photographs or slide artwork are used. Our generated pairs and chimney, device, colour and style extensions are separately labelled. We show raw cosine; the reference notebook’s optional scale of 10 is a display choice, not proof of a probability. - Computerphile — How AI Understands Images
Around 1:28–4:00: fixed categories versus descriptions and a shared representation. Around 13:15–15:50: downstream zero-shot matching. We adopt the motivation-first teaching pattern, not the video’s artwork. - Yannic Kilcher — OpenAI CLIP paper explanation
Around 4:40–8:00 and 14:40–19:20: image/text pairs, the representation-learning goal and matching within a batch. This informs our gradual move from a retrieval task to a score matrix. - Data Science Gems — OpenAI CLIP
Around 8:38–12:00: zero-shot classification and wording ambiguity. Around 27:38: examples across visual domains. We use this to motivate testing domains and prompts, while checking technical claims against the paper. - Supabase — Image Search in Python with OpenAI CLIP
Around 6:54–11:10: query a gallery using text; a nearest result can still be irrelevant. Our small browser gallery uses direct dot products instead of a database. Search workflow also documented in the linked Supabase tutorial. - Roboflow — CLIP, T-SNE, and UMAP
From 10:40: image embeddings for dataset exploration; from 17:22: similarity and possible duplicates. Our app implements nearest-image inspection. It does not claim that similar embeddings prove duplication, or that a 2D projection preserves all distances.
Image provenance
- Persian_98.jpg — Oxford-IIIT Pet, original image owners; Parkhi et al. CC BY-SA 4.0. Same bytes as Vision I.
- astronaut.jpg — NASA photograph of Eileen Collins. Public domain, via scikit-image data; exported JPEG quality 95.
- chelsea.jpg — Stefan van der Walt. CC0. Exported from scikit-image data, JPEG quality 95.
- chest-xray.png — Stillwaterising, Wikimedia Commons, CC0 1.0. Public chest radiograph; no diagnostic label is inferred by this lab.
- chimney-pair.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- coffee.jpg — Rachel Michetti, courtesy of Pikolo Espresso Bar. CC0. Exported from scikit-image data, JPEG quality 95.
- device-pair.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- hat-pair.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- mug-pair.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- newfoundland_31.jpg — Oxford-IIIT Pet, original image owners; Parkhi, Vedaldi, Zisserman and Jawahar. Distributed via timm/oxford-iiit-pet, revision 089695c834a7deb60505b7cc506672db1c31a6aa. CC BY-SA 4.0. Same bytes as Vision I.
- photo-sketch-pair.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- pug_57.jpg — Oxford-IIIT Pet, original image owners; Parkhi et al. CC BY-SA 4.0. Same bytes as Vision I.
- remote-sensing-scenes.png — Generated for this lesson with the built-in OpenAI image tool on 29 September 2026. Synthetic, not a real person, site survey or patient. See the prompt record.
- rocket.jpg — SpaceX DSCOVR launch. Public domain, via scikit-image data; exported JPEG quality 95.
Paired panels are cropped at their midpoint during model preprocessing and displayed separately. The original generated images are preserved. Generated medical anatomy and device placement are approximate and are not clinical evidence. Synthetic remote-sensing scenes have no real coordinates or measured ground resolution.
What was actually run?
OpenAI CLIP ViT-B/32 in the executed notebook; a pinned q8 ONNX conversion in the live browser and recorded CPU runs. Recorded measurements are explicitly labelled. Browser canvas and Python/PIL or Node image resizing can yield small score differences.
The separate tiny training experiment uses 270 generated geometric images for training and 90 held-out images, with both encoders initialized randomly. Its architecture and dataset differ from CLIP; it demonstrates the loss learning a shared space, not healthcare or satellite accuracy.
Recorded CPU measurements · Browser measurements · Image checksums · Generation prompts