CLIP LABBefore / after training

Sources & credits

This lab teaches existing ideas with newly written exercises and code. Source ideas, public photographs, generated pictures and measured results are distinguished below.

Ideas and teaching references

Review method: public captions and targeted timestamped sections, plus the primary paper and code. We do not claim uninterrupted audiovisual viewing. Full transcripts and third-party notebooks are not redistributed.

Image provenance

Paired panels are cropped at their midpoint during model preprocessing and displayed separately. The original generated images are preserved. Generated medical anatomy and device placement are approximate and are not clinical evidence. Synthetic remote-sensing scenes have no real coordinates or measured ground resolution.

What was actually run?

OpenAI CLIP ViT-B/32 in the executed notebook; a pinned q8 ONNX conversion in the live browser and recorded CPU runs. Recorded measurements are explicitly labelled. Browser canvas and Python/PIL or Node image resizing can yield small score differences.

The separate tiny training experiment uses 270 generated geometric images for training and 90 held-out images, with both encoders initialized randomly. Its architecture and dataset differ from CLIP; it demonstrates the loss learning a shared space, not healthcare or satellite accuracy.

Recorded CPU measurements · Browser measurements · Image checksums · Generation prompts