← Lecture

FROM CLIP TO VISION-LANGUAGE MODELS

VLM Lab

Keep the question fixed. Change the evidence. Inspect the answer.

Choose live WebGPU inference or the saved measured run.

SmolVLM-256M-Instruct · WebGPU · Web Worker · no paid API. First live load downloads about 575 MB and uses the browser cache. Recorded outputs need no network.

Selected visual evidence

Raw model output

No run selected
No model output yet.

Reveal reference answer

Show what happened · computation steps

These are observable runtime events and documented system stages, not hidden reasoning or attention maps. Image encoding, context assembly and prefill occur inside the first model forward pass; the library does not time them separately.

    Technical metadata
    Model not loaded.