FROM CLIP TO VISION-LANGUAGE MODELS
VLM Lab
Keep the question fixed. Change the evidence. Inspect the answer.
SmolVLM-256M-Instruct · WebGPU · Web Worker · no paid API. First live load downloads about 575 MB and uses the browser cache. Recorded outputs need no network.
What animal is shown?
Exactly the same prompt and decoding settings; only the image changes. Blank means a white image through the same processor.
Compare p(yt | y<t, I, P). Changes in answers show image sensitivity; they do not establish reliable grounding.
Raw model output
No run selectedNo model output yet.
Reveal reference answer
Show what happened · computation steps
These are observable runtime events and documented system stages, not hidden reasoning or attention maps. Image encoding, context assembly and prefill occur inside the first model forward pass; the library does not time them separately.
Technical metadata
Model not loaded.