4,387 children's line drawings of 12 object categories, produced by kids
aged 4–9 at four international sites — ■ San Jose (USA),
■ Beijing (China), ■ New Delhi (India),
■ Kisumu (Kenya). Each drawing sits in a 2-D map of its
CLIP (ViT-B/32) embedding. Hover a point for the drawing; click to pin and replay the strokes.
Tablet-site drawings re-rendered crisply from raw strokes; Kisumu is scanned on paper.
Mean cosine distance between sites' category-means — how
differently the four contexts depict each object.
Recognizability rises with age
Mean CLIP target similarity by age, per site (recomputed from
the visible set). This is the model-based measure — cosine similarity to the
target category's CLIP text embedding — not a behavioral 12-AFC score.
Category structure by site (RDMs)
12×12 representational dissimilarity — cosine
distance between category-mean CLIP embeddings. Hover a cell to compare that pair
across all four sites.
similardistinct
Darker (viridis) = smaller distance; brighter = more distinct.