Learning Words from a Child’s-Eye View
Bootstrapping word–object learning from a child’s-eye view
Overview
The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.
These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.
Children learn the meanings of words from remarkably little data — a few hundred million words heard by age five, embedded in rich multimodal experience. Large models need three to five orders of magnitude more (Frank 2023). One appealing route to closing that gap is to learn word–object mappings directly from the visual and linguistic stream a child actually experiences. Vong et al. (2024) showed this is possible in principle: a minimal CLIP-style model (CVCL) trained on 61 hours of headcam video and transcripts from a single child (SAYCam-S) learned to map words to their visual referents well above chance.
Our lab’s attempts to reproduce that success on other corpora — in particular our own, much larger BabyView corpus (Long et al. 2025) — failed. Trained naively on BabyView, contrastive vision–language models sit near chance, and in the end-to-end setting the loss can even climb during training (Lin et al. 2026). The consensus diagnosis across several groups (and our own symbolic modeling) is that the problem is referential alignment: in naturalistic egocentric data, the things being talked about are usually not the things on screen. Effective learning moments are rare and buried in noise.
This book documents a small, fast prototype built to iterate on that problem. It uses a frozen-vision probe — cheap enough to run dozens of controlled experiments in an afternoon — to ask a concrete question: if the problem is referential noise, does concentrating the aligned moments rescue learning?
| Chapter | What it covers |
|---|---|
| Background | The data gap, CVCL, why BabyView breaks contrastive learning, and the alignment diagnosis. |
| The data | BabyView 2025.2 — scale (36 children, ~1,390 h, 1.1M utterance–frame pairs), what the speech looks like, and the CLIP vs. Gemini annotations of alignment. |
| The pipeline | The modeling infrastructure — the frozen two-tower model, the three bootstrapping mechanisms (region / frame / utterance), and the out-of-sample Konkle evaluation. |
| Bootstrapping experiments | The systematic ladder: what region-, frame-, and utterance-bootstraps buy against oracle toplines, on all data, with error bars. |
| Cues | Re-aiming the cues children actually use — prosody, discourse, gaze, pose — at referent identification; every one is null, for a concrete reason. |
| Scaling and data distribution | How each rung scales with amount of data, and within- vs. across-child: alignment is ~90× the data, diversity buys more. |
| What got learned | The best organic model, item by item: distinctive common nouns; the bottleneck is the learning signal, not the vision encoder. |
| How much does the vision encoder matter? | Swapping the encoder — off-the-shelf DINOv3 dominates, BabyView-trained encoders transfer worse; the referential gap shrinks with a stronger encoder but never closes. |
| Architectures linking language and vision | Soft attention and a from-scratch generative captioner vs. contrastive hard-max: nothing beats the plain two-tower. The architecture isn’t the bottleneck either. |
| Conclusion | What referential alignment sets a ceiling on, and the three ways the boundary might move. |
The question in one figure
The core intuition is that egocentric speech and vision are only intermittently about the same thing. A cheap off-the-shelf signal — the CLIP cosine similarity between a frame and the utterance spoken over it — separates the moments where they align from the moments where they don’t:
The prototype tests whether training on the top row — and its analogues — is enough to learn word–object mappings where training on everything is not.
The model in one paragraph
The object of learning is a joint embedding space in which an image of a thing and the word for that thing land near each other. A frozen DINOv2 vision encoder maps a frame to a 768-d vector; a small bag-of-words text encoder maps an utterance to a vector in the same space; a symmetric contrastive (InfoNCE) objective pulls co-occurring (frame, utterance) pairs together and pushes mismatched pairs apart. Because the vision encoder is frozen, “training” is just a linear projection plus word embeddings — so each run is minutes, not hours, and the experiments isolate the effect of which pairs the model sees rather than of optimization or capacity. The next chapter walks through every piece; the code lives in src/.