7  What got learned

ImportantAI-generated — not for publication

The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.

WarningPreliminary & unvetted analysis

These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.

The chapters so far measured how much a learner reaches (the ladder) and why it cannot reach further (cues, scaling). This one opens the best organically-bootstrapped model — region-MIL trained on the full corpus (all 1.14M pairs, including the held-out 20%; the definitive model of Chapter 6), no oracle — and asks what it actually learned, and where its failures come from. The answer sharpens the whole story: the bottleneck is the referential learning signal, not the vision encoder.

7.1 The item landscape

Broken down by category, accuracy is a broad, continuous distribution from perfect to below chance (Figure 7.1), not a uniform level.

Figure 7.1: Per-category 4AFC across 176 Konkle categories (test + dev), sorted. Green ≥ 50%, red near or below chance.

What the model learns are the visually distinctive, common, unambiguous concrete nouns — apple, bike, butterfly, turtle, umbrella, microwave, fan, crib — all at or near 100%. What it fails are the rare (domino, lei, trophy), the small or visually diffuse (glove, feather, coin, cigarette, candy), and the polysemous or ambiguous (flag, bullet, microphone). This is the pattern Vong et al. (2024) report for CVCL. One sanity check worth noting: many words the old noisy detector eval scored near chance (it put cat at 14%) sit high here — the clean benchmark confirms the model genuinely learned them, and the old numbers were eval noise.

7.2 What predicts an item’s accuracy

Two candidate explanations: a word is learned because it is heard often (frequency), or because its referent is visually distinctive (the frozen features separate it). We can test both directly.

Figure 7.2: Left: the frozen-feature prototype 4AFC — pick the image closest to a category’s mean DINOv2 embedding, with no learned text — is at or above 99% for every category. Right: model accuracy tracks training frequency, not vision.

The result is stark. A vision-only prototype — classifying each Konkle image by nearest category-centroid in the frozen DINOv2 space, using no language at all — scores median 100%, minimum 99%: the frozen features separate every one of the 176 categories near-perfectly. So visual distinctiveness cannot explain the model’s failures, and indeed it barely does — model accuracy correlates with the vision prototype at only ρ = 0.19. What it does correlate with, twice as strongly, is training frequency (ρ = 0.41): the model learns the words it hears aligned often and misses the rare ones — for most categories the object is fully separable in the features yet the model scores far below what a prototype would.

7.3 What it confuses

When the model errs, does it err sensibly? We take the learned categories (≥50%), build a symmetric confusion matrix — how much each word prefers every other category’s images — and order it by hierarchical clustering over the confusion pattern (Figure 7.3). The residual confusions are anything but random: they fall into clean block-diagonal clusters that respect real category boundaries:

  • a kitchen / tableware block — bowl, fork, knife, spoon, plate, kettle, pitcher, microwave, stove, toaster, sink, fridge;
  • an animals-and-nature block — bird, butterfly, cat, dog, turtle, tree, leaves, mushroom, nest, rock;
  • a tools-and-wheeled-vehicles block — bike, train, tractor, scooter, ladder, hammer, broom, wheelbarrow;
  • a clothing block — hat, jacket, pants, shirt, shoe, suit; and small furniture (bench, crib, dresser, fireplace) and round/flat (clock, watch, frisbee, rug) blocks.

The model has not memorized isolated words; it has built a representation organized by kind, in which its few residual mistakes blur an object into its neighbors — a spoon into the other tableware, a dog into the other animals — never into a random distractor. That such clean semantic geometry emerges from a bag-of-words text tower over frozen features says the frozen space already carries rich category structure, and the learner, where it has enough data, reads it faithfully.

Figure 7.3: Confusion RDM among the learned categories, rows/cols ordered by hierarchical clustering over the confusion pattern. Bright blocks on the diagonal are clusters of mutually-confusable categories — kitchen items, animals, vehicles, clothing — the model’s residual errors respecting category boundaries.

7.4 An external check: the LEVANTE vocabulary benchmark

Everything so far is measured on the Konkle photo set, against proxies for difficulty (frequency, a vision prototype). Does the same story hold on an independent benchmark with real children to compare against? LEVANTE-bench (Alvin Wei Ming Tan et al. 2026) pairs cognitive tasks from the LEVANTE child-assessment project with data from ~1,500 children aged 5–12. Its vocabulary task is a clean 4AFC — a target word, four candidate pictures, pick the referent — the same shape as our Konkle eval, and each item carries an IRT difficulty estimated from the children. We score our definitive model on it exactly as Alvin W. M. Tan et al. (2024) scored contrastive models: embed the word, embed the four images, pick the one whose region grid the word matches best. No prompting, no retraining.

The bank is a wide-range test running from carrot and turtle up to aesthete, gesticulate, and mammalogy. Of 159 test items the model can even attempt 100 — the rest name concepts that simply never occur in a child’s-eye corpus. On those 100 it scores 42.3 ± 0.5% (3 seeds; chance 25), which places this tiny frozen learner in the range of the smallest generative VLMs LEVANTE-bench evaluates — models three to four orders of magnitude larger.

Scored fairly across all 159 items — its 42.3% on the 100 words it knows, plus chance on the 59 out-of-vocabulary ones it cannot attempt — the model reaches ~36% (Figure 7.4): just above chance, level with a 3.1B generative VLM (TinyLLaVA), and far below both children (~72–82% by age) and the frontier models (near ceiling). The point is not that a 0.5M-parameter frozen probe wins, but where it lands on that curve.

Figure 7.4: LEVANTE-bench vocabulary, all 159 items. Our frozen contrastive probe (~36%; real accuracy on the 100 known words, chance on the rest) against generative VLMs prompted for the answer, with children 5–12 as the amber reference band. Model numbers are from local LEVANTE-bench runs.

The item-by-item comparison to children is the revealing part, and it is a weak alignment with a clear cause (Figure 7.5). Across the 76 items with a child difficulty estimate, the model’s per-item confidence is only faintly related to how hard the item is for children (Spearman \(\rho \approx -0.13\), and unstable across seeds). What it tracks instead — more than twice as strongly, and robustly — is the word’s frequency in the model’s own BabyView input (\(\rho = +0.34\); accuracy climbs 25 → 40 → 58% across frequency terciles). The two are not independent: words frequent in one child’s video also tend to be the words children find easy (\(\rho = -0.42\)), so the model does align with children a little — but only through frequency, and its single-corpus counts are a noisy stand-in for the shared developmental ordering.

Figure 7.5: On the LEVANTE vocabulary 4AFC (seed-mean over 3 runs; 76 items with a child difficulty estimate; green = the model chose the correct image). A: the model’s confidence is only weakly related to children’s item difficulty. B: it tracks the word’s frequency in the model’s own BabyView training data much more strongly. The model knows the words it heard often, not the words children find easy — though the two partly overlap.

This is the chapter’s thesis confirmed on outside data, against real children rather than proxies: the model has learned exactly the words its input made available often, and its competence is governed by that idiosyncratic exposure rather than by anything about how children acquire words. It is also consistent with LEVANTE-bench’s own finding that even billion-parameter VLMs align only modestly with children at the item level — our learner sits at the small-model end of that same weak-alignment regime.

7.5 The bottleneck is the signal, not the encoder

This settles a question the conclusion of the mechanistic chapters left open. The frozen vision features are not the limit — they already separate every object we test, so unfreezing the encoder would not help; there is no visual information the model is failing to see. The failures are learning failures: words the model has not heard, aligned, often enough. The item distribution is a frequency-and-alignment distribution wearing a visual mask.

It also closes the loop with Chapter 6. There we found alignment is worth ~90× the data and the baseline never plateaus; here we see why at the level of individual words — each word needs enough aligned exposures to be learned, the frozen features are always ready to support it, and so the growth of the vocabulary across development is gated by how often each referent is named in a moment the learner can see. The child’s-eye view contains the information; what is scarce is the alignment that would let a learner use it.

Tan, Alvin W. M., Sunny Yu, Bria Long, Wanjing Anya Ma, Tonya Murray, Rebecca D. Silverman, Jason D. Yeatman, and Michael C. Frank. 2024. “DevBench: A Multimodal Developmental Benchmark for Language Learning.” In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track.
Tan, Alvin Wei Ming, David Cardinal, Tania Lorido-Botrán, Laura Bravo-Sánchez, Sunny Yu, and Michael C. Frank. 2026. “LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks.” arXiv Preprint arXiv:2606.05497.
Vong, Wai Keen, Wentao Wang, A. Emin Orhan, and Brenden M. Lake. 2024. “Grounded Language Acquisition Through the Eyes and Ears of a Single Child.” Science 383 (6682): 504–11.