8 How much does the vision encoder matter?
The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.
These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.
Everything so far has run on a single frozen off-the-shelf vision encoder — DINOv2-base. Chapter 7 argued that this encoder is not the bottleneck: its frozen features already separate every Konkle category near-perfectly, so the failures are learning failures, not seeing failures. That is an argument from one encoder. This chapter tests it directly by swapping the encoder — including encoders trained on BabyView itself — and asking whether an in-domain or simply stronger encoder moves the ceiling.
This is also the natural home for a question our lab’s earlier work raised. A BabyView-trained DINOv2 (from the CVCL-replication line, Lin et al. (2026)) gave notably low scaling; the intuition that an encoder tuned to the child’s own visual world should help word learning has, so far, not paid off. Here we quantify it.
8.1 The comparison
Five encoders embed the exact same 877,802 training frames — the topline set of Chapter 4 — so the encoder is the only thing that varies. Two are off-the-shelf (DINOv2-base, our workhorse, and DINOv3-ViT-B/16); four are trained on BabyView by our collaborators: DINOv3-ViT-L, a V-JEPA2-ViT-L video model, and two sizes of ZWM (170M and 1B), the reconstruction-trained world model of the LRAS line (Lee et al. 2025). On each encoder we train the same contrastive head (bag-of-words text, InfoNCE) and evaluate the same out-of-sample Konkle 4AFC, exactly as in DevBench’s contrastive protocol (Tan et al. 2024). For a like-for-like image summary across models that lack a CLS token, all encoders are read out as mean-pooled patch tokens (whole-frame) or a 4×4 region grid (region-MIL).
8.2 Off-the-shelf wins; in-domain training transfers worse
The result is unambiguous (Figure 8.1). The newer off-the-shelf encoder, DINOv3-B, is the best of all (whole-frame 70.8, region-MIL 72.6) — well above our DINOv2 workhorse (61.3 / 65.3), which is itself far above every BabyView-trained encoder: DINOv3-L 41.0, ZWM-1B 35.6, ZWM-170M 31.6, V-JEPA2-L 30.8 — the last three barely above chance (25).
Two controls make this clean. First, capacity is not the story: ZWM-1B (947M) is barely better than ZWM-170M (170M), and both trail the ~300M ViT-Ls, which trail off-the-shelf DINOv2-base (~86M). Second, the low BabyView scores are not an artifact of the anisotropic, un-centered features those reconstruction models produce (a real hazard flagged by whoever cached them): re-running every encoder with the dataset mean subtracted before the projection moves nothing (all within 0.5 points). The deficit is real transfer failure — encoders tuned on blurry, cluttered, egocentric child video represent clean centered object photos worse than encoders trained on web images. That lines up exactly with the CVCL-replication line’s low-scaling result: in-domain visual pretraining, on this benchmark, hurts.
The single cleanest comparison is DINOv3-L-BabyView (41.0) vs. DINOv3-B-off-the-shelf (70.8) — same architecture family, differing only in training data — and in-domain loses by 30 points.
8.3 A correction to the region-MIL rung
This comparison is what fixed the region-MIL rung of Chapter 4. An earlier draft of that ladder reported region-MIL as a +9.7 free gain over whole-frame — but that baseline was the DINOv2 CLS token (52.9), a weak object-recognition readout for DINOv2. Measured against a mean-pooled whole-frame baseline — the fair “no spatial selection” summary, and what every encoder here uses — DINOv2 whole-frame is already 61.3, so region-MIL’s honest gain is only +4.0 (the value the ladder now carries). For the stronger DINOv3-B it is +1.8; for the BabyView encoders, ≈0 (Figure 8.1). Most of the apparent “region-MIL magic” was CLS being a poor summary, not spatial selection being powerful — and the genuine region gain shrinks as the base encoder gets stronger, because a good encoder’s mean-pooled features already localize what a word needs.1
8.4 The extended ladder: the referential gap is robust to the encoder
If the encoder is not the bottleneck, then climbing the oracle rungs — filtering to aligned pairs, supplying the referent label (Section 4.4) — should still buy a large, encoder-independent gain. We built the full ladder (whole-frame → region-MIL → alignment-filter → referent-label) for all three DINOs (Figure 8.2).
| Rung | DINOv3-L (BV) | DINOv2 (OTS) | DINOv3-B (OTS) |
|---|---|---|---|
| whole-frame (meanpatch) | 41.0 | 61.3 | 70.8 |
| + region-MIL | 41.7 | 65.3 | 72.6 |
| + alignment filter (oracle) | 45.7 | 70.3 | 75.4 |
| + referent label (oracle ceiling) | 51.1 | 81.5 | 84.6 |
Three things fall out, and they refine the “encoder doesn’t matter” slogan into something more exact.
The ceiling moves. A better encoder raises the whole ladder, oracle rungs included — the clean-label ceiling climbs from 81.5 (DINOv2) to 84.6 (DINOv3-B). The “~81 ceiling” of Chapter 4 is DINOv2-specific.
The referential gap shrinks but persists. The headroom from the free base to the label ceiling is +20 for DINOv2, +14 for DINOv3-B, +10 for DINOv3-L. A stronger encoder gets closer to its ceiling on free features alone — but even the best off-the-shelf encoder we tested leaves ~+14 of label-topline headroom that only the referent oracle reaches. The referential-alignment bottleneck is not an artifact of a weak encoder; it survives a large encoder upgrade.
In-domain pretraining loses even with the oracle. BabyView DINOv3-L’s clean-label ceiling (51.1) sits below off-the-shelf DINOv2’s plain free region-MIL (65.3). Handed the correct referent word for every pair, the child’s-eye-trained encoder still cannot match a weaker web-trained encoder that never sees the label. There is no readout of these BabyView features from which a linear learner recovers clean-object identity as well as off-the-shelf features do.
8.5 What the encoder does and doesn’t change
Chapter 7 claimed the encoder is not the bottleneck; this chapter makes that precise. The encoder changes the level (base 31–71) and even the ceiling (51–85), and a stronger encoder narrows the referential gap. But the gap is robust: it persists across a 30-point spread of encoders, and in-domain pretraining — the one intervention that most directly targets “see the child’s world better” — makes everything worse. What a stronger encoder buys is a higher floor under the same ceiling-shaped problem; what it does not buy is the referential alignment that the oracle rungs supply. The bottleneck is where Chapter 7 put it: in the learning signal, not the eyes.
The one non-monotone cell — ZWM-170M region-MIL (28.9) below its whole-frame (31.6) — is a readout mismatch, not a real drop: its cached grid is a mid-block layer while its whole-frame vector is a later layer. It is the only row not measured at a matched layer.↩︎