8  How much does the vision encoder matter?

ImportantAI-generated — not for publication

The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.

WarningPreliminary & unvetted analysis

These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.

Everything so far has run on a single frozen off-the-shelf vision encoder — DINOv2-base. Chapter 7 argued that this encoder is not the bottleneck: its frozen features already separate every Konkle category near-perfectly, so the failures are learning failures, not seeing failures. That is an argument from one encoder. This chapter tests it directly by swapping the encoder — including encoders trained on BabyView itself — and asking whether an in-domain or simply stronger encoder moves the ceiling.

This is also the natural home for a question our lab’s earlier work raised. A BabyView-trained DINOv2 (from the CVCL-replication line, Lin et al. (2026)) gave notably low scaling; the intuition that an encoder tuned to the child’s own visual world should help word learning has, so far, not paid off. Here we quantify it.

8.1 The comparison

Five encoders embed the exact same 877,802 training frames — the topline set of Chapter 4 — so the encoder is the only thing that varies. Two are off-the-shelf (DINOv2-base, our workhorse, and DINOv3-ViT-B/16); four are trained on BabyView by our collaborators: DINOv3-ViT-L, a V-JEPA2-ViT-L video model, and two sizes of ZWM (170M and 1B), the reconstruction-trained world model of the LRAS line (Lee et al. 2025). On each encoder we train the same contrastive head (bag-of-words text, InfoNCE) and evaluate the same out-of-sample Konkle 4AFC, exactly as in DevBench’s contrastive protocol (Tan et al. 2024). For a like-for-like image summary across models that lack a CLS token, all encoders are read out as mean-pooled patch tokens (whole-frame) or a 4×4 region grid (region-MIL).

8.2 Off-the-shelf wins; in-domain training transfers worse

The result is unambiguous (Figure 8.1). The newer off-the-shelf encoder, DINOv3-B, is the best of all (whole-frame 70.8, region-MIL 72.6) — well above our DINOv2 workhorse (61.3 / 65.3), which is itself far above every BabyView-trained encoder: DINOv3-L 41.0, ZWM-1B 35.6, ZWM-170M 31.6, V-JEPA2-L 30.8 — the last three barely above chance (25).

Figure 8.1: Whole-frame vs. region-MIL Konkle 4AFC for six vision encoders (3 seeds), sorted by whole-frame. The off-the-shelf encoders dominate; every BabyView-trained encoder transfers far worse to clean object photos, and a bigger BabyView model (ZWM-1B) helps only marginally over a smaller one.

Two controls make this clean. First, capacity is not the story: ZWM-1B (947M) is barely better than ZWM-170M (170M), and both trail the ~300M ViT-Ls, which trail off-the-shelf DINOv2-base (~86M). Second, the low BabyView scores are not an artifact of the anisotropic, un-centered features those reconstruction models produce (a real hazard flagged by whoever cached them): re-running every encoder with the dataset mean subtracted before the projection moves nothing (all within 0.5 points). The deficit is real transfer failure — encoders tuned on blurry, cluttered, egocentric child video represent clean centered object photos worse than encoders trained on web images. That lines up exactly with the CVCL-replication line’s low-scaling result: in-domain visual pretraining, on this benchmark, hurts.

The single cleanest comparison is DINOv3-L-BabyView (41.0) vs. DINOv3-B-off-the-shelf (70.8) — same architecture family, differing only in training data — and in-domain loses by 30 points.

8.3 A correction to the region-MIL rung

This comparison is what fixed the region-MIL rung of Chapter 4. An earlier draft of that ladder reported region-MIL as a +9.7 free gain over whole-frame — but that baseline was the DINOv2 CLS token (52.9), a weak object-recognition readout for DINOv2. Measured against a mean-pooled whole-frame baseline — the fair “no spatial selection” summary, and what every encoder here uses — DINOv2 whole-frame is already 61.3, so region-MIL’s honest gain is only +4.0 (the value the ladder now carries). For the stronger DINOv3-B it is +1.8; for the BabyView encoders, ≈0 (Figure 8.1). Most of the apparent “region-MIL magic” was CLS being a poor summary, not spatial selection being powerful — and the genuine region gain shrinks as the base encoder gets stronger, because a good encoder’s mean-pooled features already localize what a word needs.1

8.4 The extended ladder: the referential gap is robust to the encoder

If the encoder is not the bottleneck, then climbing the oracle rungs — filtering to aligned pairs, supplying the referent label (Section 4.4) — should still buy a large, encoder-independent gain. We built the full ladder (whole-frame → region-MIL → alignment-filter → referent-label) for all three DINOs (Figure 8.2).

Figure 8.2: The bootstrapping ladder across three vision encoders, worst to best. Each line’s total rise (base → referent-label ceiling) is the referential headroom: +20 for DINOv2, +14 for off-the-shelf DINOv3-B, +10 for BabyView DINOv3-L. It shrinks with a stronger encoder but never closes — even the best encoder leaves ~+14 that only the referent-label oracle reaches.
Rung DINOv3-L (BV) DINOv2 (OTS) DINOv3-B (OTS)
whole-frame (meanpatch) 41.0 61.3 70.8
+ region-MIL 41.7 65.3 72.6
+ alignment filter (oracle) 45.7 70.3 75.4
+ referent label (oracle ceiling) 51.1 81.5 84.6

Three things fall out, and they refine the “encoder doesn’t matter” slogan into something more exact.

The ceiling moves. A better encoder raises the whole ladder, oracle rungs included — the clean-label ceiling climbs from 81.5 (DINOv2) to 84.6 (DINOv3-B). The “~81 ceiling” of Chapter 4 is DINOv2-specific.

The referential gap shrinks but persists. The headroom from the free base to the label ceiling is +20 for DINOv2, +14 for DINOv3-B, +10 for DINOv3-L. A stronger encoder gets closer to its ceiling on free features alone — but even the best off-the-shelf encoder we tested leaves ~+14 of label-topline headroom that only the referent oracle reaches. The referential-alignment bottleneck is not an artifact of a weak encoder; it survives a large encoder upgrade.

In-domain pretraining loses even with the oracle. BabyView DINOv3-L’s clean-label ceiling (51.1) sits below off-the-shelf DINOv2’s plain free region-MIL (65.3). Handed the correct referent word for every pair, the child’s-eye-trained encoder still cannot match a weaker web-trained encoder that never sees the label. There is no readout of these BabyView features from which a linear learner recovers clean-object identity as well as off-the-shelf features do.

8.5 What the encoder does and doesn’t change

Chapter 7 claimed the encoder is not the bottleneck; this chapter makes that precise. The encoder changes the level (base 31–71) and even the ceiling (51–85), and a stronger encoder narrows the referential gap. But the gap is robust: it persists across a 30-point spread of encoders, and in-domain pretraining — the one intervention that most directly targets “see the child’s world better” — makes everything worse. What a stronger encoder buys is a higher floor under the same ceiling-shaped problem; what it does not buy is the referential alignment that the oracle rungs supply. The bottleneck is where Chapter 7 put it: in the learning signal, not the eyes.

Lee, Wanhee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Honglin Chen, Khai Loong Aw, and Daniel L. K. Yamins. 2025. “3D Scene Understanding Through Local Random Access Sequence Modeling.” arXiv Preprint arXiv:2504.03875.
Lin, Rust, Villar Corrales, Michael C. Frank, Emmanuel Dupoux, et al. 2026. “EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data.” arXiv Preprint arXiv:2605.19130.
Tan, Alvin W. M., Sunny Yu, Bria Long, Wanjing Anya Ma, Tonya Murray, Rebecca D. Silverman, Jason D. Yeatman, and Michael C. Frank. 2024. “DevBench: A Multimodal Developmental Benchmark for Language Learning.” In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track.

  1. The one non-monotone cell — ZWM-170M region-MIL (28.9) below its whole-frame (31.6) — is a readout mismatch, not a real drop: its cached grid is a mid-block layer while its whole-frame vector is a later layer. It is the only row not measured at a matched layer.↩︎