4  Bootstrapping experiments

ImportantAI-generated — not for publication

The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.

WarningPreliminary & unvetted analysis

These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.

The question of this book is whether word–object mapping can be bootstrapped from naturalistic data — whether a learner can climb from naive contrastive training toward the ceiling using only signal it could plausibly have. This chapter maps that climb with a single, systematic design: train on all of BabyView (36 children, 911k pairs on an 80% video split), evaluate on the out-of-sample Konkle benchmark (Section 3.7), and measure what each mechanism contributes against oracle toplines. Everything is 3 seeds; the primary number is Konkle test-60.

4.1 Three bootstraps

Every mechanism here is a form of self-supervised selection — deciding what to attend to without being told the answer — at a different grain (Section 3.8):

  • Region (MIL max over frame cells): where in the frame is the referent.
  • Frame (MIL max over a ±2 s window): which moment shows it best.
  • Utterance (EM reweighting of pairs by self-estimated alignment): which pairs are aligned.

The oracles that bound them come from the Gemini annotation of Chapter 2 — a per-pair alignment score and, where a visible object is named, its referent word. Three successively stronger interventions build on those two labels, and it helps to state up front what each one physically does to the training data (the mechanics and exact numbers are in Section 4.4). All three operate on the ~9% of pairs Gemini marks as referentially aligned (~84k of 911k):

  1. Alignment filter — throw away the other 91% and train only on the aligned pairs, keeping their natural utterances. This isolates which pairs.
  2. Word selection — on those same aligned pairs, replace the whole utterance with just the referent noun, for the ~63% of them where that noun is actually spoken in the utterance. This isolates which word names the object.
  3. Vision-binding — additionally supply the referent noun for the remaining ~37%, where it is never said aloud and so could only be recovered by seeing the object. This isolates the referent when it is unspoken.

Each is an oracle because it uses a label — alignment, or the referent word — that a learner has no access to at training time.

4.2 The ladder

Figure 4.1: The bootstrapping ladder. Green rungs are free — achievable with no labels — and chain cumulatively from a mean-pooled whole-frame baseline; blue rungs are oracles — upper bounds that require alignment or referent labels we do not have at training time. The oracle rungs are an alternative branch out of the region-MIL baseline (65.3), not stacked on frame-MIL: the guide line marks where they leave the free chain, so the alignment filter’s +4.6 is measured from region-MIL. Perfect labels cannot exceed ~81 on these frozen features.

The Δ column below reads the same way: free rungs are measured from the rung above, the oracle branch from region-MIL. The from column makes the base explicit.

Rung Konkle Δ from kind
chance 25 — — —
pure learning (mean-pooled whole-frame) 61.3 +36.3 chance free
+ region MIL 65.3 +4.0 pure learning free
+ frame MIL (±2 s) 65.5 +0.2 region MIL free
+ alignment filter 69.9 +4.6 region MIL oracle
+ word selection 74.0 +4.1 filter oracle
+ vision-binding 81.5 +7.5 word selection oracle
ceiling ~81 — — unlearnable from these features

Two facts frame everything. The free gains are modest. Naive contrastive learning on a mean-pooled whole-frame already reaches 61.3; region-MIL — scoring the utterance against the best-matching cell — adds only +4.0, and frame-MIL — selecting the best moment in a window — adds essentially nothing (+0.2, within seed noise). Self-supervised spatial and temporal selection buys a little, then stops.1

Almost all the reachable climb is oracle. From region-MIL’s 65.3 to the clean-label ceiling of 81.5 is +16.2, and every step of it needs information a learner cannot access at training time — which pairs are aligned (+4.6), which spoken word is the referent (+4.1), and which object is named when no word is spoken (+7.5). No self-supervised mechanism we tried climbs those oracle rungs, and none comes close: the free route tops out at ~65, a clear +4.6 below even the first oracle rung. The referential moment is not something free selection recovers — the bottleneck is the referential signal itself.

4.3 What the bootstraps buy

Figure 4.2: The three self-supervised bootstraps. Only region selection is a real (if small) free gain; frame selection and the utterance EM bootstrap are both null within seed noise.

Region MIL (+4.0) is the largest free gain — and a modest one. Scoring the utterance against the best-matching cell of a frame, with no detector and no labels, does implicit referent localization, and against the mean-pooled whole-frame baseline it is worth +4.0. It is real (positive in all three seeds), but it is not the +9.7 an earlier draft reported, and Chapter 8 shows it shrinks toward zero for stronger encoders whose mean-pooled features already localize what a word needs.

Frame MIL (+0.2) is null. Selecting the best moment in a ±2 s window, on top of region selection, adds nothing in the clean rig (65.3 → 65.5, within seed noise). “Which moment” does not help once “which region” is handled: when there is a referential frame at all, it sits close enough to the utterance midpoint that a wider temporal reach adds no signal the model can use. (We also tried widening the window to ±5 s; the story does not change.)

Utterance EM is null. The self-supervised EM loop — reweighting pairs by the model’s own max-region alignment, the direct attempt to climb the filter rung without the oracle — does essentially nothing (both a whole-frame and a region variant land inside the seed noise, sd ≈ 2). This is a firm version of “the bootstrap does not ignite”: pair-reweighting by the model’s own alignment estimate is indistinguishable from zero, and region MIL has already absorbed most of what it targets.

So of the three self-supervised bootstraps only region selection buys anything at all, and only a little. Which-region helps a bit; which-moment and which-pairs, on their own, do not. None comes near the oracle rungs.

4.4 The label decomposition

Here are the mechanics behind the three oracle interventions previewed at the start of the chapter. They are built to be nested, so their gains add up cleanly. We start from the 83,794 referentially-aligned pairs — the ~9% of the 910,975 training pairs to which Gemini assigned a visible referent — and hold that pair set fixed. The only thing that changes across the three rungs is the text each pair is trained against:

  • filtered-natural — the aligned pairs with their natural utterance unchanged (“look, is that a doggie?”). This is the alignment filter of the main ladder, restated on the referent-bearing set.
  • referent word where spoken — for each pair, if the referent noun actually appears in the utterance, we throw the rest of the sentence away and train on the bare noun (“dog”); otherwise we leave the utterance untouched. The referent turns out to be spoken in 63% of these aligned pairs, so this rewrites 63% of them and leaves 37% as natural text.
  • referent word always — we supply the bare referent noun for all the pairs, including the 37% where it was never said.

The three differ only in text, on identical frames and identical pairs, so each Δ isolates one thing (Figure 4.3).

Figure 4.3: Where the label headroom lives: filtering to referential pairs, selecting the spoken referent word, and — the largest single piece — supplying the referent when it is not spoken at all.
step Konkle Δ isolates
region-MIL baseline 65.3 — —
filtered-natural 69.9 +4.6 which pairs
referent word where spoken 74.0 +4.1 which word
referent word always 81.5 +7.5 referent when unspoken (vision)

Why is the referent spoken in as many as 63% of aligned pairs, when only ~9% of all pairs are aligned in the first place? Because the 63% is conditional on alignment: it is 63% of the ~9%, not 63% overall. And it is high for exactly the reason those pairs were kept — Gemini calls a pair aligned when the caregiver is talking about a visible object, and most of the time naming that object is why they are talking about it, so its noun is right there in the transcript. The other 37% are still genuinely aligned — the caregiver is attending to the visible object — but they refer to it without the noun: deixis and pronouns (“look at that!”, “pick it up”), or a description (“is it yummy?”). Those pairs carry a real referential moment but no referential word.

That split is what makes the last two rungs mean different things. Supplying the noun for the spoken 63% (word selection, +4.1) mostly buys disambiguation — pulling the referent out of a sentence that also contains other content words. Supplying it for the unspoken 37% (vision-binding, +7.5) is the single largest piece, because for those pairs the noun is new information: it cannot be read off the language at all, only off the image. This is why Chapter 5 aims accessible cues at identifying the referent — including visually, when it isn’t said — rather than merely filtering frames.

4.5 Synthesis

The unaided learner reaches 65 — a mean-pooled base plus the one small free gain, region selection (+4). Frame selection and utterance reweighting add nothing, so free self-supervised selection stops there, a clear +4.6 below even the first oracle rung. Everything from 65 to the 81 ceiling — the +16 that actually matters — is oracle: which pairs are aligned, which spoken word is the referent, and which object is named when no word is spoken. The bottleneck is not optimization, not scale, and not which region or which moment — the selection mechanisms are nearly spent. It is which referent: the object being named, especially when it is never spoken. The next chapter asks whether the cues children actually use can supply any of it.


  1. An earlier draft of this ladder reported these free gains as +9.7 (region) and +2.4 (frame). Both were inflated: the region gain was measured against a CLS whole-frame baseline (52.9), and CLS is a weak object readout for DINOv2 — against the fairer mean-pooled baseline (61.3) region-MIL is only +4.0. Chapter 8 traces this across encoders and shows the region gain shrinks further for stronger ones; the numbers here are the clean, single-rig values.↩︎