2  The data

ImportantAI-generated — not for publication

The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.

WarningPreliminary & unvetted analysis

These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.

Everything in this book is learned from BabyView 2025.2, a corpus of head-mounted camera recordings from young children’s daily lives, paired with transcribed speech. This chapter describes what the corpus contains, why referential alignment is the central problem it poses, and the two annotations — CLIP and Gemini — we use to estimate which moments of speech are actually about something visible.

2.1 Scale

At 1 frame per second, aligned to transcribed utterances, the corpus gives:

children 36
recordings (videos) 8,188
utterance–frame pairs 1,145,312
total footage ~1,390 h (lower bound)
training vocabulary (≥5 occurrences) 14,888 words

This is roughly two orders of magnitude more data than SAYCam-S (the single-child corpus behind CVCL), and it spans many homes rather than one — a difference that matters for both amount and diversity (Chapter 6).

2.2 What the speech looks like

Child-directed speech in the wild is short and mostly not referential. Utterances have a median of 3 tokens (mean 4.2; Figure 2.1), and — as the annotations below quantify — the overwhelming majority of them are not naming anything visible on screen at that moment. This is the core obstacle: a naive contrastive learner that pairs every utterance with its frame is training on a signal that is ~90% noise.

The data is also extremely uneven across children: the median child contributes ~19,600 utterance–frame pairs, but the interquartile range runs from ~2,700 to ~43,500, and the extremes span 253 to 134,686. Who contributes how much is itself a variable we return to in Chapter 6.

Figure 2.1: Descriptive statistics of BabyView 2025.2 and its annotations. Left to right: utterance–frame pairs per child (highly skewed); utterance length in tokens (short, heavy-tailed); the Gemini alignment score (log scale — a large spike at 0 and a spread tail); and Gemini alignment vs. CLIP’s whole-frame score (a near-shapeless cloud — the two disagree substantially).

2.3 Two annotations of alignment

Because most speech is non-referential, we need a way to estimate, for each (utterance, frame) pair, how strongly the utterance refers to a concrete object visible in the frame. We use two, and the difference between them matters throughout the book.

CLIP (clip_score_max) scores the whole frame against the utterance with a pretrained CLIP model. It is cheap and requires no extra infrastructure, but its distribution is extremely compressed — mean 0.224, sd 0.015 — so the entire aligned/misaligned distinction lives in a razor-thin band, which is why a fixed threshold (e.g. 0.24) is so delicate. And because it scores the whole frame, it is easily fooled — it scores on-screen text and topical keywords highly even when nothing is depicted.

Gemini (alignment, 0–100, plus a referent noun) is a per-pair judgment from Gemini 2.5 Flash on Vertex AI: does this utterance name a concrete object visible in this frame, and which one? It is an IRB-approved annotation pass over the full corpus (frames are human-subjects data and never leave the approved pipeline). Its distribution is bimodal and decisive: 91% of pairs score 0 (correctly — most speech is not referential), while 9.2% clear ≥50 and 6.0% clear ≥80, each carrying a candidate referent word. This matches the intuition that only a small minority of moments are genuine naming events.

The two annotations substantially disagree — Spearman ρ ≈ 0.14 over the full corpus (ρ ≈ 0.22 on the referential tail), i.e. a lot of independent signal (Figure 2.1, right). Inspection of the disagreements favors Gemini: it recovers referents CLIP misses — shared book reading (“Llama Llama” → book), toys, and objects on TV screens — and it rejects CLIP’s keyword false-positives (abstract talk about a “high chair,” deictic “get it”). A direct test confirms it: at matched count, training on a Gemini-selected set is at least as good as a CLIP-selected one on the clean eval, and its referent labels open a much larger headroom (Chapter 4).

Throughout the rest of the book, Gemini is the gold definition of alignment (and its referent the gold label), with CLIP kept as the cheaper, weaker comparison and as the annotation behind the preliminary experiments of the appendices.

2.4 The 2026.1 consolidated release

The chapters above describe the 2025.2 working release the experiments were developed on. For the paper-phase experiments, everything moves to 2026.1: 16,298 videos from 50 children (2,700 hours; ages 2.9–54 months at recording, median 17.0), a strict superset of 2025.2 (8,566 shared videos, byte-identical; 11,735 new). All layers were consolidated under one root with a single key convention, and every claim below is generated from the committed diagnostics bundle (diagnostics/2026.1/, extracted on-cluster by src/make_diagnostics.py).

Annotation layers. Each layer covers the release nearly completely, and every layer joins on the same keys (video, and utterance or frame within video):

Figure 2.2: Fraction of the 16,298 release videos covered by each data layer. Transcripts, referent, and language cover the 16,008 videos with any speech; pose covers all but 10 sub-second clips; embeddings cover every training-pair frame.
  • Transcripts: 1,838,288 ASR utterances (16,008 videos have speech).
  • Referent annotation (Gemini alignment 0–100 + referent word): 1,838,061 of 1,838,134 utterance–frame training pairs scored (99.996%).
  • Language annotation (two independent Gemini passes + agreement): all 1,838,288 utterances; 99.4% pass agreement, 99.05% English among agreed utterances.
  • Pose: 8.3M person detections. 65% of frames contain at least one detection, but only 43% contain a social partner (a detection with a visible face or body): a fifth of all detections are hands-only — largely the wearing child’s own hands entering the frame — so analyses of social presence should use the face-or-body criterion, not raw detections.
  • Embeddings: DINOv3-B 4×4 region grid for all 1,745,489 training-pair frames.

Recording effort and age. Contribution is uneven — the largest contributor provides 7.1% of all hours — which motivates the child-diversity analyses of Chapter 6:

Figure 2.3: Hours per child (left: sorted). Right: recorded hours by child age, and each child’s cumulative hours across age — the corpus is densest in the second year of life.
Figure 2.4: Hours recorded per child, sorted.

Video lengths and silent videos. Videos are typically ~10 minutes (median 10.5 min); 428 videos (2.6%) are under a minute but carry only 0.13% of utterances, so no length exclusion is applied. 298 videos (1.8%) have no transcribed speech at all — these are not short clips (median 8.7 min): they are real recordings without talk near the microphone. Rate figures below exclude videos under 2 minutes, whose tiny denominators otherwise produce implausible per-hour estimates (e.g. a 1.1 s clip with one utterance ≈ 3,300/hour).

Figure 2.5: Distribution of video lengths.

Speech density is stable across videos and children (median ≈ 700 utterances/hour), with no outlier children that would indicate transcription failures:

Figure 2.6: Speech density per video (left) and per child against median age (right); videos ≥ 2 minutes.

Change with age. Speech density rises with child age (roughly 550 utterances/hour around 12 months to ~950 by 30 months) while mean utterance length is comparatively flat (~4–5 words); per-child trajectories (thin lines) track the corpus trend. The single child recorded beyond 48 months (S02170002; 8 videos, 1.0 h) carries an inconsistent registry entry — an S02-range (preschool-series) id, but dataset = BV-main and release = 2026.1, while every other S02 subject is tagged 2025.1.preschool or unreleased. It has been excluded from the working release (outputs/excluded_post_release.tsv; Airtable fix requested), leaving 16,298 videos from 50 children — the counts used throughout:

Figure 2.7: Per-video speech measures across age (dots), per-child 3-month median trajectories (thin lines), and the corpus median (heavy line).

Reported vs measured language — and an ASR artifact. An earlier draft of this section claimed the household survey sharply under-reported English, based on transcript-level language labels. That claim was wrong, and the error is instructive: Whisper auto-translates non-English speech, rendering Korean or Japanese caregiver talk as fluent English text, so transcript-based measurement reads ~90% English even for households speaking almost none. An audio-level re-measurement (Gemini language ID on 47,826 sampled 60-s windows, bypassing transcripts entirely; annotations/language/audio/) settles it: the survey correlates r = 0.91 with audio-measured English, while the transcript-based measure correlates only r = 0.46 — the family reports were approximately right all along. Ten children fall below 80% audio-measured English. The English training filter is therefore audio-based and video-level: videos with < 50% audio-measured English are excluded (1,258 videos, 7.7% of pairs; plus 0.9% of utterances the transcript layer itself flags as non-English), which keeps bilingual households’ genuinely-English recordings. This matters beyond our corpus: any project measuring language environments from Whisper transcripts will overestimate English by construction.

Figure 2.8: Per-child measured % English against the household survey. Grey: transcript-based measurement, inflated to ~90–100% for every child by Whisper auto-translation. Red: audio-based measurement, which tracks the survey (r = 0.91); children below 80% audio-English are labelled. Dashed line is identity.

Annotation distributions. Long utterances are real, not segmentation failures: 0.97% of utterances have ≥20 words, and inspection shows storybook reading and adult conversation. Persons-per-frame includes the 35% of frames with no detected person:

Figure 2.9: Distributions over the annotation layers: Gemini alignment scores (most speech is non-referential), persons per frame (including zero), and utterance length.