2 The data
The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.
These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.
Everything in this book is learned from BabyView 2025.2, a corpus of head-mounted camera recordings from young children’s daily lives, paired with transcribed speech. This chapter describes what the corpus contains, why referential alignment is the central problem it poses, and the two annotations — CLIP and Gemini — we use to estimate which moments of speech are actually about something visible.
2.1 Scale
At 1 frame per second, aligned to transcribed utterances, the corpus gives:
| children | 36 |
| recordings (videos) | 8,188 |
| utterance–frame pairs | 1,145,312 |
| total footage | ~1,390 h (lower bound) |
| training vocabulary (≥5 occurrences) | 14,888 words |
This is roughly two orders of magnitude more data than SAYCam-S (the single-child corpus behind CVCL), and it spans many homes rather than one — a difference that matters for both amount and diversity (Chapter 6).
2.2 What the speech looks like
Child-directed speech in the wild is short and mostly not referential. Utterances have a median of 3 tokens (mean 4.2; Figure 2.1), and — as the annotations below quantify — the overwhelming majority of them are not naming anything visible on screen at that moment. This is the core obstacle: a naive contrastive learner that pairs every utterance with its frame is training on a signal that is ~90% noise.
The data is also extremely uneven across children: the median child contributes ~19,600 utterance–frame pairs, but the interquartile range runs from ~2,700 to ~43,500, and the extremes span 253 to 134,686. Who contributes how much is itself a variable we return to in Chapter 6.
2.3 Two annotations of alignment
Because most speech is non-referential, we need a way to estimate, for each (utterance, frame) pair, how strongly the utterance refers to a concrete object visible in the frame. We use two, and the difference between them matters throughout the book.
CLIP (clip_score_max) scores the whole frame against the utterance with a pretrained CLIP model. It is cheap and requires no extra infrastructure, but its distribution is extremely compressed — mean 0.224, sd 0.015 — so the entire aligned/misaligned distinction lives in a razor-thin band, which is why a fixed threshold (e.g. 0.24) is so delicate. And because it scores the whole frame, it is easily fooled — it scores on-screen text and topical keywords highly even when nothing is depicted.
Gemini (alignment, 0–100, plus a referent noun) is a per-pair judgment from Gemini 2.5 Flash on Vertex AI: does this utterance name a concrete object visible in this frame, and which one? It is an IRB-approved annotation pass over the full corpus (frames are human-subjects data and never leave the approved pipeline). Its distribution is bimodal and decisive: 91% of pairs score 0 (correctly — most speech is not referential), while 9.2% clear ≥50 and 6.0% clear ≥80, each carrying a candidate referent word. This matches the intuition that only a small minority of moments are genuine naming events.
The two annotations substantially disagree — Spearman ρ ≈ 0.14 over the full corpus (ρ ≈ 0.22 on the referential tail), i.e. a lot of independent signal (Figure 2.1, right). Inspection of the disagreements favors Gemini: it recovers referents CLIP misses — shared book reading (“Llama Llama” → book), toys, and objects on TV screens — and it rejects CLIP’s keyword false-positives (abstract talk about a “high chair,” deictic “get it”). A direct test confirms it: at matched count, training on a Gemini-selected set is at least as good as a CLIP-selected one on the clean eval, and its referent labels open a much larger headroom (Chapter 4).
Throughout the rest of the book, Gemini is the gold definition of alignment (and its referent the gold label), with CLIP kept as the cheaper, weaker comparison and as the annotation behind the preliminary experiments of the appendices.
2.4 The 2026.1 consolidated release
The chapters above describe the 2025.2 working release the experiments were developed on. For the paper-phase experiments, everything moves to 2026.1: 16,298 videos from 50 children (2,700 hours; ages 2.9–54 months at recording, median 17.0), a strict superset of 2025.2 (8,566 shared videos, byte-identical; 11,735 new). All layers were consolidated under one root with a single key convention, and every claim below is generated from the committed diagnostics bundle (diagnostics/2026.1/, extracted on-cluster by src/make_diagnostics.py).
Annotation layers. Each layer covers the release nearly completely, and every layer joins on the same keys (video, and utterance or frame within video):
- Transcripts: 1,838,288 ASR utterances (16,008 videos have speech).
- Referent annotation (Gemini alignment 0–100 + referent word): 1,838,061 of 1,838,134 utterance–frame training pairs scored (99.996%).
- Language annotation (two independent Gemini passes + agreement): all 1,838,288 utterances; 99.4% pass agreement, 99.05% English among agreed utterances.
- Pose: 8.3M person detections. 65% of frames contain at least one detection, but only 43% contain a social partner (a detection with a visible face or body): a fifth of all detections are hands-only — largely the wearing child’s own hands entering the frame — so analyses of social presence should use the face-or-body criterion, not raw detections.
- Embeddings: DINOv3-B 4×4 region grid for all 1,745,489 training-pair frames.
Recording effort and age. Contribution is uneven — the largest contributor provides 7.1% of all hours — which motivates the child-diversity analyses of Chapter 6:
Video lengths and silent videos. Videos are typically ~10 minutes (median 10.5 min); 428 videos (2.6%) are under a minute but carry only 0.13% of utterances, so no length exclusion is applied. 298 videos (1.8%) have no transcribed speech at all — these are not short clips (median 8.7 min): they are real recordings without talk near the microphone. Rate figures below exclude videos under 2 minutes, whose tiny denominators otherwise produce implausible per-hour estimates (e.g. a 1.1 s clip with one utterance ≈ 3,300/hour).
Speech density is stable across videos and children (median ≈ 700 utterances/hour), with no outlier children that would indicate transcription failures:
Change with age. Speech density rises with child age (roughly 550 utterances/hour around 12 months to ~950 by 30 months) while mean utterance length is comparatively flat (~4–5 words); per-child trajectories (thin lines) track the corpus trend. The single child recorded beyond 48 months (S02170002; 8 videos, 1.0 h) carries an inconsistent registry entry — an S02-range (preschool-series) id, but dataset = BV-main and release = 2026.1, while every other S02 subject is tagged 2025.1.preschool or unreleased. It has been excluded from the working release (outputs/excluded_post_release.tsv; Airtable fix requested), leaving 16,298 videos from 50 children — the counts used throughout:
Reported vs measured language — and an ASR artifact. An earlier draft of this section claimed the household survey sharply under-reported English, based on transcript-level language labels. That claim was wrong, and the error is instructive: Whisper auto-translates non-English speech, rendering Korean or Japanese caregiver talk as fluent English text, so transcript-based measurement reads ~90% English even for households speaking almost none. An audio-level re-measurement (Gemini language ID on 47,826 sampled 60-s windows, bypassing transcripts entirely; annotations/language/audio/) settles it: the survey correlates r = 0.91 with audio-measured English, while the transcript-based measure correlates only r = 0.46 — the family reports were approximately right all along. Ten children fall below 80% audio-measured English. The English training filter is therefore audio-based and video-level: videos with < 50% audio-measured English are excluded (1,258 videos, 7.7% of pairs; plus 0.9% of utterances the transcript layer itself flags as non-English), which keeps bilingual households’ genuinely-English recordings. This matters beyond our corpus: any project measuring language environments from Whisper transcripts will overestimate English by construction.
Annotation distributions. Long utterances are real, not segmentation failures: 0.97% of utterances have ≥20 words, and inspection shows storybook reading and adult conversation. Persons-per-frame includes the 35% of frames with no detected person: