5  Cues

ImportantAI-generated — not for publication

The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.

WarningPreliminary & unvetted analysis

These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.

The decomposition of Chapter 4 relocates the prize. Earlier social-cue work aimed accessible cues at filtering frames, and found little — but filtering is the small rung (~+5, and region-MIL already takes most of it). The open headroom is in the referent label (~+16), split into word selection (+4.1, the referent when it is spoken) and vision-binding (+7.5, the referent when it is not). This chapter re-aims cues at referent identification — which word, and which object — measured against the Gemini referent oracle on Konkle (dev-tuned, Vong-60 test), with error bars.

The mechanism throughout is a weighted bag-of-words or a region prior: instead of pooling uniformly, weight each word (or region) by a per-item cue so the referent dominates.

5.1 Condition 0: do the cues even predict alignment?

Before asking whether a cue can recover a rung, a blunter question: does any accessible cue predict referential alignment at all — does it correlate with the Gemini gold? A cue needs to correlate with alignment at roughly ρ ≈ 0.3 (a selection AUC ~0.65) to drive a useful filter — below that, any reweighting is swamped by noise. Nothing we can read comes close.

cue grain predicts alignment
discourse continuity utterance ρ = 0.14 (best overall)
combined language predictor utterance ρ = 0.11
prosody (energy range) word ρ = 0.05
person present utterance AUC 0.55 (best social)
caregiver speaking utterance AUC 0.54
pose gesture (point / show) utterance AUC 0.53
child’s hand present utterance AUC 0.53
visible face utterance AUC 0.51
caregiver gaze direction (pitch/yaw) utterance AUC ~0.50
adult hand / reach / pointing utterance AUC ~0.50

The best single cue — discourse continuity — reaches only ρ ≈ 0.14, about half the ignition bar. The best social cue is merely “is a person present” (AUC 0.55; mean alignment 8.5 with a person vs 5.1 without) — a real whisper, but nowhere near useful; a visible face, a pointing gesture, and gaze direction add nothing beyond it, and the sharper we make them (gesture geometry, head direction from raw keypoints) the more clearly the ceiling holds. This is the foundation of every null below: the cues do not predict which moments are referential, so — on any axis, whether used to select pairs, weight words, or steer region attention — they cannot do better than chance. At 1 fps, in this stream, the signal that would tell a learner when reference is happening is simply not legible.

5.2 Language cues: word selection

Can an accessible speech cue pick the referent word inside the utterance, recovering the word rung (uniform 71.6 → oracle t15 74.0, a +2.4 headroom at best epoch)? We tested four, each as a per-word pooling weight on the referent-bearing pairs (region-MIL, Konkle test-60, 3 seeds):

Cue Mechanism Result
speaker filter drop the child’s own utterances null — drop-child 63.5 ≤ random-drop control 64.3
noun bias static content-noun weighting (spacy_pos) null at convergence (+2 transient, gone by the final epoch)
discourse newness down-weight recently-repeated words null (−0.8, within noise)
prosody per-word RMS energy from the audio null (−0.6, within noise)

Every language cue is null. Two things explain it. First, the contrastive objective already does implicit word-selection: the referent word co-occurs with its object across many pairs and so rises on its own, while function words co-occur with everything and wash out — which is why the noun bias adds nothing at convergence and why a discourse-repetition cue, by penalizing repeated (hence often referential) words, does not help. Second, and more tellingly, the headroom is small to begin with: measured cleanly at best epoch, uniform bag-of-words already reaches 71.6, within +2.4 of the referent-word oracle (74.0). The learner has essentially already solved word-selection on its own; no cue recovers the residual.

The conclusion is clean: word selection is not a cue-accessible lever. The speech side of referent identification is already handled by the learner. What remains is the harder half.

5.3 Vision cues: binding the unspoken referent

The genuinely open rung is vision-binding (+7.5) — the ~37% of aligned moments where the referent is never spoken (“look at that”). Here the object must be identified from vision, and this is a different kind of problem: an accessible cue can supply the referent’s location (gaze, salience, a caregiver’s point, the child’s hands) but not its name — naming an unspoken object is the word-learning problem itself, solvable only by cross-situational structure or an external labeler (which is exactly what the Gemini oracle is). So the question is whether pointing region attention at the right object on these deictic pairs lets the learner bind it to words it has seen elsewhere.

We tested the most obvious cue from the developmental literature — the caregiver’s hand (a pointing/holding proxy). From the full pose annotation (per-frame body-part bounding boxes over 6.5M person-detections), we take the largest person’s hand location as a per-pair region prior that biases region-MIL’s cell selection toward the indicated object (referent-bearing pairs, 70% with a hand target, 3 seeds):

Konkle
uniform (no prior) 70.0
+ caregiver-hand region prior 69.8 (−0.2)

Null. Steering attention to the caregiver’s hand does not improve referent identification. The reasons are the ones anticipated: region-MIL already selects the best-matching cell on its own, so forcing it toward the hand adds nothing where the referent is nameable; and on the deictic pairs — the ones this cue is meant to rescue — there is no content word for the attended object to bind to. A location cue cannot supply a name.

The region axis is where the pose cue had the least room, so we also tested it on the axis where it might have had the most — the utterance filter (does a gesture mark a referential pair?). It fails there too, for the deeper reason of Section 5.1: a sharper pointing/showing gesture predicts Gemini alignment at ρ ≈ 0.03 (AUC 0.53), and caregiver gaze direction at AUC ~0.50 — both chance. A cue that does not predict when reference happens cannot filter for it, pick the frame for it, or steer attention to it. The one social signal with any purchase is bare person presence (AUC 0.55), and it is far too weak to bootstrap.

5.4 The shape of the result

Across both halves — every language cue (speaker, noun bias, discourse newness, prosody) and the canonical vision cue (caregiver-hand pose) — no accessible cue recovers any oracle rung. This is not for want of the right signal in the data; the pose is faithful and the prosody is real. It is that the two oracle levers are already claimed: word selection is done implicitly by the contrastive objective (the referent word co-occurs with its object and rises on its own), and vision-binding requires naming the unspoken object, which no location cue provides and which is the word-learning problem itself. In this frozen-feature regime, the ~19 points of label headroom are real but not accessible-cue-recoverable — the honest boundary of what these data and this probe can reach.