9 Architectures linking language and vision
The code and prose in this book were written by Claude Code (an AI coding agent) under human direction. They have not been line-by-line verified by a human and must not be published or cited as-is.
These are exploratory analyses of preliminary, in-progress work. Results and numbers have not been independently reviewed or validated and require checking before any use.
Chapter 8 showed the vision encoder is not the bottleneck. But every result in this book has used a single architecture: a CLIP-style contrastive two-tower, in which a word and a frame are each crushed to one vector and aligned by a dot product, with a hard max-over-regions (MIL) sneaking in a single bit of spatial selection. That is a strong, rigid inductive bias. A natural worry — and a natural next idea — is that the contrastive objective itself is what caps us: perhaps a more flexible, attention-based or generative model, of the kind that has reduced alignment problems elsewhere (LLaVA-style captioners, the LRAS token models of Lee et al. (2025)), would extract more from the same noisy pairs.
This chapter tests that, holding the encoder (frozen DINOv2 region features) and the data (the same 877k pairs) fixed so that architecture is the only variable. We change it along its two axes — how a word aligns to the image (Exp 1), and what the model is trained to do (Exp 2) — and ask whether either beats the plain contrastive two-tower.
9.1 Softening the alignment (Exp 1): hard selection is optimal
Our MIL is a hard max: a word’s score is its similarity to the single best-matching region, winner-take-all. We can relax that to a temperature-controlled softmax over regions — a soft, weighted average — where the max is exactly the τ→0 limit. And we can go further and let the model learn the pooling, with a cross-attention head (the word queries the regions, keys/values are learned projections). Both keep the contrastive loss; only the pooling changes.
Softening does not help — it monotonically hurts (Figure 9.1). Hard max is best (66.1); raising the temperature walks it down (65.3 → 63.1 → 61.6 as τ goes 0.05 → 0.15 → 0.50); and the learned attention pooling is worst of all (60.2). This is a clean, slightly counterintuitive result: a word refers to one object in one place, so winner-take-all is exactly the right bias — averaging over regions, even a learned text-conditioned average, just readmits distractor regions. The “flexibility” we hoped would help is actively counterproductive. The rigid hard-max is not a crude hack we tolerate; it is the correct model of referential selection.
9.2 Changing the objective (Exp 2): generative ties, does not beat
The deeper change is to drop contrastive discrimination for generation. We train a small, from-scratch autoregressive captioner (no pretrained language model): the frozen region features are the cross-attention memory, and a tiny Transformer decoder generates the utterance token by token. Alignment now lives in the decoder’s cross-attention — each generated word attends over the regions and pulls in the visual information it needs — learned purely from the next-token loss. Evaluation is generative too, following DevBench’s continuation-scoring (Tan et al. 2024): for a Konkle item we pick the image under which the target word is most probable, \(\arg\max_i P(\text{word} \mid \text{image}_i)\).
It works — and it ties the contrastive baseline: 62.1, squarely in the contrastive range, neither beating nor really trailing it (Figure 9.1). Two things are worth noting. First, a generative model with cross-attention does learn clean word–object grounding from the child data alone, with no language prior — a real positive result about what the architecture can do. Second, it does so despite spending most of its loss on non-referential function words (“look at the …”), which the contrastive setup ignores entirely — yet the extra machinery buys nothing over a dot product. (LLaVA-style models do beat this, but by importing a large pretrained language prior — a different question, about how much external knowledge substitutes for referential alignment, not about the architecture per se.)
9.3 The architecture is not the bottleneck either
Across both axes the verdict is the same: nothing beats the plain contrastive two-tower. Softening its alignment hurts, learning the alignment hurts more, and swapping its objective for generation merely ties. Combined with Chapter 8 — where a 30-point spread of encoders left the referential gap intact — this closes the argument Chapter 7 opened. It is not the eyes (the encoder), and it is not the objective or the alignment mechanism (the architecture). Every lever we can pull on the model leaves the same referential ceiling in place. What remains, as the only thing that ever moves it, is the oracle information about which pairs are aligned and which word is the referent — that is, the referential signal in the data. The bottleneck was never the machine; it is what the machine is given to learn from.