Visual grounding helps language models learn concrete object properties, but this benefit doesn't show up in standard benchmarks—suggesting we need better evaluation methods to measure what models actually learn from multimodal information.
This paper tests whether giving a small language model visual grounding for words (like showing it images of 'banana') before training helps it learn language better. The author finds that visual initialization leaves a lasting imprint on the model, especially for object-property knowledge, but this advantage is invisible to most standard benchmarks that test grammar and abstract reasoning.