Geometric Consistency in Multimodal Visual Media: A Longitudinal Analysis of Spatial Reasoning and Scaling Behavior
Main Article Content
Abstract
Vision–language models have made impressive progress in interpreting and describing visual scenes, yet they still struggle with the kind of precise geometric reasoning needed to understand how objects relate and move in space. In this work, we introduce the Geometric Consistency Score (GCS) — a new way to measure how consistently models reason about spatial relationships. GCS formalizes geometric inference through logical constraints, enabling a quantitative view of transitive spatial coherence. We evaluate existing multimodal architectures across fourteen diverse benchmarks covering dynamic and static scenes, indoor and outdoor settings, and both real and synthetic data. This large-scale analysis leads to the first scaling laws for geometric reasoning in vision–language systems. Notably, transformer models improve only sub-linearly in spatial precision (α=0.31±0.04), far slower than their gains in semantic understanding (α=0.78±0.06), and their performance drops sharply under occlusion (IoU 62.1%→18.3%, p<0.001). Our ablation studies reveal that conventional contrastive pre-training makes models largely permutation-invariant to spatial order — embeddings remain highly similar (>85%) even when object positions are shuffled. To address this limitation, we adapt Group Relative Policy Optimization with a geometry-aware reward shaped by GCS. This approach raises geometric reasoning accuracy to 40.8% (95% CI: [38.4, 43.2]) compared to 13.3% for standard models. While a substantial gap remains to human performance (κ=0.89 across 50 participants), our findings highlight both the promise and current limits of multimodal architectures in spatial reasoning. To foster progress, we release the GCS Evaluation Toolkit, enabling standardized and transparent benchmarking of geometric consistency in visual understanding systems.


