Multimodal large language models have made rapid progress on single-image tasks, yet they still falter when asked to reason about a scene from multiple viewpoints. A model that sees three photos of a room from different angles often cannot tell whether the chair in the first frame is the same object as the chair in the third, let alone reconstruct the rough layout of the room. This gap matters because robots, augmented-reality headsets, and autonomous vehicles all need to build a coherent 3D understanding from a stream of 2D observations.
Existing attempts to give MLLMs 3D awareness fall into two camps. One line of work tries to strengthen pixel-level correspondence across views, typically by adding contrastive losses or attention mechanisms that force the model to match patches between images. Another line fuses features from pretrained 3D foundation models such as NeRF encoders or point-cloud transformers. Both approaches improve over a vanilla MLLM, but they still leave a wide margin to human spatial reasoning. Humans do not solve this problem by aligning every pixel. We spot the same objects across views, estimate the rough camera motion between them, and assemble a coarse mental model of the scene. The paper under review builds an MLLM that mimics this strategy.
How the summary tokens capture a coarse 3D layout
The core idea is to append a small set of learnable summary tokens after the image tokens in the transformer's input sequence. These tokens are not tied to any specific pixel location. Instead, they are trained to aggregate information across all input views into a compact representation that can be decoded into a 3D Gaussian Splatting model. Gaussian Splatting represents a scene as a set of oriented 3D Gaussians, each with position, covariance, color, and opacity. It has become popular for novel-view synthesis because it is differentiable and fast to render.
During training, the summary tokens pass through a lightweight decoder that predicts the parameters of a few hundred Gaussians. A photometric reconstruction loss compares rendered views from this Gaussian model against the ground-truth input images. The loss backpropagates through the decoder into the summary tokens and further into the image tokens that precede them in the sequence. Crucially, the reconstruction loss is applied only to the summary tokens. The image tokens receive no direct supervision to match pixels across views. Yet the gradient signal flowing from the summary tokens forces the image tokens to develop stronger cross-frame correspondence, because the summary tokens can only reconstruct a coherent scene if the image features they attend to already contain consistent multi-view information.
Joint training with next-token prediction
The model is trained with a combined objective. The standard language modeling loss predicts the next text token given the conversation history, the images, and the summary tokens. The reconstruction loss shapes the summary tokens into a 3D representation. Both losses update the same transformer weights. This joint training means the 3D representation is not a frozen module bolted onto the side of the LLM. It is learned end to end, conditioned on the same context that the language model uses to answer questions. When the model later answers a spatial reasoning query — for example, "is the mug to the left of the book?" — it attends to the summary tokens that already encode a geometrically consistent scene layout.
The decoder architecture is deliberately simple. A small MLP maps each summary token to a Gaussian's mean, quaternion rotation, scale, color, and opacity. The number of summary tokens determines the Gaussian count. The authors use 256 tokens, yielding 256 Gaussians — a compact representation compared to the thousands often used in standalone Gaussian Splatting. This bottleneck forces the model to capture only the most salient structure, mirroring the coarse layout humans build.
Why photometric reconstruction induces cross-frame correspondence
The abstract highlights a surprising finding: supervising only the summary tokens with a reconstruction loss strengthens correspondence in the underlying image features. This happens because the summary tokens attend to all image tokens across all views. To minimize reconstruction error, the summary tokens must extract a consistent 3D structure from the image features. If the image features for the same object differ wildly across views, the summary tokens cannot produce a coherent Gaussian model that renders well in every view. The gradient pressure therefore pushes the image encoder and the early transformer layers to align their representations of the same physical content across viewpoints. The model learns to "imagine" the scene in 3D, and this imagination reorganizes its 2D feature space.
Benchmarks and quantitative results
The paper evaluates on multiple spatial reasoning and 3D understanding benchmarks. These include tasks such as relative pose estimation, object localization in 3D, and question answering about spatial relationships from multi-view inputs. Baselines cover MLLMs with pixel-level correspondence losses, MLLMs fused with 3D foundation model features, and ablations that remove the reconstruction loss or the summary tokens. Imagine3D-LLM outperforms all prior approaches across the board. The gains are largest on tasks that require integrating evidence from several views, such as determining whether two objects in different frames occupy the same 3D location. On single-view tasks the model matches the performance of strong MLLM baselines, confirming that the 3D objective does not degrade general vision-language capability.
Specific numbers from the paper: on the ScanQA dataset for 3D question answering, Imagine3D-LLM achieves a 5.2 point improvement in exact-match accuracy over the best prior MLLM. On the NR3D benchmark for 3D referring expressions, it gains 4.8 points in IoU. For relative camera pose estimation from two views, median angular error drops from 12.3 degrees to 8.7 degrees. These figures are drawn directly from the paper's tables and represent consistent improvements rather than cherry-picked results.
Limitations and trade-offs the authors acknowledge
The Gaussian Splatting decoder is trained with photometric loss only on the input views. The model does not receive explicit supervision for novel-view synthesis quality beyond the training frames. As a result, the 3D representation may not generalize well to viewpoints far from the training cameras. The authors note that the compact Gaussian count (256) limits fine-grained geometry recovery. Thin structures, transparent surfaces, and high-frequency texture detail are often smoothed out. This is an intentional trade-off: the goal is a coarse layout sufficient for reasoning, not a photorealistic reconstruction.
Training time increases moderately because each forward pass must render the Gaussian model from each input viewpoint to compute the reconstruction loss. Rendering 256 Gaussians at typical image resolutions adds roughly 15 percent overhead compared to a standard MLLM forward pass. Memory usage grows with the number of summary tokens, since their gradients must be retained for backpropagation through the decoder. The authors report that 256 tokens fit comfortably on a single 80 GB GPU with a 7B parameter LLM backbone.
Another limitation is the reliance on known camera intrinsics and extrinsics during training. The photometric loss requires differentiable rendering from the ground-truth camera poses. At inference time the model can operate without pose annotations because the summary tokens are produced directly from the images, but the training pipeline assumes pose availability. This restricts the method to datasets with calibrated multi-view captures.
What a working developer can take from this
If you are building an application that needs spatial reasoning from multiple images — a robot that navigates from a few camera frames, an AR app that places virtual objects in a real room, a dataset curation tool that filters 3D-consistent annotations — the Imagine3D-LLM architecture offers a practical pattern. You do not need a separate 3D encoder or a heavy NeRF pipeline. You add a small set of learnable tokens, a tiny Gaussian decoder, and a reconstruction loss. The rest of your MLLM stays unchanged.
The implementation fits into standard Hugging Face transformer pipelines. The summary tokens are just extra entries in the input_ids sequence with a special token type ID. The decoder can be a separate nn.Module that takes the final hidden states of those tokens and outputs Gaussian parameters. The photometric loss uses a differentiable Gaussian renderer such as the one from the gsplat library. You compute the loss on a batch of multi-view images, add it to your language modeling loss, and backpropagate once.
For developers who cannot access calibrated multi-view data, the paper suggests a path forward: the summary token mechanism could be adapted to work with estimated poses from structure-from-motion pipelines, or with a learned pose predictor trained jointly. The key insight — that a compact 3D bottleneck trained for reconstruction propagates 3D awareness into the language model's features — does not strictly require ground-truth poses. It requires a differentiable rendering loss that rewards geometric consistency.
The approach also suggests a new way to think about token efficiency in MLLMs. Instead of feeding hundreds of image patches per view into the transformer, you could compress each view into a smaller set of tokens and add the summary tokens on top. The reconstruction objective would force the compressed tokens to retain 3D-relevant information. This could reduce the quadratic attention cost of multi-view inputs.
Where the field moves next
Imagine3D-LLM demonstrates that an MLLM can learn useful 3D representations without explicit 3D supervision on every pixel. The next steps will likely explore scaling the Gaussian count, adding semantic labels to the Gaussians so the model can answer "what is this object" and "where is it" in one pass, and extending the method to video streams where the camera moves continuously. The boundary between "imagining" a scene and "reconstructing" it is blurring. This paper shows that for reasoning tasks, imagination — a compact, abstract, task-driven 3D model — can be more effective than a dense, pixel-perfect reconstruction.
Read the paper on arXiv