Most unified multimodal models split along a fundamental divide: some treat both language and images as discrete tokens, others pair discrete language prediction with continuous image generation. The token-based approach creates a visual quantization loss, while the hybrid setup demands separate objectives and sampling steps for each modality. A fully continuous alternative sidesteps these compromises and permits a shared generative process, yet has seen limited exploration for multimodal pretraining. Multimodal Flow directly addresses this gap.

The core idea organizes text blocks and images as ordered continuous hyperchunks. Textual token order stays intact, and visual spatial structure remains readable by the model. A single chunk-causal flow backbone learns one vector field over these hyperchunks through flow matching. Cross-modal signals flow through a joint attention module, while each modality also passes through its own feed-forward networks. During training the system predicts several target chunks in parallel; at inference time it generates hyperchunks one after another.

Hyperchunk construction and ordering

Text enters the model as a sequence of embedding vectors, kept in the order they appear. Image patches form their own ordered set, preserving two-dimensional layout when the chunks are arranged. Both modalities live in the same continuous space, so the flow backbone can move between them without conversion layers. The ordering constraint means the model knows which dimension corresponds to language and which to vision, even though no discrete token IDs appear.

Flow matching over shared hyperchunks

The backbone treats the combined language‑vision representation as a single trajectory. Flow matching finds a vector field that pushes a source distribution toward a target distribution across the hyperchunks. Because the same field serves both modalities, the model learns a joint embedding geometry rather than separate ones. The matching objective operates on the continuous vectors, eliminating the need for discrete token predictors or per‑modality loss terms.

Joint attention and modality‑specific feed‑forward streams

Cross‑modal interaction happens through a shared attention layer that can look at language tokens and image patches simultaneously. This lets the model align words with visual concepts without relying on a separate cross‑modal projector. Each modality also runs through its own feed‑forward network, preserving specialization so that language‑style patterns and image‑style patterns are not blurred together.

Training and inference procedure

At training time the model samples multiple future hyperchunks and predicts them in one step, which stabilizes the flow matching signal. At inference the generation proceeds hyperchunk by hyperchunk, each new chunk conditioned on the previous ones. This sequential factorization mirrors how a reader might scan a document word by word while also absorbing visual layout.

MF‑1 and empirical results

The authors instantiated MF‑1 and pretrained it on multimodal data at three scales: 0.6 billion, 1.2 billion, and 1.6 billion parameters. With only 150 billion pretraining tokens the model scored an average of 82.8 on GenEval and DPG‑Bench, and 75.3 on VQAv2, MMBench, and POPE. These numbers place MF‑1 on par with unified models that were trained on substantially more data. When data, optimization budget, and parameter count are matched, MF‑1 surpasses representative hybrid and discrete systems, confirming that the continuous chunk‑based flow approach can be both efficient and effective.

Constraints and trade‑offs

The paper notes that continuous representations require careful design of the chunk boundaries so that spatial information does not degrade. The sequential generation at inference introduces a dependency on prior chunks, which can slow production compared with fully parallel token‑wise decoders. Additionally, the flow matching objective depends on the quality of the source distribution; poor initialization can limit the reach of the learned vector field.

Practical takeaways

For engineers building multimodal systems, the Multimodal Flow framework suggests that a single continuous backbone can replace disjoint discrete and image‑generation pipelines. The hyperchunk organization makes it possible to reason about language and vision together without converting one to the other's native format. Models trained with flow matching show strong zero‑shot performance on vision‑language benchmarks even when trained on an order of magnitude fewer tokens than competing systems. The released code and pretrained checkpoints provide a starting point for anyone who wants to experiment with fully continuous multimodal generation.

Read the paper on arXiv