Grounding Executable Visual Semantics in Unified Multimodal Models

Published:

VLMs that emit graphics code (TikZ, manim, SVG, matplotlib, slide DSLs) treat the program purely as text: the model never observes what its own code renders to, and any signal about visual fidelity must arrive through a text-only proxy. Unified models — architectures that natively generate both text and image tokens in a single model — remove this restriction, since the rendered output can sit in the same training sequence as the program that produced it. We propose to test whether a continued mid-training stage on (code, rendered image) sequences, initialized from the released BAGEL 7B checkpoint, improves a unified model’s ability to (i) generate code that renders to a target image, (ii) predict what a given program will render to, and (iii) reason about images by referring to their underlying executable code structure (e.g., using a program as an intermediate representation when inputting an image). The mid-training corpus mixes prompt-conditioned code-image pairs (<prompt><code><image>), render traces (e.g. <prompt><code chunk 1: draw (0,0) circle (1cm);><image A: one circle><code chunk 2: draw (2,0) circle (1cm);><image B: two circles> <code chunk 3: node at (1,0) hello;><image C: two circles + hello>), and prompt-free <code>, <image> & <image>,<code> pairs. We use TikZ as the primary domain for its determinism and existing data scale, extend to SVG, matplotlib, manim, and slide DSLs to test cross-DSL transfer, and follow mid-training with a renderer-as-verifier RL stage. Throughout, we compare against matched-compute baselines: the same BAGEL checkpoint continued on the same code, but with the loss on rendered images masked out, so the model sees the same data but receives no supervision on producing images.