The Repurposing Problem in ML Pipeline Figures
Machine learning papers rely on pipeline figures to communicate model architectures, data flows, and training procedures. A single diagram must serve multiple output formats: two-column conference pages, 16:9 presentation slides, portrait research posters, 1:1 social media teasers, and 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph. When a connection breaks silently during resizing, the figure no longer represents the actual method. Researchers currently handle this by manually redrawing diagrams for each canvas, a tedious process that introduces errors.
The core difficulty is that flowchart relayout is not simple image resizing. The spatial arrangement of blocks and edges encodes semantic relationships. A block stretched to fit a wide canvas may obscure its label. An edge rerouted automatically may connect the wrong components. Existing automation approaches fail in predictable ways: image-to-image models stretch content and reject extreme ratios, text-to-image systems invent blocks that were not in the source, and parse-then-render tools misroute edges. This paper formulates aspect-ratio-adaptive flowchart relayout as a distinct task and proposes an agentic pipeline to solve it.
Why Existing Methods Fall Short
Three families of prior work address scientific figure generation, but each fails at relayout for specific reasons. Image-to-image models like Nano Banana Pro and GPT Image 2.0 operate in pixel space. They uniformly scale the entire raster, which distorts embedded text and block proportions. They also refuse aspect ratios outside approximately 3:1 to 1:3, making them unusable for phone or poster formats.
Text-to-image agentic systems such as PaperBanana and SciFig take a different approach. They use a planner agent to interpret a methodology description, then generate a new figure from scratch. This introduces hallucination: blocks present in the source disappear, and labels not in the source appear. The generated figure reflects the agent's understanding of the method, not the actual source diagram.
Parse-then-render systems like AutoFigure-Edit attempt to recover structure from the raster, then re-render it. They extract blocks and edges but lose the precise connectivity information during parsing. When laying out for a new aspect ratio, edges connect to incorrect blocks or cross in misleading ways. The output may look plausible but misrepresents the computational graph.
None of these approaches produces an editable output in a standard diagram format. Researchers cannot open the result in draw.io to fix a misplaced label or adjust a connection. The paper argues that relayout requires all three properties simultaneously: structural fidelity to the source graph, zero hallucination, and arbitrary aspect ratio support with editable output.
The Agentic Pipeline Architecture
The proposed method decomposes the task into three stages: Parse, Style, and Layout. Each stage pairs a main agent with a critic agent. The critic combines deterministic constraint checks with vision-language model (VLM) visual feedback. This dual verification ensures connectivity is explicitly validated rather than assumed correct.
Parse Stage: Raster to Structured XML
The Parse stage converts the input raster flowchart into mxGraph XML. The main agent receives the source image plus spatial priors from SAM 3 (Segment Anything Model 3), which provides instance-level bounding boxes, category predictions, and spatial attributes for each component. This grounding reduces the VLM's spatial reasoning errors. Embedded images and icons are cropped and stored as assets with XML placeholders for later reinjection.
The Parse Critic compares the rendered XML output against the source image. It checks connector logic, structural consistency, and visual correspondence. If discrepancies exist, it issues correction instructions. The iteration continues until the critic passes or a maximum iteration limit is reached (three iterations for Parse). The output is structured XML encoding every block, edge, and group with explicit source and target references for each connector.
Style Stage: Visual Fidelity Without VLM Color Hallucination
The Style stage refines the parsed XML to match the visual appearance of the source. A critical design decision here: VLMs are unreliable at color perception. Instead of asking the agent to infer colors, a deterministic PIL-based Color Extraction Tool samples dominant colors from the source pixels at the block positions recorded in the XML. The Style Agent then applies these exact colors to fill, stroke, text, and arrow properties while keeping block text, edge references, and group membership immutable.
The Style Critic verifies visual correspondence by comparing the rendered styled XML against the source, using the extracted color assignments as ground truth. It runs for up to two iterations. This separation of color extraction from style application prevents the common failure where a VLM assigns plausible but incorrect colors to blocks.
Layout Stage: Directed-Graph Relayout with Connectivity Guarantees
The Layout stage rearranges the styled graph for the target aspect ratio. The Layout Agent extracts the directed graph from the XML and performs hierarchical relayout using a DAG-based layering algorithm. It processes containers first, then arranges components within and across containers to preserve connectivity and keep related blocks spatially proximate.
The Layout Critic is the most complex. It runs deterministic checks for three constraint categories: no same-level components overlap, all original edge relationships are preserved (source and target blocks still connected in the same direction), and all components remain within canvas boundaries. It also performs visual comparison between the rendered layout and the source to catch routing issues that deterministic checks miss. This stage iterates up to five times, the highest budget in the pipeline, reflecting the difficulty of layout under arbitrary aspect ratios.
Editable Output Representation
All intermediate and final representations use draw.io-compatible mxGraph XML. This format explicitly encodes blocks, edges, and group hierarchies. Each edge stores source and target block references, so connectors update automatically when blocks move. Researchers can open the output directly in draw.io and edit fonts, borders, colors, and positions while connector relationships remain intact. This is a practical advantage over SVG or raster outputs where manual edits break structure.
FlowchartRelayoutBench: Benchmark Design
The authors curated 100 flowcharts from oral-level papers at top-tier venues (ECCV, CVPR, NeurIPS, ICLR, ICCV, ICML) over three years. The set spans three structural complexity tiers: 25 easy, 40 medium, 35 hard. Average complexity is 15.5 nodes and 13.6 edges per figure, increasing with tier. Each source is evaluated at five target ratios (9:16, 2:3, 1:1, 3:2, 16:9), yielding 500 relayout tasks.
Evaluation uses a four-metric VLM-as-a-judge protocol with Gemini 3.1 Pro, validated against human judgments. Relationship Preservation checks if connector logic (source, target, direction) matches the source. Hallucination-free Rate verifies no elements were added or omitted. Layout Quality ranks candidates by canvas utilization and arrangement quality, penalizing trivial scaling or rotation. Style Similarity ranks by preservation of block colors, text styles, and connector styles. Hierarchical scoring prioritizes content correctness: candidates are first ranked by passed Content Fidelity checks (Relationship + Hallucination), then by Visual Quality.
Experimental Results
The full pipeline achieves 68.6% Content Fidelity (both Relationship Preservation and Hallucination-free Rate pass), versus 11.2% for PaperBanana, 24.8% for AutoFigure-Edit, 41.4% for Nano Banana Pro, and 40.2% for GPT Image 2.0. Relationship Preservation alone reaches 74.2%, Hallucination-free Rate 90.6%. On Visual Quality, the method ranks competitively: Layout Quality and Style Similarity scores place it near the top despite the XML reconstruction introducing a modest style gap compared to pixel-preserving image-to-image models.
The Overall hierarchical score is 81.85, 1.15x the next best (Nano Banana Pro and GPT Image 2.0 at 71.10). A user study with 30 participants evaluating 250 outputs confirms the VLM judgments: the method wins on Relationship Preservation and Hallucination-free Rate, with strong pairwise win rates on Layout Quality. Image-to-image models score higher on Style Similarity due to direct pixel preservation, but users rate the editable XML output as visually competitive.
Agreement between VLM and human judgments is significant across all metrics: Phi coefficient 0.671 for Relationship, 0.593 for Hallucination, Spearman rank correlation 0.620 for Layout, 0.757 for Style (all p<0.001).
Ablation Study Findings
Removing all critic agents drops Content Fidelity from 76.7% to 46.7% and Overall from 87.04 to 61.85 on a 30-figure subset. Removing the Parse Critic or Layout Critic individually also degrades Content Fidelity. Removing the Style Stage primarily hurts Style Similarity and lowers Overall to 76.30. These results confirm the complementary roles of stage-specific critics.
Replacing the Layout stage with Graphviz DOT reduces Overall to 35.19. Conventional graph layout lacks target-canvas awareness, often leaving large unused areas and failing to preserve container-structure relationships. Direct prompting without multi-stage decomposition is far worse: one-shot VLM-to-XML generation achieves only 33.3% Content Fidelity with Gemini 3.1 Pro and 30.0% with GPT-5.5. Even direct prompting of individual stages underperforms the full pipeline, confirming that task-specific decomposition with iterative critique is necessary.
Limitations and Trade-offs
The authors acknowledge several limitations. Style fidelity has a persistent gap: XML reconstruction cannot perfectly replicate raster typography and border rendering. The method depends on Gemini 3.1 Pro Preview, a closed-source VLM, limiting reproducibility for teams without API access. Convergence is not guaranteed; some cases hit the iteration limit without passing the critic. Edge routing failures persist in highly complex diagrams with overlapping or ambiguous connections, particularly in the hard complexity tier. Runtime and API cost scale with iteration count, which may be prohibitive for large diagrams or production workflows.
Practical Implications for Developers
For a researcher preparing a paper, this pipeline means drawing a pipeline figure once and getting valid, editable outputs for every target canvas. The mxGraph XML opens in draw.io where any element can be adjusted without regenerating the diagram. For tool builders, the structured XML provides a lossless interchange format that preserves connectivity semantics. The benchmark and evaluation protocol give the community a standardized way to measure progress on aspect-ratio-adaptive relayout.
The critic architecture demonstrates a pattern applicable beyond flowcharts: pair generative agents with critics that combine deterministic verification (cheap, exhaustive, guaranteed) with VLM visual feedback (expensive, semantic, flexible). This catches errors that either approach alone would miss. The deterministic color extraction in the Style stage shows how to bypass known VLM weaknesses by offloading specific perceptions to classical computer vision.
Where This Leaves the Field
Aspect-ratio-adaptive flowchart relayout is now a formulated task with a benchmark and a working solution that outperforms prior work by a wide margin on structural fidelity. The 68.6% Content Fidelity leaves room for improvement, particularly on hard diagrams. The closed-source VLM dependency and style gap are the most tractable next targets. Open-weight VLMs could replace Gemini 3.1 Pro with some performance trade-off. Learned color and typography transfer could narrow the style gap. The critic framework itself is extensible to other diagram types where connectivity semantics matter: architecture diagrams, circuit schematics, and process flows.
Read the paper on arXiv