Image priors × layout priorsUnder reviewIEEE TVCG · 2026

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

Shirong Yang, Bo Yang, Ying Cao

IEEE Transactions on Visualization and Computer Graphics · 2026

A joint diffusion framework that learns how pretrained image and layout models communicate instead of asking one modality to follow the other.

PROJECT / 02Image priors × layout priors
Four wide-format examples jointly generated as background images and foreground layouts by InterIL.
Project teaser · click to inspect the full-resolution figure.

Sequential generation hides the dependency.

A background image and its foreground layout are not independent outputs: composition, empty space, visual saliency, and semantic intent constrain both. Sequential pipelines generate one modality first and force the second to adapt to a decision that can no longer change.

InterIL models the two modalities in a single generative process. The aim is not to replace strong pretrained priors, but to teach them when and what to exchange while their predictions are still evolving.

InterIL architecture with frozen image and layout backbones joined by a trainable communication module.
Figure 01Only the communication pathway is learned; the image and layout backbones remain frozen and exchange information throughout denoising.

Freeze the priors; learn the interaction.

InterIL connects frozen image and layout diffusion backbones with a lightweight bidirectional communication module. The module is trained to move useful signals across the two denoising trajectories while retaining the knowledge already encoded by each backbone.

At inference time, guidance can adjust the balance between image fidelity and layout preference without retraining the two generators.

InterIL examples showing coordinated image and layout generation across multiple design prompts.
Figure 02Joint samples illustrate how image composition and foreground regions can evolve together across different content and aspect requirements.

What this project adds.

  1. 01

    Introduces a joint image-layout generation paradigm for design templates rather than a fixed sequential pipeline.

  2. 02

    Learns explicit bidirectional interaction between frozen pretrained image and layout diffusion priors.

  3. 03

    Supports inference-time guidance for different image-layout preferences without updating the backbones.

Under review

Public arXiv preprint; manuscript under review at IEEE TVCG.

This page describes ongoing research and does not imply acceptance or publication.

Next project / 03Prior-Informed Design: Text-to-Design Generation with Multimodal Priors