Multimodal priors × editable designUnder reviewACM TOG · 2026

Prior-Informed Design: Text-to-Design Generation with Multimodal Priors

Shirong Yang, Ying Cao

ACM Transactions on Graphics · 2026

A data-efficient text-to-design framework that starts from a multimodal prior and learns to correct it in an editable design latent space.

PROJECT / 03Multimodal priors × editable design
A grid of layered graphic designs generated by PriDe across square, landscape, and portrait formats.
Project teaser · click to inspect the full-resolution figure.

Structured design data is the scarce resource.

A finished graphic design contains far more than a flattened image: editable layers, geometry, typography, visual assets, and semantic relationships must all remain coherent. Large collections with this full structure are difficult to obtain.

PriDe asks how much design knowledge can be transferred from models trained on abundant image and language data. It uses frozen text-to-image and language priors to make structured text-to-design learning possible with 10K layered examples.

PriDe shared layout VAE encoding editable design elements into a token-level latent representation.
Figure 01The shared design latent preserves element category, geometry, order, typography, imagery, SVG information, and canvas attributes instead of collapsing the artifact into pixels.

Model the correction, not the whole design from zero.

The system first represents each layered design in a shared token-level latent space. A multimodal prior proposes useful imagery and semantic structure; the trainable generator then models a residual that transforms that prior into a coherent, editable target design.

Intermediate prior features are injected into a lightweight generator so the learned component can focus on graphic-design-specific organization rather than relearning broad visual and linguistic knowledge.

Examples comparing multimodal prior designs with final structured outputs from PriDe.
Figure 02PriDe learns a design-specific residual: it keeps useful content from the multimodal prior while correcting composition, hierarchy, and typography in the final editable design.

What this project adds.

  1. 01

    Studies data efficiency as a first-class problem for structured graphic design generation.

  2. 02

    Combines frozen text-to-image and language-model priors in a prior-informed text-to-design framework.

  3. 03

    Uses residual generative modeling and feature injection to produce layered designs with only 10K training examples.

Under review

Manuscript under review; no public preprint or code release yet.

This page describes ongoing research and does not imply acceptance or publication.

Next project / 01Content-Edited Layout Generation for Graphic Design