Prior-Informed Design: Text-to-Design Generation with Multimodal Priors
ACM Transactions on Graphics · 2026
A data-efficient text-to-design framework that starts from a multimodal prior and learns to correct it in an editable design latent space.

01 / Research problem
Structured design data is the scarce resource.
A finished graphic design contains far more than a flattened image: editable layers, geometry, typography, visual assets, and semantic relationships must all remain coherent. Large collections with this full structure are difficult to obtain.
PriDe asks how much design knowledge can be transferred from models trained on abundant image and language data. It uses frozen text-to-image and language priors to make structured text-to-design learning possible with 10K layered examples.

02 / Approach
Model the correction, not the whole design from zero.
The system first represents each layered design in a shared token-level latent space. A multimodal prior proposes useful imagery and semantic structure; the trainable generator then models a residual that transforms that prior into a coherent, editable target design.
Intermediate prior features are injected into a lightweight generator so the learned component can focus on graphic-design-specific organization rather than relearning broad visual and linguistic knowledge.

03 / At a glance
What this project adds.
- 01
Studies data efficiency as a first-class problem for structured graphic design generation.
- 02
Combines frozen text-to-image and language-model priors in a prior-informed text-to-design framework.
- 03
Uses residual generative modeling and feature injection to produce layered designs with only 10K training examples.
04 / Publication status
Under review
Manuscript under review; no public preprint or code release yet.
This page describes ongoing research and does not imply acceptance or publication.