From research to creation · Fewer missed details

The bird is there. Where is the spaghetti?

An official comparison separates a good-looking image from a fulfilled request. See how Diffusion Controller steers generation, and what remains between this research and everyday tools.

Google Research presented this work on September 29; the paper first appeared on March 7. This is a look at its evidence and method, not a newly released image feature.

In this article3 chapters

Even with bird and pasta, check the action

The base image has a bluejay but no spaghetti; the other two include pasta. Yet pasta beside a bird is not the same as a bird eating it. Our reading of the figure: composition, required objects and the requested action need separate checks. One demonstration cannot establish reliability across prompts.

Google's comparison of the pretrained model, LoRA and the proposed method: cat, bluejay and lizard prompts.
Original Google Research figure. Pretrained is the base model, LoRA a fine-tuning method, and Ours the paper's method. The middle prompt asks for a bluejay eating spaghetti; the base result contains only the bird. Enlarge to compare. · Open full-size image
Sources and further reading

It steers generation, with access requirements

One Diffusion Controller approach trains a small network to correct generation while freezing the base model. The experiments use Stable Diffusion 1.4. Even the restricted-access version needs intermediate denoising outputs; a prompt-in, image-out API is insufficient. These results do not establish support in Nano Banana or GPT Image, or a success rate for everyday requests.

Sources and further reading

Three checks for your own image

For an illustration of a dog in a red scarf watching rain by a window, check the objects, then the scarf's color, then position, gaze and weather. Liking the style does not settle those checks. Naming the missing detail makes the next attempt easier to assess than asking for a nicer image. This is our editorial suggestion, not a reproduction of the research.

An image must connect objects and actions; a multi-shot video must preserve those relationships over time. Our earlier pianist-to-violinist example follows that problem across a cut.