Most AI image generators want a sentence. Whisk wanted three pictures. You supplied a subject, a scene and a style as separate reference images, and it produced something that borrowed from each — sidestepping prompt-writing entirely.
That trade had a cost as well as a benefit, and this guide covers both. Whisk was retired on 30 April 2026 and folded into Flow; the mechanic survived the move, the interface did not.
What Whisk actually was
Whisk was an experiment from Google Labs built around image-based prompting. Instead of describing what you wanted, you dropped in photos defining three components, and it blended them into something new.
It was never built for pixel-accurate editing. It was built for fast visual exploration — roughing out concepts for stickers, pins, merchandise, mood boards, anything where you wanted twenty variations rather than one exact result. Judged as a precision tool it disappoints. Judged as a way to think out loud in pictures, it made sense.
It was free, ran in the browser, and needed only a Google account. Availability was restricted — for much of its life it was US-only, which is worth knowing if you read older tutorials and wondered why you could never open it.
The three inputs
- Subject — the main focus: a person, animal, object, vehicle, or any central element. This is what the image is of.
- Scene — the environment the subject sits in, from landscapes and streets to abstract or invented spaces. This is where it is.
- Style — the artistic treatment, palette and mood. This is what it looks like.
Separating the three is what made Whisk different. In a text generator, all three are tangled in one sentence, and changing the style often drags the subject along with it. Here you could swap one input and hold the other two still.
Why that differed from DALL-E, Midjourney and Stable Diffusion
Text-based tools ask you to convert a mental picture into words, then hope the model converts it back the same way. Every step in that round trip loses something. Describing "a vintage motorcycle in a cyberpunk cityscape, impressionist" is a lot of work, and the result frequently is not what you pictured.
Whisk removed the round trip. You showed it a motorcycle, showed it a city, showed it a painting.
The cost is control. Words are precise in a way reference images are not. A prompt can specify "three people, left-handed, at dusk"; a photo cannot be argued with. So the comparison is not about which produces better images — it is about whether your bottleneck is describing things or finding things. There is a fuller comparison here.
Choosing reference images
This is where results were won or lost. The interface was simple; the inputs were everything.
Subject images work best with sharp focus, even lighting, and clear separation from the background. Busy backgrounds, several competing objects, or heavy filtering all confuse the extraction. For people, a neutral pose with clear features beats an action shot. Clean product photography is close to ideal.
Scene images need to be interesting without being loud. Beaches, forests, mountains, tidy interiors and uncluttered streets all work. Scenes packed with small detail or lit from several conflicting directions tend to fight the subject rather than hold it.
Style images should commit to one look. A single art movement, one photographic treatment, one consistent grade. A reference that mixes several aesthetics produces a muddle, because there is no single signal to copy.
The most reliable predictor of a good result was whether the three references shared compatible lighting and colour. Three images that already look like they belong in the same world combine well. Three that do not, will not.
The workflow
Upload the subject first and read the description Whisk generates back at you. That text is the tell — it shows what the model thinks it is looking at. If the description is wrong, the output will be too, and no amount of iterating on the other inputs will rescue it.
Repeat for the scene, then the style, checking each interpretation as you go. There was also an "Inspire Me" option that filled slots with suggestions, useful when you were stuck rather than when you had something specific in mind.
Before generating, look at all three together and ask whether they belong in one picture. Then generate. Output typically took somewhere in the region of 30 to 90 seconds depending on load.
You could add optional text guidance alongside the images — "the robot is running", "use a pastel palette" — which was the escape hatch for the things pictures cannot say.
What to expect from the results
Whisk prioritised the essence of your references over faithful reproduction. Exact height, build, hairstyle and skin tone were frequently approximated rather than reproduced. That was a design decision, not a defect: the tool extracted characteristics in order to remix them.
It does mean Whisk was a poor choice whenever a specific person or product had to look exactly like itself. When a result was close but wrong, a Refine mode let you adjust, and you could edit the underlying prompt that Gemini had written from your images — which was often faster than swapping references.
Where this leaves you now
The interface described above no longer exists. In Flow, the same idea appears as Ingredients — reference images that hold a subject consistent across shots — though it is now pointed at video rather than single stills, and image generation runs on Nano Banana rather than Imagen 3.
What carries over is the judgement, not the buttons. Pick clean references. Check what the model says it sees before you generate. Keep lighting and colour compatible across inputs. Accept approximation, or use a different tool. None of that changed.
The full breakdown of the move is here, including what happened to libraries and credits.