Edit one image with several visual references

One picture rarely contains every instruction an image edit needs. The product photograph has the exact bottle shape. A campaign still has the lighting. A sketch has the camera angle. A brand sheet has the colours. Asking an editor to infer all four from one sentence turns a precise art direction into a lottery, while flattening the references into a collage makes the model guess where one source ends and the next begins.
Creator lets those pictures stay separate. Attach them as a reference deck, say what each one contributes, and run one edit. The newest reference is the picture being edited; the earlier ones travel as additional visual context. More importantly, the finished turn records the whole deck, so reopening it from History still shows what the model actually saw.
Start by assigning each reference one job
A multi-reference edit works best when the sources do not compete. Pick one image for identity or geometry, another for light, another for composition, and say those roles in the instruction. Keep the bottle, use the second image's amber rim light, and frame it like the third image is much easier to settle than combine these. The latter asks the model to make three decisions you already made.
The ordering has one concrete consequence in Creator. The newest attached image is primary: it is the source sent as the image being edited. Every earlier image is additional context. If the shape that must survive is in the first card and the mood board is last, the model is being asked to edit the mood board into the product. Reorder by removing and re-attaching, or simply attach the identity image last.
This is a complement to multi-step image editing, not a replacement. Use several references when one operation needs several inputs; use a chain when you want to approve one change before asking for the next.
Build the deck without stopping to file things
References can arrive through the file picker, the shared asset library, the Cast picker, a clipboard paste, or a drag anywhere onto the Creator page. Images and videos are accepted by the composer; a video reference is reduced to the frame actually sent when the step is recorded, so History shows the visual input rather than pretending the whole clip was inspected frame by frame.
The transfer starts as soon as a reference is attached, while you are still writing. Pressing Send usually posts a staged upload id instead of putting the entire file transfer between the click and the first progress update. Remove the card before sending and the staged upload is cancelled; send it and the card moves out of the composer into the user turn. The next prompt begins with an empty deck because references belong to the instruction that consumed them, not to every idea that follows.
The small deck always exposes the latest card and expands to show the complete set. Each card keeps its own aspect instead of being forced into a square, which matters when the difference between a portrait pose and a landscape composition is part of the instruction.
Pick a model that can accept the deck
Attaching the first image changes the job from text-to-image to image edit, and the model picker changes with it. A text-only image generator is not offered for a request that will go to the edit endpoint. Remove the last reference and the job becomes image generation again. If a model selected a moment ago no longer matches the job, Creator returns the picker to Auto image model instead of waiting for the server to reject the request after you press Send.
Reference count matters too. Models that publish a max_refs ceiling disappear when the deck exceeds it, and an already selected model is demoted to Auto at that moment. A model with no declared ceiling stays visible: absence means unknown, not incompatible, so the picker does not erase useful options on the strength of missing catalogue data. The same filtering applies to Try again on a failed card and Vary on a finished one; each re-run gets a model that fits its own inputs.
If the short list does not contain the editor you need, Add more to the list… opens the account catalogue directly on image-edit models. The choice returns to Creator through the same browser, and the picker sees it without a reload.
Write an instruction that separates identity from style
The useful prompt names what must be preserved before it names what should change. That order does not create a magic weighting rule, but it makes the acceptance test readable to both you and the model. A good product-edit instruction might be: Use the newest image as the exact bottle and label. Take only the low amber side light from reference one and the waist-high camera angle from reference two. Keep the white sweep background. Do not copy the props or typography from either reference.
That last sentence is as important as the positive direction. A style reference contains accidental facts — a red table, a hand, a headline — and without boundaries those can leak into the output. Say light only, pose only or palette only. When a face or character must remain stable, say which card owns identity and which details are non-negotiable. Creator templates and batches can save that structure as a repeatable starting point when a team makes the same kind of edit every week.
A worked example: one product, two art directions
Start with a clean phone photograph of a ceramic coffee bag. Attach a magazine still with hard blue light first, then a sketch of a low three-quarter composition, and attach the product photograph last so it becomes primary. Write: Keep the exact bag shape, label and cream paper texture from the newest image. Use only the hard blue edge light from the first reference. Use the low three-quarter framing and empty right-hand space from the sketch. Put the bag on dark slate; do not include the sketch's cup or handwritten notes.
Set the result count to four. One submit creates four sibling edits from the same deck, letting you compare how strongly each take followed the light without rebuilding the instruction four times. The live estimate reflects the four generations before the request is sent. Keep the closest result, open it, and use Vary with a different compatible edit model if the composition is right but the material is not.
For a second campaign direction, do not overwrite the first. Return to the same primary photograph, replace only the lighting reference, and send a new turn. Both routes remain visible, which is the practical benefit of reproducible creative workflows: the rejected art direction is still evidence, not a file named final-old-2.png.
Every source comes back when the thread reopens
Multi-image input is only useful if the record does not quietly simplify it afterwards. Creator stores the primary input beside the step as {step}-input.png. Each additional image is stored as {step}-input-{n}.png, and their ordered names ride on the step's detail record. Reopening the session returns the extra URLs first and the primary last — the same order the composer showed before Send.
That sounds like housekeeping until a client asks why the label changed. A history view showing one of three inputs would make the result look as though the model invented the other two. The complete turn now carries the instruction, model, result and every visual source. Old single-reference steps have no extra field and reopen unchanged rather than failing a newer reader.
The storage is per step by design. A four-take batch records its inputs on each sibling, so any card can survive on its own after the others are deleted or excluded. That costs more storage than a pointer shared between siblings, but it avoids reference-counted deletion and keeps a node self-describing. It does not add generation credits; only the model calls do.
Vary means vary the same edit, not generate a stranger
A variation must run on the picture the original result was made from. Sending only the old words would turn an edit into a fresh text-to-image request, which is how a preserved face becomes a different person while the button still says Vary. Creator re-sends the recorded reference set and hangs the new result under the card it came from.
The same rule holds after a reload. If a reference exists only as a stored URL, Creator fetches it back into a file the first time a re-run needs it and caches that file for a second send. A missing source fails loudly with That reference is no longer available instead of silently dropping into text-only generation. This is the difference between provenance as a caption and provenance as working material: the record can actually be used to repeat the operation.
Compared with a collage or a giant prompt
A collage is convenient when the model accepts only one image, but it destroys useful structure. Every source is rescaled, borders become visual content, and the model has to infer whether the upper-left panel owns the face, the palette or both. Separate references preserve their original resolution and let the instruction name their roles. The trade-off is model compatibility: a declared one-reference model is available for the collage and correctly unavailable for a three-card deck.
A giant text prompt avoids that compatibility problem but asks language to describe facts an image already states exactly. Warm cinematic lighting cannot specify the precise falloff you liked, and a paragraph describing a person's jacket is weaker identity evidence than the photograph. Use words for decisions and images for appearance.
Exporting every source to another editor works too, but the record splits across tools. The shared-library workflow described in an AI workflow with no export step keeps the sources, edit and later handoff on one shelf.
The limits are deliberate
The server accepts attachments up to 100 MB on this path and bounds the multi-image provider call. The primary source is always first-class; at most eight additional decodable images ride beside it as visual inputs. Anything beyond that does not expand the provider call indefinitely. In practice, a deck of two to four purposeful references is stronger than nine sources with overlapping roles.
Not every runtime can carry several pictures. A single-source runtime receives the primary and is not told in the prompt that it saw images it never received. Catalogue filtering prevents known count mismatches; Auto routing is the safest default when compatibility is uncertain. Cost follows the generations, not the attachment count: four takes are four edits, while saving and reopening their sources uses no credits. Lowest-cost model routing is useful when the deck matters more than the brand name on the model.
Frequently asked questions
Which reference image is treated as the main source?
The newest card in the Creator deck is primary: it is the image the request edits. Earlier cards are additional visual context. If identity, geometry or exact typography must survive, attach that source last and describe the earlier cards with narrow roles such as lighting, pose or palette. The ordering is visible before Send and restored in the same order from History.
How many image references can I use in one edit?
The server bounds the provider input at the primary image plus eight additional decodable pictures. Individual models may publish a smaller max_refs limit; Creator filters those models as the deck grows and returns an incompatible selection to Auto. More is rarely better: two or three sources with explicit, non-overlapping jobs usually produce a more controllable result than a crowded mood board.
Do reference images remain after I close Creator?
Yes. The turn stores the primary input and every additional visual input beside the resulting step. Reopen the session from History and the complete deck appears above the instruction in its original order. This is also what allows Vary or a retry to resend the actual sources rather than rebuilding the edit from words alone.
Can I use a video as a reference?
Yes. Creator accepts image or video reference media. The edit path works from a representative frame for a video source, and the step records the frame that was actually sent as an image. That makes History honest about the model's visual input: it shows the inspected frame, not a whole clip the image editor never consumed frame by frame.
Does adding more references cost more credits?
Credits are charged for generation requests, not for storing reference cards. A single edit with several compatible inputs is one generation; setting the result count to four creates four paid sibling edits. The composer shows the estimate before Send. Keeping the deck in History and reopening it later costs no generation credits, though each sibling stores its own copy of the inputs.
Open Creator, attach the mood and composition first, attach the image that must survive last, and give every card one job.


