Best AI image editing models with reference images

Runbo Li
Runbo Li
·
· 9 min read
Best AI Image Editing Models With Reference Images

Quick answer

For AI image editing with reference images, start with GPT Image, Nano Banana, FLUX.2 or MAI-Image-2.6 when you need to change an image through instructions. Use Photoshop or Midjourney’s current Editor when selecting the area to change is central to the job. For a quick browser workflow, Magic Hour’s AI Image Editor lets you upload a source image and describe the edit.

The key question is what must stay unchanged: a person’s face, a product’s shape, the background, or the visual style. Those requirements call for different controls. A reference image guides a model; it does not guarantee that every logo, facial feature or label survives generation.

Edit an image with a current AI model

Upload an image, describe the exact change, and compare the result against your source in Magic Hour.

Try AI Image Editor

Magic Hour publishes this guide and includes its own product in the comparison. Treat the recommendations as editorial guidance from a vendor, and verify the linked first-party product details and your own output requirements before choosing a tool.

Best reference-image editing options compared

Model or workflow

Why to evaluate it

Reference or editing control

Important distinction

Magic Hour AI Image Editor

Start an edit in a browser and continue into other media tools

Source image plus a prompt; model-specific options in the full workflow

Magic Hour is a platform, not a single model or one universal input limit

OpenAI GPT Image 2.5

Instruction-based editing and repeated revisions

Image inputs, editing instructions and supported masking workflows

ChatGPT Images 2.5 and API Flare/Sunburst are separate product surfaces

Google Nano Banana family

Combine reference subjects, objects and scenes

Multiple image inputs with limits specific to each named model

Nano Banana 2, Pro and Lite have different reference capabilities

Black Forest Labs FLUX.2

Combine multiple sources with explicit image roles

Multi-reference editing; documented API and playground limits differ

Select the exact variant and host before assuming its controls or price

Microsoft MAI-Image-2.6

Evaluate another current model for image editing

Image edits and multi-reference capabilities; Foundry access is in preview

Check the deployed model’s endpoint, output limits and quota

Midjourney Edit Model

Combine prompt edits with masks, layers and style references

Current V8 editing workflow supports multiple references and local edits

Older descriptions of Midjourney as only a style-reference tool are incomplete

Photoshop with Firefly Fill & Expand

Replace or place a reference object in a selected area

Selection, reference image, Object/Whole image and Swap/Place intent

Photoshop’s editing workflow is different from every model inside Firefly

Stable Diffusion with compatible adapters

Build a custom pipeline around pose, edges or image conditioning

IP-Adapter, ControlNet and inpainting, depending on the pipeline

Requires compatible models and setup; not one guaranteed-fidelity preset

This guide was reviewed on September 13, 2026. It compares documented capabilities and practical selection criteria. Magic Hour publishes it and is included; the table does not assign invented image-quality scores.

For a separate quantitative view, use the dated AI image editing model leaderboard. It combines several public benchmark sources; use this guide for workflow fit and the leaderboard for source-by-source model scores.

Choose the reference control that matches the job

What you need to preserve

Start with

Why it matters

Most of the original photograph

A local selection, mask or manual composite

Limits the area you ask the model to reinterpret

A person or product across a new scene

Subject references with a clear description of what may change

Gives the model identity cues while allowing a new composition

Color, lighting and visual treatment

A style reference

Style alone does not specify the exact person or product

Pose, depth or layout

Structural references or a compatible control pipeline

Provides information that a descriptive prompt may leave ambiguous

A label, logo or other exact text

Original pixels or an editable text layer

Avoids depending on generation to reproduce legally or commercially important details

For example, changing the background behind a bottle is different from creating a new scene that happens to contain a similar bottle. If the actual packaging must remain exact, preserve the product cutout and build the scene around it.

1. Magic Hour: a practical browser starting point

Magic Hour’s image editor takes an uploaded image and a description of the change. Its product page lists editing options and models, including Qwen Edit, GPT Image 2, Nano Banana models and FLUX.2 Klein. Check the current selector for the model and settings available to your account.

The public tool offers three guest edits per day without signup and watermark-free image downloads. Free outputs are for personal use; paid plans permit commercial use. A free browser edit does not establish the same model access, reference count or allowance as a paid or API workflow.

Try a precise instruction: “Remove the chair on the left. Keep the person, camera angle, wall color and remaining furniture unchanged.” Compare the output with the original, especially around the edit boundary.

Choose Magic Hour when the workflow may continue into image-to-video after the image is approved. For an integration, use the model reference rather than assuming every available model accepts the same inputs.

2. OpenAI GPT Image: conversational and instruction-based edits

OpenAI introduced ChatGPT Images 2.5 on September 8, 2026. ChatGPT adds sketch, template and image-comment workflows, while the API offers GPT-Image-2.5 Flare and Sunburst. These are options to evaluate for focused changes and successive revisions.

The image-generation guide documents image editing, image inputs and masking. Confirm the selected model’s supported controls and dimensions. A ChatGPT subscription does not pay for separate API requests.

Describe the requested change and the features to preserve. If you are combining images, explain each image’s role: source subject, target scene, style or composition. Avoid treating a reference upload as an implicit instruction to copy every element in it.

OpenAI’s launch describes improved editing fidelity, but that is a provider claim, not evidence that every product label or person will remain exact in your files. Inspect the details that matter to the final use.

3. Nano Banana: match the exact model to the references

Google’s Gemini image guide distinguishes Nano Banana 2, Nano Banana Pro, Nano Banana 2 Lite and the original Nano Banana. Their reference capabilities differ.

The current guide permits up to 14 reference images across the Gemini 3 image models, with different categories of supported references. It describes Nano Banana 2, or Gemini 3.1 Flash Image, with up to 10 object references and four character references. Nano Banana Pro, or Gemini 3 Pro Image, lists up to six object references, five character references and three style references within the 14-image total. Do not transfer those figures to every Nano Banana interface or variant.

Use this family when you need to specify several source elements. Name their roles clearly: “Use the person from image 1, the jacket from image 2 and the setting from image 3.” Start with only the references needed to remove ambiguity; extra images can introduce conflicting cues.

For a single product, one clear view plus a close view of an important detail can be more useful than a large set of unrelated inspiration images. The downloaded result still needs inspection.

4. FLUX.2: explicit multi-reference composition

Black Forest Labs’ FLUX.2 editing documentation describes text-guided edits and combining multiple source images. It distinguishes up to eight references through the documented API workflow from up to ten in the playground, with variant-specific capabilities.

The family includes max, pro, flex and klein options. Choose the specific model and host before comparing costs or assuming the same controls are available. A capability in Black Forest Labs’ playground is not automatically exposed by every app that offers FLUX.

A useful brief assigns one role to each source: “Keep the room layout from image 1. Replace only the sofa with the sofa in image 2. Match the room’s existing light and perspective.” Review the sofa’s proportions and the surrounding floor after the edit.

This is a practical option for multi-source scene building. It should not be described as universally better or worse at identity preservation without a comparison using the same inputs and requirements.

5. MAI-Image-2.6: another current model to evaluate

Microsoft made MAI-Image-2.6 and the faster-oriented Flash variant available in Microsoft Foundry public preview on September 4, 2026. The announcement describes multi-image reference editing, web grounding and dynamic aspect ratios.

The Foundry documentation provides the actual generation and edit endpoints, input requirements, output format and deployment quotas. Follow those endpoint details rather than assuming a feature’s announcement describes every available client or setting.

Include MAI-Image-2.6 in a model evaluation if you are choosing an API for reference edits. Preview availability and output limits matter alongside visual quality. There is no reason to choose it solely because a general leaderboard places it ahead by a few points.

What the independent benchmark does—and does not—show

The Artificial Analysis editing leaderboard ranks results through blind user preferences for edits to the same input and instruction. In the September 10 snapshot, MAI-Image-2.6 scored 1122 and GPT Image 2 at high quality scored 1117, with overlapping reported 95% confidence intervals. That supports including both in a shortlist; it does not establish a decisive quality difference or perfect reference fidelity.

That GPT Image 2 result is not an evaluation of the newly released GPT Image 2.5 models. Aggregate preference also does not tell you whether a particular bottle label, facial feature or catalog requirement will pass review.

6. Midjourney: use its current editing workflow

Midjourney’s August 27 Edit Model announcement introduced a V8.2 editing workflow with instruction-based edits, up to four image references, inpainting and outpainting. It also describes personalization, moodboards and style-reference support.

The current web Editor supports uploaded images, selections and layers. This makes older “Discord-only” or “style reference only” descriptions misleading. Older Omni Reference instructions are not a substitute for checking the current Edit Model workflow.

Choose it when you want to combine art direction with direct image edits. Use a selection for a local change and a reference for the specific visual information you want to introduce. Review the result rather than assuming that reference weight alone locks identity.

7. Photoshop: select the area, then specify the reference

Adobe’s Reference Image instructions give a concrete workflow: create a selection, choose Generative Fill, add a reference image and select Firefly Fill & Expand.

The controls distinguish an Object reference from a Whole image reference, and swapping the selected area from placing an object into it. Those choices help clarify whether you want to replace something or add a new item while retaining the setting.

Use Photoshop when selections, layers and manual repair are essential to the deliverable. For exact packaging, keep the original label or product layer available instead of repeatedly asking generation to recreate it. Adobe’s partner-model workflows have their own controls, so do not assume every Firefly option behaves identically.

8. Stable Diffusion with adapters: a custom control pipeline

For technical users, a compatible Stable Diffusion pipeline can combine image conditioning, structural controls and inpainting. Hugging Face’s IP-Adapter documentation explains image-conditioned generation and combinations with ControlNet.

IP-Adapter uses image information to guide a result; ControlNet can add a structural condition such as pose, edges or depth. The exact combination depends on the base model, adapter, pipeline and weights. An adapter built for one model family is not automatically compatible with another.

Choose this route when the control requirement justifies setup and maintenance. Keep the model versions and settings with the project. Self-hosting also requires suitable hardware and applicable model licenses; it is not automatically the cheapest way to finish a small batch.

A reference-editing brief you can reuse

Use this structure to make the task clear:

Source: which image is the starting photograph? References: what should each additional image contribute? Change: what is the one main edit? Preserve: which details must remain? Output: what dimensions, format and intended use are required?

Image 1 is the original backpack photo. Image 2 is a lighting reference only. Replace the background with a gray studio wall and match image 2’s soft side lighting. Preserve the backpack’s shape, pockets, straps, stitching, color and logo. Do not add accessories.

This is an example instruction, not a tested output or a guarantee. If a logo must be exact, retain it from the source image or an approved graphic asset.

Check the result before making a batch

Review area

What to compare with the source

If it fails

Face or character

Eye spacing, face shape, hair, expression and distinguishing details

Reuse the approved source; narrow the edit or switch the workflow

Product

Silhouette, proportions, seams, label and color

Preserve the product layer or cutout instead of regenerating it

Background

Lines, perspective, reflections and unwanted scene changes

Use a smaller selection or simplify conflicting references

Text

Every required letter, number and symbol

Add or restore editable text or original pixels

Delivery

Dimensions, transparency, watermark and usage conditions

Correct the export or plan requirement before producing more images

Judge the result at the size where it will be used, and zoom in on important details. If you compare models, use the same source, instruction and acceptance criteria, then record billable attempts and correction time. Our image-generation pricing guide explains how to compare the cost of accepted outputs.

For repeatable Seedream prompts and a practical workflow, use the Seedream editing guide.

Separate preservation from transformation

Direct answer: The best reference-image editor changes the requested region while preserving everything else. Evaluate instruction following and collateral damage separately.

Five-edit benchmark. Change an object, background, color, pose detail and text-adjacent region. Keep source images, prompts and output size fixed. Use portraits, products and complex scenes.

Review. Score requested change, identity and geometry preservation, text and logo integrity, edge quality, color consistency and unexpected edits outside the target area.

Decision rule. Use precise editors for localized production work. Use more generative systems when reinterpretation is acceptable. One model can excel at creative transformation while failing strict preservation.

Try the workflow: Open the matching Magic Hour tool. Use the same representative input and acceptance criteria before comparing results.

Continue learning

Continue learning: how AI image generation works, best AI image generators, and best AI image editors.

Frequently asked questions

A model’s reference support does not guarantee exact preservation. Use references for guidance and inspect the result. When a region must remain pixel-exact, preserve that region from the original and edit or composite around it.

No. A style reference guides the visual treatment, such as color, texture or lighting. A character or subject reference supplies information about a particular person or object. Some workflows combine them, but the controls and limits depend on the model.

The instruction and references may conflict, or the workflow may give more influence to one input than you intended. State each image’s role, remove unnecessary references and narrow the requested change. Use a local selection when only one area should be edited.

Start with one source image in Magic Hour’s AI Image Editor. For product work, use the product-photo workflow to decide which details should be preserved manually before creating variations or turning the result into video.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Illustrated denoising sequence resolving colored noise into a flower, glass vase, and architectural scene
Recommended next
How AI image generation works: a simple guide

Learn how AI image generators turn prompts and reference images into pictures, how training differs from generation, and why every output needs review.

Collage of the best AI image generator logos
10 best AI image generators: features, costs and examples
The Best AI Image Editors
6 best AI image editors in 2026: choose the right one
AI Image Gen 3
AI image generator pricing: plans, APIs and real costs
AI Batch Image Gen
8 best AI batch image editors (2026): automate bulk edits