29 open-source image generation model families to run locally in 2026

Runbo Li
Runbo Li
·
· 33 min read
29 open-source image generation models field guide — hero

Quick answer

Choose a local image model by its task, checkpoint license and hardware requirements. Base, Turbo, Edit, Dev and quantized releases can differ in inputs, quality and permitted use. Compare the exact checkpoint rather than a family name alone. Downloadable weights give deployment control, but setup, inference cost and output review remain your responsibility.

A quiet glass observatory above a cloud layer at blue hour, warm interior lights, one telescope, cinematic realistic lighting, balanced composition, fine architectural detail, no people, text, or logos.

Controlled prompt used for both outputs: A windswept alpine valley at blue hour, small illuminated observatory on a ridge, layered atmospheric depth, realistic rock and snow texture, cinematic wide landscape rendered as a square composition.

Controlled same-prompt blog comparison output for post 407 generated with z-image-turbo
Z Image Turbo
z image turbo: first output from the controlled prompt above, September 17, 2026, 1K. One pair illustrates differences; it does not establish an overall model ranking.
Controlled same-prompt blog comparison output for post 407 generated with flux-schnell
Flux Schnell
flux schnell: first output from the controlled prompt above, September 17, 2026, 1K. One pair illustrates differences; it does not establish an overall model ranking.

Test the brief before running a model stack

Generate a representative image in the browser first, then compare control, output, hardware, licensing, and maintenance against a local workflow.

Try AI Image Generator

Magic Hour publishes this guide. A model family is included only when it has downloadable weights or a verifiable local deployment route and a distinct current use case; hosted-only announcements are excluded. Capability, benchmark, hardware, and license claims are tied to the exact checkpoint and linked source rather than generalized across a family.

Magic Hour publishes this guide and includes its own product in the comparison. Treat the recommendations as editorial guidance from a vendor, and verify the linked first-party product details and your own output requirements before choosing a tool.

What should you compare open-weight image models on?

  • Capabilities: Match the checkpoint to text-to-image generation, image editing, text rendering, or multi-reference personalization.
  • Released checkpoint: Confirm that the exact weights are downloadable. An announcement or hosted variant cannot stand in for a local release.
  • Benchmark evidence: Compare exact checkpoints under the same protocol. Artificial Analysis derives Quality Elo from user votes on serverless API outputs, while ImageBench V1 uses 192 prompts and publishes local runs. Neither score is a substitute for testing your prompts on your hardware.
  • Parameter count: Use model size as a planning input, then size the real runtime from published VRAM, precision, resolution, offload, and quantization requirements.
  • Terms: Read the license attached to the exact checkpoint. Code, weights, gates, commercial rights, and territory limits can differ even within one family.
  • Local runner: Prefer a publisher-supported route such as Hugging Face Diffusers, an official PyTorch repository, SGLang, vLLM-Omni, or ComfyUI.

This roster consolidates related checkpoints under their upstream family and excludes fine-tunes, repackages, quantized mirrors, UI layers, and hosted-only systems. License terms can change, so verify the exact checkpoint’s current license before commercial use.

Quick comparison

Family

Recommended for

Variants

How to run it locally

Recommended Hardware

FLUX.1

General generation and Kontext editing · hosted in Magic Hour

schnell, dev, Kontext dev

Official PyTorch code, Diffusers, ComfyUI

No publisher minimum found

FLUX.2

Generation and editing on consumer GPUs · hosted in Magic Hour

klein 4B and 9B; Base variants; dev

Official PyTorch code and Diffusers

klein 4B: approximately 13GB VRAM

Z-Image

Fast local photorealism · hosted in Magic Hour

Base, Turbo

Diffusers and ComfyUI

Turbo: 16GB consumer VRAM

Stable Diffusion

Mature ecosystem and broad style work

SDXL Base/Refiner; SD3.5 Large, Turbo, Medium

Diffusers, official PyTorch repos, ComfyUI

SD3.5 Medium: 9.9GB excluding text encoders; SDXL Base test: 11.47GB–28.09GB

Qwen-Image

Text rendering and precise editing

Qwen-Image, 2512, Edit-2511, Layered

Diffusers and ComfyUI

No publisher minimum found

HunyuanImage 3.0

Knowledge-rich, large-scale generation

Base, Instruct, Distil

Official PyTorch repo

Base: at least 3×80GB; Instruct: at least 8×80GB

GLM-Image

Text-heavy and knowledge-intensive images

GLM-Image

Diffusers or SGLang

More than 80GB on one GPU or multiple GPUs

ERNIE-Image

Standard and fast bilingual generation

Standard, Turbo

Diffusers

24GB consumer GPU

HiDream-O1-Image

Generation, editing, and references

Full, Dev, Dev-2604

Official PyTorch repo

CUDA GPU; no publisher minimum found

DeepGen 1.0

Compact unified generation and editing

SFT, RL

Diffusers

No publisher minimum found

JoyAI-Image

Spatial and camera-aware editing

Edit, Edit-Plus

Official PyTorch repo and ComfyUI

CUDA GPU; no publisher minimum found

Krea 2

Illustration and varied styles

Raw, Turbo

Official PyTorch repo and ComfyUI

No publisher minimum found

Ideogram 4

Typography and structured design

9.3B nf4, fp8

Official PyTorch repo and ComfyUI

No publisher minimum found

Boogu-Image-0.1

Generation and editing at several memory tiers

Base, Turbo, Edit, Edit-Turbo

Official PyTorch repo and ComfyUI

12GB–80GB, depending on configuration

Cosmos 3

Frontier-scale image and world generation

Super-Text2Image, four-step variant

Diffusers, vLLM-Omni, SGLang

Multi-GPU data-center class

LongCat-Image

Bilingual generation, text, and editing

Base, Dev, Edit, Edit-Turbo

Diffusers and ComfyUI

About 17GB with CPU offload

FIBO

Structured, repeatable art direction

FIBO generation and editing releases

Diffusers and ComfyUI

No publisher minimum found

Infinity

Fast autoregressive image generation

2B, 8B

Official PyTorch repo and Docker

No publisher minimum found

Lumina-Image 2.0

Efficient 1024px generation

2.6B checkpoint

Diffusers, official PyTorch repo, ComfyUI

CPU offload supported; no minimum found

Sana

Efficient high-resolution generation

0.6B, 1.6B, 4.8B, Sprint

Diffusers, SGLang, ComfyUI

Quantized 4K path within about 8GB

OmniGen2

Generation, editing, and in-context composition

OmniGen2

Official PyTorch repo and ComfyUI

About 17GB native; offload options below that

BAGEL

Unified generation, editing, and understanding

7B-active/14B-total MoT

Official PyTorch repo

NF4 path for 12GB–32GB

CogView4

Chinese/English text and high-resolution generation

CogView4-6B

Diffusers and CogKit

13GB–35GB at 1024px, depending on offload

Kandinsky 5

Russian/English generation and editing

T2I Lite, Image Editing

Official PyTorch repo and Diffusers

No publisher minimum found

F-Lite

Copyright-safe, SFW generation

Standard, Texture, 7B

Diffusers and ComfyUI

At least 24GB VRAM

Janus-Pro

Unified understanding and compact generation

1B, 7B

Official PyTorch repo and ComfyUI

No publisher minimum found

Kolors

Chinese/English generation and portraits

Base, IP-Adapter, ControlNet, inpainting

Diffusers and ComfyUI

No publisher minimum found

AuraFlow

Literal prompt following under Apache 2.0

v0.1–v0.3

Diffusers and ComfyUI

No publisher minimum found

PixArt-Σ

Efficient high-resolution generation

512px, 1024px, 2K

Diffusers and official PyTorch repo

No publisher minimum found

FLUX.1

FLUX.1 official GitHub repository

Black Forest Labs’ FLUX.1 is a 12B family spanning fast generation with schnell, higher-quality generation with dev, and editing through Kontext. It remains a strong local baseline for literal prompt following and complex scenes, but practitioners also describe a narrow stylistic range, recurring faces, and overly smooth photographic skin; schnell trades more detail for speed.

Choose schnell when permissive commercial terms and speed matter. Evaluate dev or Kontext when their quality or editing capabilities justify the extra restrictions.

Use it on MagicHour

To use FLUX.1 without a local GPU, select flux-schnell in Magic Hour’s AI Image Generator or pass model="flux-schnell" to POST /v1/ai-image-generator.

It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with 1–4 images per job.

If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

FLUX.2

FLUX.2 official GitHub repository

FLUX.2 combines generation and editing, with klein aimed at local use. Early users praise its speed, positional prompting, style range, and edit mode, while 4B testers report unwanted changes and inconsistent anatomy or text.

  • Benchmark: The Artificial Analysis arena recorded klein 4B at Elo 1,057 and rank 26, and klein Base 4B at Elo 968 and rank 37 on July 30, 2026. These are not local-hardware runs.

Use it on MagicHour

To use FLUX.2 without a local GPU, select flux-2-klein in Magic Hour’s AI Image Generator or pass model="flux-2-klein" to POST /v1/ai-image-generator.

It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with one image per job. Pro, Flex, and Max are hosted offerings rather than additional local families.

Z-Image

Z-Image official GitHub repository

Tongyi-MAI’s Z-Image is a compact 6B family whose Turbo checkpoint is popular for fast, crisp photorealism on modest hardware. Users praise its quality-to-speed ratio but report limited seed diversity and weaker handling of multi-person scenes, left-versus-right instructions, and uncommon objects. Base is slower but often preferred for stylistic exploration.

Choose Turbo for a fast consumer-GPU path and Base when its slower, less-distilled behavior better suits your style.

Use it on MagicHour

To use Z-Image without a local GPU, select z-image-turbo in Magic Hour’s AI Image Generator or pass model="z-image-turbo" to POST /v1/ai-image-generator.

Stable Diffusion

Stable Diffusion official GitHub repository

Stability AI’s local family includes the mature SDXL ecosystem and the newer SD3.5 Large, Large Turbo, and Medium checkpoints. Practitioners value SD3.5’s style range and creativity but often find FLUX more reliable for people and long prompts; hands, feet, and anatomy remain recurring complaints. SDXL is older, yet local users still value its fine-tune, ControlNet, and face-workflow ecosystem.

  • Benchmark: The Artificial Analysis arena recorded SD3.5 Large Turbo at Elo 1,024, Large at 1,023, Medium at 946, and SDXL 1.0 at 876 on July 30, 2026.

Choose this family when ecosystem depth, fine-tuning, and control tooling matter as much as base-model leaderboard position.

Qwen-Image

Qwen-Image official GitHub repository

Alibaba’s Qwen team built Qwen-Image around text rendering, prompt understanding, and precise editing. Local users particularly value Edit for preserving faces, poses, and scene structure, while realistic edits can look plasticky or airbrushed; testers report that 2512 improves skin, hands, and small details.

Choose the generation checkpoint for typography-heavy prompts and the Edit or Layered releases when structural preservation matters.

HunyuanImage 3.0

HunyuanImage 3.0 official GitHub repository

Tencent’s HunyuanImage 3.0 is an unusually large native multimodal family for knowledge-rich prompts, text rendering, and complex composition. One early practitioner found that expanded prompts could produce striking results and stylized lettering, but simple prompts were more average and Base felt unfinished. Its footprint keeps it outside ordinary desktop workflows.

Choose it only when the model’s semantic and typography strengths justify a multi-GPU deployment and its territory terms fit your use.

GLM-Image

GLM-Image official GitHub repository

Z.ai’s GLM-Image combines an autoregressive semantic stage with a diffusion decoder, targeting text-heavy and knowledge-intensive images. Early testers found it creative and promising for layout, but ordinary text-to-image output can feel underbaked and slow beside lighter models.

Choose GLM-Image for dense text and semantic layout rather than as the easiest general-purpose local generator.

ERNIE-Image

ERNIE-Image official GitHub repository

Baidu’s ERNIE-Image is an 8B text-to-image family with standard and Turbo checkpoints. Community tests praise clean illustration, coherent backgrounds, and solid text, while reporting prompt-expander drift, camera-angle misses, diagonal artifacts, and an overprocessed HDR look.

Choose standard for maximum quality or Turbo when its eight-step path matters more than the small benchmark gap.

HiDream-O1-Image

HiDream-O1-Image official GitHub repository

HiDream’s O1 family is an 8B native multimodal model for generation, editing, and multi-reference personalization. Its broad workflow is attractive, but early testing is mixed: users report lost detail in distilled Dev outputs and plastic skin or artifact-like texture in some full-model results.

Choose HiDream when one local family must cover generation, editing, and subject-driven personalization.

DeepGen 1.0

DeepGen 1.0 official GitHub repository

DeepGen Team’s DeepGen 1.0 is a compact 5B research model that unifies generation, editing, and reasoning-oriented control. Local-model users like the unified design, but early discussion flags its fixed 512×512 experiments and immature UI and fine-tuning ecosystem.

Choose DeepGen when compact unified generation and editing matter more than a mature local ecosystem.

JoyAI-Image

JoyAI-Image official GitHub repository

JD’s JoyAI-Image is an edit-first family with Edit and Edit-Plus checkpoints plus text-to-image generation. Early testers found it strong at spatial changes and alternate camera angles, but saw facial-detail drift and distortion at extreme views. Adoption has also been slowed by a heavy runtime and rough early ComfyUI workflows.

  • Benchmark: The independent KRIS-Bench leaderboard scores JoyAI-Image-Edit at 63.44 overall across 1,267 knowledge-reasoning edits.

Choose JoyAI when spatially aware editing and camera changes are more important than lightweight deployment.

Krea 2

Krea 2 official GitHub repository

Krea 2 is built around style diversity, with Raw for exploration and Turbo for speed. The team says it prioritized illustration and varied styles over realism; practitioners praise anatomy, animals, and wide compositions while reporting weaker photography, occasional prompt misses, VAE grid texture, and aggressive safety behavior.

Choose Raw for style exploration and Turbo when iteration speed matters.

Ideogram 4

Ideogram 4 official GitHub repository

Ideogram’s first open-weight family is a design-focused 9.3B generator for multilingual typography, structured JSON prompts, bounding boxes, palettes, and native 2K output. Practitioners like its direct layout control, but criticize missing full-precision weights and restrictive safety behavior.

Choose Ideogram 4 for local, non-commercial typography and layout work.

Boogu-Image-0.1

Boogu-Image-0.1 official GitHub repository

The Boogu Project’s first family covers Base, Turbo, Edit, and Edit-Turbo. A 192-prompt practitioner test found high detail and good overall fidelity but a strong symmetry bias and weak composition; edit users report failures on complex changes and artifacts in small faces, limbs, eyes, and text.

Choose Boogu when you want generation and editing variants with documented memory tradeoffs.

Cosmos 3

Cosmos 3 official GitHub repository

NVIDIA’s Cosmos 3 is a 64B omnimodal world-model family with a specialized text-to-image checkpoint. It leads the current open-weights arena, but local practitioners stress that its reasoner-plus-generator footprint is impractical on ordinary desktops.

  • Benchmark: The Artificial Analysis arena recorded Cosmos3-Super-Text2Image at Elo 1,218 and rank 1, and the four-step version at 1,181 and rank 3 on July 30, 2026.

Choose Cosmos 3 only when frontier-scale quality is worth data-center-class infrastructure.

LongCat-Image

LongCat-Image official GitHub repository

Meituan’s 6B LongCat-Image family covers bilingual generation, typography, and editing. Users like its product and face preservation, but report slow Base inference, uneven editing quality, and occasional oversaturated or plastic output.

Choose LongCat for bilingual text rendering or editing when 17GB-class offloaded inference is acceptable.

FIBO

FIBO official Hugging Face model card

Bria AI’s 8B FIBO is built for structured, repeatable art direction: a VLM expands an idea into a long JSON description that controls lighting, camera, composition, and color. Practitioners praise its prompt adherence and texture, but report weaker hands and note that the local release lacks broad native ComfyUI support.

Choose FIBO when reproducible, programmatic visual control matters more than free-form prompting.

Infinity

Infinity official GitHub repository

FoundationVision’s Infinity is a bitwise autoregressive generator rather than a diffusion model. Its 2B and 8B checkpoints target fast 1024px synthesis and strong composition. Practitioners like the prompt adherence of autoregressive models, but some see little improvement in the familiar “AI-generated” look and note that the larger 20B release is still absent.

Choose Infinity when you want to evaluate an autoregressive alternative to local diffusion pipelines.

Lumina-Image 2.0

Lumina-Image 2.0 official GitHub repository

Alpha-VLLM’s Lumina-Image 2.0 is a compact 2.6B 1024px generator with strong prompt adherence and an open fine-tuning path. Practitioners find it capable for composition, but describe unstable anatomy, uneven detail, and a demanding setup relative to mature SDXL workflows.

Choose Lumina when an Apache-licensed, fine-tunable 2.6B base is more important than ecosystem maturity.

Sana

Sana official GitHub repository

NVIDIA’s Sana family uses linear attention and aggressive latent compression for efficient high-resolution generation. It is attractive for small models and fast Sprint checkpoints, although practitioners characterize Sprint as extremely fast but visibly lower quality.

Choose Sana when efficiency, resolution, and fine-tuning openness matter more than arena rank.

OmniGen2

OmniGen2 official GitHub repository

VectorSpaceLab’s OmniGen2 unifies text-to-image generation, instruction editing, and in-context composition. Users like its ability to combine references, but report slow local runs, inconsistent edits, oversaturation, and installation friction.

Choose OmniGen2 when multi-image composition and unified editing matter more than speed.

BAGEL

BAGEL official GitHub repository

ByteDance Seed’s BAGEL is a 7B-active, 14B-total Mixture-of-Transformer-Experts model that combines understanding, generation, and editing. Practitioners praise its ambitious unified workflow, but report blurry outputs, aggressive safety behavior, anatomy failures, and heavy local memory use.

Choose BAGEL to experiment with a single multimodal model across understanding, generation, and editing.

CogView4

CogView4 official GitHub repository

Z.ai’s CogView4 is a 6B bilingual generator designed for Chinese and English prompts and native high-resolution output. Community interest centers on its text accuracy and prompt length, while users warn that its large text encoder and custom stack make it heavier than the 6B headline suggests.

Choose CogView4 for bilingual text-heavy work when you can accommodate its text encoder.

Kandinsky 5

Kandinsky 5 official GitHub repository

Kandinsky Lab’s current family includes 6B text-to-image and editing checkpoints with strong Russian concepts and typography, plus a related video stack. One community workflow praises its skin texture and image-to-image results, but calls the family an underrated underdog and needs offload workarounds to fit it on an 8GB GPU.

Choose Kandinsky 5 for Russian-language concepts, typography, or a shared image-and-video research stack.

F-Lite

F-Lite official GitHub repository

Freepik and fal built F-Lite from licensed, SFW data, with Standard, Texture, and smaller 7B releases. That provenance is its clearest differentiator; practitioners welcome the licensing story but note limited text rendering and malformed anatomy in some outputs.

Choose F-Lite when licensed-data provenance and commercial safety review are central to the evaluation.

Janus-Pro

Janus-Pro official GitHub repository

DeepSeek’s Janus-Pro combines visual understanding and image generation in 1B and 7B checkpoints. Practitioners praise the 1B release’s literal prompt adherence and speed, but its hard-coded 384px output and weaker aesthetic finish make it more of a compact research tool than a production image model.

Choose Janus-Pro for compact multimodal experiments rather than maximum-resolution image work.

Kolors

Kolors official GitHub repository

Kuaishou’s Kolors is a bilingual latent-diffusion family for Chinese and English generation, portraits, IP-Adapter workflows, ControlNet, and inpainting. Practitioners value its Chinese handling and portraits, but prompt comparisons find weaker complex-scene adherence than newer models.

Choose Kolors for bilingual portrait and control workflows when its model terms fit the deployment.

AuraFlow

AuraFlow official Hugging Face model card

Fal’s AuraFlow is a fully Apache-licensed flow-based generator built around literal prompt adherence. Practitioners found v0.2 unusually good at complex composition, but also described weak aesthetics, anatomy problems, and a clip-art-like look; v0.3 improved cohesion for some prompts while regressing adherence for others.

Choose AuraFlow when permissive terms and literal composition matter more than polished default aesthetics.

PixArt-Σ

PixArt-Σ official GitHub repository

The PixArt team’s PixArt-Σ is a compact diffusion-transformer family for 512px, 1024px, and 2K generation. Practitioners praise its prompt understanding and low training cost, while reporting weaker default aesthetics, anatomy errors, and a need for refinement.

  • Benchmark: Z.ai’s shared table reports PixArt-α at DPG-Bench 71.11; this older sibling score should not be treated as a PixArt-Σ result.

Choose PixArt-Σ when you want a compact, permissively licensed base for high-resolution generation and fine-tuning.

How to choose your first local test

  1. Pick the capability: generation, typography, editing, or multi-reference composition.
  1. Select the exact released checkpoint rather than the family name alone.
  1. Confirm the checkpoint’s current access and commercial terms.
  1. Match its published GPU configuration to your machine.
  1. Compare exact releases under one benchmark protocol.
  1. Run your own prompts, resolutions, edit inputs, and seeds before production use.

For consumer hardware, the publisher figures point to FLUX.2 klein 4B at approximately 13GB, Z-Image Turbo at 16GB, ERNIE-Image at 24GB, LongCat-Image around 17GB with offload, and OmniGen2 around 17GB natively. Sana and Boogu publish lower-memory configurations through quantization and offload. HunyuanImage 3.0, GLM-Image, and Cosmos 3 belong in the multi-GPU or greater-than-80GB class.

Is FLUX.1 better than Stable Diffusion 3.5?

The live arena places FLUX.1 dev at Elo 1,029, narrowly above SD3.5 Large Turbo at 1,024 and Large at 1,023, but that does not make it a universal winner. FLUX.1 is usually the stronger prompt-following baseline; Stable Diffusion retains a deeper fine-tuning and control ecosystem. Compare their exact checkpoints on your prompts, then include license and hardware fit in the decision.

If local model operations fall outside your workflow, Magic Hour’s AI Image Generator exposes selectable flux-schnell, flux-2-klein, and z-image-turbo model IDs through its hosted product. The AI Image Generator API reference documents model selection and job creation.

Evaluate open models as systems, not names

Direct answer: An open image model is only one part of the workflow. Compare license, weights, inference code, hardware, quantization, control adapters, editing support and maintenance before choosing a stack.

Reproducible benchmark. Record model and checkpoint hash, sampler, steps, guidance, seed, resolution, hardware, precision and runtime. Use prompts covering people, products, typography, composition and editing. Publish failures alongside selected outputs.

Operational cost. Include setup, downloads, GPU time, idle capacity, storage, moderation, updates and engineering support. Hosted inference can be cheaper for intermittent demand; local deployment can be better when privacy or sustained utilization justifies it.

License gate. Read the model card and license for the exact weights and adapters. “Open source,” “open weights” and “source available” are not interchangeable commercial permissions.

Try the workflow: Open the matching Magic Hour tool. Use the same representative input and acceptance criteria before comparing results.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Build Flexible Video Pipelines Without Lock-In or Heavy Infrastructure
Recommended next
6 video AI APIs and open-weight deployment options

Compare fal.ai, Replicate, Magic Hour, Runway, LTX 2.5 and MiniMax H3 by model access, deployment, job handling, licensing and total production cost.

Best Open Source AI Video Generation Models
2 open-source AI video models to evaluate now: LTX 2.5 and MiniMax H3
openai chatgpt 4o cover
GPT-4o image generation review: what changed since 2025
AI UGC Ad Generators (2026)
6 best AI UGC ad generators by workflow (2026)
Best AI Image Editing Models With Reference Images
Best AI image editing models with reference images