Magic Hour
  • Pricing
  • API
Magic Hour
  • Create New
  • Pricing
  • API
Try Magic HourLogin
Video

Create your first video in under 60 seconds

Generate or edit video, image, and audio - free to start.
Start Creating Free
No credit card requiredFree daily creditsNo signup required
Join the Discord
Video
Company
PricingAboutBlogAPIAll ToolsTemplatesAI ModelsPrivacy PolicyTerms of ServiceRefund Policy
Video Products
AI UGC Ad GeneratorAI Video EditorAI Video ExpanderAI Video ExtenderAI Video UpscalerAnimationAudio-to-VideoCharacter ReplaceFace Swap VideoImage-to-VideoLip SyncMusic Video GeneratorSubtitle GeneratorTalking PhotoText-to-VideoVideo ColorizerVideo-to-Video
Image Products
AI Clothes ChangerAI Face EditorAI GIF GeneratorAI Headshot GeneratorAI Image EditorAI Image ExpanderAI Image GeneratorAI Image UpscalerAI Influencer GeneratorAI Meme GeneratorAI Selfie GeneratorAI Storyboard GeneratorBackground RemoverBody SwapFace Swap PhotoHead SwapPhoto ColorizerQR Code GeneratorCharactersMoodboards
Audio Products
AI Music GeneratorVideo-to-AudioVoice ChangerVoice ClonerVoice Generator
Support
CommunityFAQContact UsStatus
Social
Instagram
X
TikTok
Facebook
YouTube
LinkedIn
support@magichour.ai
Backed byCombinator

© 2026 Magic Hour AI, Inc.

Back to Blog
  1. Blog
  2. Top Choice
  3. 29 Open-Source Image Generation Model Families to Run Locally in 2026
Top Choice

29 Open-Source Image Generation Model Families to Run Locally in 2026

Runbo Li
Runbo Li
·
CEO of Magic Hour
Aug 05, 2026
(Updated Aug 05, 2026)
· 33 min read
AI Summary:
ChatGPTClaudeGeminiPerplexity
29 open-source image generation models field guide — hero

Contents

The best local image model is the one whose capabilities, checkpoint terms, and hardware requirements fit your workflow. Treat Base, Turbo, Edit, Dev, quantized, and revision checkpoints as choices within a family rather than separate models.

This guide compares 29 families with downloadable weights and a publisher-documented local inference path: FLUX.1, FLUX.2, Z-Image, Stable Diffusion, Qwen-Image, HunyuanImage 3.0, GLM-Image, ERNIE-Image, HiDream-O1-Image, DeepGen 1.0, JoyAI-Image, Krea 2, Ideogram 4, Boogu-Image-0.1, Cosmos 3, LongCat-Image, FIBO, Infinity, Lumina-Image 2.0, Sana, OmniGen2, BAGEL, CogView4, Kandinsky 5, F-Lite, Janus-Pro, Kolors, AuraFlow, and PixArt-Σ.

What should you compare open-weight image models on?

  • Capabilities: Match the checkpoint to text-to-image generation, image editing, text rendering, or multi-reference personalization.
  • Released checkpoint: Confirm that the exact weights are downloadable. An announcement or hosted variant cannot stand in for a local release.
  • Benchmark evidence: Compare exact checkpoints under the same protocol. Artificial Analysis derives Quality Elo from user votes on serverless API outputs, while ImageBench V1 uses 192 prompts and publishes local runs. Neither score is a substitute for testing your prompts on your hardware.
  • Parameter count: Use model size as a planning input, then size the real runtime from published VRAM, precision, resolution, offload, and quantization requirements.
  • Speed: Keep step count and latency attached to the tested configuration. For example, Z-Image Turbo’s sub-second claim uses H800 GPUs, while its 16GB figure describes consumer-device fit.
  • Terms: Read the license attached to the exact checkpoint. Code, weights, gates, commercial rights, and territory limits can differ even within one family.
  • Local runner: Prefer a publisher-supported route such as Hugging Face Diffusers, an official PyTorch repository, SGLang, vLLM-Omni, or ComfyUI.

This roster consolidates related checkpoints under their upstream family and excludes fine-tunes, repackages, quantized mirrors, UI layers, and hosted-only systems. License terms can change, so verify the exact checkpoint’s current license before commercial use.

Quick comparison

Family

Recommended for

Variants

How to run it locally

Recommended Hardware

FLUX.1

General generation and Kontext editing · hosted in Magic Hour

schnell, dev, Kontext dev

Official PyTorch code, Diffusers, ComfyUI

No publisher minimum found

FLUX.2

Generation and editing on consumer GPUs · hosted in Magic Hour

klein 4B and 9B; Base variants; dev

Official PyTorch code and Diffusers

klein 4B: approximately 13GB VRAM

Z-Image

Fast local photorealism · hosted in Magic Hour

Base, Turbo

Diffusers and ComfyUI

Turbo: 16GB consumer VRAM

Stable Diffusion

Mature ecosystem and broad style work

SDXL Base/Refiner; SD3.5 Large, Turbo, Medium

Diffusers, official PyTorch repos, ComfyUI

SD3.5 Medium: 9.9GB excluding text encoders; SDXL Base test: 11.47GB–28.09GB

Qwen-Image

Text rendering and precise editing

Qwen-Image, 2512, Edit-2511, Layered

Diffusers and ComfyUI

No publisher minimum found

HunyuanImage 3.0

Knowledge-rich, large-scale generation

Base, Instruct, Distil

Official PyTorch repo

Base: at least 3×80GB; Instruct: at least 8×80GB

GLM-Image

Text-heavy and knowledge-intensive images

GLM-Image

Diffusers or SGLang

More than 80GB on one GPU or multiple GPUs

ERNIE-Image

Standard and fast bilingual generation

Standard, Turbo

Diffusers

24GB consumer GPU

HiDream-O1-Image

Generation, editing, and references

Full, Dev, Dev-2604

Official PyTorch repo

CUDA GPU; no publisher minimum found

DeepGen 1.0

Compact unified generation and editing

SFT, RL

Diffusers

No publisher minimum found

JoyAI-Image

Spatial and camera-aware editing

Edit, Edit-Plus

Official PyTorch repo and ComfyUI

CUDA GPU; no publisher minimum found

Krea 2

Illustration and varied styles

Raw, Turbo

Official PyTorch repo and ComfyUI

No publisher minimum found

Ideogram 4

Typography and structured design

9.3B nf4, fp8

Official PyTorch repo and ComfyUI

No publisher minimum found

Boogu-Image-0.1

Generation and editing at several memory tiers

Base, Turbo, Edit, Edit-Turbo

Official PyTorch repo and ComfyUI

12GB–80GB, depending on configuration

Cosmos 3

Frontier-scale image and world generation

Super-Text2Image, four-step variant

Diffusers, vLLM-Omni, SGLang

Multi-GPU data-center class

LongCat-Image

Bilingual generation, text, and editing

Base, Dev, Edit, Edit-Turbo

Diffusers and ComfyUI

About 17GB with CPU offload

FIBO

Structured, repeatable art direction

FIBO generation and editing releases

Diffusers and ComfyUI

No publisher minimum found

Infinity

Fast autoregressive image generation

2B, 8B

Official PyTorch repo and Docker

No publisher minimum found

Lumina-Image 2.0

Efficient 1024px generation

2.6B checkpoint

Diffusers, official PyTorch repo, ComfyUI

CPU offload supported; no minimum found

Sana

Efficient high-resolution generation

0.6B, 1.6B, 4.8B, Sprint

Diffusers, SGLang, ComfyUI

Quantized 4K path within about 8GB

OmniGen2

Generation, editing, and in-context composition

OmniGen2

Official PyTorch repo and ComfyUI

About 17GB native; offload options below that

BAGEL

Unified generation, editing, and understanding

7B-active/14B-total MoT

Official PyTorch repo

NF4 path for 12GB–32GB

CogView4

Chinese/English text and high-resolution generation

CogView4-6B

Diffusers and CogKit

13GB–35GB at 1024px, depending on offload

Kandinsky 5

Russian/English generation and editing

T2I Lite, Image Editing

Official PyTorch repo and Diffusers

No publisher minimum found

F-Lite

Copyright-safe, SFW generation

Standard, Texture, 7B

Diffusers and ComfyUI

At least 24GB VRAM

Janus-Pro

Unified understanding and compact generation

1B, 7B

Official PyTorch repo and ComfyUI

No publisher minimum found

Kolors

Chinese/English generation and portraits

Base, IP-Adapter, ControlNet, inpainting

Diffusers and ComfyUI

No publisher minimum found

AuraFlow

Literal prompt following under Apache 2.0

v0.1–v0.3

Diffusers and ComfyUI

No publisher minimum found

PixArt-Σ

Efficient high-resolution generation

512px, 1024px, 2K

Diffusers and official PyTorch repo

No publisher minimum found

FLUX.1

FLUX.1 official GitHub repository

Black Forest Labs’ FLUX.1 is a 12B family spanning fast generation with schnell, higher-quality generation with dev, and editing through Kontext. It remains a strong local baseline for literal prompt following and complex scenes, but practitioners also describe a narrow stylistic range, recurring faces, and overly smooth photographic skin; schnell trades more detail for speed.

  • Benchmark: As checked July 30, 2026, the Artificial Analysis open-weights arena recorded FLUX.1 dev at Elo 1,029 and rank 32, and schnell at Elo 1,000 and rank 35. The KRIS-Bench editing leaderboard scores Kontext dev at 49.54 overall.
  • Size and hardware: Schnell is a 12B model; the publisher gives no universal minimum VRAM figure.
  • Run locally: Use the official PyTorch repository, Hugging Face Diffusers, or ComfyUI.
  • Access and terms: Schnell uses Apache 2.0 and permits commercial use; dev uses the FLUX.1 dev Non-Commercial License. Kontext dev uses the same non-commercial family of terms.

Choose schnell when permissive commercial terms and speed matter. Evaluate dev or Kontext when their quality or editing capabilities justify the extra restrictions.

Use it on MagicHour

To use FLUX.1 without a local GPU, select flux-schnell in Magic Hour’s AI Image Generator or pass model="flux-schnell" to POST /v1/ai-image-generator.

It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with 1–4 images per job.

If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

FLUX.2

FLUX.2 official GitHub repository

FLUX.2 combines generation and editing, with klein aimed at local use. Early users praise its speed, positional prompting, style range, and edit mode, while 4B testers report unwanted changes and inconsistent anatomy or text.

  • Benchmark: The Artificial Analysis arena recorded klein 4B at Elo 1,057 and rank 26, and klein Base 4B at Elo 968 and rank 37 on July 30, 2026. These are not local-hardware runs.
  • Size and hardware: Klein 4B has 4B parameters and a publisher-stated footprint of approximately 13GB VRAM.
  • Run locally: Use BFL’s official PyTorch code or Diffusers; ComfyUI supports the family.
  • Access and terms: Klein 4B and its Base checkpoint are Apache 2.0. Check dev-family terms separately.

Use it on MagicHour

To use FLUX.2 without a local GPU, select flux-2-klein in Magic Hour’s AI Image Generator or pass model="flux-2-klein" to POST /v1/ai-image-generator.

It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with one image per job. Pro, Flex, and Max are hosted offerings rather than additional local families.

If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

Z-Image

Z-Image official GitHub repository

Tongyi-MAI’s Z-Image is a compact 6B family whose Turbo checkpoint is popular for fast, crisp photorealism on modest hardware. Users praise its quality-to-speed ratio but report limited seed diversity and weaker handling of multi-person scenes, left-versus-right instructions, and uncommon objects. Base is slower but often preferred for stylistic exploration.

  • Benchmark: ImageBench V1 scored Turbo at 49.2 overall with 18.1-second latency and Base at 47.4 overall with 130.7-second latency. The Artificial Analysis arena recorded Turbo at Elo 1,102 and rank 18 on July 30, 2026.
  • Size and hardware: The family has 6B parameters; Turbo fits within 16GB consumer VRAM, while the sub-second claim uses H800 GPUs.
  • Run locally: Use Diffusers or ComfyUI.
  • Access and terms: The official repository code is Apache 2.0. Verify the license attached to the exact weight file before commercial deployment.

Choose Turbo for a fast consumer-GPU path and Base when its slower, less-distilled behavior better suits your style.

Use it on MagicHour

To use Z-Image without a local GPU, select z-image-turbo in Magic Hour’s AI Image Generator or pass model="z-image-turbo" to POST /v1/ai-image-generator.

It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with 1–4 images per job.

If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

Stable Diffusion

Stable Diffusion official GitHub repository

Stability AI’s local family includes the mature SDXL ecosystem and the newer SD3.5 Large, Large Turbo, and Medium checkpoints. Practitioners value SD3.5’s style range and creativity but often find FLUX more reliable for people and long prompts; hands, feet, and anatomy remain recurring complaints. SDXL is older, yet local users still value its fine-tune, ControlNet, and face-workflow ecosystem.

  • Benchmark: The Artificial Analysis arena recorded SD3.5 Large Turbo at Elo 1,024, Large at 1,023, Medium at 946, and SDXL 1.0 at 876 on July 30, 2026.
  • Released checkpoints: Stability AI publishes SD3.5 Large, Large Turbo, and Medium weights, while SDXL provides Base and Refiner weights.
  • Checkpoint sizes and hardware: SD3.5 Large and Large Turbo are 8.1B models; the publisher does not state a minimum VRAM figure for either. SD3.5 Medium is 2.5B and requires 9.9GB VRAM for the model alone, excluding its text encoders. For SDXL Base, Hugging Face measured 28.09GB unoptimized, 21.72GB at fp16, and 11.47GB with VAE slicing plus sequential CPU offload while generating a batch of four 1024px images on an A100; the separate Refiner adds another pipeline stage, and no universal minimum is published.
  • Run locally: Use Hugging Face Diffusers, Stability AI’s official PyTorch repositories, or ComfyUI.
  • Access and terms: SDXL uses CreativeML Open RAIL++-M. SD3.5 checkpoints use Stability AI’s Community License, which should be checked at download time.

Choose this family when ecosystem depth, fine-tuning, and control tooling matter as much as base-model leaderboard position.

Qwen-Image

Qwen-Image official GitHub repository

Alibaba’s Qwen team built Qwen-Image around text rendering, prompt understanding, and precise editing. Local users particularly value Edit for preserving faces, poses, and scene structure, while realistic edits can look plasticky or airbrushed; testers report that 2512 improves skin, hands, and small details.

  • Benchmark: ImageBench V1 scored Qwen-Image-2512 at 58.6 overall, 50.0% capability, 67.1 estimated preference, and 80.2-second average latency. The Artificial Analysis arena recorded the original Qwen Image at Elo 1,064 and rank 24 on July 30, 2026.
  • Size and hardware: The original release is a 20B MMDiT model; the publisher gives no universal minimum VRAM figure.
  • Run locally: Use Hugging Face Diffusers or ComfyUI.
  • Access and terms: The family is Apache 2.0.

Choose the generation checkpoint for typography-heavy prompts and the Edit or Layered releases when structural preservation matters.

HunyuanImage 3.0

HunyuanImage 3.0 official GitHub repository

Tencent’s HunyuanImage 3.0 is an unusually large native multimodal family for knowledge-rich prompts, text rendering, and complex composition. One early practitioner found that expanded prompts could produce striking results and stylized lettering, but simple prompts were more average and Base felt unfinished. Its footprint keeps it outside ordinary desktop workflows.

  • Benchmark: The Artificial Analysis arena recorded HunyuanImage 3.0 at Elo 1,125 and Instruct at 1,119 on July 30, 2026. A shared publisher table reports GenEval 0.72, DPG-Bench 86.10, and WISE 0.57.
  • Size and hardware: Base has 80B total and 13B active parameters and needs at least 3×80GB; Instruct needs at least 8×80GB.
  • Run locally: Use Tencent’s official PyTorch repository.
  • Access and terms: Tencent’s community license excludes the EU, UK, and South Korea.

Choose it only when the model’s semantic and typography strengths justify a multi-GPU deployment and its territory terms fit your use.

GLM-Image

GLM-Image official GitHub repository

Z.ai’s GLM-Image combines an autoregressive semantic stage with a diffusion decoder, targeting text-heavy and knowledge-intensive images. Early testers found it creative and promising for layout, but ordinary text-to-image output can feel underbaked and slow beside lighter models.

  • Benchmark: The Artificial Analysis arena recorded Elo 1,049 and rank 28 on July 30, 2026. Z.ai reports DPG-Bench 84.78, LongText-Bench 0.966 average, and CVTG-2K word accuracy 0.9116.
  • Size and hardware: Current optimization requires one GPU with more than 80GB or multiple GPUs.
  • Run locally: Use Diffusers or SGLang.
  • Access and terms: The checkpoint and repository use Apache 2.0.

Choose GLM-Image for dense text and semantic layout rather than as the easiest general-purpose local generator.

ERNIE-Image

ERNIE-Image official GitHub repository

Baidu’s ERNIE-Image is an 8B text-to-image family with standard and Turbo checkpoints. Community tests praise clean illustration, coherent backgrounds, and solid text, while reporting prompt-expander drift, camera-angle misses, diagonal artifacts, and an overprocessed HDR look.

  • Benchmark: The Artificial Analysis arena recorded ERNIE Image at Elo 1,166 and Turbo at 1,163 on July 30, 2026. Baidu reports GenEval 0.8856 without prompt enhancement for standard and 0.8667 for Turbo.
  • Size and hardware: The family has 8B DiT parameters and runs in a publisher-documented 24GB consumer-GPU configuration.
  • Run locally: Use Hugging Face Diffusers.
  • Access and terms: The official code and weights use Apache 2.0.

Choose standard for maximum quality or Turbo when its eight-step path matters more than the small benchmark gap.

HiDream-O1-Image

HiDream-O1-Image official GitHub repository

HiDream’s O1 family is an 8B native multimodal model for generation, editing, and multi-reference personalization. Its broad workflow is attractive, but early testing is mixed: users report lost detail in distilled Dev outputs and plastic skin or artifact-like texture in some full-model results.

  • Benchmark: The Artificial Analysis arena recorded Dev-2604 at Elo 1,189 and rank 2, Full at 1,116, and Dev at 1,080 on July 30, 2026. The official model card reports GenEval 0.90, DPG-Bench 89.83, and HPSv3 10.37.
  • Size and hardware: The family has 8B parameters and requires a CUDA-capable GPU.
  • Run locally: Use HiDream’s official PyTorch repository.
  • Access and terms: Code and models use the MIT License.

Choose HiDream when one local family must cover generation, editing, and subject-driven personalization.

DeepGen 1.0

DeepGen 1.0 official GitHub repository

DeepGen Team’s DeepGen 1.0 is a compact 5B research model that unifies generation, editing, and reasoning-oriented control. Local-model users like the unified design, but early discussion flags its fixed 512×512 experiments and immature UI and fine-tuning ecosystem.

  • Benchmark: The publisher reports DeepGen RL at GenEval 0.87, DPG-Bench 87.90, UniGenBench 75.74, and GEdit-EN 7.17.
  • Size and hardware: The model has 5B parameters: a 3B VLM and 2B DiT.
  • Run locally: Use Hugging Face Diffusers.
  • Access and terms: The official checkpoint uses Apache 2.0.

Choose DeepGen when compact unified generation and editing matter more than a mature local ecosystem.

JoyAI-Image

JoyAI-Image official GitHub repository

JD’s JoyAI-Image is an edit-first family with Edit and Edit-Plus checkpoints plus text-to-image generation. Early testers found it strong at spatial changes and alternate camera angles, but saw facial-detail drift and distortion at extreme views. Adoption has also been slowed by a heavy runtime and rough early ComfyUI workflows.

  • Benchmark: The independent KRIS-Bench leaderboard scores JoyAI-Image-Edit at 63.44 overall across 1,267 knowledge-reasoning edits.
  • Size and hardware: The official runner requires a CUDA-capable GPU; no publisher minimum VRAM figure is given.
  • Run locally: Use JD’s official PyTorch repository or ComfyUI.
  • Access and terms: The repository and released weights use Apache 2.0.

Choose JoyAI when spatially aware editing and camera changes are more important than lightweight deployment.

Krea 2

Krea 2 official GitHub repository

Krea 2 is built around style diversity, with Raw for exploration and Turbo for speed. The team says it prioritized illustration and varied styles over realism; practitioners praise anatomy, animals, and wide compositions while reporting weaker photography, occasional prompt misses, VAE grid texture, and aggressive safety behavior.

  • Benchmark: ImageBench V1 scored Krea 2 Turbo at 56.5 overall, 58% capability, 55.2 estimated preference, and 73.2-second average latency.
  • Run locally: Use Krea’s official PyTorch repository or ComfyUI.
  • Access and terms: Raw and Turbo are gated and use the Krea 2 Community License.

Choose Raw for style exploration and Turbo when iteration speed matters.

Ideogram 4

Ideogram 4 official GitHub repository

Ideogram’s first open-weight family is a design-focused 9.3B generator for multilingual typography, structured JSON prompts, bounding boxes, palettes, and native 2K output. Practitioners like its direct layout control, but criticize missing full-precision weights and restrictive safety behavior.

  • Benchmark: The Artificial Analysis arena recorded the Quality checkpoint at Elo 1,175 and rank 4 on July 30, 2026. Ideogram cites a blind typography test in which 10 designers selected Ideogram 4 first 47.9% of the time and rated it 3.55/5 for real client work.
  • Size and hardware: The nf4 and fp8 releases contain a 9.3B model.
  • Run locally: Use the official PyTorch repository or ComfyUI.
  • Access and terms: The gated weights use the Ideogram 4 Non-Commercial License.

Choose Ideogram 4 for local, non-commercial typography and layout work.

Boogu-Image-0.1

Boogu-Image-0.1 official GitHub repository

The Boogu Project’s first family covers Base, Turbo, Edit, and Edit-Turbo. A 192-prompt practitioner test found high detail and good overall fidelity but a strong symmetry bias and weak composition; edit users report failures on complex changes and artifacts in small faces, limbs, eyes, and text.

  • Benchmark: ImageBench V1 scored Boogu-Image-0.1-Turbo at 61.9 overall, 57% capability, 66.6 estimated preference, and 9.9-second average latency.
  • Size and hardware: Published configurations span 12GB to 80GB VRAM, depending on resolution, offload, and quantization.
  • Run locally: Use the official PyTorch repository or ComfyUI.
  • Access and terms: The family uses Apache 2.0.

Choose Boogu when you want generation and editing variants with documented memory tradeoffs.

Cosmos 3

Cosmos 3 official GitHub repository

NVIDIA’s Cosmos 3 is a 64B omnimodal world-model family with a specialized text-to-image checkpoint. It leads the current open-weights arena, but local practitioners stress that its reasoner-plus-generator footprint is impractical on ordinary desktops.

  • Benchmark: The Artificial Analysis arena recorded Cosmos3-Super-Text2Image at Elo 1,218 and rank 1, and the four-step version at 1,181 and rank 3 on July 30, 2026.
  • Size and hardware: Cosmos3-Super and its text-to-image checkpoint are 64B and use multi-GPU configurations.
  • Run locally: Use NVIDIA’s Cosmos framework, Diffusers, vLLM-Omni, or SGLang.
  • Access and terms: Code and models use OpenMDW 1.1.

Choose Cosmos 3 only when frontier-scale quality is worth data-center-class infrastructure.

LongCat-Image

LongCat-Image official GitHub repository

Meituan’s 6B LongCat-Image family covers bilingual generation, typography, and editing. Users like its product and face preservation, but report slow Base inference, uneven editing quality, and occasional oversaturated or plastic output.

  • Benchmark: Meituan reports GenEval 0.87, DPG-Bench 86.80, and WISE 0.65. The Artificial Analysis arena recorded Elo 1,032 and rank 30 on July 30, 2026.
  • Size and hardware: The dense DiT has 6B parameters and CPU-offloaded generation needs about 17GB VRAM.
  • Run locally: Use Diffusers or ComfyUI.
  • Access and terms: The family uses Apache 2.0.

Choose LongCat for bilingual text rendering or editing when 17GB-class offloaded inference is acceptable.

FIBO

FIBO official Hugging Face model card

Bria AI’s 8B FIBO is built for structured, repeatable art direction: a VLM expands an idea into a long JSON description that controls lighting, camera, composition, and color. Practitioners praise its prompt adherence and texture, but report weaker hands and note that the local release lacks broad native ComfyUI support.

  • Benchmark: The Artificial Analysis arena recorded Elo 1,068 and rank 22 on July 30, 2026.
  • Size and hardware: FIBO has 8B parameters; the publisher gives no universal minimum VRAM figure.
  • Run locally: Use the official Diffusers pipeline; Bria also supplies ComfyUI integrations.
  • Access and terms: Weights are non-commercial; commercial deployment requires Bria licensing.

Choose FIBO when reproducible, programmatic visual control matters more than free-form prompting.

Infinity

Infinity official GitHub repository

FoundationVision’s Infinity is a bitwise autoregressive generator rather than a diffusion model. Its 2B and 8B checkpoints target fast 1024px synthesis and strong composition. Practitioners like the prompt adherence of autoregressive models, but some see little improvement in the familiar “AI-generated” look and note that the larger 20B release is still absent.

  • Benchmark: The publisher reports 8B GenEval 0.79 and DPG-Bench 86.6. The Artificial Analysis arena recorded Infinity 8B at Elo 1,045 and rank 29 on July 30, 2026.
  • Size and hardware: Official releases include 2B and 8B checkpoints; the publisher gives no universal minimum VRAM figure.
  • Run locally: Use the official PyTorch notebooks or Docker workflow.
  • Access and terms: The project uses the MIT License.

Choose Infinity when you want to evaluate an autoregressive alternative to local diffusion pipelines.

Lumina-Image 2.0

Lumina-Image 2.0 official GitHub repository

Alpha-VLLM’s Lumina-Image 2.0 is a compact 2.6B 1024px generator with strong prompt adherence and an open fine-tuning path. Practitioners find it capable for composition, but describe unstable anatomy, uneven detail, and a demanding setup relative to mature SDXL workflows.

  • Benchmark: The Artificial Analysis arena recorded Elo 968 and rank 36 on July 30, 2026.
  • Size and hardware: The checkpoint has 2.6B parameters and supports CPU offload.
  • Run locally: Use Diffusers, the official PyTorch repository, or ComfyUI.
  • Access and terms: The project uses Apache 2.0; its Gemma 2 text encoder requires Hugging Face access.

Choose Lumina when an Apache-licensed, fine-tunable 2.6B base is more important than ecosystem maturity.

Sana

Sana official GitHub repository

NVIDIA’s Sana family uses linear attention and aggressive latent compression for efficient high-resolution generation. It is attractive for small models and fast Sprint checkpoints, although practitioners characterize Sprint as extremely fast but visibly lower quality.

  • Benchmark: The Artificial Analysis arena recorded Sana Sprint 1.6B at Elo 934 and rank 40 on July 30, 2026. NVIDIA reports Sana 1.5 1.6B at GenEval 0.82 and DPG 84.5.
  • Size and hardware: Releases span 0.6B, 1.6B, and 4.8B. NVIDIA documents 4K inference within about 8GB using tiling, offload, and quantization.
  • Run locally: Use Diffusers, SGLang, the official repository, or ComfyUI.
  • Access and terms: The code is Apache 2.0; model weights use NVIDIA’s Open Model License.

Choose Sana when efficiency, resolution, and fine-tuning openness matter more than arena rank.

OmniGen2

OmniGen2 official GitHub repository

VectorSpaceLab’s OmniGen2 unifies text-to-image generation, instruction editing, and in-context composition. Users like its ability to combine references, but report slow local runs, inconsistent edits, oversaturation, and installation friction.

  • Benchmark: The Artificial Analysis arena recorded Elo 905 and rank 43 on July 30, 2026.
  • Size and hardware: Native inference needs about 17GB VRAM; CPU offload cuts that roughly in half, and slower sequential offload can run below 3GB.
  • Run locally: Use the official PyTorch repository or ComfyUI.
  • Access and terms: The project uses Apache 2.0.

Choose OmniGen2 when multi-image composition and unified editing matter more than speed.

BAGEL

BAGEL official GitHub repository

ByteDance Seed’s BAGEL is a 7B-active, 14B-total Mixture-of-Transformer-Experts model that combines understanding, generation, and editing. Practitioners praise its ambitious unified workflow, but report blurry outputs, aggressive safety behavior, anatomy failures, and heavy local memory use.

  • Benchmark: The publisher reports GenEval 0.82 without rewriting and 0.88 with rewriting, plus WISE 0.52 and 0.70 respectively. The Artificial Analysis arena recorded Elo 900 and rank 45 on July 30, 2026.
  • Size and hardware: BAGEL has 7B active and 14B total parameters; official guidance provides NF4 for 12GB–32GB, INT8 for 22GB–32GB, and full precision above 32GB.
  • Run locally: Use ByteDance’s official PyTorch repository.
  • Access and terms: The family uses Apache 2.0.

Choose BAGEL to experiment with a single multimodal model across understanding, generation, and editing.

CogView4

CogView4 official GitHub repository

Z.ai’s CogView4 is a 6B bilingual generator designed for Chinese and English prompts and native high-resolution output. Community interest centers on its text accuracy and prompt length, while users warn that its large text encoder and custom stack make it heavier than the 6B headline suggests.

  • Benchmark: Z.ai reports DPG-Bench 85.13 and Chinese text F1 0.6168.
  • Size and hardware: At 1024px and batch four, the publisher reports 35GB normally, 20GB with CPU offload, and 13GB with offload plus a four-bit text encoder.
  • Run locally: Use Diffusers or the CogKit toolkit.
  • Access and terms: The project uses Apache 2.0.

Choose CogView4 for bilingual text-heavy work when you can accommodate its text encoder.

Kandinsky 5

Kandinsky 5 official GitHub repository

Kandinsky Lab’s current family includes 6B text-to-image and editing checkpoints with strong Russian concepts and typography, plus a related video stack. One community workflow praises its skin texture and image-to-image results, but calls the family an underrated underdog and needs offload workarounds to fit it on an 8GB GPU.

  • Performance: The publisher reports 13-second latency for the 6B T2I Lite checkpoint on an H100 80GB and a number-one open-source LMArena ranking for the related Video Pro model; treat the latency as a server-GPU result, not a consumer estimate.
  • Run locally: Use Kandinsky’s official PyTorch repository or Diffusers.
  • Access and terms: The project uses the MIT License.

Choose Kandinsky 5 for Russian-language concepts, typography, or a shared image-and-video research stack.

F-Lite

F-Lite official GitHub repository

Freepik and fal built F-Lite from licensed, SFW data, with Standard, Texture, and smaller 7B releases. That provenance is its clearest differentiator; practitioners welcome the licensing story but note limited text rendering and malformed anatomy in some outputs.

  • Size and hardware: The primary checkpoints are 10B and the publisher recommends at least 24GB VRAM.
  • Run locally: Use Diffusers, the official repository, or ComfyUI.
  • Access and terms: The weights use CreativeML Open RAIL-M.

Choose F-Lite when licensed-data provenance and commercial safety review are central to the evaluation.

Janus-Pro

Janus-Pro official GitHub repository

DeepSeek’s Janus-Pro combines visual understanding and image generation in 1B and 7B checkpoints. Practitioners praise the 1B release’s literal prompt adherence and speed, but its hard-coded 384px output and weaker aesthetic finish make it more of a compact research tool than a production image model.

  • Benchmark: The Artificial Analysis arena recorded Janus-Pro at Elo 711 and rank 48 on July 30, 2026. Z.ai’s shared table reports Janus-Pro-7B at DPG-Bench 84.19.
  • Run locally: Use DeepSeek’s official PyTorch repository or ComfyUI.
  • Access and terms: Code is MIT; weights use the DeepSeek Model License.

Choose Janus-Pro for compact multimodal experiments rather than maximum-resolution image work.

Kolors

Kolors official GitHub repository

Kuaishou’s Kolors is a bilingual latent-diffusion family for Chinese and English generation, portraits, IP-Adapter workflows, ControlNet, and inpainting. Practitioners value its Chinese handling and portraits, but prompt comparisons find weaker complex-scene adherence than newer models.

  • Benchmark: Kuaishou reports MPS 10.3, human overall satisfaction 3.59, visual appeal 3.99, and text faithfulness 4.17 on its KolorsPrompts evaluation.
  • Run locally: Use Diffusers, the official repository, or ComfyUI.
  • Access and terms: Code is Apache 2.0; the weights are open for academic research, while commercial use requires registration or separate permission.

Choose Kolors for bilingual portrait and control workflows when its model terms fit the deployment.

AuraFlow

AuraFlow official Hugging Face model card

Fal’s AuraFlow is a fully Apache-licensed flow-based generator built around literal prompt adherence. Practitioners found v0.2 unusually good at complex composition, but also described weak aesthetics, anatomy problems, and a clip-art-like look; v0.3 improved cohesion for some prompts while regressing adherence for others.

  • Benchmark: The publisher’s model collection reports GenEval above 0.7 for the AuraFlow v0.x series.
  • Run locally: Use Hugging Face Diffusers or ComfyUI.
  • Access and terms: The model card and weights use Apache 2.0.

Choose AuraFlow when permissive terms and literal composition matter more than polished default aesthetics.

PixArt-Σ

PixArt-Σ official GitHub repository

The PixArt team’s PixArt-Σ is a compact diffusion-transformer family for 512px, 1024px, and 2K generation. Practitioners praise its prompt understanding and low training cost, while reporting weaker default aesthetics, anatomy errors, and a need for refinement.

  • Benchmark: Z.ai’s shared table reports PixArt-α at DPG-Bench 71.11; this older sibling score should not be treated as a PixArt-Σ result.
  • Size and hardware: The main DiT is approximately 0.6B parameters. Community workflows report local use around 6GB VRAM, but hardware needs vary with the T5 text encoder and offload.
  • Run locally: Use Diffusers or the official PyTorch repository.
  • Access and terms: The project uses Apache 2.0.

Choose PixArt-Σ when you want a compact, permissively licensed base for high-resolution generation and fine-tuning.

How to choose your first local test

  1. Pick the capability: generation, typography, editing, or multi-reference composition.
  1. Select the exact released checkpoint rather than the family name alone.
  1. Confirm the checkpoint’s current access and commercial terms.
  1. Match its published GPU configuration to your machine.
  1. Compare exact releases under one benchmark protocol.
  1. Run your own prompts, resolutions, edit inputs, and seeds before production use.

For consumer hardware, the publisher figures point to FLUX.2 klein 4B at approximately 13GB, Z-Image Turbo at 16GB, ERNIE-Image at 24GB, LongCat-Image around 17GB with offload, and OmniGen2 around 17GB natively. Sana and Boogu publish lower-memory configurations through quantization and offload. HunyuanImage 3.0, GLM-Image, and Cosmos 3 belong in the multi-GPU or greater-than-80GB class.

Is FLUX.1 better than Stable Diffusion 3.5?

The live arena places FLUX.1 dev at Elo 1,029, narrowly above SD3.5 Large Turbo at 1,024 and Large at 1,023, but that does not make it a universal winner. FLUX.1 is usually the stronger prompt-following baseline; Stable Diffusion retains a deeper fine-tuning and control ecosystem. Compare their exact checkpoints on your prompts, then include license and hardware fit in the decision.

If local model operations fall outside your workflow, Magic Hour’s AI Image Generator exposes selectable flux-schnell, flux-2-klein, and z-image-turbo model IDs through its hosted product. The AI Image Generator API reference documents model selection and job creation.

Runbo Li
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
See all articles
Magic Hour
Create content faster with AI.
Trusted by 3M+ creators and teams

Related Posts

Build Flexible Video Pipelines Without Lock-In or Heavy Infrastructure
Videos
7 Best Open-Source-Friendly Video AI APIs in 2026 (Build Faster Without Lock-In)
Jan 21, 2026
Best Open Source AI Video Generation Models
Videos
Best Open Source AI Video Generation Models
Jan 01, 2026
openai chatgpt 4o cover
Images
GPT-4o Image Generation Review: The Best AI Image Generator Yet?
Mar 25, 2025
AI UGC Ad Generators (2026)
Top Choice
Best AI UGC Ad Generators (2026): Scripts, Avatars, and Ad-Ready Exports
May 15, 2026
Best AI Image Editing Models With Reference Images
ImagesTop Choice
Best AI Image Editing Models With Reference Images (2026): Keep Identity and Style While You Edit
Mar 03, 2026
AI Subtitle Generators (2026)
Top ChoiceVideos
Best AI Subtitle Generators (2026): Fast Captions, Styling, and Translation
May 08, 2026