

The best local image model is the one whose capabilities, checkpoint terms, and hardware requirements fit your workflow. Treat Base, Turbo, Edit, Dev, quantized, and revision checkpoints as choices within a family rather than separate models.
This guide compares 29 families with downloadable weights and a publisher-documented local inference path: FLUX.1, FLUX.2, Z-Image, Stable Diffusion, Qwen-Image, HunyuanImage 3.0, GLM-Image, ERNIE-Image, HiDream-O1-Image, DeepGen 1.0, JoyAI-Image, Krea 2, Ideogram 4, Boogu-Image-0.1, Cosmos 3, LongCat-Image, FIBO, Infinity, Lumina-Image 2.0, Sana, OmniGen2, BAGEL, CogView4, Kandinsky 5, F-Lite, Janus-Pro, Kolors, AuraFlow, and PixArt-Σ.
This roster consolidates related checkpoints under their upstream family and excludes fine-tunes, repackages, quantized mirrors, UI layers, and hosted-only systems. License terms can change, so verify the exact checkpoint’s current license before commercial use.
Family | Recommended for | Variants | How to run it locally | Recommended Hardware |
|---|---|---|---|---|
General generation and Kontext editing · hosted in Magic Hour | schnell, dev, Kontext dev | Official PyTorch code, Diffusers, ComfyUI | No publisher minimum found | |
Generation and editing on consumer GPUs · hosted in Magic Hour | klein 4B and 9B; Base variants; dev | Official PyTorch code and Diffusers | ||
Fast local photorealism · hosted in Magic Hour | Base, Turbo | Diffusers and ComfyUI | ||
Mature ecosystem and broad style work | SDXL Base/Refiner; SD3.5 Large, Turbo, Medium | Diffusers, official PyTorch repos, ComfyUI | SD3.5 Medium: 9.9GB excluding text encoders; SDXL Base test: 11.47GB–28.09GB | |
Text rendering and precise editing | Qwen-Image, 2512, Edit-2511, Layered | Diffusers and ComfyUI | No publisher minimum found | |
Knowledge-rich, large-scale generation | Base, Instruct, Distil | Official PyTorch repo | ||
Text-heavy and knowledge-intensive images | GLM-Image | Diffusers or SGLang | ||
Standard and fast bilingual generation | Standard, Turbo | Diffusers | ||
Generation, editing, and references | Full, Dev, Dev-2604 | Official PyTorch repo | ||
Compact unified generation and editing | SFT, RL | Diffusers | No publisher minimum found | |
Spatial and camera-aware editing | Edit, Edit-Plus | Official PyTorch repo and ComfyUI | ||
Illustration and varied styles | Raw, Turbo | Official PyTorch repo and ComfyUI | No publisher minimum found | |
Typography and structured design | 9.3B nf4, fp8 | Official PyTorch repo and ComfyUI | No publisher minimum found | |
Generation and editing at several memory tiers | Base, Turbo, Edit, Edit-Turbo | Official PyTorch repo and ComfyUI | ||
Frontier-scale image and world generation | Super-Text2Image, four-step variant | Diffusers, vLLM-Omni, SGLang | ||
Bilingual generation, text, and editing | Base, Dev, Edit, Edit-Turbo | Diffusers and ComfyUI | ||
Structured, repeatable art direction | FIBO generation and editing releases | Diffusers and ComfyUI | No publisher minimum found | |
Fast autoregressive image generation | 2B, 8B | Official PyTorch repo and Docker | No publisher minimum found | |
Efficient 1024px generation | 2.6B checkpoint | Diffusers, official PyTorch repo, ComfyUI | CPU offload supported; no minimum found | |
Efficient high-resolution generation | 0.6B, 1.6B, 4.8B, Sprint | Diffusers, SGLang, ComfyUI | ||
Generation, editing, and in-context composition | OmniGen2 | Official PyTorch repo and ComfyUI | ||
Unified generation, editing, and understanding | 7B-active/14B-total MoT | Official PyTorch repo | ||
Chinese/English text and high-resolution generation | CogView4-6B | Diffusers and CogKit | ||
Russian/English generation and editing | T2I Lite, Image Editing | Official PyTorch repo and Diffusers | No publisher minimum found | |
Copyright-safe, SFW generation | Standard, Texture, 7B | Diffusers and ComfyUI | ||
Unified understanding and compact generation | 1B, 7B | Official PyTorch repo and ComfyUI | No publisher minimum found | |
Chinese/English generation and portraits | Base, IP-Adapter, ControlNet, inpainting | Diffusers and ComfyUI | No publisher minimum found | |
Literal prompt following under Apache 2.0 | v0.1–v0.3 | Diffusers and ComfyUI | No publisher minimum found | |
Efficient high-resolution generation | 512px, 1024px, 2K | Diffusers and official PyTorch repo | No publisher minimum found |

Black Forest Labs’ FLUX.1 is a 12B family spanning fast generation with schnell, higher-quality generation with dev, and editing through Kontext. It remains a strong local baseline for literal prompt following and complex scenes, but practitioners also describe a narrow stylistic range, recurring faces, and overly smooth photographic skin; schnell trades more detail for speed.
Choose schnell when permissive commercial terms and speed matter. Evaluate dev or Kontext when their quality or editing capabilities justify the extra restrictions.
To use FLUX.1 without a local GPU, select flux-schnell in Magic Hour’s AI Image Generator or pass model="flux-schnell" to POST /v1/ai-image-generator.
It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with 1–4 images per job.
If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

FLUX.2 combines generation and editing, with klein aimed at local use. Early users praise its speed, positional prompting, style range, and edit mode, while 4B testers report unwanted changes and inconsistent anatomy or text.
To use FLUX.2 without a local GPU, select flux-2-klein in Magic Hour’s AI Image Generator or pass model="flux-2-klein" to POST /v1/ai-image-generator.
It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with one image per job. Pro, Flex, and Max are hosted offerings rather than additional local families.
If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

Tongyi-MAI’s Z-Image is a compact 6B family whose Turbo checkpoint is popular for fast, crisp photorealism on modest hardware. Users praise its quality-to-speed ratio but report limited seed diversity and weaker handling of multi-person scenes, left-versus-right instructions, and uncommon objects. Base is slower but often preferred for stylistic exploration.
Choose Turbo for a fast consumer-GPU path and Base when its slower, less-distilled behavior better suits your style.
To use Z-Image without a local GPU, select z-image-turbo in Magic Hour’s AI Image Generator or pass model="z-image-turbo" to POST /v1/ai-image-generator.
It is available on every tier, including free, from 5 credits per image at 640px, 1K, or 2K, with 1–4 images per job.
If you're new to MagicHour API, use coupon FIRSTAPI to get 10% off. Sign up here.

Stability AI’s local family includes the mature SDXL ecosystem and the newer SD3.5 Large, Large Turbo, and Medium checkpoints. Practitioners value SD3.5’s style range and creativity but often find FLUX more reliable for people and long prompts; hands, feet, and anatomy remain recurring complaints. SDXL is older, yet local users still value its fine-tune, ControlNet, and face-workflow ecosystem.
Choose this family when ecosystem depth, fine-tuning, and control tooling matter as much as base-model leaderboard position.

Alibaba’s Qwen team built Qwen-Image around text rendering, prompt understanding, and precise editing. Local users particularly value Edit for preserving faces, poses, and scene structure, while realistic edits can look plasticky or airbrushed; testers report that 2512 improves skin, hands, and small details.
Choose the generation checkpoint for typography-heavy prompts and the Edit or Layered releases when structural preservation matters.

Tencent’s HunyuanImage 3.0 is an unusually large native multimodal family for knowledge-rich prompts, text rendering, and complex composition. One early practitioner found that expanded prompts could produce striking results and stylized lettering, but simple prompts were more average and Base felt unfinished. Its footprint keeps it outside ordinary desktop workflows.
Choose it only when the model’s semantic and typography strengths justify a multi-GPU deployment and its territory terms fit your use.

Z.ai’s GLM-Image combines an autoregressive semantic stage with a diffusion decoder, targeting text-heavy and knowledge-intensive images. Early testers found it creative and promising for layout, but ordinary text-to-image output can feel underbaked and slow beside lighter models.
Choose GLM-Image for dense text and semantic layout rather than as the easiest general-purpose local generator.

Baidu’s ERNIE-Image is an 8B text-to-image family with standard and Turbo checkpoints. Community tests praise clean illustration, coherent backgrounds, and solid text, while reporting prompt-expander drift, camera-angle misses, diagonal artifacts, and an overprocessed HDR look.
Choose standard for maximum quality or Turbo when its eight-step path matters more than the small benchmark gap.

HiDream’s O1 family is an 8B native multimodal model for generation, editing, and multi-reference personalization. Its broad workflow is attractive, but early testing is mixed: users report lost detail in distilled Dev outputs and plastic skin or artifact-like texture in some full-model results.
Choose HiDream when one local family must cover generation, editing, and subject-driven personalization.

DeepGen Team’s DeepGen 1.0 is a compact 5B research model that unifies generation, editing, and reasoning-oriented control. Local-model users like the unified design, but early discussion flags its fixed 512×512 experiments and immature UI and fine-tuning ecosystem.
Choose DeepGen when compact unified generation and editing matter more than a mature local ecosystem.

JD’s JoyAI-Image is an edit-first family with Edit and Edit-Plus checkpoints plus text-to-image generation. Early testers found it strong at spatial changes and alternate camera angles, but saw facial-detail drift and distortion at extreme views. Adoption has also been slowed by a heavy runtime and rough early ComfyUI workflows.
Choose JoyAI when spatially aware editing and camera changes are more important than lightweight deployment.

Krea 2 is built around style diversity, with Raw for exploration and Turbo for speed. The team says it prioritized illustration and varied styles over realism; practitioners praise anatomy, animals, and wide compositions while reporting weaker photography, occasional prompt misses, VAE grid texture, and aggressive safety behavior.
Choose Raw for style exploration and Turbo when iteration speed matters.

Ideogram’s first open-weight family is a design-focused 9.3B generator for multilingual typography, structured JSON prompts, bounding boxes, palettes, and native 2K output. Practitioners like its direct layout control, but criticize missing full-precision weights and restrictive safety behavior.
Choose Ideogram 4 for local, non-commercial typography and layout work.

The Boogu Project’s first family covers Base, Turbo, Edit, and Edit-Turbo. A 192-prompt practitioner test found high detail and good overall fidelity but a strong symmetry bias and weak composition; edit users report failures on complex changes and artifacts in small faces, limbs, eyes, and text.
Choose Boogu when you want generation and editing variants with documented memory tradeoffs.

NVIDIA’s Cosmos 3 is a 64B omnimodal world-model family with a specialized text-to-image checkpoint. It leads the current open-weights arena, but local practitioners stress that its reasoner-plus-generator footprint is impractical on ordinary desktops.
Choose Cosmos 3 only when frontier-scale quality is worth data-center-class infrastructure.

Meituan’s 6B LongCat-Image family covers bilingual generation, typography, and editing. Users like its product and face preservation, but report slow Base inference, uneven editing quality, and occasional oversaturated or plastic output.
Choose LongCat for bilingual text rendering or editing when 17GB-class offloaded inference is acceptable.

Bria AI’s 8B FIBO is built for structured, repeatable art direction: a VLM expands an idea into a long JSON description that controls lighting, camera, composition, and color. Practitioners praise its prompt adherence and texture, but report weaker hands and note that the local release lacks broad native ComfyUI support.
Choose FIBO when reproducible, programmatic visual control matters more than free-form prompting.

FoundationVision’s Infinity is a bitwise autoregressive generator rather than a diffusion model. Its 2B and 8B checkpoints target fast 1024px synthesis and strong composition. Practitioners like the prompt adherence of autoregressive models, but some see little improvement in the familiar “AI-generated” look and note that the larger 20B release is still absent.
Choose Infinity when you want to evaluate an autoregressive alternative to local diffusion pipelines.

Alpha-VLLM’s Lumina-Image 2.0 is a compact 2.6B 1024px generator with strong prompt adherence and an open fine-tuning path. Practitioners find it capable for composition, but describe unstable anatomy, uneven detail, and a demanding setup relative to mature SDXL workflows.
Choose Lumina when an Apache-licensed, fine-tunable 2.6B base is more important than ecosystem maturity.

NVIDIA’s Sana family uses linear attention and aggressive latent compression for efficient high-resolution generation. It is attractive for small models and fast Sprint checkpoints, although practitioners characterize Sprint as extremely fast but visibly lower quality.
Choose Sana when efficiency, resolution, and fine-tuning openness matter more than arena rank.

VectorSpaceLab’s OmniGen2 unifies text-to-image generation, instruction editing, and in-context composition. Users like its ability to combine references, but report slow local runs, inconsistent edits, oversaturation, and installation friction.
Choose OmniGen2 when multi-image composition and unified editing matter more than speed.

ByteDance Seed’s BAGEL is a 7B-active, 14B-total Mixture-of-Transformer-Experts model that combines understanding, generation, and editing. Practitioners praise its ambitious unified workflow, but report blurry outputs, aggressive safety behavior, anatomy failures, and heavy local memory use.
Choose BAGEL to experiment with a single multimodal model across understanding, generation, and editing.

Z.ai’s CogView4 is a 6B bilingual generator designed for Chinese and English prompts and native high-resolution output. Community interest centers on its text accuracy and prompt length, while users warn that its large text encoder and custom stack make it heavier than the 6B headline suggests.
Choose CogView4 for bilingual text-heavy work when you can accommodate its text encoder.

Kandinsky Lab’s current family includes 6B text-to-image and editing checkpoints with strong Russian concepts and typography, plus a related video stack. One community workflow praises its skin texture and image-to-image results, but calls the family an underrated underdog and needs offload workarounds to fit it on an 8GB GPU.
Choose Kandinsky 5 for Russian-language concepts, typography, or a shared image-and-video research stack.

Freepik and fal built F-Lite from licensed, SFW data, with Standard, Texture, and smaller 7B releases. That provenance is its clearest differentiator; practitioners welcome the licensing story but note limited text rendering and malformed anatomy in some outputs.
Choose F-Lite when licensed-data provenance and commercial safety review are central to the evaluation.

DeepSeek’s Janus-Pro combines visual understanding and image generation in 1B and 7B checkpoints. Practitioners praise the 1B release’s literal prompt adherence and speed, but its hard-coded 384px output and weaker aesthetic finish make it more of a compact research tool than a production image model.
Choose Janus-Pro for compact multimodal experiments rather than maximum-resolution image work.

Kuaishou’s Kolors is a bilingual latent-diffusion family for Chinese and English generation, portraits, IP-Adapter workflows, ControlNet, and inpainting. Practitioners value its Chinese handling and portraits, but prompt comparisons find weaker complex-scene adherence than newer models.
Choose Kolors for bilingual portrait and control workflows when its model terms fit the deployment.

Fal’s AuraFlow is a fully Apache-licensed flow-based generator built around literal prompt adherence. Practitioners found v0.2 unusually good at complex composition, but also described weak aesthetics, anatomy problems, and a clip-art-like look; v0.3 improved cohesion for some prompts while regressing adherence for others.
Choose AuraFlow when permissive terms and literal composition matter more than polished default aesthetics.

The PixArt team’s PixArt-Σ is a compact diffusion-transformer family for 512px, 1024px, and 2K generation. Practitioners praise its prompt understanding and low training cost, while reporting weaker default aesthetics, anatomy errors, and a need for refinement.
Choose PixArt-Σ when you want a compact, permissively licensed base for high-resolution generation and fine-tuning.
For consumer hardware, the publisher figures point to FLUX.2 klein 4B at approximately 13GB, Z-Image Turbo at 16GB, ERNIE-Image at 24GB, LongCat-Image around 17GB with offload, and OmniGen2 around 17GB natively. Sana and Boogu publish lower-memory configurations through quantization and offload. HunyuanImage 3.0, GLM-Image, and Cosmos 3 belong in the multi-GPU or greater-than-80GB class.
The live arena places FLUX.1 dev at Elo 1,029, narrowly above SD3.5 Large Turbo at 1,024 and Large at 1,023, but that does not make it a universal winner. FLUX.1 is usually the stronger prompt-following baseline; Stable Diffusion retains a deeper fine-tuning and control ecosystem. Compare their exact checkpoints on your prompts, then include license and hardware fit in the decision.
If local model operations fall outside your workflow, Magic Hour’s AI Image Generator exposes selectable flux-schnell, flux-2-klein, and z-image-turbo model IDs through its hosted product. The AI Image Generator API reference documents model selection and job creation.
