IMAGE
AI image generation and editing: open models, tools, workflows, comparisons and practical implementation.

Open Source AI Image Generators: Models, Tools, and Local Workflows
Open-source AI image generators compared across FLUX, Stable Diffusion, editing, ComfyUI workflows, licensing, local hardware and practical use cases.

Jaaz: An Open-Source, Local-First Canvas for AI Images and Video
Jaaz is an open-source Canva AI alternative and local-first creative canvas for AI images and video, with an AI agent that can use local ComfyUI or cloud models. License unverified.

Stable Virtual Camera: Stability AI’s Model for New Camera Angles
Stable Virtual Camera is a 1.3B diffusion model for novel view synthesis, generating 3D-consistent camera views from one or more input images with controllable target cameras. License unverified.

ReVersion: Teaching Diffusion Models a Relation From Examples
ReVersion is a diffusion-based relation inversion method that learns visual relationships from a few example images and applies the learned relation to new subjects and scenes. License unverified.

VTP: MiniMax’s Research on Scalable Visual Tokenizers
VTP explores scalable visual tokenizer pre-training for image generation, combining representation learning and reconstruction to improve generative model scaling. Code and weights available; license unverified.

RedInk: Self-Hosted Generator for Xiaohongshu Image Posts
RedInk turns one sentence into a multi-page Xiaohongshu-style image post using Gemini and Nano Banana Pro via your own API keys. Docker install. License unverified.

PromptEnhancer: Tencent Hunyuan’s Prompt Rewriter for Image Models
PromptEnhancer (CVPR 2026) rewrites rough prompts into structured ones for text-to-image and image editing. 7B and 32B models, GGUF options. License unverified.

DeepSeek Ships an Experimental Vision Model, Claims It Rivals Opus 4.8 on Agent Tasks
DeepSeek quietly released DeepSeek-V4-Flash-Vision-Exp, a multimodal variant the company says closes the gap with Anthropic's Opus 4.8 on agent benchmarks that require visual understanding.
