Open Source Hugging Face Blog

Rebuilding AUTOMATIC1111 with Gradio Workflow

Gradio WorkflowStable DiffusionHugging Face SpacesAUTOMATIC1111

Following a previous post that built five small gr.Workflow graphs, the team now walks through Workflow1111, which rebuilds most of AUTOMATIC1111's feature set as one workflow canvas. The canvas is a graph of eleven media pipelines built from seventy-three nodes, bringing together state-of-the-art models for text-to-image, hi-resolution fix, image-to-image, prompt-matrix grids, VLM interrogate, detection-to-inpaint masks, ControlNet-style annotators, background removal, PNG Info storing, and image-to-video. Anyone can run the pipelines by signing in with a Hugging Face account or supplying an access token, with model calls drawing on their own quota; users can also duplicate the Space and rewire it for their own use case.

Every pipeline is composed from the same four operator kinds introduced in the earlier post: fn (a Python function), model (a model called through InferenceClient), space (another Gradio Space), and dataset (a row from a Hub dataset). Each node wraps one operator, and the operator's inputs and outputs become the ports that edges connect to.

Pipeline specifics: the core text-to-image node offers the controls expected from A1111's txt2img tab — negative prompt, steps, CFG, seed, width and height, plus a model_id field for choosing the checkpoint. The prompt first passes through a prompt-builder fn node that appends the selected style preset and cleans up the text, then into a model node that calls the checkpoint through Inference Providers, while a post-process fn node writes the generation parameters into the PNG's metadata on the way out — the same metadata the PNG Info pipeline later reads back. Hi-resolution fix, which in A1111 upscales the txt2img output and runs a second denoising pass, here becomes a two-node detour: the text-to-image result goes into a FLUX.1-Kontext model node with the refine instruction "enhance fine detail and micro-texture, keep the composition identical," returning a sharper, larger image. That same Kontext node doubles as the image-to-image tab: upload an image, describe the change, and it returns the edited image.

Another pipeline turns a rough prompt such as "A lighthouse in a storm" into a detailed one by sending it to a Qwen3-4B model node, with a small fn node converting the reply into a clean tag list capped at forty items (e.g. "stormy sea, wet rocks, dramatic composition, low angle shot, volumetric lighting, ominous tone"); any diffusion model node can be connected to that output to render the image. The post notes that unlike ComfyUI this requires no custom node, since in a Gradio workflow the LLM and the diffusion model are both ordinary model operators on the same canvas.

Read original →

← Back to home