MMagic 中的 Stable Diffusion 文生图与图像修复实战:配置解析、推理调用与 ToMe 加速指南
媒体生成计算机视觉深度学习人工智能大模型【免费下载链接】mmagicOpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic : Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.项目地址https://gitcode.com/gh_mirrors/mm/mmagic点击查看免费下载本文以 OpenMMLab 的 MMagic 工具箱中configs/stable_diffusion/下的官方配置与文档为主线系统讲解如何在 MMagic 中完成 Stable Diffusion v1.5 的文生图Text2Image与图像修复Inpainting推理并深入解析 ToMeToken Merging加速插件的配置参数、底层实现与实测性能。读完本文你将能够基于仓库内配置文件快速搭建推理脚本掌握StableDiffusion/StableDiffusionInpaint两个模型的调用方式并学会用tomesd_cfg一键为扩散模型提速。Stable Diffusion 在 MMagic 中的定位Stable Diffusion 是一种以 CLIP 文本编码器输出的文本嵌入为条件condition的潜空间扩散模型latent diffusion model可以从自然语言描述直接生成图像。它建立在 CVPR22 工作High-Resolution Image Synthesis with Latent Diffusion Models之上官方代码与 HuggingFace diffusers 均有实现。MMagic 在 configs/stable_diffusion/README.md 中说明引入该算法的目的方便社区在同一框架下学习、对比不同的 text2image 方法。与原生 diffusers 管线一致MMagic 中的 Stable Diffusion 由以下组件构成见 stable-diffusion_ddim_denoisingunet.pyVAEAutoencoderKL负责图像 ↔ 潜空间的编解码UNetUNet2DConditionModel在潜空间内执行条件去噪文本编码器ClipWrapperclip_typehuggingface加载text_encoder子目录输出条件嵌入TokenizerCLIP tokenizertokenizer字段直接给模型名/路径调度器EditDDIMScheduler训练与测试均使用 DDIM 采样。注意MMagic 的这套配置直接以from_pretrained方式复用 HuggingFace diffusers 的权重与网络结构注释中明确写着 Use DiffuserWrapper!因此网络骨架细节UNet 各层的 channels、heads 等不再像自研DenoisingUnet那样逐项手写——被注释掉的DenoisingUnet/EditAutoencoderKL配置块保留了可选的纯 MMagic 实现方案方便对照学习。三种配置与预训练权重准备configs/stable_diffusion/目录下提供三个模型配置配置任务关键差异stable-diffusion_ddim_denoisingunet.pyText2Image标准 v1.5无加速stable-diffusion_ddim_denoisingunet-tomesd_5e-1.pyText2Image追加tomesd_cfgdict(ratio0.5)stable-diffusion_ddim_denoisingunet-inpaint.pyInpainting模型类型为StableDiffusionInpaint权重源为runwayml/stable-diffusion-inpainting三者共用同一套调度器配置diffusion_scheduler dict( typeEditDDIMScheduler, variance_typelearned_range, beta_end0.012, beta_schedulescaled_linear, beta_start0.00085, num_train_timesteps1000, set_alpha_to_oneFalse, clip_sampleFalse)这些参数刻画了 v1.5 官方权重的噪声调度num_train_timesteps1000表示训练时使用 1000 步扩散beta_schedulescaled_linear、beta_start0.00085、beta_end0.012对应原版噪声强度曲线set_alpha_to_oneFalse与clip_sampleFalse保证与 diffusers 原始 pipeline 行为一致。模型使用 Stable Diffusion v1.5 的权重包含 vae、unet、clip 三部分。可以按 README 提供的方式下载git lfs install git clone https://huggingface.co/runwayml/stable-diffusion-v1-5下载后把配置中的from_pretrained指向权重目录即可例如# config.model.unet.from_pretrained /path/to/your/stable-diffusion-v1-5 # config.model.vae.from_pretrained /path/to/your/stable-diffusion-v1-5不手动指定时代码会以runwayml/stable-diffusion-v1-5为标识自动从 HuggingFace Hub 拉取权重。修复模型对应使用runwayml/stable-diffusion-inpainting。文生图快速开始README 给出的最小文生图示例quick start如下from mmengine import MODELS, Config from torchvision import utils from mmengine.registry import init_default_scope init_default_scope(mmagic) config configs/stable_diffusion/stable-diffusion_ddim_denoisingunet.py config Config.fromfile(config).copy() # change the pretrained_model_path if you have downloaded the weights manually # config.model.unet.from_pretrained /path/to/your/stable-diffusion-v1-5 # config.model.vae.from_pretrained /path/to/your/stable-diffusion-v1-5 StableDiffuser MODELS.build(config.model) prompt A mecha robot in a favela in expressionist style StableDiffuser StableDiffuser.to(cuda) image StableDiffuser.infer(prompt)[samples][0] image.save(robot.png)流程要点init_default_scope(mmagic)将默认注册表作用域切换为mmagic使MODELS.build能解析StableDiffusion等 MMagic 注册的类Config.fromfile(config).copy()读取并复制配置便于在内存中修改from_pretrained指向MODELS.build(config.model)构建完整模型对象内部会依次构建unet、vae、text_encoder、tokenizer、scheduler、test_scheduler见 stable_diffusion.py 的__init__infer(prompt)返回{samples: [...]}samples[0]是 PIL 图像。infer 核心参数详解从 StableDiffusion.infer 的签名与 docstring 可以整理出推理阶段最常用的参数参数默认值说明prompt必填str或str列表引导生成的文本height/width自动默认unet_sample_size * vae_scale_factor即 512必须能被 8 整除check_inputs会校验num_inference_steps50去噪步数越多通常质量越高、耗时越长guidance_scale7.5无分类器引导CFG强度 1.0时启用双路前向条件 无条件negative_promptNone负向提示词仅在启用 CFG 时生效num_images_per_prompt1每个 prompt 生成的图像数量eta0.0DDIM 的 η 参数仅对 DDIMScheduler 生效取值 [0, 1]seed1推理随机种子infer内部调用set_random_seedreturn_typeimageimage返回 PIL 列表numpy返回[N, C, H, W]数组tensor返回解码器输出张量latentsNone预生成的噪声潜变量可复现同一批潜变量下的多次生成推理的底层调用链结合 stable_diffusion.pyinfer内部按标准扩散管线依次执行默认height/width并做输入校验check_inputs_encode_prompttokenizer 编码 → CLIP 文本编码器前向 → 按num_images_per_prompt复制嵌入若启用 CFG则将无条件嵌入与条件嵌入torch.cat成一批一次前向同时得到两种预测test_scheduler.set_timesteps(num_inference_steps)生成时间步序列prepare_latents按(batch, latent_channels, H//scale, W//scale)采样高斯噪声并乘上scheduler.init_noise_sigma去噪循环unet预测噪声 → CFG 合并noise_pred_uncond guidance_scale * (noise_pred_text - noise_pred_uncond)→test_scheduler.step迭代decode_latents以1 / 0.18215缩放潜变量后交给 VAE 解码最终经output_to_pil归一化并转为 PIL 图像。另外StableDiffusion.train_step展示了 v1.5 的微调训练路径VAE 编码图像得到潜变量 → 加噪 → CLIP 编码 prompt → UNet 预测噪声 → MSE 损失FP32 计算并支持noise_offset_weight与v_prediction两种预测类型。图像修复InpaintingMMagic 的修复实现是独立的StableDiffusionInpaint模型stable_diffusion_inpaint.py它继承自StableDiffusion并重写infer。README 提供完整示例import mmcv from mmengine import MODELS, Config from mmengine.registry import init_default_scope from PIL import Image init_default_scope(mmagic) config configs/stable_diffusion/stable-diffusion_ddim_denoisingunet-inpaint.py config Config.fromfile(config).copy() # change the pretrained_model_path if you have downloaded the weights manually # config.model.unet.from_pretrained /path/to/your/stable-diffusion-inpainting # config.model.vae.from_pretrained /path/to/your/stable-diffusion-inpainting StableDiffuser MODELS.build(config.model) prompt a mecha robot sitting on a bench img_url https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo.png # noqa mask_url https://raw.githubusercontent.com/CompVis/latent-diffusion/main/data/inpainting_examples/overture-creations-5sI6fQgYIuo_mask.png # noqa image Image.fromarray(mmcv.imread(img_url, channel_orderrgb)) mask Image.fromarray(mmcv.imread(mask_url)).convert(L) StableDiffuser StableDiffuser.to(cuda) image StableDiffuser.infer( prompt, image, mask )[samples][0] image.save(inpaint.png)修复推理的关键差异在于infer额外接收image与mask_imagePIL Image / ndarray / Tensor 均可由prepare_mask_and_masked_image统一预处理缩放至目标尺寸、二值化 mask 0.5置 0 0.5置 1、图像归一化到[-1, 1]并生成masked_image image * (mask 0.5)prepare_mask_latents将 mask 插值到潜空间尺寸并用 VAE 编码 masked image 得到masked_image_latentsCFG 开启时 mask 与 masked image latents 都会按批次复制两份输入 UNet 时若num_channels_unet 9v1.5-inpainting 的默认情况会把[latents, mask, masked_image_latents]沿通道维拼接后送入网络stable_diffusion_inpaint.py 中的 denoising loop。需要注意修复模型当前只实现infer其val_step/test_step/train_step均直接raise NotImplementedError属于纯推理用途。使用 ToMe 加速 Stable DiffusionMMagic 集成了 tomesd——基于 ToMeToken Merging源自Token Merging: Your ViT but Faster开发的扩散模型 token 合并加速工具。启用方式极其简单在model配置里追加一个tomesd_cfg即可例如 stable-diffusion_ddim_denoisingunet-tomesd_5e-1.pymodel dict( typeStableDiffusion, unetunet, vaevae, enable_xformersFalse, text_encoderdict( typeClipWrapper, clip_typehuggingface, pretrained_model_name_or_pathstable_diffusion_v15_url, subfoldertext_encoder), tokenizerstable_diffusion_v15_url, schedulerdiffusion_scheduler, test_schedulerdiffusion_scheduler, tomesd_cfgdict( ratio0.5))唯一的硬性前置条件是torch 1.12.1ToMe 的 token 合并依赖torch.Tensor.scatter_reduce()源码 tome_utils.py 中会显式检查版本并抛出ImportError提示升级。运行 demo 前请务必确认环境满足该要求。tomesd_cfg 参数速查表README 对每个参数都给出了明确的语义与建议取值参数类型含义与建议ratiofloat要合并的 token 比例如 0.4 表示总 token 减少 40%。上限为1 - 1/(sx*sy)默认上限 0.75通常建议 0.5越大加速越多、画质损失越大max_downsampleint对“下采样不超过该值”的层应用 ToMe。1 表示仅作用于无下采样的层8 表示所有层取值应从 1/2/4/8 中选择推荐 1、2sx, sy(int, int)计算 dst 集合时的步长步长越大可合并的 token 越多默认 (2, 2) 在多数场景表现良好二者不必整除图像尺寸use_randbool计算 dst 集合时是否允许随机扰动默认 True若出现异常伪影可尝试关闭merge_attnbool是否对自注意力层合并 token推荐开启merge_crossattnbool是否对交叉注意力层合并 token不推荐merge_mlpbool是否对 MLP 层合并 token尤其不推荐加速效果评测脚本README 提供了一个端到端对比脚本先关掉tomesd_cfg置None跑基线再分别以ratio0.5、0.75复测统计生成 100 张图像的耗时import time import numpy as np from mmengine import MODELS, Config from mmengine.registry import init_default_scope init_default_scope(mmagic) _device 0 work_dir /path/to/your/work_dir config configs/stable_diffusion/stable-diffusion_ddim_denoisingunet-tomesd_5e-1.py config Config.fromfile(config).copy() # # change the pretrained_model_path if you have downloaded the weights manually # config.model.unet.from_pretrained /path/to/your/stable-diffusion-v1-5 # config.model.vae.from_pretrained /path/to/your/stable-diffusion-v1-5 # w/o tomesd config.model.tomesd_cfg None StableDiffuser MODELS.build(config.model).to(fcuda:{_device}) prompt A mecha robot in a favela in expressionist style # inference time evaluation params size 512 ratios [0.5, 0.75] samples_perprompt 5 t time.time() for i in range(100//samples_perprompt): image StableDiffuser.infer(prompt, heightsize, widthsize, num_images_per_promptsamples_perprompt)[samples][0] if i 0: image.save(f{work_dir}/wo_tomesd.png) print(fGenerating 100 images with {samples_perprompt} images per prompt, without ToMe speed-up, time used : {time.time() - t}s) for ratio in ratios: # w/ tomesd config.model.tomesd_cfg dict(ratioratio) sd_model MODELS.build(config.model).to(fcuda:{_device}) t time.time() for i in range(100//samples_perprompt): image sd_model.infer(prompt, heightsize, widthsize, num_images_per_promptsamples_perprompt)[samples][0] if i 0: image.save(f{work_dir}/w_tomesd_ratio_{ratio}.png) print(fGenerating 100 images with {samples_perprompt} images per prompt, merging ratio {ratio}, time used : {time.time() - t}s)实测性能对比README 官方数据README 给出了在单张 RTX 3090、torch 2.0.0cu118环境下的推理耗时对比512 尺寸、每 prompt 生成 5 张、共 100 张模型xformersRatioSize / 每 prompt 图数耗时秒stable_diffusion_v1.5-tomesdw/ow/o tome / 0.5 / 0.75512 / 5542.20 / 427.65↓21.1%/ 393.05↓27.5%stable_diffusion_v1.5-tomesdw/w/o tome / 0.5 / 0.75512 / 5541.64 / 428.53↓20.9%/ 396.38↓26.8%README 的结论是开启xformers时加速比例略有下降但tomesd仍能显著降低推理耗时当image_size与num_images_per_prompt较大时特别推荐启用tomesd——因为此时相似 token 更多合并收益更大。xformers的开关由模型配置中的enable_xformers控制仓库配置默认False底层由 set_xformers 对 diffusers 的注意力模块调用set_use_memory_efficient_attention_xformers。ToMe 加速的源码级原理MMagic 对 ToMe 的封装位于 tome_utils.py入口是StableDiffusion.__init__中调用的set_tomesd(self, **self.tomesd_cfg)model_utils.py。整体机制是在运行时对 UNet 的 Transformer block 做就地 patch二分软匹配bipartite soft matching核心函数bipartite_soft_matching_random2d把每张特征图按(sy, sx)网格划分 src/dst 集合用余弦相似度贪心找出最相似的一批 token 对随后通过scatter_reduce做均值合并use_randFalse时退化为固定取左上角no_rand避免随机扰动层级过滤build_merge根据tome_info[size]与当前层分辨率计算downsample倍数只有downsample max_downsample的层才真正执行合并并通过merge_attn/merge_crossattn/merge_mlp三个开关分别决定自注意力、交叉注意力、MLP 三段是否套用 merge/unmerge 函数forward 注入build_mmagic_tomesd_block/build_mmagic_wrapper_tomesd_block生成继承自原 block 类的ToMeBlock在其forward中按“norm → merge → attention → unmerge 残差”的顺序插入合并逻辑同时add_tome_cfg_hook注册 forward pre-hook 以捕获输入分辨率奇数 batch 保护若x.shape[0]为奇数CFG 下条件/无条件无法成对use_rand会被强制关闭以避免伪影。测试用例 test_stable_diffusion_tomesd.py 使用小尺寸DenoisingUnetimage_size128、base_channels32验证了tomesd_cfg配置能被正确构建与推理test_stable_diffusion.py 与 test_stable_diffusion_inpaint.py 则分别覆盖文生图与修复模型的构建和infer流程可作为自定义配置时的回归参考。扩展到更多 Stable Diffusion 生态模型tomesd_cfg同样适用于仓库中其他基于 Stable Diffusion 的模型。README 明确指出可以评估 DreamBooth 与 ControlNet 等模型的加速效果从源码结构看set_tomesd工具被 dreambooth.py、controlnet.py、animatediff.py、textual_inversion.py、stable_diffusion_xl.py 等多个编辑器复用意味着配置一行tomesd_cfg即获得 ToMe 加速这一模式在 MMagic 中是通用的。此外StableDiffusion相关组件还被DiffusersWrapper风格配置体系承载用户可参考 diffusers_pipeline 了解与 diffusers 生态的对接方式。总结在 MMagic 中使用 Stable Diffusion v1.5核心路径只有三步下载权重 → 用Config.fromfile读取 configs/stable_diffusion/ 下的配置并MODELS.build→ 调用infer输出 PIL 图像。在此基础上通过tomesd_cfg一行配置即可引入 ToMe token 合并加速在 512 分辨率、多图生成的场景下可获得约 20%27% 的推理耗时下降官方 README 数据。无论是文生图、图像修复还是在此基础上扩展的 DreamBooth、ControlNet 等模型这套配置驱动 注册表构建 infer 推理的范式都保持一致。注本文引用的性能数据来自仓库 configs/stable_diffusion/README.md 在单卡 RTX 3090、torch 2.0.0cu118 下的官方记录不同硬件、驱动与 PyTorch 版本下结果会有所差异实际加速效果请以本地复测为准。赞分享媒体生成计算机视觉深度学习人工智能大模型【免费下载链接】mmagicOpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic : Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.项目地址https://gitcode.com/gh_mirrors/mm/mmagic点击查看免费下载相关推荐MMagic 中的 Disco Diffusion 实战CLIP-Guided Diffusion 图文生成原理、配置与调参指南MMagic 中的 Disco Diffusion 实战CLIP Guided Diffusion 图文生成原理、配置与调参指南 本文以 MMagic 对 D媒体生成计算机视觉深度学习人工智能大模型Stable Diffusion v2 完整实践指南从文本到图像、深度条件生成到图像修复的推理与配置解析Stable Diffusion v2 完整实践指南从文本到图像、深度条件生成到图像修复的推理与配置解析 Stable Diffusion v2 是本仓库S计算机视觉媒体生成深度学习基础模型mmagic 中的 Disco Diffusion 实战CLIP 引导扩散的文本生图原理、配置与推理全解析mmagic 中的 Disco Diffusion 实战CLIP 引导扩散的文本生图原理、配置与推理全解析 Disco DiffusionDD是最早让“用媒体生成计算机视觉深度学习人工智能大模型上一篇go-smtp与标准库对比为什么选择更现代的SMTP解决方案下一篇Kite革命性Kubernetes仪表板如何简化集群管理10个必知核心功能创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考