Table of contents :#table-of-contents What is NeMo Automodel? :#what-is-nemo-automodel Supported diffusion models :#supported-diffusion-models What this collaboration unlocks :#what-this-collaboration-unlocks A look at the fine-tuning workflow :#a-look-at-the-fine-tuning-workflow 1. Pre-encode the dataset :#1-pre-encode-the-dataset 2. Launch training with the existing FLUX YAML :#2-launch-training-with-the-existing-flux-yaml 3. Generate from the fine-tuned checkpoint :#3-generate-from-the-fine-tuned-checkpoint 4. Performance :#4-performance Other Finetuned/LoRA examples :#other-finetunedlora-examples Try it today :#try-it-today Coming next: Pythonic recipe APIs :#coming-next-pythonic-recipe-apis Resources :#resources A joint post from NVIDIA and Hugging Face. Special thanks to Sayak Paul from Hugging Face for their contributions to the integration work and for co-authoring this blog.
Diffusion models power some of the most exciting open-source releases of the last two years — such as FLUX.1-dev:https://huggingface.co/black-forest-labs/FLUX.1-dev for text-to-image and Wan 2.1:https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers and HunyuanVideo:https://huggingface.co/hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v for text-to-video. The 🤗 Diffusers:https://github.com/huggingface/diffusers library has become the de facto home for these models, giving researchers and builders a single, consistent interface for inference, adaptation, and pipeline composition.
In addition, training and fine-tuning diffusion models are also on the rise, requiring utilities that offer memory-efficient sharding, latent caching, multiresolution bucketing, and configurations that scale gracefully from one GPU to hundreds.
To cater to these technical demands, we offer the NVIDIA NeMo Automodel:https://github.com/NVIDIA-NeMo/Automodel open-source library. Today, we're highlighting the collaboration between NVIDIA and Hugging Face that brings production-grade, distributed diffusion training to any Diffusers-format model on the Hugging Face Hub — with no checkpoint conversion and no model rewrites for any new model. The integration is documented in the Diffusers training guide:https://docs.nvidia.com/nemo/automodel/recipes-e2e-examples/diffusion-fine-tuning and is fully open source under Apache 2.0.
NeMo Automodel is an open-source PyTorch DTensor-native training library, part of the NVIDIA NeMo framework, built around two design principles that matter for the Diffusers ecosystem:
AutoModel currently supports flow-matching models only. Under the hood, it uses flow matching as the training objective, with latent-space training (via pre-encoded VAE outputs) and multiresolution bucketed dataloading to accelerate throughput.
NeMo Automodel integration ships with ready-to-use fine-tuning recipes for the open diffusion models below. The list reflects the recipes currently in examples/diffusion/finetune .
For Diffusers users, the practical gains break down into a few concrete capabilities.
No checkpoint conversion. Pretrained weights from the Hub work out of the box. There's no separate "training format" to convert to, then convert back. Your fine-tuned checkpoint loads directly into a DiffusionPipeline for inference, or back to the Hub for sharing. Downstream tools — quantization, compilation, LoRA adapters, custom samplers — all keep working.
Fast path to new model support. When a new diffusion model lands in Diffusers, enabling it in NeMo Automodel takes a small, contained code addition — a data preprocessing handler and a model adapter — rather than a full custom training script. The rest of the recipe stack (FSDP2, bucketed dataloading, checkpointing, generation) carries over unchanged, and the same YAML-driven workflow applies.
Full and parameter-efficient fine-tuning. Both full fine-tuning and LoRA-style PEFT are supported, so you can choose between maximum quality (full FT on a large cluster) or maximum efficiency (LoRA on a single node). The same recipe structure handles both.
Scalable training that goes beyond what built-in scripts offer. NeMo Automodel adds sharding schemes such as FSDP2, tensor, context, and pipeline parallelisms, multi-node orchestration (SLURM today, Kubernetes coming), and multiresolution bucketing. These capabilities make training larger models like FLUX.1-dev (12B) and HunyuanVideo (13B) possible.
In this section, we walk through the typical workflow for fine-tuning any of the supported models. The recommended way to install Automodel is the NeMo Automodel Docker container ( nvcr.io/nvidia/nemo-automodel:26.06 ), which ships with PyTorch, TransformerEngine, and other CUDA-compiled dependencies pre-built. Alternatively, install with pip3 install nemo-automodel or from source ( pip3 install git+https://github.com/NVIDIA-NeMo/Automodel.git ); see the installation guide:https://docs.nvidia.com/nemo/automodel/latest/get-started/installation for all options.
This guide walks through a full-transformer fine-tune of FLUX.1-dev on the 78-card Rider–Waite tarot dataset:https://huggingface.co/datasets/multimodalart/1920-raider-waite-tarot-public-domain, then generating from the resulting checkpoint. It reuses the checked-in YAML configs and applies run-specific settings as command-line overrides, so no new config files are required.
The diffusion recipe consumes cached VAE latents and text embeddings instead of encoding source images during every training step. Stream the 78 Rider–Waite images directly from Hugging Face and distribute preprocessing across all visible GPUs:
The captions already contain the trtcrd trigger token. With this pixel budget and the dataset's portrait aspect ratio, preprocessing assigns the samples to the 384×640 bucket used by the showcase run.
For image training, preprocessing produces .pt cache files and sharded metadata:
Use examples/diffusion/finetune/flux_t2i_flow.yaml directly. The YAML already selects FLUX.1-dev, full transformer fine-tuning, the FLUX flow-matching adapter, an effective batch size of 32, and eight-way FSDP2.
Supply the tarot-specific paths and settings as command-line overrides:
The run produces checkpoints at steps 50, 100, 150, and 200. The final checkpoint is labeled epoch_66_step_199 ; the label is zero-based even though it represents the completed 200th optimizer step.
Use the existing FLUX generation YAML and point model.checkpoint at the complete training checkpoint:
Include trtcrd to invoke the learned tarot style. For a control comparison, keep the seed and scene fixed but omit the trigger:
At step 200, the triggered astronaut prompts retain their requested content while acquiring a cream, red, and black vintage palette, heavy ink contours, flat color fields, aged-paper tones, and allegorical card composition. The untriggered astronaut remains photographic, demonstrating that the learned effect is substantially associated with trtcrd rather than replacing the base model globally.
All measurements were collected on one node with 8× NVIDIA H100 80GB GPUs. Results are means ± sample standard deviation over three steady-state 10-step windows.
Each sample is one 49-frame video clip.
The results from fine-tuning and LoRA showcase the power of NeMo Automodel for domain specialization. For instance, fine-tuning the Wan 2.1 model on a Ghibli video dataset successfully adapted the output style, demonstrated by a noticeable change in a flower's appearance compared to the baseline.
We also observed the distinct impact of using LoRA, where applying the adapter to Wan 2.1 caused the video to adopt a characteristic Ghibli style, particularly visible in the highlighting of characters' eyes.
These examples, including those for FLUX.2, confirm that users can achieve both maximum quality via full fine-tuning and maximum efficiency via LoRA-style PEFT, tailoring the output to specific stylistic domains.
Learn more about the integration and find more fine-tuning examples in the NeMo Automodel documentation:https://docs.nvidia.com/nemo/automodel/recipes-e2e-examples/diffusion-fine-tuning
YAML is a strong fit for reproducible configuration, especially for teams that want files they can check in, review, and reuse, but many teams also need a programmatic interface.
In an upcoming NeMo Automodel release, we plan to surface the diffusion recipes through a fully typed Pythonic API as well. Users will be able to compose the same model, data, optimizer, PEFT/LoRA, parallelism, checkpointing, and generation pieces directly from Python.
The Pythonic path is intended to make the recipes easier to use from existing training code, notebooks, and experiment workflows, and to offer a first-class Pythonic interface alongside the YAML quick-start path.