arlaz's picture
Initial Multidiff Modular export
d8bf41d
|
Raw History Blame Contribute Delete
7.3 kB

Examples

This folder contains runnable generation examples.

  • example.py: local checkout/editable-install path using root block.py.
  • example_remote.py: Hub path using ModularPipeline.from_pretrained(..., trust_remote_code=True).
  • prompt.csv plus masks-*: toy regional prompting inputs.

Install, export, and Hugging Face setup are covered in the root README.md. The commands below assume they are run from the repo root.

Basic Local Run

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a dense renaissance fresco" \
  --height 2048 \
  --width 2048 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --dtype bfloat16 \
  --device cuda \
  --output output.png

Remote Hub Run

This is the easiest end-to-end test of the published repo:

uv run python examples/example_remote.py \
  --repo-id arlaz/modular-flux2-multidiffusion \
  --prompt "a dense renaissance fresco" \
  --height 1024 \
  --width 1024 \
  --height-generation 512 \
  --width-generation 512 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --num-inference-steps 1 \
  --dtype bfloat16 \
  --device cuda \
  --output remote_smoke_1024.png

Use --local-files-only only after the upstream Flux.2 components are already cached.

Large Tiled Run

--batch-size controls the maximum number of window or regional work items sent through one denoiser forward pass. 1 preserves sequential behavior.

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a majestuous renaissance fresco, with iridiscent light, caustic lights" \
  --height 8192 \
  --width 8192 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 1024 \
  --window-stride-width 1024 \
  --window-stride-height-offset 256 \
  --window-stride-width-offset 256 \
  --batch-size 4 \
  --dtype bfloat16 \
  --device cuda \
  --output output.png \
  --local-files-only \
  --enable-tiling \
  --enable-slicing

Panorama Run

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a continuous renaissance fresco panorama" \
  --height 4096 \
  --width 4096 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --panorama-width \
  --panorama-height \
  --dtype bfloat16 \
  --device cuda \
  --output panorama.png

Inspect the wrapped seams:

uv run multidiff-modular inspect-panorama panorama.png
uv run python scripts/inspect_panorama.py panorama.png

Regional Prompting

Regional prompting is enabled when --prompt points to a CSV and --masks points to a mask folder.

CSV format:

mask,prompt
00.png,"a snowy mountain"
01.png,"a forest at sunset"

Mask rules:

  • filenames in the CSV must exist inside --masks;
  • masks must match --height and --width;
  • grayscale masks are accepted;
  • non-binary grayscale masks warn and are used as fractional weights;
  • masks must cover every packed latent cell.

Toy folders:

  • masks-1024, masks-2048, masks-4096: dense binary regions.
  • masks-1024-ramp, masks-2048-ramp, masks-4096-ramp: ramped grayscale regions.
uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt examples/prompt.csv \
  --masks examples/masks-4096-ramp \
  --height 4096 \
  --width 4096 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --batch-size 4 \
  --dtype bfloat16 \
  --device cuda \
  --output regional-ramp.png

Img2Img

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a painterly architectural fresco" \
  --image-img2img input.png \
  --strength 0.75 \
  --height 2048 \
  --width 2048 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --dtype bfloat16 \
  --device cuda \
  --output img2img.png

Image Conditioning

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a fresco guided by the conditioning image" \
  --image-conditioning conditioning.png \
  --height 2048 \
  --width 2048 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --dtype bfloat16 \
  --device cuda \
  --output conditioned.png

Quantization

Pass --quantization without a value to use float8wo, or pass one of none, float8wo, int8wo, int4wo, or float8dyn.

uv run python examples/example.py \
  --base-model black-forest-labs/FLUX.2-klein-4B \
  --prompt "a dense renaissance fresco" \
  --height 2048 \
  --width 2048 \
  --height-generation 1024 \
  --width-generation 1024 \
  --window-stride-height 512 \
  --window-stride-width 512 \
  --quantization float8wo \
  --vae-quantization none \
  --dtype bfloat16 \
  --device cuda \
  --output quantized.png

Argument Reference

Model and loading:

  • --base-model: upstream Flux.2 repo id or local model directory for example.py.
  • --repo-id: exported Hub repo id or local exported repo path for example_remote.py.
  • --lora-path, --lora_path: optional LoRA directory or file path.
  • --local-files-only: disable network fetches and use cached/local files only.

Inputs:

  • --prompt: text prompt, or a CSV file with mask,prompt headers.
  • --masks: mask folder used with prompt CSV files.
  • --image-img2img: image path for img2img initialization.
  • --image-conditioning: image path for Flux.2 image conditioning.
  • --strength: img2img renoising strength. Default: 1.0.

Canvas and windowing:

  • --height, --width: final canvas size in pixels.
  • --height-generation, --width-generation: local denoising window size.
  • --window-stride-height, --window-stride-width: sliding-window stride.
  • --window-stride-height-offset, --window-stride-width-offset: per-step stride offset.
  • --panorama-width, --panorama-height: wrap width and/or height.
  • --weighting-type: none, linear, or cosine.
  • --weighting-range: blend ramp width scalar.

Inference:

  • --num-inference-steps: denoising steps.
  • --guidance-scale: optional guider override.
  • --terra-scale: optional Terra LoRA scale.
  • --seed: generator seed. Default: 42.
  • --num-images-per-prompt: output count. Default: 1.
  • --batch-size: maximum window/regional work items per denoiser forward pass.

Runtime:

  • --dtype: float16, bfloat16, or float32.
  • --device: torch device string. Default: cuda.
  • --output: output path. Images larger than 4096 * 4096 pixels save as BigTIFF automatically.
  • --allow-tf32: enable CUDA TF32 settings.
  • --compile: compile repeated transformer blocks with torch.compile.
  • --quantization: TorchAO strategy for transformer, text encoder, and VAE.
  • --transformer-quantization, --text-encoder-quantization, --vae-quantization: component-specific overrides.
  • --enable-tiling: enable VAE tiling.
  • --enable-slicing: enable VAE slicing.