SceneWorks's picture
Add SANA MLX mirror (verbatim repackage of NVIDIA diffusers checkpoint into SanaPipeline::from_snapshot layout)
eef346e verified
Raw History Blame Contribute Delete
2.45 kB
SANA 1600M 1024px — SceneWorks MLX mirror
=========================================
This repository is a re-hosted, un-gated mirror of NVIDIA's SANA 1600M 1024px
diffusers checkpoint, repackaged into the directory layout the SceneWorks native
MLX worker (mlx-gen-sana) loads via `SanaPipeline::from_snapshot`.
Upstream model:
Efficient-Large-Model/Sana_1600M_1024px_diffusers
https://huggingface.co/Efficient-Large-Model/Sana_1600M_1024px
LICENSE
-------
SANA is distributed under the NVIDIA Open Model License Agreement (see the
bundled `LICENSE` file). It is a NON-COMMERCIAL license: the model and its
outputs are for research and evaluation use only. By downloading or using these
weights you agree to the terms of the NVIDIA Open Model License.
This mirror carries the upstream license text verbatim and adds no additional
grant. Generations produced with this model are for non-commercial / research
use only.
CONTENTS
--------
transformer/diffusion_pytorch_model.safetensors
The SANA 1.6B Linear-DiT trunk, copied verbatim from the upstream
`transformer/diffusion_pytorch_model.safetensors` (F16, dtype preserved —
no conversion or re-quantization). 396 tensors.
vae/diffusion_pytorch_model.safetensors
The 32x deep-compression DC-AE (f32c32) autoencoder, copied verbatim from
the upstream `vae/diffusion_pytorch_model.safetensors` (F32, dtype
preserved — the DC-AE decode path runs in f32). 375 tensors.
text_encoder/gemma-2-2b-it.safetensors
text_encoder/tokenizer.json
The gemma-2-2b-it Complex-Human-Instruction (CHI) caption encoder, sourced
from the un-gated SceneWorks/gemma-2-2b-it mirror (BF16) and merged from its
two shards into a single `model.*`-keyed safetensors so the worker can load
it with `SanaTextEncoder::from_snapshot`. gemma-2-2b-it is distributed by
Google under the Gemma Terms of Use; see https://huggingface.co/SceneWorks/gemma-2-2b-it
for that license and notice. It is bundled here only so the SANA snapshot is
self-contained; the weights are unmodified.
ATTRIBUTION
-----------
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
NVIDIA / MIT HAN Lab. Project: https://nvlabs.github.io/Sana/
Paper: https://arxiv.org/abs/2410.10629
No weights were altered: this mirror is a verbatim repackage (dtype-preserving)
of the upstream tensors into the SceneWorks snapshot layout.