--- pipeline_tag: text-to-speech tags: - mlx - tts - maya1 - apple-silicon - snac base_model: maya-research/maya1 library_name: mlx --- # maya1-mlx-bf16 MLX bfloat16 conversion of [maya-research/maya1](https://huggingface.co/maya-research/maya1) for text-to-speech on Apple Silicon. Converted using [mlx-lm](https://github.com/ml-explore/mlx-examples/tree/main/llms/mlx_lm). > **Note:** This is the full-precision (bfloat16) variant. It offers the highest quality but runs below real-time on most hardware. Consider the [8-bit](https://huggingface.co/hashb/maya1-mlx-8bit) or [4-bit](https://huggingface.co/hashb/maya1-mlx-4bit) variants for faster inference. ## Benchmarks (M4 Max, 36GB) | Variant | Size | Tokens/s | Real-time factor | |---------|------|----------|-----------------| | **bf16** | **6.2 GB** | **~51 tok/s** | **0.50x** | | 8-bit | 3.3 GB | ~91 tok/s | 0.82x | | 4-bit | 1.8 GB | ~108 tok/s | 1.58x | ## Quick Start Requires macOS with Apple Silicon and [uv](https://docs.astral.sh/uv/). ```bash # From a text file uv run tts.py input.txt -o output.wav # From stdin echo "Hello world" | uv run tts.py - -o hello.wav # With a custom voice description uv run tts.py input.txt -o output.wav -d "Deep male voice, British accent, slow pacing." ``` ## CLI Options | Option | Default | Description | |--------|---------|-------------| | `input` | *(required)* | Input text file path, or `-` for stdin | | `-o`, `--output` | `output.wav` | Output WAV file path | | `-d`, `--description` | Calm/clear female voice | Voice description prompt | | `-m`, `--model` | `.` | Path to MLX model directory | | `--max-chars` | `200` | Max characters per chunk | | `--max-tokens` | `2048` | Max tokens per chunk | | `--temperature` | `0.4` | Sampling temperature | | `--top-p` | `0.9` | Top-p sampling | | `--repetition-penalty` | `1.1` | Repetition penalty | ## Requirements - macOS with Apple Silicon (M1 or later) - [uv](https://docs.astral.sh/uv/) (dependencies are declared inline via PEP 723)