Inflect v2 CoreML (Micro + Nano)
Fixed-shape CoreML conversion of Inflect-Micro-v2 (9.36M params) and Inflect-Nano-v2 (3.97M params) โ ultra-tiny VITS-family end-to-end English TTS, 24 kHz mono, fp16.
Converted by FluidInference (mobius);
conversion pipeline, parity checks, and benchmark harness live in the
models/tts/inflect-v2/coreml/ directory of the mobius repo.
Contents
micro/ 9.4M params, ~19 MB fp16 total per graph set
encoder.mlpackage tokens+mask -> m_p, logs_p, logw [t_text=512]
synthesizer_f{256,384,512,640,768,896,1024,2048}.mlpackage
z_p+mask -> waveform [bucket frames x 256 samples]
upstream_config.json
nano/ 4.0M params, ~8 MB fp16 per graph set (same layout)
symbols.json keithito-style symbol table (178 symbols, add_blank interspersal)
LICENSE Apache-2.0 (upstream)
Inference pipeline
Two deterministic models with everything stochastic or dynamically shaped on the host:
- espeak-ng
en-usphonemization (with stress) โ symbol ids โ intersperse blanks (pad_id=0) โ pad to 512. encoder: โm_p,logs_p,logw.- Host:
w = ceil(exp(logw))/speed; repeat-expandm_p/logs_ptoy_lenframes;z_p = m_p + randn * exp(logs_p) * noise_scale(default 0.667). - Pick the smallest synthesizer bucket โฅ
y_len; zero-padz_p, mask valid frames. synthesizer: reverse coupling flow + HiFiGAN โ waveform; trim toy_len * 256samples.
Buckets are fixed-shape for ANE compatibility; at 128-frame granularity padding overhead is within ~5% of exact-shape synthesis.
Benchmarks (M5 Pro, macOS 26.6, GPU, MiniMax-English 100 phrases)
| Variant | Synth p50 / p95 | Agg RTFx | WER* | CER* |
|---|---|---|---|---|
| Micro | 25.7 / 35.5 ms | 245ร | 1.43% | 0.44% |
| Nano | 13.2 / 16.9 ms | 460ร | 1.92% | 0.57% |
* Parakeet TDT v3 roundtrip, same scoring path as the FluidAudio TTS benchmarks. fp16 parity vs the PyTorch reference: audio correlation > 0.9999, predicted durations bit-identical.
Notes
- ANE: the synthesizer's waveform-rate tensors exceed the ANE width limit (W โค 65536) above the 256-frame bucket; f256 reaches 77% ANE residency, larger buckets run on GPU (which is faster at every size on M-series).
- English only, one voice per variant, no cloning. Sample rate 24 kHz.
- Upstream weights and frontend: Apache-2.0, ยฉ the Inflect authors. This repo redistributes converted weights under the same license.
- Downloads last month
- -
Model tree for FluidInference/inflect-v2-coreml
Base model
owensong/Inflect-Micro-v2