2026-06-16 19:50:23 +00:00
2026-06-16 19:50:23 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00
2026-06-16 19:32:09 +00:00

license, language, tags, pipeline_tag, library_name
license language tags pipeline_tag library_name
apache-2.0
en
text-to-speech
tts
speech-synthesis
pytorch
tiny-tts
experimental
text-to-speech pytorch

Inflect-Nano-v1

Inflect-Nano-v1 is an experimental ultra-small English TTS stack. It is built to test how far a sub-5M-parameter text-to-speech system can be pushed with a compact non-autoregressive acoustic model and a small neural vocoder.

This is not a production-quality or SOTA TTS model. It is a research/demo release: small, local, and runnable, but still audibly limited.

Quick Facts

Item Value
Total inference parameters 4.632M
Acoustic model 3.465M
Vocoder generator 1.167M
Language English
Voice single Mark-style synthetic male voice
Sample rate 24 kHz
Acoustic output 80-bin mel spectrogram
Vocoder custom Snake-activation HiFi-GAN-style generator
Training source synthetic Qwen3-TTS Mark-style teacher data

Audio Examples

These are unseen/OOD stress prompts, not hand-picked training rows.

Prompt Audio
Wait, are you actually being for real now? I can't believe it!
Sophia sent me 43 pictures of her doing stuff... interesting.
Please say chrysanthemum, thoroughly, proprietary, and rural without rushing through the middle syllables.
No, seriously, did Jordan leave the receipt in Albuquerque, or did Priya move it to Worcester?
The Wi-Fi password is Q7-Delta-9921, but please do not say the dash like a minus sign.
I appreciate the honesty, but that explanation sounded weirdly dramatic for a Tuesday morning.
Could you whisper the first part, then brighten up when you say, 'we finally solved it'?
The dermatologist, the anesthesiologist, and the statistician all disagreed about February.

Install

git clone https://huggingface.co/owensong/Inflect-Nano-v1
cd Inflect-Nano-v1
pip install -r requirements.txt

The text frontend uses TinyTTS-style English G2P and may download bert-base-uncased tokenizer files on first run.

Usage

python inference.py \
  --text "Wait, are you actually being for real now? I can't believe it!" \
  --out sample.wav

CPU example:

python inference.py \
  --device cpu \
  --text "Please say chrysanthemum, thoroughly, proprietary, and rural clearly." \
  --out sample_cpu.wav

Optional controls:

python inference.py \
  --text "No, seriously, did Jordan leave the receipt in Albuquerque?" \
  --length-scale 1.03 \
  --pitch-scale 1.00 \
  --energy-scale 1.00 \
  --out sample_controlled.wav

Gradio demo:

python app.py

Architecture

Inflect-Nano-v1 is a two-part TTS stack:

text
-> TinyTTS-style normalization + G2P phoneme/tone/lang IDs
-> compact FastSpeech-style acoustic model
-> 80-bin mel spectrogram
-> Snake V2Mid HiFi-GAN-style vocoder
-> 24 kHz waveform

Acoustic Model

The acoustic model is a small non-autoregressive FastSpeech-style network. It predicts duration, energy, brightness, and pitch, then expands token states into frame states and decodes mels.

Main config:

{
  "hidden": 168,
  "encoder_layers": 5,
  "decoder_layers": 6,
  "decoder_ff_mult": 3,
  "kernel_size": 7,
  "speaker_dim": 64,
  "dropout": 0.08,
  "n_mels": 80,
  "sample_rate": 24000,
  "max_frames": 1400,
  "postnet_scale": 0.1,
  "use_frame_pitch": true,
  "abs_frame_bins": 512
}

Acoustic parameter split:

total acoustic: 3.465M
encoder:        1.292M
decoder:        1.211M
postnet:        0.276M
local context:  0.226M
frame GRU:      0.128M
heads/embeds/projections: remainder

Vocoder

The vocoder is a custom Snake-activation HiFi-GAN-style generator.

Main config:

{
  "variant": "snake_v2mid",
  "sample_rate": 24000,
  "n_fft": 1024,
  "hop_size": 256,
  "win_size": 1024,
  "num_mels": 80,
  "fmax": 12000.0,
  "upsample_rates": [8, 8, 2, 2],
  "upsample_kernel_sizes": [16, 16, 4, 4],
  "upsample_initial_channel": 144,
  "resblock_kernel_sizes": [3, 7, 11],
  "resblock_dilation_sizes": [[1, 3, 5], [1, 3, 5], [1, 3, 5]],
  "activation": "snake"
}

The vocoder was trained with HiFi-GAN-style adversarial losses and multi-resolution spectral pressure. Discriminators are training-only and are not included in inference.

Training Data

The final acoustic model was trained primarily on synthetic Mark-style teacher speech.

Known final broad training mix:

mixed_80k_v2_40_v1_20_old_20
rows: 80,000
estimated duration: ~125.2 hours
composition:
  40k newer generalization v2 examples
  20k generalization v1 examples
  20k older Mark anchor examples

The practical release voice is one speaker:

qwen3_mark

Training Recipe

The final F checkpoint is a staged recovery candidate, not a single monolithic run.

High-level acoustic lineage:

1. Mark-focused acoustic base
2. mixed 80k generalization training
3. predictor-exposure heads training
4. robust prosody bridge
5. recovery phases A/B/C
6. short predictor-tail cleanup

Known acoustic continuation after the earlier Mark base:

mixed80k broad training:      ~32k selected checkpoint
predictor exposure:            9k steps
robust prosody bridge:       1.2k steps
recovery phase A:             3k steps
recovery phase B:             3k steps
recovery phase C:           2.5k steps
predictor tail:             0.8k steps
known continuation total:   ~51.5k steps

Acoustic losses included mel reconstruction, MSE, delta/acceleration losses, duration loss, energy loss, brightness loss, pitch loss, predicted-prosody exposure, and robust-prosody exposure.

Final selected files in this repo:

weights/inflect_nano_v1_acoustic.pt
weights/inflect_nano_v1_vocoder.pt

Limitations

This model is intentionally tiny and has clear quality limits:

  • Unseen text can stumble or sound unstable.
  • The voice can sound robotic, buzzy, or artifacted.
  • Long or unusual prompts are less reliable.
  • It inherits habits from synthetic Qwen3-TTS teacher data.
  • It is not a voice cloning model.
  • It is not multilingual.
  • It is not suitable for production accessibility, safety, or high-quality narration use.

Use this as:

An experimental 4.63M-parameter English TTS model exploring the quality/size tradeoff for ultra-small local speech synthesis.

Do not present it as SOTA or production-quality.

License

Apache-2.0. The repo includes TinyTTS text frontend code; its license is included as TINY_TTS_LICENSE.

S
Description
No description provided
Readme Apache-2.0 30 MiB
Languages
Python 99.3%
Jupyter Notebook 0.7%