{"ID":22952742,"CreatedAt":"2026-09-17T02:12:05.498442134Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.18856","arxiv_id":"2609.18856","title":"GrainSpeech: Less Context, More Detail for Compact Speech Synthesis","abstract":"Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance supervision restores fine-scale variation but degrades predicted quality, motivating a Mel-specific formulation with axis-specific gradients, overlapping local statistics, and log-domain variance matching. GrainSpeech contains only 264.8K parameters and achieves 17.9x real-time Mel generation on a microcontroller (MCU), while attaining UTMOS scores comparable to substantially larger models with less than 1.5% of their parameters. Source code and demos are available at https://github.com/lab-emi/GrainSpeech.","short_abstract":"Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction....","url_abs":"https://arxiv.org/abs/2609.18856","url_pdf":"https://arxiv.org/pdf/2609.18856v1","authors":"[\"Zitao Liang\",\"Chang Gao\"]","published":"2026-09-16T15:59:20Z","proceeding":"eess.AS","tasks":"[\"eess.AS\",\"cs.AI\",\"cs.SD\",\"eess.SP\"]","methods":"[]","has_code":false,"code_links":[{"ID":639779,"CreatedAt":"2026-09-17T02:12:05.498442134Z","UpdatedAt":"2026-09-17T02:12:05.498442134Z","DeletedAt":null,"paper_id":22952742,"paper_url":"https://arxiv.org/abs/2609.18856","paper_title":"GrainSpeech: Less Context, More Detail for Compact Speech Synthesis","repo_url":"https://github.com/lab-emi/GrainSpeech","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"github_stars":0}]}
