{"ID":429846,"CreatedAt":"2026-03-04T20:58:37Z","UpdatedAt":"2026-03-04T20:58:37Z","DeletedAt":null,"paper_url":"https://paperswithcode.com/paper/wenet-2-0-more-productive-end-to-end-speech","arxiv_id":"2203.15455","title":"WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit","abstract":"Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model. To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates. (1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-to-left attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage. (2) We introduce an n-gram based language model and a WFST-based decoder into WeNet 2.0, promoting the use of rich text data in production scenarios. (3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios. (4) We design a unified IO to support large-scale data for effective model training. In summary, the brand-new WeNet 2.0 achieves up to 10\\% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features.","short_abstract":"Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model.","url_abs":"https://arxiv.org/abs/2203.15455v2","url_pdf":"https://arxiv.org/pdf/2203.15455v2.pdf","authors":"[\"BinBin Zhang\", \"Di wu\", \"Zhendong Peng\", \"Xingchen Song\", \"Zhuoyuan Yao\", \"Hang Lv\", \"Lei Xie\", \"Chao Yang\", \"Fuping Pan\", \"Jianwei Niu\"]","published":"2022-03-29T00:00:00Z","tasks":"[\"Decoder\", \"Language Modelling\", \"speech-recognition\", \"Speech Recognition\"]","methods":"[]","has_code":false,"code_links":[{"ID":301948,"CreatedAt":"2026-03-04T21:00:12Z","UpdatedAt":"2026-03-04T21:00:12Z","DeletedAt":null,"paper_id":429846,"paper_url":"https://paperswithcode.com/paper/wenet-2-0-more-productive-end-to-end-speech","paper_title":"WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit","repo_url":"https://github.com/wenet-e2e/wenet","is_official":true,"mentioned_in_paper":true,"mentioned_in_github":true,"framework":"pytorch","github_stars":0},{"ID":417334,"CreatedAt":"2026-03-04T21:00:12Z","UpdatedAt":"2026-03-04T21:00:12Z","DeletedAt":null,"paper_id":429846,"paper_url":"https://paperswithcode.com/paper/wenet-2-0-more-productive-end-to-end-speech","paper_title":"WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit","repo_url":"https://github.com/leonwlw/wenet","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"framework":"pytorch","github_stars":0},{"ID":448812,"CreatedAt":"2026-03-04T21:00:12Z","UpdatedAt":"2026-03-04T21:00:12Z","DeletedAt":null,"paper_id":429846,"paper_url":"https://paperswithcode.com/paper/wenet-2-0-more-productive-end-to-end-speech","paper_title":"WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit","repo_url":"https://github.com/mobvoi/wenet","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"framework":"pytorch","github_stars":0}]}
