Modern ASR models treat transcription style (literal vs. intent) as an uncontrolled latent variable, leading to unstable decoding, with up to 60% of WER re...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
Modern ASR models treat transcription style (literal vs. intent) as an uncontrolled latent variable,
leading to unstable decoding, with up to 60% of WER resulting from style mismatches. By training task tokens of a coverage-aware decoder on parallel literal/intent transcription pairs, training only in English can improve the F1 for disfluent German from 10% in zero-shot to 79%. Supervised cross-attention fine-tuning achieves word-level timestamps on disfluent speech that surpass forced alignment baselines for the first time, and the new task 'verbatimize' can improve rare word recall from 6.8% to 96.1%.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Modern ASR models treat transcription style (literal vs. intent) as an uncontrolled latent variable, leading to unstable decoding, with up to 60% of WER re...