Generally speaking, spoken term detection system will degrade significantly because of mismatch between acoustic model and spontaneous speech. This paper presents an improved spoken term detection strategy, which integrated with a novel phoneme confusion matrix and an improved word-level minimum classification error (MCE) training method. The first technique is presented to improve spoken term detection rate while the second one is adopted to reject false accepts. On mandarin conversational telephone speech (CTS), the proposed methods reduce the equal error rate (EER) by 8.4% in relative.
Chuanxu Wang. Pengyuan Zhang. "Optimization of Spoken Term Detection System." J. Appl. Math. 2012 (SI08) 1 - 8, 2012. https://doi.org/10.1155/2012/548341