Article detail · 2008
Spoken Term Detection for Turkish Broadcast News
- Year
- 2008
- Type
- conference-paper
Abstract
OpenAlex · English
In this paper, we present a baseline spoken term detection (STD) system for Turkish broadcast news. The agglutinative structure of Turkish causes a high out-of-vocabulary (OOV) rate and increases word error rate (WER) in automatic speech recognition. Several approaches are attempted to reduce this negative effect on the STD system. Sub-word units are used to handle the OOV queries and lattice-based indexing is used to obtain different operating points and handle high WER cases. A recently proposed method for setting term specific thresholds is also evaluated and extended to allow us to choose an operating point suitable for our needs. Best results are obtained by using a cascade of word and sub-word lattice indices with term-thresholding.
Topics
Citations
OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.
79 citations
OpenAlex cited_by_count (cache / database)
14 publications in the local catalog that cite this work (OpenAlex reference match; not the full global list).
- Retrieval and browsing of spoken content 2008
- Lattice Indexing for Spoken Term Detection 2011
- Turkish Broadcast News Transcription and Retrieval 2009
- Effect of pronounciations on OOV queries in spoken term detection 2009
- Performance Analysis and Improvement of Turkish Broadcast News Retrieval 2012
- Automatic sign segmentation from continuous signing via multiple sequence alignment 2009
- Speech and sliding text aided sign retrieval from hearing impaired sign news videos 2008
- Speech and sliding text aided sign retrieval from hearing impaired sign news videos 2008
- Web derived pronunciations for spoken term detection 2009
- Distance metric learning for posteriorgram based keyword search 2017