Skip to content
akaturk Academic measurement

Article detail · 2023

Effect of number and position of frames in speaker age estimation

Sigma Journal of Engineering and Natural Sciences

YÖKSİS OpenAlex ISSN 1304-7191 DOI 10.14744/sigma.2023.00036 Citations 1 Open access · diamond SJR Q4 JCR Q3

10.14744/sigma.2023.00036

YÖKSİS YÖKSİS article record

OpenAlex OpenAlex enrichment (abstract, citations, topics)

Abstract

OpenAlex record

English (OpenAlex)

With the invention of powerful processing devices as well as lucrative capabilities in the first two decades of the 21 st century, machine learning algorithms will soon be able to predict speaker age with higher accuracy or much lower error rate.It is an age-old quest for the human society to profile individuals remotely which basically includes age.Speaker age estimation has been treated in quite few perspectives.However, most of these approaches fail to show the effect of utterance length, aka number of frames on speaker age estimation.We present a detailed analysis on the effect of number of frames and position of frames for speaker age estimation using four magnitude-based and one phase-based spectral feature sets.The optimal speech duration for this objective is investigated.In addition, the mismatch between the training and test utterance duration is explored.The magnitude-based features are mainly derived from filter bank analysis.After the filter-bank analysis, an i-vector is generated for each utterance.Least squares support vector regression (LSSVR) is employed for speaker age estimation.In the experiments, the aGender database which consists of utterances from four age groups of German speakers is used.Increasing number of frames in the training and test increases the age estimation accuracy.This can be associated with the notion that more data helps the estimation process.Concerning position, the frames located at the centre of utterances tend to offer better results for both genders.The backend algorithms offer the best performance when the utterance length of training and test sets are equal for longer speech segments, otherwise training with medium length utterances and testing with longer ones offers better estimation performance especially for the female dataset.

OpenAlex enrichment

Topics

  • Speech Recognition and Synthesis

Type: article Speech Recognition and Synthesis

Index information

WoS (JCR) and Scopus (SJR) quartiles by ISSN and publication year. · 2023

Scopus (SJR) / WoS (JCR)

Sigma Journal of Engineering and Natural Sciences

Scopus (SJR) Q4 0,178 Year 2023
WoS (JCR) Q3 JIF 0,6 Year 2023

Universities

  • İZMİR DEMOKRASİ ÜNİVERSİTESİ

Authors

  1. Mohammed Muntaz Osman
  2. OSMAN BÜYÜK İZMİR DEMOKRASİ ÜNİVERSİTESİ
  3. ALİ TANGEL