Automatic Detection of Speech Sound Disorder in Child Speech Using Posterior-based Speaker Representations

by   Si-Ioi Ng, et al.

This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The disorder can be detected by focally analyzing the phonemes or the words elicited by the child subject. In the present study, instead of attempting to detect individual phone- and word-level errors, we propose to extract a subject-level representation from a long utterance that is constructed by concatenating multiple test words. The speaker verification approach, and posterior features generated by deep neural network models, are applied to derive various types of holistic representations. A linear classifier is trained to differentiate disordered speech in normal one. On the task of detecting SSD in Cantonese-speaking children, experimental results show that the proposed approach achieves improved detection performance over previous method that requires fusing phone-level detection results. Using articulatory posterior features to derive i-vectors from multiple-word utterances achieves an unweighted average recall of 78.2 of 78.0


page 1

page 2

page 3

page 4


Automatic Detection of Phonological Errors in Child Speech Using Siamese Recurrent Autoencoder

Speech sound disorder (SSD) refers to the developmental disorder in whic...

Detection of Consonant Errors in Disordered Speech Based on Consonant-vowel Segment Embedding

Speech sound disorder (SSD) refers to a type of developmental disorder i...

Phone Duration Modeling for Speaker Age Estimation in Children

Automatic inference of important paralinguistic information such as age ...

Deep segmental phonetic posterior-grams based discovery of non-categories in L2 English speech

Second language (L2) speech is often labeled with the native, phone cate...

VOP Detection for Read and Conversation Speech using CWT Coefficients and Phone Boundaries

In this paper, we propose a novel approach for accurate detection of the...

Persian Typographical Error Type Detection using Many-to-Many Deep Neural Networks on Algorithmically-Generated Misspellings

Digital technologies have led to an influx of text created daily in a va...

Deep Learning for Prominence Detection in Children's Read Speech

Expressive reading, considered the defining attribute of oral reading fl...

Please sign up or login with your details

Forgot password? Click here to reset