Computer Science Faculty Publications

Automatic Lip-Synchronized Video-Self-Modeling Intervention for Voice Disorders

Ju Shen, University of DaytonFollow
Changpeng Ti, University of Kentucky
Sen-ching S. Cheung, University of Kentucky
Ravi R. Patel, University of Kentucky

Document Type

Conference Paper

Publication Date

10-2012

Publication Source

IEEE International Conference EHealth Networking Application & Service

Abstract

Video self-modeling (VSM) is a behavioral intervention technique in which a learner models a target behavior by watching a video of him- or herself. In the field of speech language pathology, the approach of VSM has been successfully used for treatment of language in children with Autism and in individuals with fluency disorder of stuttering. Technical challenges remain in creating VSM contents that depict previously unseen behaviors. In this paper, we propose a novel system that synthesizes new video sequences for VSM treatment of patients with voice disorders. Starting with a video recording of a voice-disorder patient, the proposed system replaces the coarse speech with a clean, healthier speech that bears resemblance to the patient's original voice. The replacement speech is synthesized using either a text-to-speech engine or selecting from a database of clean speeches based on a voice similarity metric. To realign the replacement speech with the original video, a novel audiovisual algorithm that combines audio segmentation with lip-state detection is proposed to identify corresponding time markers in the audio and video tracks. Lip synchronization is then accomplished by using an adaptive video re-sampling scheme that minimizes the amount of motion jitter and preserves the spatial sharpness. Experimental evaluations on a dataset with 31 subjects demonstrate the effectiveness of the proposed techniques.

Inclusive pages

244-249

ISBN/ISSN

978-1-4577-2040-6

Comments

Permission documentation is on file.

Copyright

Publisher

IEEE

Place of Publication

Beijing, China

Peer Reviewed

yes

eCommons Citation

Shen, Ju; Ti, Changpeng; Cheung, Sen-ching S.; and Patel, Ravi R., "Automatic Lip-Synchronized Video-Self-Modeling Intervention for Voice Disorders" (2012). Computer Science Faculty Publications. 54.
https://ecommons.udayton.edu/cps_fac_pub/54

Link to Full Text

COinS

Computer Science Faculty Publications

Automatic Lip-Synchronized Video-Self-Modeling Intervention for Voice Disorders

Document Type

Publication Date

Publication Source

Abstract

Inclusive pages

ISBN/ISSN

Comments

Copyright

Publisher

Place of Publication

Peer Reviewed

eCommons Citation

ENTER SEARCH TERMS

Contribute Work

SelectedWorks

Browse

Contribute Work

Browse

Links

Computer Science Faculty Publications

Automatic Lip-Synchronized Video-Self-Modeling Intervention for Voice Disorders

Author(s)

Document Type

Publication Date

Publication Source

Abstract

Inclusive pages

ISBN/ISSN

Comments

Copyright

Publisher

Place of Publication

Peer Reviewed

eCommons Citation

Share

ENTER SEARCH TERMS

Contribute Work

SelectedWorks

Browse

Contribute Work

Browse

Links