Diarizace řečníků a syntéza přeloženého hlasu ve videu

Abstract

This master’s thesis focuses on speech data processing, with particular emphasis on speaker diariza tion and voice cloning as key technologies for the automated localization of audiovisual content. The theoretical part summarizes the principles of speaker diarization, modern approaches to speech syn thesis and voice cloning, and current deep learning-based architectures. The practical part presents the design and implementation analysis of a system that integrates transcription, speech alignment, diarization, translation, and voice synthesis into a single processing pipeline. The thesis also in cludes an evaluation of the proposed solution using state-of-the-art metrics for assessing diarization quality and overall speech processing performance. The work thus connects a theoretical overview of current methods with the practical implementation of a system for multilingual video localization.

Description

Delayed publication

Available after

Subject(s)

speaker diarization, voice cloning, speech synthesis, speech data processing, video localization, deep learning, automated dubbing

Citation