Aplikace modelů umělé inteligence pro vytvoření interagujícího avatara

Abstract

This bachelor’s thesis focuses on the design and implementation of a demonstration application of an interactive avatar utilizing artificial intelligence models for bidirectional voice communication with a user. The work concentrates on the integration of speech-to-text (ASR) technologies, natural language processing using large language models with the application of retrieval-augmented generation (RAG), and text-to-speech (TTS) synthesis for response generation. The thesis also includes the design of a system architecture supporting efficient data flow between individual application components. The practical part presents the implementation of an application featuring a 2D/3D avatar capable of basic visual articulation and simple motion responses. Furthermore, available technologies are analyzed in terms of quality, latency, and availability. The resulting solution is experimentally evaluated with respect to speech recognition accuracy, the quality of generated responses, system response time, and overall user experience.

Description

Delayed publication

Available after

Subject(s)

artificial intelligence, interactive avatar, speech recognition, speech synthesis, large language models, retrieval-augmented generation, natural language processing, voice communication

Citation