Докторська школа імені родини Юхименків
Permanent URI for this collection
Browse
Browsing Докторська школа імені родини Юхименків by Author "Ivaniuk, Andrii"
Now showing 1 - 2 of 2
Results Per Page
Sort Options
Item Guided inverse problems(Національний університет "Києво-Могилянська академія", 2024) Ivaniuk, Andrii; Kravchuk, Oleg; Kriukova, GalynaThe given work proposes a novel approach for solving inverse problems in machine learning leveraging Physics-Guided Neural Networks (PGNNs). This method incorporates domain knowledge through an additional inverse problem, leading to significant improvements in model performance and accuracy.Item Latent diffusion model for speech signal processing(2024) Ivaniuk, AndriiTopicality. The development of generative models for audio synthesis, including text-to-speech (TTS), text-to-music, and text-to-audio applications, largely depends on their ability to handle complex and varied input data. This paper centers on latent diffusion modeling, a versatile approach that leverages stochastic processes to generate high-quality audio outputs. Key goals. This study aims to evaluate the efficacy of latent diffusion modeling for TTS synthesis on the EmoV-DB dataset, which features multi-speaker recordings across five emotional states, and to contrast it with other generative techniques. Research methods. We applied latent diffusion modeling to TTS synthesis specifically and evaluated its performance using metrics that assess intelligibility, speaker similarity, and emotion preservation in the generated audio signal. Results. The study reveals that while the proposed model demonstrates decent efficiency in maintaining speaker characteristics, it is outperformed by the discrete autoregressive model: xTTS v2 in all assessed metrics. Notably, the researched model exhibits deficiencies in emotional classification accuracy, suggesting potential misalignment between the emotional intents encoded by the embeddings and those expressed in the speech output. Conclusions. The findings suggest that further refinement of the encoder's ability to process and integrate emotional data could enhance the performance of the latent diffusion model. Future research should focus on optimizing the balance between speaker and emotion characteristics in TTS models to achieve a more holistic and effective synthesis of human-like speech.