'Voice Conversion' paper candidate 2305.04476

github-actions[bot] commented 1 year ago

Please check whether this paper is about 'Voice Conversion' or not.

article info.

title: AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment
summary: The speech-to-singing (STS) voice conversion task aims to generate singing samples corresponding to speech recordings while facing a major challenge: the alignment between the target (singing) pitch contour and the source (speech) content is difficult to learn in a text-free situation. This paper proposes AlignSTS, an STS model based on explicit cross-modal alignment, which views speech variance such as pitch and content as different modalities. Inspired by the mechanism of how humans will sing the lyrics to the melody, AlignSTS: 1) adopts a novel rhythm adaptor to predict the target rhythm representation to bridge the modality gap between content and pitch, where the rhythm representation is computed in a simple yet effective way and is quantized into a discrete space; and 2) uses the predicted rhythm representation to re-align the content based on cross-attention and conducts a cross-modal fusion for re-synthesize. Extensive experiments show that AlignSTS achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://alignsts.github.io.
id: http://arxiv.org/abs/2305.04476v1

judge

Write [vclab::confirmed] or [vclab::excluded] in comment.

tarepan commented 9 months ago

[vclab::confirmed]

github-actions[bot] commented 9 months ago

Thunk you very much for contribution! Your judgement is refrected in arXivSearches.json, and is going to be used for VCLab's activity. Thunk you so much.

tarepan / VoiceConversionLab

'Voice Conversion' paper candidate 2305.04476 #463

article info.

judge