RSSAmplifier

Blog

LESS IS MORE

LESS IS MORE

r9y9.github.ioRSS feed ↗100 posts

Latest posts

Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control

Preprint: arXiv:2409.17452 (Submitted to ICASSP 2025 ) Abstract Figure Comparison of different systems English TTS Japanese TTS Style controllability samples Pitch control (en) Speed control (en) Pitch control (ja) Speed control (ja) Abstract We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description…

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection

Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment

SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark

Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data

2023年の振り返り/Looking Back at 2023

はじめに 新年明けましておめでとうございます。今年もよろしくお願いします。ブログを書かなくなって久しいですが、久しぶりになんか書いてみます。自分用のメモみたいなものですが、よろしければご覧ください。 2023年 タイムライン Googleカレンダーを見ながら思い出せる範囲で書いてみます 1月 静岡旅行に行きました。 海鮮丼 が美味しかったです 家事代行サービスの利用を始めました。月2回の頻度で今も続けていますが、水回りの掃除がかなり楽になりました 2月 ICASSP 2023 論文アクセプト 4/4 3月 年度末で仕事が忙しかったような気がします 4月 博士研究の中間発表をしました。タイトル: Neural Singing Voice Synthesis Towards Unified and Controllable Voice…

PromptTTS++: Controlling Speaker Identity in Prompt-based Text-to-Speech using Natural Language Descriptions

Enhancing Multilingual TTS with Voice Conversion based Data Augmentation and Posterior Embedding

Electrolaryngeal Speech Intelligibility Enhancement Through Robust Linguistic Encoders

A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023

Preprint: arXiv:2310.05203 (Accepted to ASRU 2023 ) This page provides audio samples of our singing voice conversion system (denoted as T13; the Nagoya University system) for The Singing Voice Conversion Challenge 2023 . Task 1: In-domain SVC Target: IDM1 Target: IDF1 Task 2: Cross-domain SVC Target: CDM1 Target: CDF1 Task 1: In-domain SVC Target: IDM1 Sample: 30013 Source Target T13 Sample: 30017…

NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit

Preprint: arXiv:2210.15987 (Submitted to ICASSP 2023 ) Updates Authors Abstract Systems Samples SVS A/S Mixed demo Sample 1 Sample 2 Sample 3 Bonus samples Japanese Mandarin Error cases Unstable pitch of DiffSinger Acknowledgments Appendix Note pitch distribution References Updates 2023/02/27: Added error cases to address reviewer’s comments. See Error cases . 2022/11/27: Added samples of…

Non-parallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs

Period VITS: Variational Inference With Explicit Pitch Modeling For End-to-End Emotional Speech Synthesis

Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform

DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation Learning

TTS-by-TTS 2: Data-selective Augmentation for Neural Speech Synthesis Using Ranking Support Vector Machine with Variational Autoencoder

Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation

A Unified Accent Estimation Method Based on Multi-Task Learning for Japanese Text-to-Speech

Language Model-Based Emotion Prediction Methods for Emotional Speech Synthesis Systems

企業における音声合成の研究開発 / Research and development for TTS in industry @名古屋工業大学

恩師である酒向先生 ( http://sakoweb.net/joomla3/ ) にお声がけいただき、母校の名古屋工業大学で講義をさせていただきました。

Hugo Academic を使ってウェブサイトをアップデートしました

Summary 新: https://github.com/r9y9/website トップページ ブログ一覧 デモページ一覧 旧: https://github.com/r9y9/blog + https://github.com/r9y9/demos-src これまで、ブログとデモページ 1 をそれぞれ個別に Hugo で管理していましたが、それらを統合して一つの Hugo site として管理するように変更しました。 内部的に色々変わっていますが、URL はほぼ変わらないようにしています。 はじめに 2年振りくらいにブログを書いています。大したことを書くわけではないのですが、ウェブサイト(ブログ含む)を大幅にアップデートしたので、その記録を残しておきます。 Why なぜ大幅にアップデートをしたのか、理由は以下のとおりです。…

国際会議Interspeech2021参加報告 / Report on Participation in Interspeech2021 @SLP研究会

Abstract (ja) 2021 年 8 月 30 日から 9 月 3 日にかけてチェコ・ブルノおよびオンラインのハイブリッド形式で Interspeech2021 が開催された.ここでは,会議概要や最新の技術動向,注目の発表について報告する.

ESPnet2-TTS: Extending the Edge of TTS Research

ttslearn: Library for Pythonで学ぶ音声合成 (Text-to-speech with Python)

Voicing-Aware Parallel WaveGAN for High-Quality Speech Synthesis

Submitted to IEEE signal processing letters Authors Abstract TTS samples M1 (male) M2 (male) F1 (female) F2 (female) Acknowledgements Authors Ryuichi Yamamoto (LINE Corp.) Min-Jae Hwang (Search Solutions Inc.) Eunwoo Song (NAVER Corp.) Abstract This letter proposes a voicing-aware Parallel Wave- GAN (VA-PWG) vocoder for a neural text-to-speech (TTS) system. To generate a high-quality speech…

High-fidelity Parallel WaveGAN with Multi-band Harmonic-plus-Noise Model

Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis

ここまで来た音声技術・今後の展望 / Current progress on speech technologies and its future prospects @ LINE DEV DAY 2020

Abstract (ja)…

Parallel WaveGAN: GPUを利用した高速かつ高品質な音声合成 / Parallel WaveGAN: Fast and High-Quality GPU Text-to-Speech @ LINE DEV DAY 2020

Abstract (ja) コンピュータによってテキストから人間の声を合成する技術は、テキスト音声合成と呼ばれます。LINE CLOVAのスマートスピーカーを初めとするユーザとのリアルタイムのインタラクションが必要なサービスでは、音声合成システムには合成品質が高いことだけでなく、高速に音声を生成できることが求められます。本セッションでは、高速かつ高品質な音声合成を実現するために、NAVERとLINEで共同で開発したGPUベースの音声合成の研究成果について発表します。従来の方法では、品質が良くても合成速度が遅い、合成速度は速い一方でモデルの学習に多大な時間がかかるなどの問題がありました。我々はそのような問題に対してどのようにアプローチしたのか、音声信号処理のトップカンファレンスICASSP 2020に採択された論文の内容を元に、近年の関連分野の発展を交えて紹介します。

Improved Parallel WaveGAN with perceptually weighted spectrogram loss

TTS-by-TTS: TTS-driven Data Augmentation for Fast and High-Quality Speech Synthesis

Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators

Preprint: arXiv:2010.14151 (accepted to ICASSP 2021 ) Table of contents Analysis/synthesis samples (Japanese) Text-to-speech samples (Japanese) Bonus: analysis/synthesis samples for CMU ARCTIC (English) Authors Ryuichi Yamamoto (LINE Corp.) Eunwoo Song (NAVER Corp.) Min-Jae Hwang (Search Solutions Inc.) Jae-Min Kim (NAVER Corp.) Abstract This paper proposes voicing-aware conditional discriminators…

NNSVS: Pytorchベースの研究用歌声合成ライブラリ

Summary コード: https://github.com/r9y9/nnsvs Discussion: https://github.com/r9y9/nnsvs/issues/1 Demo on Google colab 春が来た 春が来た どこに来た。 山に来た 里に来た、野にも来た。花がさく 花がさく どこにさく。山にさく 里にさく、野にもさく。 Your browser does not support the audio element. NNSVS はなに? Neural network-based singing voice synthesis library for research 研究用途を目的とした、歌声合成エンジンを作るためのオープンソースのライブラリを作ることを目指したプロジェクトです。このプロジェクトについて、考えていることをまとめておこうと思います。…

Neural text-to-speech with a modeling-by-generation excitation vocoder

End-to-End 音声合成の研究を加速させるツールキット ESPnet-TTS / ESPnet-TTS: A toolkit to accelerate research on end-to-end speech synthesis @ ASJ 2020s

ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram

Preprint: arXiv:1910.11480 (accepted to ICASSP 2020 ) Audio samples (Japanese) Audio samples (English) Japanese samples were used in the subjective evaluations reported in our paper. Authors Ryuichi Yamamoto (LINE Corp.) Eunwoo Song (NAVER Corp.) Jae-Min Kim (NAVER Corp.) Abstract We propose Parallel WaveGAN 1 , a distillation-free, fast, and small-footprint waveform generation method using a…

Probability Density Distillation with Generative Adversarial Networks for High-Quality Parallel Waveform Generation

Preprint: arXiv:1904.04472 , Published version: ISCA Archive Interspeech 2019 Authors Ryuichi Yamamoto (LINE Corp.) Eunwoo Song (NAVER Corp.) Jae-Min Kim (NAVER Corp.) Abstract This paper proposes an effective probability density distillation (PDD) algorithm for WaveNet-based parallel waveform generation (PWG) systems. Recently proposed teacher-student frameworks in the PWG system have…

LJSpeech は価値のあるデータセットですが、ニューラルボコーダの品質比較には向かないと思います

LJSpeech Dataset: https://keithito.com/LJ-Speech-Dataset/ まとめ 最近いろんな研究で LJSpeech が使われていますが、合成音の品質を比べるならクリーンなデータセットを使ったほうがいいですね。でないと、合成音声に含まれるノイズがモデルの限界からくるノイズなのかコーパスの音声が含むノイズ(LJSpeechの場合リバーブっぽい音)なのか区別できなくて、公平に比較するのが難しいと思います。 例えば、LJSpeechを使うと、ぶっちゃけ WaveGlow がWaveNetと比べて品質がいいかどうかわかんないですよね… 1 . 例えば最近のNICT岡本さんの研究 ( 基本周波数とメルケプストラムを用いたリアルタイムニューラルボコーダに関する検討 ) を引用すると、実際にクリーンなデータで実験すれば(Noise shaping…

WN-based TTSやりました / Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions [arXiv:1712.05884]

Summary Thank you for coming to see my blog post about WaveNet text-to-speech. Your browser does not support the audio element. 論文リンク: https://arxiv.org/abs/1712.05884 オンラインデモ: Tacotron2: WaveNet-based text-to-speech demo コード r9y9/wavenet_vocoder , Rayhane-mamah/Tacotron-2 音声サンプル: https://r9y9.github.io/wavenet_vocoder/ 三行まとめ 自作WaveNet ( WN ) と既存実装Tacotron 2 (WNを除く) を組み合わせて、英語TTSを作りました…

WaveNet vocoder をやってみましたので、その記録です / WaveNet: A Generative Model for Raw Audio [arXiv:1609.03499]

Summary コード: https://github.com/r9y9/wavenet_vocoder 音声サンプル: https://r9y9.github.io/wavenet_vocoder/ 三行まとめ Local / global conditioning を最低要件と考えて、WaveNet を実装しました DeepVoice3 / Tacotron2 の一部として使えることを目標に作りました PixelCNN++ の旨味を少し拝借し、16-bit linear PCMのscalarを入力として、(まぁまぁ)良い22.5kHzの音声を生成させるところまでできました Tacotron2 は、あとはやればほぼできる感じですが、直近では僕の中で優先度が低めのため、しばらく実験をする予定はありません。興味のある方はやってみてください。 音声サンプル 左右どちらかが合成音声です^^…

An open-source implementation of WaveNet vocoder

An open-source implementation of Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning

【108 話者編】Deep Voice 3: 2000-Speaker Neural Text-to-Speech / arXiv:1710.07654 [cs.SD]

Summary 論文リンク: arXiv:1710.07654 コード: https://github.com/r9y9/deepvoice3_pytorch VCTK: https://datashare.ed.ac.uk/handle/10283/2950 音声サンプルまとめ: https://r9y9.github.io/deepvoice3_pytorch/ 三行まとめ arXiv:1710.07654: Deep Voice 3: 2000-Speaker Neural Text-to-Speech を読んで、複数話者の場合のモデルを実装しました 論文のタイトル通りの2000話者とはいきませんが、 VCTK を使って、108 話者対応の英語TTSモデルを作りました(学習時間1日くらい)…

Interactive C++: Jupyter上で対話的にC++を使う方法の紹介 [Jupyter Advent Calendar 2017]

Jupyter Advent Calendar 2017 21日目の記事です。 C++をJupyterで使う方法はいくつかあります。この記事では、僕が試したことのある以下の4つの方法について、比較しつつ紹介したいと思います。 root/cling 付属のカーネル root/root 付属のカーネル xeus-cling Keno/Cxx.jl をIJuliaで使う まとめとして、簡単に特徴などを表にまとめておきますので、選ぶ際の参考にしてください。詳細な説明は後に続きます。 cling ROOT xeus-cling Cxx.jl + IJulia C++インタプリタ実装 C++ C++ C++ Julia + C++ (Tab) Code completion ○ ○ ○ x Cインタプリタ △ 1 △ △ ○ %magics x %%cpp, %%jsroot, その他 x △ 2…

ニューラルネットの学習過程の可視化を題材に、Jupyter + Bokeh で動的な描画を行う方法の紹介 [Jupyter Advent Calendar 2017]

Line https://bokeh.pydata.org/en/latest/docs/reference/models/glyphs/line.html VBar https://bokeh.pydata.org/en/latest/docs/reference/models/glyphs/vbar.html HBar https://bokeh.pydata.org/en/latest/docs/reference/models/glyphs/hbar.html ImageRGBA https://bokeh.pydata.org/en/latest/docs/reference/models/glyphs/image_rgba.html ImageRGBA…

【単一話者編】Deep Voice 3: 2000-Speaker Neural Text-to-Speech / arXiv:1710.07654 [cs.SD]

Summary 論文リンク: arXiv:1710.07654 コード: https://github.com/r9y9/deepvoice3_pytorch 三行まとめ arXiv:1710.07654: Deep Voice 3: 2000-Speaker Neural Text-to-Speech を読んで、単一話者の場合のモデルを実装しました(複数話者の場合は、今実験中です ( deepvoice3_pytorch/#6 ) arXiv:1710.08969 と同じく、RNNではなくCNNを使うのが肝です 例によって LJSpeech Dataset を使って、英語TTSモデルを作りました(学習時間半日くらい)。論文に記載のハイパーパラメータでは良い結果が得られなかったのですが、 arXiv:1710.08969 のアイデアをいくつか借りることで、良い結果を得ることができました。…

Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention. [arXiv:1710.08969]

Summary 論文リンク: arXiv:1710.08969 コード: https://github.com/r9y9/deepvoice3_pytorch 三行まとめ arXiv:1710.08969: Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention. を読んで、実装しました RNNではなくCNNを使うのが肝で、オープンソースTacotronと同等以上の品質でありながら、 高速に (一日程度で) 学習できる のが売りのようです。 LJSpeech Dataset を使って、英語TTSモデルを作りました(学習時間一日くらい)。完全再現とまではいきませんが、大まかに論文の主張を確認できました。 前置き 本当は DeepVoice3…

日本語 End-to-end 音声合成に使えるコーパス JSUT の前処理 [arXiv:1711.00354]

Summary コーパス配布先リンク: JSUT (Japanese speech corpus of Saruwatari Lab, University of Tokyo) - Shinnosuke Takamichi (高道 慎之介) 論文リンク: arXiv:1711.00354 三行まとめ 日本語End-to-end音声合成に使えるコーパスは神、ありがとうございます クリーンな音声であるとはいえ、冒頭/末尾の無音区間は削除されていない、またボタンポチッみたいな音も稀に入っているので注意 僕が行った無音区間除去の方法(Juliusで音素アライメントを取って云々)を記録しておくので、必要になった方は参考にどうぞ。ラベルファイルだけほしい人は連絡ください JSUT とは ツイート引用:…