* Create emo_gen.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * update server.py, fix bugs in func get_text() and infer(). (#52) * Extract get_text() and infer() from webui.py. (#53) * Extract get_text() and infer() from webui.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * add emo emb * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * init emo gen * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * init emo * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * init emo * Delete bert/bert-base-japanese-v3 directory * Create .gitkeep * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Create add_punc.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix bug in bert_gen.py (#54) * Update README.md * fix bug in models.py (#56) * 更新 models.py * Fix japanese cleaner (#61) * 初步,睡觉明天继续写( * 好好好放错分支了,熬夜是大忌 * [pre-commit.ci] pre-commit autoupdate (#55) * [pre-commit.ci] pre-commit autoupdate updates: - [github.com/pre-commit/pre-commit-hooks: v4.4.0 → v4.5.0](https://github.com/pre-commit/pre-commit-hooks/compare/v4.4.0...v4.5.0) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Create tokenizer_config.json * update preprocess_text.py:过滤一个音频匹配多个文本的情况 (#57) * update preprocess_text.py:过滤音频不存在的情况 (#58) * 修复日语cleaner和bert * better * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Sora <atri@suzakuintsubaki.com> * Apply Code Formatter Change * Add config.yml for global configuration. (#62) * Add config.yml for global configuration. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix bug in webui.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Rename config.yml to default_config.yml. Add ./config.yml to gitignore. * Add config.py to parse config.yml * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Update webui.py (#65) * Update webui.py: 1. Add auto translation from Chinese to Japanese. 2. Start to use config.py in webui.py to set config instead of using the command line. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix (#68) * 加上ー * fix * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Update infer.py and webui.py. Supports loading and inference models of 1.1.1 version. (#66) * Update infer.py and webui.py. Supports loading and inference models of 1.1.1 version. * Update config.json * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix bug in translate.py (#69) * Supports loading and inference models of 1.1、1.0.1、1.0 version. (#70) * Supports loading and inference models of 1.1、1.0.1、1.0 version. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Delete useless file in OldVersion --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Update japanese.py (#71) Handling JA long pronunciations * 使用配置文件配置bert_gen.py, preprocess_text.py, resample.py (#72) * Update bert_gen.py, preprocess_text.py, resample.py. Support using config.yml in these files. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update bert_gen.py * Update bert_gen.py, fix bug. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Delete bert/bert-base-japanese-v3 directory * Create config.json * Create tokenizer_config.json * Create vocab.txt * Update server.py. 支持多版本多模型 (#76) * Update server.py. 支持多版本多模型 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Dev webui (#77) * 申请pr (#75) * 2023/10/11 update 界面优化 * Update webui.py 翻译英文页面为中文 * Update train_ms.py 单卡训练 * 加入图片 * Update extern_subprocess.py * Update asr_transcript.py * Update asr_transcript.py * Update asr_transcript.py * Update extern_subprocess.py * Update asr_transcript.py * Update asr_transcript.py * Update asr_transcript.py * Update all_process.py * Update extern_subprocess.py * Update all_process.py * Update all_process.py * Update asr_transcript.py * Update extern_subprocess.py * Update webui.py * Create re_matching.py * Update webui.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update all_process.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update all_process.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update all_process.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update asr_transcript.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Pack 'update' functions into a module * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update all_process.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update asr_transcript.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update extern_subprocess.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update README.md * Update README.md * Update README.md * Update README.md * Update README.md * Update all_process.py * Update asr_transcript.py * Update webui.py * Add files via upload * Update extern_subprocess.py * Update all_process.py * Update asr_transcript.py * Update bert_gen.py * Update extern_subprocess.py * Update preprocess_text.py * Update re_matching.py * Update resample.py * Update update_status.py * Update update_status.py * Update webui.py * Update all_process.py * Update preprocess_text.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update train_ms.py --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> Co-authored-by: innnky <67028263+innnky@users.noreply.github.com> * Delete all_process.py * Delete asr_transcript.py * Delete extern_subprocess.py --------- Co-authored-by: spicysama <122108331+AnyaCoder@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: innnky <67028263+innnky@users.noreply.github.com> * Create config.json * Create preprocessor_config.json * Create vocab.json * Delete emotional/wav2vec2-large-robust-12-ft-emotion-msp-dim/.gitkeep * Update emo_gen.py * Delete add_punc.py * add emotion_clustering.i * Apply Code Formatter Change * Update models.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update preprocess_text.py (#78) * Update preprocess_text.py. 检测重复以及不存在的音频 (#79) * Handle Janpanese long pronunciations (#80) * Handle Janpanese long pronunciations * Update japanese.py * Update japanese.py * Use unified phonemes for Japanese long vowel (#82) * Use an unified phoneme for Japanese long vowel `symbol.py` has not been updated to ensure compatibility with older version models. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * 增加一个按钮,点击后可以按句子切分,添加“|” (#81) * Update re_matching.py * Update webui.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix phonemer bug (#83) * Fix phonemer bug * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix long vowel handler bug (#84) * Fix long vowel handler bug * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * 加入整合包管理器的特性:长文本合成可以自定义句间段间停顿 (#85) * Update webui.py * Update re_matching.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Update train_ms.py * fix' * Update cleaner.py * add en * add en * Update english.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add en * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add en * add en * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add en * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * 更新 README.md * 更新 README.md * 更新 README.md * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Change phonemer to pyopenjtalk (#86) * Change phonemer to pyopenjtalk * 修改为openjtalk便于安装 --------- Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> * 更新 english.py * Fix english_bert_mock.py. (#87) * Add punctuation execptions (#88) * Add punctuation execptions * Ellipses exceptions * remove get bert * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix bug in oldVersion. (#89) * Update requirements.txt * change to large * rollback requirements.txt * Feat: Enable 1.1.1 models using fix-ver infer. (#91) * Feat: Enable 1.1.1 models using fix-ver infer. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Add Japanese accent (high-low) (#90) * Add punctuation execptions * Ellipses exceptions * Add Japanese accent * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Do not replace iteration mark (#92) * Add punctuation execptions * Ellipses exceptions * Add Japanese accent * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Do not replace iteration mark --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix: fix import error in oldVersion (#93) * Refactor: reusing model loading in webui.py and server.py. (#94) * Feat: Enable using config.yml in train_ms.py (#96) * 更新 emo_gen.py * Change emo_gen.py (#97) * Fix emo_gen bugs * Add multiprocess * Fix queue (#98) * Fix emo_gen bugs * Add multiprocess * Del var * Fix queue * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix training bugs (#99) * Updatge cluster notebook * Fix train * Fix filename * Update infer.py (#100) * Update infer.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Add reference audio (#101) * Add reference audio * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update * Update * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> * Fix: fix 1.1.1-fix (#102) * Fix infer bug (#103) * Feat: Add server_fastapi.py. (#104) * Feat: Add server_fastapi.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix: Update requirements.txt. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix: requirements.txt. (#105) * Swith to deberta-v3-large (#106) * Swith to deberta-v3-large * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Feat: Update config.py. (#107) * Feat: Update config.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Dev fix (#108) * fix bugs when deploying * fix bugs when deploying * fix bugs when deploying * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Revert "Dev fix (#108)" (#109) This reverts commit 685e18a10498d602b1a9a26079340d11925646f0. * Dev fix (#110) * fix bugs when deploying * fix bugs when deploying * fix bugs when deploying * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix fixed bugs * fix fixed bugs * fix fixed bug 3 * fix fixed bug 4 * fix fixed bug 5 * fix * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Add emo vec quantizer (#111) Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> * Clean req and gitignore (#112) * Clean req and gitignore * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Switch to deberta-v2-large-japanese (#113) * Switch to deberta-v2-large-japanese * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix emo bugs (#114) * Fix english (#115) * Remove emo (#117) * Don't train codebook * Remove emo * Update * Update * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Merge dev into no-emo (#122) * [pre-commit.ci] pre-commit autoupdate (#95) * [pre-commit.ci] pre-commit autoupdate updates: - [github.com/astral-sh/ruff-pre-commit: v0.0.292 → v0.1.1](https://github.com/astral-sh/ruff-pre-commit/compare/v0.0.292...v0.1.1) - [github.com/psf/black: 23.9.1 → 23.10.0](https://github.com/psf/black/compare/23.9.1...23.10.0) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Don't train codebook (#116) * Update requirements.txt * Update english_bert_mock.py * Fix: server_fastapi.py (#118) * Fix: server_fastapi.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Fix: don't print debug logging. (#119) * Fix: don't print debug logging. * Feat: support emo_gen config * Fix config * Apply Code Formatter Change * 更新,修正bug (#121) * Feat: Update infer.py preprocess_text.py server_fastapi.py. * Fix resample.py. Maintain same directory structure in out_dir as in_dir. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update resample.py * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * Update server_fastapi.py to no-emo ver * Update config.py, no emo config --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: OedoSoldier <31711261+OedoSoldier@users.noreply.github.com> Co-authored-by: Stardust·减 <star_dust_chen@foxmail.com> Co-authored-by: Stardust-minus <Stardust-minus@users.noreply.github.com> * Update train_ms.py * Update latest version info (#124) --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: jiangyuxiaoxiao <atri@suzakuintsubaki.com> Co-authored-by: AkitoLiu <39857739+Akito-UzukiP@users.noreply.github.com> Co-authored-by: Stardust-minus <Stardust-minus@users.noreply.github.com> Co-authored-by: OedoSoldier <31711261+OedoSoldier@users.noreply.github.com> Co-authored-by: spicysama <122108331+AnyaCoder@users.noreply.github.com> Co-authored-by: innnky <67028263+innnky@users.noreply.github.com> Co-authored-by: YYuX-1145 <138500330+YYuX-1145@users.noreply.github.com>
400 lines
11 KiB
Python
400 lines
11 KiB
Python
# Convert Japanese text to phonemes which is
|
||
# compatible with Julius https://github.com/julius-speech/segmentation-kit
|
||
import re
|
||
import unicodedata
|
||
|
||
from transformers import AutoTokenizer
|
||
|
||
from text import punctuation, symbols
|
||
|
||
from num2words import num2words
|
||
|
||
import pyopenjtalk
|
||
import jaconv
|
||
|
||
|
||
def kata2phoneme(text: str) -> str:
|
||
"""Convert katakana text to phonemes."""
|
||
text = text.strip()
|
||
if text == "ー":
|
||
return ["ー"]
|
||
elif text.startswith("ー"):
|
||
return ["ー"] + kata2phoneme(text[1:])
|
||
res = []
|
||
prev = None
|
||
while text:
|
||
if re.match(_MARKS, text):
|
||
res.append(text)
|
||
text = text[1:]
|
||
continue
|
||
if text.startswith("ー"):
|
||
if prev:
|
||
res.append(prev[-1])
|
||
text = text[1:]
|
||
continue
|
||
res += pyopenjtalk.g2p(text).lower().replace("cl", "q").split(" ")
|
||
break
|
||
# res = _COLON_RX.sub(":", res)
|
||
return res
|
||
|
||
|
||
def hira2kata(text: str) -> str:
|
||
return jaconv.hira2kata(text)
|
||
|
||
|
||
_SYMBOL_TOKENS = set(list("・、。?!"))
|
||
_NO_YOMI_TOKENS = set(list("「」『』―()[][]"))
|
||
_MARKS = re.compile(
|
||
r"[^A-Za-z\d\u3005\u3040-\u30ff\u4e00-\u9fff\uff11-\uff19\uff21-\uff3a\uff41-\uff5a\uff66-\uff9d]"
|
||
)
|
||
|
||
|
||
def text2kata(text: str) -> str:
|
||
parsed = pyopenjtalk.run_frontend(text)
|
||
|
||
res = []
|
||
for parts in parsed:
|
||
word, yomi = replace_punctuation(parts["orig"]), parts["pron"].replace("’", "")
|
||
if yomi:
|
||
if re.match(_MARKS, yomi):
|
||
if len(word) > 1:
|
||
word = [replace_punctuation(i) for i in list(word)]
|
||
yomi = word
|
||
res += yomi
|
||
sep += word
|
||
continue
|
||
elif word not in rep_map.keys() and word not in rep_map.values():
|
||
word = ","
|
||
yomi = word
|
||
res.append(yomi)
|
||
else:
|
||
if word in _SYMBOL_TOKENS:
|
||
res.append(word)
|
||
elif word in ("っ", "ッ"):
|
||
res.append("ッ")
|
||
elif word in _NO_YOMI_TOKENS:
|
||
pass
|
||
else:
|
||
res.append(word)
|
||
return hira2kata("".join(res))
|
||
|
||
|
||
def text2sep_kata(text: str) -> (list, list):
|
||
parsed = pyopenjtalk.run_frontend(text)
|
||
|
||
res = []
|
||
sep = []
|
||
for parts in parsed:
|
||
word, yomi = replace_punctuation(parts["orig"]), parts["pron"].replace("’", "")
|
||
if yomi:
|
||
if re.match(_MARKS, yomi):
|
||
if len(word) > 1:
|
||
word = [replace_punctuation(i) for i in list(word)]
|
||
yomi = word
|
||
res += yomi
|
||
sep += word
|
||
continue
|
||
elif word not in rep_map.keys() and word not in rep_map.values():
|
||
word = ","
|
||
yomi = word
|
||
res.append(yomi)
|
||
else:
|
||
if word in _SYMBOL_TOKENS:
|
||
res.append(word)
|
||
elif word in ("っ", "ッ"):
|
||
res.append("ッ")
|
||
elif word in _NO_YOMI_TOKENS:
|
||
pass
|
||
else:
|
||
res.append(word)
|
||
sep.append(word)
|
||
return sep, [hira2kata(i) for i in res], get_accent(parsed)
|
||
|
||
|
||
def get_accent(parsed):
|
||
labels = pyopenjtalk.make_label(parsed)
|
||
|
||
phonemes = []
|
||
accents = []
|
||
for n, label in enumerate(labels):
|
||
phoneme = re.search(r"\-([^\+]*)\+", label).group(1)
|
||
if phoneme not in ["sil", "pau"]:
|
||
phonemes.append(phoneme.replace("cl", "q").lower())
|
||
else:
|
||
continue
|
||
a1 = int(re.search(r"/A:(\-?[0-9]+)\+", label).group(1))
|
||
a2 = int(re.search(r"\+(\d+)\+", label).group(1))
|
||
if re.search(r"\-([^\+]*)\+", labels[n + 1]).group(1) in ["sil", "pau"]:
|
||
a2_next = -1
|
||
else:
|
||
a2_next = int(re.search(r"\+(\d+)\+", labels[n + 1]).group(1))
|
||
# Falling
|
||
if a1 == 0 and a2_next == a2 + 1:
|
||
accents.append(-1)
|
||
# Rising
|
||
elif a2 == 1 and a2_next == 2:
|
||
accents.append(1)
|
||
else:
|
||
accents.append(0)
|
||
return list(zip(phonemes, accents))
|
||
|
||
|
||
_ALPHASYMBOL_YOMI = {
|
||
"#": "シャープ",
|
||
"%": "パーセント",
|
||
"&": "アンド",
|
||
"+": "プラス",
|
||
"-": "マイナス",
|
||
":": "コロン",
|
||
";": "セミコロン",
|
||
"<": "小なり",
|
||
"=": "イコール",
|
||
">": "大なり",
|
||
"@": "アット",
|
||
"a": "エー",
|
||
"b": "ビー",
|
||
"c": "シー",
|
||
"d": "ディー",
|
||
"e": "イー",
|
||
"f": "エフ",
|
||
"g": "ジー",
|
||
"h": "エイチ",
|
||
"i": "アイ",
|
||
"j": "ジェー",
|
||
"k": "ケー",
|
||
"l": "エル",
|
||
"m": "エム",
|
||
"n": "エヌ",
|
||
"o": "オー",
|
||
"p": "ピー",
|
||
"q": "キュー",
|
||
"r": "アール",
|
||
"s": "エス",
|
||
"t": "ティー",
|
||
"u": "ユー",
|
||
"v": "ブイ",
|
||
"w": "ダブリュー",
|
||
"x": "エックス",
|
||
"y": "ワイ",
|
||
"z": "ゼット",
|
||
"α": "アルファ",
|
||
"β": "ベータ",
|
||
"γ": "ガンマ",
|
||
"δ": "デルタ",
|
||
"ε": "イプシロン",
|
||
"ζ": "ゼータ",
|
||
"η": "イータ",
|
||
"θ": "シータ",
|
||
"ι": "イオタ",
|
||
"κ": "カッパ",
|
||
"λ": "ラムダ",
|
||
"μ": "ミュー",
|
||
"ν": "ニュー",
|
||
"ξ": "クサイ",
|
||
"ο": "オミクロン",
|
||
"π": "パイ",
|
||
"ρ": "ロー",
|
||
"σ": "シグマ",
|
||
"τ": "タウ",
|
||
"υ": "ウプシロン",
|
||
"φ": "ファイ",
|
||
"χ": "カイ",
|
||
"ψ": "プサイ",
|
||
"ω": "オメガ",
|
||
}
|
||
|
||
|
||
_NUMBER_WITH_SEPARATOR_RX = re.compile("[0-9]{1,3}(,[0-9]{3})+")
|
||
_CURRENCY_MAP = {"$": "ドル", "¥": "円", "£": "ポンド", "€": "ユーロ"}
|
||
_CURRENCY_RX = re.compile(r"([$¥£€])([0-9.]*[0-9])")
|
||
_NUMBER_RX = re.compile(r"[0-9]+(\.[0-9]+)?")
|
||
|
||
|
||
def japanese_convert_numbers_to_words(text: str) -> str:
|
||
res = _NUMBER_WITH_SEPARATOR_RX.sub(lambda m: m[0].replace(",", ""), text)
|
||
res = _CURRENCY_RX.sub(lambda m: m[2] + _CURRENCY_MAP.get(m[1], m[1]), res)
|
||
res = _NUMBER_RX.sub(lambda m: num2words(m[0], lang="ja"), res)
|
||
return res
|
||
|
||
|
||
def japanese_convert_alpha_symbols_to_words(text: str) -> str:
|
||
return "".join([_ALPHASYMBOL_YOMI.get(ch, ch) for ch in text.lower()])
|
||
|
||
|
||
def japanese_text_to_phonemes(text: str) -> str:
|
||
"""Convert Japanese text to phonemes."""
|
||
res = unicodedata.normalize("NFKC", text)
|
||
res = japanese_convert_numbers_to_words(res)
|
||
# res = japanese_convert_alpha_symbols_to_words(res)
|
||
res = text2kata(res)
|
||
res = kata2phoneme(res)
|
||
return res
|
||
|
||
|
||
def is_japanese_character(char):
|
||
# 定义日语文字系统的 Unicode 范围
|
||
japanese_ranges = [
|
||
(0x3040, 0x309F), # 平假名
|
||
(0x30A0, 0x30FF), # 片假名
|
||
(0x4E00, 0x9FFF), # 汉字 (CJK Unified Ideographs)
|
||
(0x3400, 0x4DBF), # 汉字扩展 A
|
||
(0x20000, 0x2A6DF), # 汉字扩展 B
|
||
# 可以根据需要添加其他汉字扩展范围
|
||
]
|
||
|
||
# 将字符的 Unicode 编码转换为整数
|
||
char_code = ord(char)
|
||
|
||
# 检查字符是否在任何一个日语范围内
|
||
for start, end in japanese_ranges:
|
||
if start <= char_code <= end:
|
||
return True
|
||
|
||
return False
|
||
|
||
|
||
rep_map = {
|
||
":": ",",
|
||
";": ",",
|
||
",": ",",
|
||
"。": ".",
|
||
"!": "!",
|
||
"?": "?",
|
||
"\n": ".",
|
||
".": ".",
|
||
"...": "…",
|
||
"···": "…",
|
||
"・・・": "…",
|
||
"·": ",",
|
||
"・": ",",
|
||
"、": ",",
|
||
"$": ".",
|
||
"“": "'",
|
||
"”": "'",
|
||
"‘": "'",
|
||
"’": "'",
|
||
"(": "'",
|
||
")": "'",
|
||
"(": "'",
|
||
")": "'",
|
||
"《": "'",
|
||
"》": "'",
|
||
"【": "'",
|
||
"】": "'",
|
||
"[": "'",
|
||
"]": "'",
|
||
"—": "-",
|
||
"−": "-",
|
||
"~": "-",
|
||
"~": "-",
|
||
"「": "'",
|
||
"」": "'",
|
||
}
|
||
|
||
|
||
def replace_punctuation(text):
|
||
pattern = re.compile("|".join(re.escape(p) for p in rep_map.keys()))
|
||
|
||
replaced_text = pattern.sub(lambda x: rep_map[x.group()], text)
|
||
|
||
replaced_text = re.sub(
|
||
r"[^\u3040-\u309F\u30A0-\u30FF\u4E00-\u9FFF\u3400-\u4DBF\u3005"
|
||
+ "".join(punctuation)
|
||
+ r"]+",
|
||
"",
|
||
replaced_text,
|
||
)
|
||
|
||
return replaced_text
|
||
|
||
|
||
def text_normalize(text):
|
||
res = unicodedata.normalize("NFKC", text)
|
||
res = japanese_convert_numbers_to_words(res)
|
||
# res = "".join([i for i in res if is_japanese_character(i)])
|
||
res = replace_punctuation(res)
|
||
return res
|
||
|
||
|
||
def distribute_phone(n_phone, n_word):
|
||
phones_per_word = [0] * n_word
|
||
for task in range(n_phone):
|
||
min_tasks = min(phones_per_word)
|
||
min_index = phones_per_word.index(min_tasks)
|
||
phones_per_word[min_index] += 1
|
||
return phones_per_word
|
||
|
||
|
||
def handle_long(sep_phonemes):
|
||
for i in range(len(sep_phonemes)):
|
||
if sep_phonemes[i][0] == "ー":
|
||
sep_phonemes[i][0] = sep_phonemes[i - 1][-1]
|
||
if "ー" in sep_phonemes[i]:
|
||
for j in range(len(sep_phonemes[i])):
|
||
if sep_phonemes[i][j] == "ー":
|
||
sep_phonemes[i][j] = sep_phonemes[i][j - 1][-1]
|
||
return sep_phonemes
|
||
|
||
|
||
tokenizer = AutoTokenizer.from_pretrained("./bert/deberta-v2-large-japanese")
|
||
|
||
|
||
def align_tones(phones, tones):
|
||
res = []
|
||
for pho in phones:
|
||
temp = [0] * len(pho)
|
||
for idx, p in enumerate(pho):
|
||
if len(tones) == 0:
|
||
break
|
||
if p == tones[0][0]:
|
||
temp[idx] = tones[0][1]
|
||
if idx > 0:
|
||
temp[idx] += temp[idx - 1]
|
||
tones.pop(0)
|
||
temp = [0] + temp
|
||
temp = temp[:-1]
|
||
if -1 in temp:
|
||
temp = [i + 1 for i in temp]
|
||
res.append(temp)
|
||
res = [i for j in res for i in j]
|
||
assert not any([i < 0 for i in res]) and not any([i > 1 for i in res])
|
||
return res
|
||
|
||
|
||
def g2p(norm_text):
|
||
sep_text, sep_kata, acc = text2sep_kata(norm_text)
|
||
sep_tokenized = [tokenizer.tokenize(i) for i in sep_text]
|
||
sep_phonemes = handle_long([kata2phoneme(i) for i in sep_kata])
|
||
# 异常处理,MeCab不认识的词的话会一路传到这里来,然后炸掉。目前来看只有那些超级稀有的生僻词会出现这种情况
|
||
for i in sep_phonemes:
|
||
for j in i:
|
||
assert j in symbols, (sep_text, sep_kata, sep_phonemes)
|
||
tones = align_tones(sep_phonemes, acc)
|
||
|
||
word2ph = []
|
||
for token, phoneme in zip(sep_tokenized, sep_phonemes):
|
||
phone_len = len(phoneme)
|
||
word_len = len(token)
|
||
|
||
aaa = distribute_phone(phone_len, word_len)
|
||
word2ph += aaa
|
||
phones = ["_"] + [j for i in sep_phonemes for j in i] + ["_"]
|
||
tones = [0] + tones + [0]
|
||
word2ph = [1] + word2ph + [1]
|
||
assert len(phones) == len(tones)
|
||
return phones, tones, word2ph
|
||
|
||
|
||
if __name__ == "__main__":
|
||
tokenizer = AutoTokenizer.from_pretrained("./bert/deberta-v2-large-japanese")
|
||
text = "hello,こんにちは、世界ー!……"
|
||
from text.japanese_bert import get_bert_feature
|
||
|
||
text = text_normalize(text)
|
||
print(text)
|
||
|
||
phones, tones, word2ph = g2p(text)
|
||
bert = get_bert_feature(text, word2ph)
|
||
|
||
print(phones, tones, word2ph, bert.shape)
|