Would a QuartzNet 15x5 pretrained on English have a notably hard time hearing nuances in Vietnamese? Will training gradually improve the pretrained Encoder layer to fix this, or is the Encoder layer too far down, gradient-wise, to notice? I'm just asking because I've been retraining QuartzNet 15x5 English for a number of days on Vietnamese and it still seems to be missing what seems to be isolated phonemic subtleties quite frequently, and my intuitive sense is that it is listening with an English "ear" that just can't distinguish the Vietnamese phonemes.
Can you think of some workarounds to improve this other than just grinding on with the training? For example:
[NeMo I 2020-10-24 10:51:06 wer:149] reference:chứ sao cũng gần mà chắc là cố gắng cưới lúc mùa hè để cho bé Dung về
[NeMo I 2020-10-24 10:51:06 wer:150] decoded :chứ sao cũng gần mà chắc là cố gắng cướii lúc mùa hè để cho bé Dung vềé
Epoch 171: 21%|███████████████████████████████████████████▋ | 99/464 [00:58<03:35, 1.69it/s, loss=8.938, v_num=4][NeMo I 2020-10-24 10:51:35 wer:148]
[NeMo I 2020-10-24 10:51:35 wer:149] reference:chửi gì chút xíu nữa đi chút xíu nữa à
[NeMo I 2020-10-24 10:51:35 wer:150] decoded :chứi vì chút xíu nữa đâi chút xíu nữa àa
Epoch 171: 32%|█████████████████████████████████████████████████████████████████▌ | 149/464 [01:28<03:07, 1.68it/s, loss=9.278, v_num=4][NeMo I 2020-10-24 10:52:05 wer:148]
[NeMo I 2020-10-24 10:52:05 wer:149] reference:chắc là không ơi cuối tháng sáu mới thi xong mà
[NeMo I 2020-10-24 10:52:05 wer:150] decoded :chắc là không ơi cuối tháng sáu mới thi xongàé
Epoch 171: 43%|███████████████████████████████████████████████████████████████████████████████████████▍ | 199/464 [01:55<02:34, 1.72it/s, loss=8.951, v_num=4][NeMo I 2020-10-24 10:52:32 wer:148]
[NeMo I 2020-10-24 10:52:32 wer:149] reference:tự nhiên ở nhà với Tư rồi vòng về cho tụi nó ăn bữa còn có một ngày à
[NeMo I 2020-10-24 10:52:32 wer:150] decoded :tự nhên ở nhà với Tư rồi vòng về cho tụi ng ăn bữa còn có một ngày àé
Would a QuartzNet 15x5 pretrained on English have a notably hard time hearing nuances in Vietnamese? Will training gradually improve the pretrained Encoder layer to fix this, or is the Encoder layer too far down, gradient-wise, to notice? I'm just asking because I've been retraining QuartzNet 15x5 English for a number of days on Vietnamese and it still seems to be missing what seems to be isolated phonemic subtleties quite frequently, and my intuitive sense is that it is listening with an English "ear" that just can't distinguish the Vietnamese phonemes.
Can you think of some workarounds to improve this other than just grinding on with the training? For example: