Skip to content

Should I modify the Quartznet 5x15 architecture for longer sentences? #1302

Description

@catskillsresearch

I have training data where some clips are up to 33 seconds long. I've tried splitting them on silence and allocating the text proportional to size but that is very ad hoc. So I prefer to keep them.

I've noticed during training that the QuartzNet 15x5 tends to more errors at the end of a phrase than at the beginning. I feel like it may be losing long-term dependencies or that those are harder to learn.

On the other hand, what I hear you saying about the parameter space is that it is independent of the max_duration parameter of the config. So same number of parameters to parse 1 second phrases as 33 second phrases.

Is that really true? Or can I get better long-term dependency by adding layers? If so, can you recommend a config which will work better for training against longer phrases?

If that would improve matters, I assume that means I have to train from scratch? I.e. if I add layers, I can't use the pretrained weights for the base Q15x5? Or if I can, can you give me a snippet of code which would show me how to add the extra layers to the pretrained model?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions