This flag is used in some of the launch files like https://github.com/NVIDIA/NeMo/blob/main/examples/nlp/language_modeling/conf/megatron_gpt_config.yaml. I know this flag causes different implementation of the model to be used, but can their differences and recommendations for when to use one vs another be explained in more detail in the doc?
This flag is used in some of the launch files like https://github.com/NVIDIA/NeMo/blob/main/examples/nlp/language_modeling/conf/megatron_gpt_config.yaml. I know this flag causes different implementation of the model to be used, but can their differences and recommendations for when to use one vs another be explained in more detail in the doc?