Qualcomm AI Engine Direct - [LLM QAT] LLM Quant-Aware Distillation (QAD)#21036
Qualcomm AI Engine Direct - [LLM QAT] LLM Quant-Aware Distillation (QAD)#21036DannyYuyang-quic wants to merge 1 commit into
Conversation
…for LLMs Training: - Add BaseTrainer, Trainer (CE), KDTrainer (knowledge distillation) - Add CrossEntropyLoss, KLDivergenceLoss, linear warm-up cosine LR scheduler - Add TrainingArgs dataclass (epochs, lr, alpha, temperature, grad_accum_steps, warmup_ratio, lr_config) Data pipeline: - Add DecoderDatasetBuilder.from_hf_source for HuggingFace SFT chat datasets - Add build_qat_dataloaders: explicit calib/train split - Add LLMTrainingCollator, make_causal_labels, make_conversation_labels - Add DataConfig train fields: train_tasks, train_hf_dataset, train_hf_limit Quantization Strategy: - Add QATStrategy: PTQ calibration pass followed by KDTrainer/Trainer fine-tuning - Branch TextDecoder.quantize on --qat: prepare_qat_pt2e + move_exported_model_to_train vs prepare_pt2e - Select qat_recipe over quant_recipe when --qat is active Quant recipe: - Add StaticLLMQATRecipe base class as a example - Add Smollm2QATQuantRecipe; add qat_recipe field to LLMModelConfig - Add 16a8w QAT qconfig Fix: - SeqMSE: unwrap FakeQuantize wrapper before extracting observer Testing: - Add test_static_llm_qat: assert QAT PPL < PTQ PPL on smollm2_135m
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21036
Note: Links to docs will display an error until the docs builds have been completed.
|
|
@psiddh Hi, With QAT, we can further push quantization down to W4 PCQ, and potentially even lower precisions (W2). Below is a comparison between PTQ and QAT on Please have a look, thanks! Experiment Results: W4 Decoder PCQ (QAT vs PTQ)QAT (W4 PCQ decoder)python examples/qualcomm/oss_scripts/llama/llama.py --build_folder build-android --device ${SERIAL_NUM} --soc_model ${SOC_MODEL} --decoder_model smollm2_135m --model_mode hybrid --max_seq_len 1024 --prompt "What dose it mean to edit written content?" --qat --train_hf_dataset "HuggingFaceTB/smol-smoltalk" --train_hf_limit 1000 --calib_hf_dataset "HuggingFaceTB/smol-smoltalk" --calib_hf_limit 1000Result[INFO 2026-07-20 12:10:52,191 llama.py:290] Device Inference Results[0]:
<|im_start|>user
What dose it mean to edit written content?<|im_end|>
<|im_start|>assistant
1. **Editing for clarity**: When you edit your writing, you aim to make it clear and concise, ensuring that your message is easy for readers to understand and retain. This involves cutting unnecessary words, phrases, or sentences, and rephrasing or reorganizing your content to make it more readable.
2. **Editing for grammar**: Grammar is another crucial aspect of writing. When you edit your writing, you check for errors in grammar, syntax, and punctuation, which can make your writing more polished and error-free.
3. **Editing for style**: The final stage of editing is where you refine your writing to make it more engaging, clear, and engaging. This involves using language that is engaging, yet clear, and engaging in the way you want it to be.
4. **Editing for impact**: The final stage of editing is where you aim to make your writing more impactful. This means that you strive to make your message clear, concise, and impactful, and to convey it in a way that resonates with your readers.
5. **Editing for consistency**: When you edit your writing, you ensure that it is consistent in terms of tone, style, and voice. This involves using the same language, vocabulary, and structure to create a consistent tone, which is essential for effective communication.
To get started, take a step back and look at your writing. Read it aloud, if possible, to get a sense of the tone, style, and voice you're using. Then, read your writing aloud to get a sense of the rhythm, cadence, and flow.
When you're on the right track, you can begin to refine your writing and make it more engaging and effective.<|im_end|>PTQ (W4 PCQ decoder)python examples/qualcomm/oss_scripts/llama/llama.py --build_folder build-android --device ${SERIAL_NUM} --soc_model ${SOC_MODEL} --decoder_model smollm2_135m --model_mode hybrid --max_seq_len 1024 --prompt "What dose it mean to edit written content?" --calib_hf_dataset "HuggingFaceTB/smol-smoltalk" --calib_hf_limit 1000Result[INFO 2026-07-20 15:12:57,609 llama.py:290] Device Inference Results[0]:
<|im_start|>user
What dose it mean to edit written content?<|im_end|>
<|im_start|>assistant
1. A sentence is usually written in the first person, i.e., "He said, she said" or "She said, she said."
2. A paragraph is usually written in the third person, i.e., "She said, he said" or "She said, she said"
3. A couple is usually written in the second person, i.e., "She said, he said" or "He said, she said"
4. A little is usually written in the third person, i.e., "She said, he said" or "She said, she said"
5. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
6. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
7. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
8. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
9. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
10. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "He said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., " |
|
@pytorchbot label "release notes: qualcomm" |
Summary
Training:
Data pipeline:
Quantization Strategy:
Quant recipe:
Fix:
CI Testing:
README:
examples/qualcomm/oss_scripts/llama/quantization_guidance.mdE2E script:
Result
Test plan
cc: @shewu-quic @haowhsu-quic @winskuo-quic @psiddh @abhinaykukkadapu