GitHub - Porcupine1/ese539_proj-energy_aware_quantization: End-to-end framework for evaluating accuracy, latency, and energy of quantized LLMs (FP32, FP16, INT8). Includes pre-tokenized SST-2 datasets, reproducible inference harness, GPU power logging via nvidia-smi, and analysis tools for energy-aware model deployment.

Branches Tags

Name		Name	Last commit message	Last commit date
Latest commit History 97 Commits
datasets		datasets
notebooks		notebooks
.gitignore		.gitignore
requirements.txt		requirements.txt

About

End-to-end framework for evaluating accuracy, latency, and energy of quantized LLMs (FP32, FP16, INT8). Includes pre-tokenized SST-2 datasets, reproducible inference harness, GPU power logging via nvidia-smi, and analysis tools for energy-aware model deployment.