M.S. in Applied AI, Nanyang Technological University
·
LLM Algorithm Intern, Huawei
shock22330@gmail.com
·
LinkedIn
·
Available for full-time roles from late Oct 2026
I work on LLM post-training: distributed RL (GRPO / GSPO / DPO), agentic pipelines, and the systems work that makes them run on large GPU/NPU clusters. Previously I built production Agent/RAG services and code-LLM data pipelines. I care about the path from alignment algorithms to stable, measurable training.
| Post-Training & RL | Languages | Systems & Serving |
|
|
|
|
| Agents & Data | DevOps & Cloud-Native | Backend |
|
|
|
|
-
LMSYS Chatbot Arena Human Preference Predictions (Kaggle)
- QLoRA SFT of Qwen3-14B + DPO on non-tie pairs (RTX 6000); stacked LLM OOF with GBDT. Log loss 0.95 (competition winner 0.96).
-
📜 "Network Security Threat Detection System Based on Knowledge Graph"
- Second author; YAC 2025 / IEEE Xplore. Multi-source KG (ATT&CK, CVE, D3FEND) on Neo4j; RAG with Weaviate + GPT-4o / Llama 3 routing.
M.S. in Applied Artificial Intelligence @ Nanyang Technological University (Nov 2025 – Oct 2026)
B.E. in Software Engineering @ Beijing Jiaotong University (Sep 2021 – Jun 2025)


