I'm a PhD graduate from Zhejiang University, working at the intersection of AI infrastructure and computer vision. I contribute to large-scale training frameworks (VeOmni, verl) while publishing research at top-tier venues (NeurIPS, ICCV, IJCAI, ECAI, IEEE TMI).
My research spans distributed training systems, vision-language models, and computational imaging.
Selected Publications:
- Dig into Multi-modal Cues for Video Retrieval with Hierarchical Alignment — IJCAI 2021 (first author)
- SinLane: Siamese Visual Transformer via Pyramid Feature Integration for Lane Detection — ECAI 2024
- A Siamese Transformer with Hierarchical Refinement for Lane Detection — NeurIPS 2024
- IFPNet: Integrated Feature Pyramid Network with Fusion Factor for Lane Detection — ICCV Workshop 2023
- PR-Net: Preference Reasoning for Personalized Video Highlight Detection — ICCV 2021
- Interactive Few-Shot Learning: Limited Supervision, Better Medical Image Segmentation — IEEE TMI 2021
- Temporal Structured Illumination and Vision-Transformer Enables Large Field-of-View Binary Snapshot Ptychography — Optics Express 2024
- Propagation Dynamics of a Spatiotemporal Vortex Pulse in the Spatial Fractional System — Photonics Research 2024
- FlowTrain: Flow-Based Decoupled Training for Industrial-Grade Vision-Language Models — arXiv 2026
📖 Full publication list: Google Scholar
I'm a contributor to open-source AI training frameworks, including VeOmni (model-centric distributed training recipe zoo) and verl / verl-omni (multimodal RL post-training frameworks).
Selected Contributions:
- VeOmni: Deterministic mapping cycles for dataset mixtures, HSDP all-reduce control, and Qwen3-VL FSDP2 compilation support
- verl-omni: Refill incomplete diffusion rollout groups, batched PickScore inference, and preserve diffusion asynchronous semantics
Research directions:
- Parallelism & memory optimization — packed sequence packing, ChunkMBS, FSDP2/hsdp integration
- Training-time validation — distributed metric aggregation and evaluation pipeline
- Hardware portability — NPU training enablement and cross-platform validation
- Multimodal reward engineering — batched reward inference and reward model integration
- Rollout & training pipeline — async semantics, rollout-train consistency, group management
- Efficient post-training — LoRA/FSDP for VLMs, distributed reward handling
Languages Python · C++ · CUDA · Shell
Frameworks PyTorch · VeOmni · verl · FSDP/FSDP2 · DeepSpeed
Distributed DDP · HSDP · FSDP · Tensor Parallel · Expert Parallel · ChunkMBS
Infrastructure Slurm · Ray · wandb · Docker
Vision/RL Diffusion Models · VLMs · Video Retrieval · Medical Imaging
- 📧 Email: w_wenzhe@163.com
- 🏠 Location: Hangzhou, China
💫 "From research papers to production training frameworks — bridging the gap between algorithms and infrastructure."

