Skip to content
View Yofuria's full-sized avatar

Highlights

  • Pro

Block or report Yofuria

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Yofuria/README.md

Xiaobo Wang | 王晓博

PhD Student @ USTC · Research Intern @ BIGAI

Reward Modeling · LLM Alignment · Continual Learning · Agent Memory

Homepage · Google Scholar · Hugging Face · Email

Unique visitor-days

About

I work on building language models that keep improving from experience through reliable reward signals, robust alignment, continual adaptation, and memory.

I am currently a PhD student at the University of Science and Technology of China and a research intern at the Beijing Institute for General Artificial Intelligence (BIGAI). I am open to research collaboration.

Selected Research

  • SAVE — On-policy feedback for reward model self-supervised improvement. [Paper] [Project]
  • UAPO — Uncertainty-aware preference optimization. EMNLP 2025. [Paper] [Data]
  • PoliCon — Evaluating LLMs on diverse political consensus objectives. ICLR 2026. [Paper] [Code]
  • ICE — Learning knowledge from self-induced contextual distributions. ICLR 2025. [Paper] [Code]
  • RAM — An ever-improving memory system that learns from communication. [Paper] [Code]

Current Projects

Pinned Loading

  1. simple-agent-lab/RSIHub simple-agent-lab/RSIHub Public

    A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.

    Python 127 6

  2. bigai-nlco/belief-context-graph bigai-nlco/belief-context-graph Public

    A context graph which is based on confidence-aware belief that drives more determinstic reasoning in long-horizon agents.

    TypeScript 165 2

  3. simple-agent-lab/simple-long-horizon-agent simple-agent-lab/simple-long-horizon-agent Public

    Simple Long Horizon Agent - A simple yet effective AI agent for learning, experimentation, and long horizon work.

    Python 33 1

  4. bigai-nlco/PoliCon bigai-nlco/PoliCon Public

    ICLR 2026 | PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives

    Python 4 1

  5. bigai-nlco/RAM bigai-nlco/RAM Public

    An innovative RAG-based framework with an ever-improving memory.

    Python 2

  6. Yofuria.github.io Yofuria.github.io Public

    Forked from RayeRen/acad-homepage.github.io

    Personal homepage of Xiaobo Wang.

    HTML