Skip to content

Monthly paper sweep — 2026-07 #3

Description

@github-actions

Monthly paper sweep — 2026-07

7 candidate(s) from the last 35 days not yet in INDEX.md / academia.md. Triage each per INTEGRATION_GUIDE.md §1 (scope) and §2 (paradigm + dedicated-file-vs-academia).

2606.31825 — Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

submitted 2026-06-30

Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult …

2606.30378 — OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

submitted 2026-06-29

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of panoramic imagery. The full 360°$\times$180° field of view of panoramas essentially supports complex global multi-step re…

2606.28164 — EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography

submitted 2026-06-26

Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreting an echocardiogram requires synthesizing complementary evidence across multiple heart views to identify abnormalities and produce structured clini…

2606.27922 — Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

submitted 2026-06-26

Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforce…

2606.27608 — Qwen-Image-2.0-RL Technical Report

submitted 2026-06-25

We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instruction-following capability of the Qwen-Image-2.0 diffusion model. To provide reliable reward signals, we…

2606.26930 — PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation

submitted 2026-06-25

Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods often favor superficial aesthetics, such as over-saturated colors, leaving critical flaws like AI artifacts and biological implausibilities unresolve…

2606.01985 — MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

submitted 2026-06-01

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-tu…

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions