Monthly paper sweep — 2026-07
7 candidate(s) from the last 35 days not yet in INDEX.md / academia.md. Triage each per INTEGRATION_GUIDE.md §1 (scope) and §2 (paradigm + dedicated-file-vs-academia).
2606.31825 — Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
submitted 2026-06-30
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult …
2606.30378 — OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
submitted 2026-06-29
Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of panoramic imagery. The full 360°$\times$180° field of view of panoramas essentially supports complex global multi-step re…
2606.28164 — EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
submitted 2026-06-26
Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreting an echocardiogram requires synthesizing complementary evidence across multiple heart views to identify abnormalities and produce structured clini…
2606.27922 — Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
submitted 2026-06-26
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforce…
2606.27608 — Qwen-Image-2.0-RL Technical Report
submitted 2026-06-25
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instruction-following capability of the Qwen-Image-2.0 diffusion model. To provide reliable reward signals, we…
2606.26930 — PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation
submitted 2026-06-25
Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods often favor superficial aesthetics, such as over-saturated colors, leaving critical flaws like AI artifacts and biological implausibilities unresolve…
2606.01985 — MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
submitted 2026-06-01
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-tu…
Monthly paper sweep — 2026-07
7 candidate(s) from the last 35 days not yet in
INDEX.md/academia.md. Triage each perINTEGRATION_GUIDE.md§1 (scope) and §2 (paradigm + dedicated-file-vs-academia).2606.31825 — Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
submitted 2026-06-30
2606.30378 — OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
submitted 2026-06-29
2606.28164 — EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
submitted 2026-06-26
2606.27922 — Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
submitted 2026-06-26
2606.27608 — Qwen-Image-2.0-RL Technical Report
submitted 2026-06-25
2606.26930 — PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation
submitted 2026-06-25
2606.01985 — MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
submitted 2026-06-01