Monthly paper sweep — 2026-08
6 candidate(s) from the last 35 days not yet in INDEX.md / academia.md. Triage each per INTEGRATION_GUIDE.md §1 (scope) and §2 (paradigm + dedicated-file-vs-academia).
2607.28509 — RefCaptioner: Multi-Reference Image-Grounded Video Captioning
submitted 2026-07-30
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference g…
2607.28464 — Can Vision-Language Models Reason about AI Edits in Images?
submitted 2026-07-30
Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Langu…
2607.28006 — MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
submitted 2026-07-30
Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention dr…
2607.27782 — RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
submitted 2026-07-30
Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rol…
2607.26432 — FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
submitted 2026-07-29
Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but s…
2607.25962 — LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
submitted 2026-07-28
Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusio…
Monthly paper sweep — 2026-08
6 candidate(s) from the last 35 days not yet in
INDEX.md/academia.md. Triage each perINTEGRATION_GUIDE.md§1 (scope) and §2 (paradigm + dedicated-file-vs-academia).2607.28509 — RefCaptioner: Multi-Reference Image-Grounded Video Captioning
submitted 2026-07-30
2607.28464 — Can Vision-Language Models Reason about AI Edits in Images?
submitted 2026-07-30
2607.28006 — MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
submitted 2026-07-30
2607.27782 — RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
submitted 2026-07-30
2607.26432 — FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
submitted 2026-07-29
2607.25962 — LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
submitted 2026-07-28