Skip to content

Monthly paper sweep — 2026-08 #4

Description

@github-actions

Monthly paper sweep — 2026-08

6 candidate(s) from the last 35 days not yet in INDEX.md / academia.md. Triage each per INTEGRATION_GUIDE.md §1 (scope) and §2 (paradigm + dedicated-file-vs-academia).

2607.28509 — RefCaptioner: Multi-Reference Image-Grounded Video Captioning

submitted 2026-07-30

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference g…

2607.28464 — Can Vision-Language Models Reason about AI Edits in Images?

submitted 2026-07-30

Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Langu…

2607.28006 — MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

submitted 2026-07-30

Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention dr…

2607.27782 — RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

submitted 2026-07-30

Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rol…

2607.26432 — FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

submitted 2026-07-29

Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but s…

2607.25962 — LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

submitted 2026-07-28

Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusio…

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions