AI Research

Latest research papers from arXiv covering machine learning, computer vision, natural language processing, and more.

arXivPDF

CoCo-IR: Contextual Composed Image Retrieval

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables u...

Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding
Aug 5, 2026
arXivPDF

Objects as Audio-Visual Modal Sound Fields

While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling app...

Zisen Shao, Zihao Wei, Derong Jin
Aug 5, 2026
arXivPDF

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and ...

Boxiu Li, Zimo Wen, Yijia Fan
Aug 5, 2026
arXivPDF

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resour...

Indraneil Paul, Falko Helm, Goran Glavaš
Aug 5, 2026
arXivPDF

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (...

Yue Zhang, Yingzhao Jian, Yunqiu Xu
Aug 5, 2026
arXivPDF

The Loss Does Not See the Basis, but Adam Does

Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mech...

Devender Singh
Aug 5, 2026
arXivPDF

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalanc...

Aniri, Jinhe Bi, Peng Liao
Aug 5, 2026
arXivPDF

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propos...

Adel Javanmard, David P. Woodruff, Vahab Mirrokni
Aug 5, 2026
arXivPDF

Chained Recursive Language Models for Multi-Iteration Reasoning

Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This becomes particularly difficult in tasks that require e...

Purbesh Mitra, Sennur Ulukus
Aug 5, 2026

Data from arXiv.org • Updated hourly