AI Research

Latest research papers from arXiv covering machine learning, computer vision, natural language processing, and more.

arXivPDF

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design ...

Hongyang Du, Lan Yan, Christian Flores
Sep 18, 2026
arXivPDF

MintAct: A Unified Visual Agent for Digital Environments

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance...

Mingfei Gao, Rui Tian, Haiming Gang
Sep 18, 2026
arXivPDF

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and composi...

Wenxue Li, Peiyan Guan, Haoyang Jiang
Sep 18, 2026
arXivPDF

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be e...

Bowen Ye, Lei Li, Shicheng Li
Sep 18, 2026
arXivPDF

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent ...

Renkai Ma, Ruyuan Wan, Xuan Lu
Sep 18, 2026
arXivPDF

Benchmarking World Models for Continual Learning on Compositional Tasks

A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent's ability to efficien...

Haoyu Zhou, Joe Watson, Anson Lei
Sep 18, 2026
arXivPDF

Available Guardrails: Certifying Selective Prediction across ML Systems

A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup...

Parivesh Priye, Yufeng Wang, Haibin Ling
Sep 18, 2026
arXivPDF

$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable i...

Yufeng Wang, Parivesh Priye, Meeshawn Marathe
Sep 18, 2026
arXivPDF

PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream ...

Erik Deinzer, Naya Baslan, Luca Paparusso
Sep 18, 2026

Data from arXiv.org • Updated hourly