Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
arXiv:2603.24257v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved inconsistencies using offline multi-view aggregation or multi-stage pipelines that decouple exploration, data association, and caption learning, with limited capacity to reason over previously observed objects. In this paper, we introduce a unified, memory-augmented Vision-Language agent that simultaneously handle — Tommaso Galliena, Stefano Rosa, Tommaso Apicella, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
View PDF
Abstract:Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved inconsistencies using offline multi-view aggregation or multi-stage pipelines that decouple exploration, data association, and caption learning, with limited capacity to reason over previously observed objects. In this paper, we introduce a unified, memory-augmented Vision-Language agent that simultaneously handles data association, object captioning, and exploration policy within a single autoregressive framework. The model processes the current RGB observation, a top-down explored map, and an object-level episodic memory serialized into object-level tokens, ensuring persistent object identity and semantic consistency across extended sequences. To train the model in a self-supervised manner, we collect a dataset in photorealistic 3D environments using a disagreement-based policy and a pseudo-captioning model that enforces consistency across multi-view caption histories. Extensive evaluation on a manually annotated object-level test set, demonstrate improvements of up to +11.86% in standard captioning scores and +7.39% in caption self-similarity over baseline models, while enabling scalable performance through a compact scene representation. Code, model weights, and data are available at this https URL.
Comments: 24 pages, 7 figures, 7 tables (including Supplementary Materials)
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2603.24257 [cs.CV]
(or arXiv:2603.24257v2 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2603.24257
arXiv-issued DOI via DataCite
Submission history
From: Tommaso Galliena [view email] [v1] Wed, 25 Mar 2026 12:52:32 UTC (5,798 KB) [v2] Mon, 30 Mar 2026 09:01:07 UTC (5,798 KB)
Sign in to highlight and annotate this article

Conversation starters
Daily AI Digest
Get the top 5 AI stories delivered to your inbox every morning.
More about
researchpaperarxivExclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ
<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxPZVppQTVFSV9BaGFBMU9hWGlGS3ltTFdJZ3ZEREdzVkxBT2pSR2VaXy1QbEFEWkIyeEJSbmJXMWpoNnJWVXJiUWtRRlh1SC00anVxOERKcHlIOU95bjdQRktMbnVsOFVkSnBnVUVIV19uOFRJOVNDM3BmSXlrd0pqNHAwOWdua0VhX1BfMWxScnlGaEFNVUlRczJMTVdfa1hSNlNLSU11d2hMTXNqWlBVdUJLNmpDajk5a3RoaW1uam1TZW1IYTB5eUd3MHZWNUFPUWIzc2VIUU9lTTVTWVhub3VKVVJFTExqa1k0NWlXMFBYOEdIWXE0RV9ZbFFGazJhZVJLUGEwNmpMWWx1X2xRYXA2LU9HbjNFZ0h4WU1ZWmhGeEdSbGZQXzRIaWR2TlpPNWJ6dTNRN1NyQmRMdVNFX3F6ay0xYWNEUlU2MDJkSGU4ZXBnLTllR0hYbTZjM0lpUjI3NklvaVpDS1hDZjBIQ01DV1dUd3F6UzVta3JtNV9UOF9MV2NrRUxsbVdZemx6aEMwU3FJcVFuQmVjVHlDNWRjU2lBWm1aVzJMd3dfNFpzV2R2VHZHSg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> <font color="#6f6f6f">WSJ</font>
Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ
<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxQamNrT0NoNFYxYTFFMUhzbWtzSVNVSHBQWVVPQ0ZOR3o1bTJTNFVuMlJLVDhXaHNhRDZGZWZjUUgtcU01WkJSM0hYSWVCQmV3ZllKbjVWcDBuX0pheTQ3QThDc254NElERTZqeXZRdE43UkJaYVlRUk03NDlMa1RlU2NUb2N6c25UMlFwQWhITVo3M3dLV1JNblZvcUxuTV9YUE04S2ZCRGFLaWZhQlNXdzdQd0dRLW5va0YzVjVkY0hyV1NyaDRLOVo2UEFxcTk5QWZhTFduUXZLUVdXN1hkbTFGeWx3YVJUMzR6eThmaExhajJOSTRIY0p0Tmt1Yy0zbE9nREJGV1hLM2xPcGpPd1RaRmFTaGZ4M09HanRGSnJSVF9yN0laang2Ui1fWTNuZjZRWEdseVNXelc3Q0d4eW82SG9DcXg2cW1UQ3pEbmYtdnZ5ZFd5ajhFV2ZWX0dSODlVZEVRVEh2LWgzTDNkRDBlNWI4U1p3cE0xQVJUY1ZKMTUyRnBDQ3RlOTg0X1J6M2hSMW9Da3hlTG54dzJ5cTRIMG5KUG45Q1c4TW41UW5XQzZwQmxmRg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> <font color="#6f6f6f">WSJ</font>
Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ
<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNRHZ4V00wTUhOaFprdE9sTTBWLWVtRzNSMjNRLXUyRWNXU3NPbVlrT1ctUk9HaHVTNmxnRzV1MWVaUHpGUk5VRXZNMll4T2ppTkVqQkhDbF9MUzJ2a2Zydm8zUVR0QzJ6aURwcS1tOVJnUUtrR0hjX1dZWXNBQkpMSUs4VGFCanBLR21ON2xrYlRDVnk4a2JjSTNmLWtlMnNmRDBVT182aElEam02UHppenFQQ2Z2QmNwMWNaRXNQMzdnckJYZnpMcEIzMmNjQUhHb3N6Wl95d09LZGVzNzhsUEFQMFJNcjVXNmpSSXlSVUp3WDFmZHVfaXBrcFdPQk4tSHpCc3hSeXFUcVVQeC0wV2gzNk1TN2phdzR1b1VKOWR1aW9vaGxNYWVwY0tJV0ItTFUtclpfZEg1a0N2elA1VHZUbVVYT3JCR093U1gyaWZWUWc5b2gxbG4zNmVLM3BmSUZGY21VM2t3RXJTdVd5dllKS0pCR3QxZ04yUmxTZzF4UWY2bFdZZ0J2SVl3eFluVVI0RGtyMmluNXN2NGYzQTNkWmgweGJ1WWNxVDNTN3BMeWEyTmgyTg?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> <font color="#6f6f6f">WSJ</font>
Knowledge Map
Connected Articles — Knowledge Graph
This article is connected to other articles through shared AI topics and tags.
More in Research Papers

Hidden Helpers: Pittsburgh’s Industrial Past Might Hold the Key to a Cleaner Future
<p> <img loading="lazy" src="https://www.cmu.edu/news/sites/default/files/styles/listings_desktop_1x_/public/2026-03/260305B_WTM_Armbruster038.jpg.webp?itok=8RGXrI_N" width="900" height="508" alt="Researchers examine soil"> </p> Pittsburgh has reinvented itself from a steel powerhouse to a hub for health care and education. But the city’s industrial past left a hidden legacy: toxic compounds like benzene and toluene in the soil. While most life can’t survive such a contamination, some microbes adapted to use the pollutants as food.
XR is XR: Rethinking MR and XR as Neutral Umbrella Terms
arXiv:2603.29939v1 Announce Type: new Abstract: The term XR is currently widely used as an expression encompassing Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). However, there is no clear consensus regarding its origin or meaning. XR is sometimes explained as an abbreviation for Extended Reality, but multiple interpretations exist regarding its etymology and formation process. This paper organizes the historical formation of terminology related to VR, AR, MR, and XR, and reexamines the context in which the term XR emerged and how it has spread. In particular, by presenting a timeline that distinguishes between the coinage of terms and the drivers of their adoption, we suggest that XR, as an umbrella term, functions not as an abbreviation of Extended Reality, but rat
Interview-Informed Generative Agents for Product Discovery: A Validation Study
arXiv:2603.29890v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are grounded in, yet approximate population-level response distributions. These findings highlight both the potential and the limits of LLM simulation in desig
Beyond Legacy OFDM: A Mobility-Adaptive Multi-Gear Framework for 6G
arXiv:2603.29721v1 Announce Type: new Abstract: While Third Generation Partnership Project (3GPP) has confirmed orthogonal frequency division multiplexing (OFDM) as the baseline waveform for sixth-generation (6G), its performance is severely compromised in the high-mobility scenarios envisioned for 6G. Building upon the GEARBOX-PHY vision, we present gear-switching OFDM (GS-OFDM): a unified framework in which the base station (BS) adaptively selects among three gears, ranging from legacy OFDM to delay-Doppler domain processing based on the channel mobility conditions experienced by the user equipments (UEs). We illustrate the benefit of adaptive gear switching for communication throughput and, finally, we conclude with an outlook on research challenges and opportunities.

Discussion
Sign in to join the discussion
No comments yet — be the first to share your thoughts!