Enhancing Spatial Cognition in MLLMs with Depth Maps and Point Cloud Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Enhancing Spatial Cognition in MLLMs with Depth Maps and Point Cloud Data Wang Zhenxing, Ruidi Qi, Ziyan Wu, Xuan Dou, Dehu Du This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8634056/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Contemporary multimodal large language models (MLLMs), particularly those integrating visual and textual modalities, have demonstrated remarkable capabilities in both image comprehension and text generation. However, current multimodal learning paradigms predominantly focus on RGB images and traditional text, often exhibiting limitations in spatial cognition. In this study, we enhance the spatial understanding of MLLMs by preprocessing raw images to extract spatial information such as depth maps and point cloud data, subsequently incorporating these into the learning process. Additionally, we employ instruction tuning techniques with comprehensive and detailed textual descriptions to enrich the model’s spatial awareness. Our experiments reveal that models trained with these enhancements surpass baseline models in tasks such as image captioning and visual question answering (VQA). Although traditional metrics such as CIDEr and ROUGE-L show improvement, they fail to capture the model's enhanced spatial reasoning abilities, necessitating complementary evaluation methods like Ref_LongCLIPScore. Empirically, we observed statistically significant improvements: a 1.5% absolute increase in Ref_LongCLIPScore (p < 0.05) and a 1.2% boost in the average accuracy of VQA tasks (p < 0.05). These gains underscore the model’s superior performance in describing spatial relationships within images. Our model weights are publicly available at huggingface.co/fisheries/wcsllava/tree/main. Biological sciences/Neuroscience Biological sciences/Psychology Social science/Psychology multimodal large language models (MLLMs) spatial cognition depth maps point cloud data visual question answering (VQA) Ref_LongCLIPScore Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction In recent studies on multimodal large-scale language models, Vision-Language Models (VLMs) have demonstrated significant capabilities in visual comprehension and text representation across various tasks, such as image captioning and Visual Question Answering (VQA).[1-3] These studies predominantly utilize standard datasets comprising images of everyday human activities alongside concise descriptions, such as COCO, Flickr8K, and VQAV2. However, two main issues persist with these datasets. Firstly, the images are conventional 2D RGB images lacking spatial information .[4-6] Secondly, the image descriptions are often simplistic, failing to capture intricate and contextual details beyond the image itself . Consequently, models must independently infer spatial information, which leads to limitations in spatial reasoning[7] and detailed textual output, with many current models producing relatively brief text of about 10-30 words. Extending the text length without sufficient training data can induce hallucinations[8,9], potentially explaining some of these phenomena. Recently, instruction tuning[10] — a supervised fine-tuning method that trains models to follow human-provided natural language instructions[11] — has emerged as a highly effective technique for enhancing multimodal large-scale language models (MLLMs).[12] By fine-tuning MLLMs with natural language instructions across diverse tasks, this approach improves the models’ visual input comprehension and precise image information representation. Instruction tuning’s strength lies in exposing models to a wide semantic context during training, thus enhancing their ability to process visual data without predefined contextual knowledge. The application of Low-Rank Adaptation (LoRA) fine-tuning has also gained traction in this domain. LoRA facilitates parameter-efficient fine-tuning by integrating low-rank matrices into large pre-trained models, thereby significantly reducing computational overhead. Moreover, monocular depth estimation is a fundamental problem with extensive applications in robotics and autonomous driving, aiming to infer depth information from flat images. Although specific visual processing modules are typically required, most current MLLMs lack such components. Nevertheless, given the robust comprehension abilities of MLLMs, they can extract requisite information, such as medical imagery. However, research on actively providing MLLMs with specialized image data to enrich their understanding of the everyday world remains limited. Simultaneously, long-form image description presents a challenging research area. The prevalent image-text contrastive learning pre-training model, CLIP, exhibits exceptional image-text matching capabilities across a wide array of tasks. However, due to architectural constraints, CLIP’s token length is limited to 77 tokens. LongCLIP represents a significant improvement, extending the token input limit beyond 300 while retaining CLIP’s original strengths. For image description tasks, the CIDEr metric evaluates consistency by computing cosine similarity between generated and reference texts. However, CIDEr faces challenges with evaluating extremely lengthy texts, as it relies on multiple reference texts whose availability and quality may diminish with increasing text length, potentially impacting evaluation accuracy and reliability. In summary, our contributions are as follows: By synergizing instruction tuning with LoRA fine-tuning, integrating depth maps and point cloud information extracted using Depth anything into instructions and input images during training, we successfully fine-tuned a model, WeCanSee-llava, based on llava1.5, with comprehensive experimental validation. In Image Textualization, we trained the model to enhance image information acquisition and foster the ability to generate extended text descriptions. Utilizing LongCLIP as an evaluation tool, we assessed the model’s outputs by calculating similarity scores with generated texts and images, employing the RefLongCLIP_score for intuitive demonstration of this similarity. Related Work 2.1 Zero-shot depth estimation In the realm of spatial cognition, a novel research endeavor known as Visual Spatial Description (VSD) has emerged[13-17], with several proposed solutions. These include the incorporation of external 3D spatial understanding modules[18-21] and the introduction of visual spatial understanding instructions during fine-tuning. In our work, we leverage additional depth estimation images, employing methods such as weakly supervised training through the collection of extensive training images, and recently developed techniques utilizing affine-invariant loss to disregard variances in depth scales and inter-dataset variations, thereby deriving relative depth information. Moreover, robust relative depth estimation models can be effectively adapted for generalized metric depth estimation by fine-tuning on metric depth data. 2.2 Ref_LongCLIP Score Ref_LongCLIP is an evaluation metric based on LongCLIP that measures image-text similarity. Its output scores are normalized to the range [0, 100], where a higher score indicates greater semantic closeness between the generated text and the image. The domain of long-form image description remains a formidable research challenge. When generating extended text, the intrinsic limitations of the CIDEr metric can result in low scores even when the semantic similarity between texts is high. To address this, CLIPScore is introduced as an evaluation metric for long text generation. It operates without requiring reference captions by calculating the cosine similarity between embeddings of images and machine-generated captions within the CLIP model. Despite the widespread use of CLIP for its powerful image-text matching capabilities, its input token length is constrained to 77 due to architectural limitations. LongCLIP significantly enhances this by extending the input token length to over 300 while maintaining the original strengths of CLIP. 2.3 Instruction Fine-tuning Instruction tuning has recently captivated attention for its potential to enhance the understanding of large language models (LLMs) regarding specific task requirements through meticulously crafted input instructions. In the multimodal domain, researchers have applied instruction tuning by utilizing small but high-quality instruction datasets to refine image-to-text multimodal models. This is exemplified by InstructPix2Pix, which enables rapid image editing in response to user instructions within seconds. Method 3.1 Visual Depth Information Extraction To effectively capture the depth information of im- ages, we utilized the Depth-Anything model as our primary image depth extraction tool. This model comprises two components: DepthNet for depth prediction and the PosNet for estimating the pose between adjacent monocular views. However, given that our focus is primarily on static images, we employed only the DepthNet module. The DepthNet encoder consists of four stages, leveraging the Cascaded Dilated Convolu- tion (CDC) and Local-Global Feature Interaction (LGFI) modules to extract rich hierarchical fea- tures. Specifically, let us consider an original im- age denoted as I.Upon Fd processing with Depth- Anything, represented as I d , we obtain results Idand T d , corresponding to the depth map and point cloud data, respectively. These can be formu- lated as follows: (Id, Td) = Fd(I) In this context, Id represents the relative depth data of objects in the image with respect to the camera, while T d provides detailed spatial and morphological information. By processing all utilized data in this manner, we derive the corresponding depth map and point cloud data for the original image. However, as the point cloud data is stored in text format, even a relatively simple representation of an image typically comprises several hundred points. Each individual point requires approximately 45 characters for storage, resulting in a textual length of about 13,000 characters, which often significantly exceeds the input capacity of most large language models. To address this, we employed the technique of Statistical Outlier Removal using Kernel Density Estimation [ 34 ] to filter the data, ultimately retaining approximately 10% of the point cloud data. Despite the retained text still being lengthy, it is within the processable range for large language models. The filtering procedure is as follows: For a given point cloud dataset P, we calculate the average distance d i for each point p i to its k-nearest neighbors. Let N i denote the neighborhood of point p i ,The average distance is calculated as: $$\:{d}_{i}=\frac{1}{k}\sum\:_{j=1}^{k}||{p}_{i}-{p}_{ij}||$$ Here, ||p i -p ij || represents the Euclidean distance. Subsequently, we compute the mean η and standard deviation σ of these average distances across the entire dataset. A threshold T is established to eliminate outliers, defined as: $$\:T={\eta\:}+{\alpha\:}\bullet\:{\sigma\:}$$ where \(\:{\alpha\:}\) is a user-defined multiplier. Points satisfying d i ≤ T are retained. Given the variability among images, the optimal k and \(\:{\alpha\:}\) may vary; however, we generally preserve approximately 10% of the point cloud data. Hence, the refined point cloud data \(\:{T}_{d}^{{\prime\:}}\) is obtained from the original T d . 3.2 Instruction Fine-tuning Our objective is to enable the multi-modal large language model (MLLM) to acquire spatial information that is challenging to discern from RGB images alone. Consequently, the design of our instructions is tailored to facilitate the learning of such spatial data. Our instructions typically prompt the model to focus on spatial comprehension, exemplified by directives such as: "Please refer to the original image and depth map, along with the provided point cloud information..." Furthermore, we conclude these instructions with the emphatic statement, "We will not provide depth maps and point cloud information subsequently," to ensure the model does not become dependent on this type of input. Subsequent experiments have demonstrated the significance of this latter statement in enhancing the model’s learning capabilities. We utilized LLaVA 1.5 as our baseline model. Building upon its instruction fine-tuning guided by GPT-4’s knowledge, we incorporated additional instructions processed through the aforementioned methodology for further refinement, employing low-rank adaptation (LoRA) techniques. In summary, our approach leverages the spatial information extraction model, Depth Anything, to obtain depth maps and point cloud data from images. After sampling the point cloud data, we integrate these components with the original image to create a cohesive visual-language instruction set. This set is then used for fine-tuning LLaVA 1.5 through visual instruction tuning, employing low-rank adaptation (LoRA) techniques. The resultant model, which we have named WCSLLaVA, serves as a comprehensive visual-language assistant, embodying the concept of "We can see everything." three distinct datasets: original images, depth images, and sampled point cloud data. During the instruc- tion fine-tuning phase, the directives we employed are illustrated on the right, where the red segments indicate placeholders for image and point cloud data. At the time of actual input, the corresponding information is directly conveyed to the multi-modal large language model (MLLM). Upon completion of this setup, we standardly appended the directive "We will not provide depth maps and point cloud information in the future" to reflect the common practice in everyday interactions of not providing precise depth and point cloud data. Model B4 METEOR ROUGE_L Ref_longCLIP Semantic Similarity WCSLLaVA 1.4 10.7 18.9 57.88 0.9997 LLaVA1.5 1.1 9.5 16.7 57.02 0.9996 BLIP2 0.3 5.9 17.3 -11.27 0.9993 Note : The Ref_LongCLIP score ranges from 0 to 100, with higher values indicating greater semantic similarity between the generated text and the image. Table 1: Performance Metrics for Different Models on COCO-long and SA1B Datasets Experiment We employed Depth-Anything-V2-Large as our spatial information extraction model, with LLaVA-1.5-7B-hf serving as the baseline model. Although the baseline model's maximum input length is sufficient to accommodate our extended text-based point cloud data, we applied the aforementioned sampling procedure to expedite the training process. 4.1 Dataset and Training Settings Our dataset consisted of the COCO-long subset, which was synthetically generated with GPT-4V, and the refined SA1B dataset, enhanced through QwenVL. This compilation, drawn from the Image-Textualization Dataset, comprised approximately 73,000 image-long text pairs, with an average text length of 328 tokens. Of these, 90% were allocated for training and 10% for evaluation. The training was conducted over approximately seven days on three 4090Ti 24GB nodes. It spanned 100 epochs with a batch size of 32. The initial learning rate was set at 1e-5 and was dynamically adjusted, tapering down to approximately 1e-8 during the final training phases. 4.2 Evaluation indicators To comprehensively assess caption quality across varying text lengths, we employ a combination of metrics, each with distinct characteristics. Traditional metrics—BLEU-4, METEOR, and ROUGE-L—primarily measure n-gram overlap and are most reliable for evaluating short captions (typically ≤ 30 words); they may underestimate longer, lexically diverse descriptions. CIDEr, though designed for image captioning, is highly sensitive to text length and lexical alignment, often penalizing semantically correct but phrasally divergent long descriptions. For robust long-form evaluation, we prioritize Ref_LongCLIP, a LongCLIP-based metric that computes image-text similarity and outputs scores normalized to the range [0, 100]. Additionally, we use Semantic Similarity to measure embedding-level correspondence between texts, complementing the lexical-based metrics. In our analysis, Ref_LongCLIP and Semantic Similarity serve as primary metrics for long-text assessment, while traditional metrics provide supplementary reference. 4.3 Experimental Result To assess whether the trained model successfully learned spatial information from images, we initially employed conventional image description evaluation metrics such as BLEU-4 (B4), METEOR, ROUGE-L, SPICE, and CIDEr. However, we observed inconsistencies, particularly with CIDEr, which occasionally scored zero despite semantic similarity in some human-acceptable text descriptions (e.g., the image used initially received a CIDEr score of zero). Given CIDEr's sensitivity to long texts, we decided to exclude certain metrics. Instead, we found that the dot product similarity between two long texts and LongCLIP's computation of long text-image similarity were more suitable for our evaluation objectives. Consequently, we adopted the LongBERT-based dot product similarity scoring and Ref_LongCLIP as primary evaluation metrics, with traditional metrics serving as supplementary. Understanding CIDEr Failures in Long-Text Evaluation.The observed zero CIDEr scores for semantically acceptable descriptions stem from its design for short captions (typically < 20 words). CIDEr relies on n-gram overlap between generated and reference texts, which diminishes as text length increases due to lexical variation and structural divergence. For instance, in Fig. 6 , the two descriptions differ in phrasing and detail yet convey equivalent semantics. To further illustrate, we provide additional examples in Appendix A, where CIDEr similarly fails to reflect human judgment. This limitation underscores the need for metrics better suited to long-form generation, such as Ref_LongCLIP. Table 2 Zero-shot Evaluation Accuracy on VQAV2 and VIZWIZ Datasets Dataset WCSLLaVA LLaVA1.5 BLIP2 VQAV2 40.3 39.63 36.9 VIZWIZ 32.65 24.4 14.2 To qualitatively illustrate the improvements of WCSLLaVA, we compare generated captions for the image in Fig. 1 across baseline models. BLIP-2 produces a concise but limited description: “Two children using a computer.” LLaVA-1.5 offers more detail: “Two children sit in front of an old computer in a room with books and other devices.” In contrast, our WCSLLaVA generates a longer, more spatially and contextually rich narrative: “In a softly decorated room, two children sit before an old-style computer... The scene evokes a nostalgic, timeless atmosphere.”[ 26 – 28 ] This comparison demonstrates that our model, enhanced with spatial information, produces more descriptive and spatially-aware captions. The BLIP2 model, which generates relatively short texts (approximately 20 characters), is penalized on Ref_LongCLIP due to the preference for longer reference texts, highlighting its limitations in producing ultra-long image descriptions. In contrast, the fine-tuned WCSLLaVA demonstrates superior performance in generating ultra-long text descriptions from images without sacrificing prior knowledge. To further validate the model's acquisition of spatial knowledge, we conducted zero-shot evaluations on the VQAV2 and VIZWIZ datasets[ 29 – 33 ], which revealed notable improvements. Despite using a straightforward accuracy metric without semantic scoring, our model outperforms the baseline under consistent conditions. Table 3 Ablation Study Results: The Effect of Spatial Data on Performance Configuration B4 METEOR ROUGE_L Ref_longCLIP Semantic Similarity Image 1.1 9.5 16.7 57.77 0.9996 Image and Depth 1.4 10.8 18.8 57.80 0.9997 Image and Pointcloud 1.4 9.8 17.9 57.87 0.9996 WCSLLaVA 1.4 10.7 18.9 57.88 0.9997 4.4 Ablation Study Our decision to incorporate both depth maps and point cloud data as additional inputs is substantiated through ablation studies. We evaluated three configurations—image-only, image-depth, and image-pointcloud—each trained with the same parameters and dataset. The results affirm that combining depth maps with point cloud data yields superior outcomes. Conclusion This study effectively demonstrates the potential to enhance the spatial cognition capabilities of multimodal large language models (MLLMs) through the integration of depth maps and point cloud data. By crafting sophisticated instructions and employing Low-Rank Adaptation (LoRA) fine-tuning techniques, we developed a model named WeCanSee-LLaVA. This model exhibits superior performance in tasks such as image captioning and visual question answering, particularly excelling in handling intricate spatial information. Our experimental results reveal that the integration of spatial data not only addresses the limitations inherent in traditional RGB images and succinct text descriptions but also significantly enhances the quality and detail of image-to-text generation, especially in producing extended textual narratives. This research offers novel insights and methodologies for advancing MLLMs’ comprehension and articulation of complex spatial relationships in multimodal tasks, highlighting a promising avenue for broad applications.[ 34 – 38 ] Limitations Although our approach has demonstrated the capability to enable MLLMs to learn spatial information from planar images, this learning process does not entirely eliminate the issue of hallucinations but mitigates the issue to a certain extent rather than eliminating it entirely. Furthermore, when generating exceptionally long texts, conventional metrics are inadequate for effectively quantifying and evaluating the generated content. The metrics currently employed are limited to basic testing, underscoring the necessity for the development of evaluation methods tailored to these specific requirements. Addressing these challenges mandates further investigation and research. Declarations 7 Funding The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Author Contribution Z.Wang and R.Q. conceived the experiments, conducted the model training and analysis, and wrote the main manuscript text. Z.Wu contributed to the data preprocessing and collection. D. Du contributed to the methodology discussion and provided technical support throughout the research. X.D. prepared Figures 1–6. All authors reviewed the manuscript. Data Availability The datasets generated and/or analysed during the current study are not publicly available due to the large volume and proprietary format of the raw depth maps and point cloud data, but are available from the corresponding author on reasonable request. References Zhu, Y. et al. ChatNav: Leveraging LLM to Zero-Shot Semantic Reasoning in Object Navigation . IEEE Trans. Circuits Syst. Video Technol. (2025). Zhao, R. et al. Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning . IEEE Rob. Autom. Lett. (2025). Wu, D. et al. Empowering natural human-robot collaboration through multimodal language models and spatial intelligence: Pathways and perspectives ( Robotics and Computer-Integrated Manufacturing, 2025). Xiong, H. et al. 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding . IEEE Trans. Multimedia (2025). Geng Liang ༆ Jianqin Yin. ViewInfer3D: 3D Visual Grounding Based on Embodied Viewpoint Inference . IEEE Rob. Autom. Lett. (2024). Zhang, W. et al. EarthGPT-X: A Spatial MLLM for Multilevel Multisource Remote Sensing Imagery Understanding With Visual Prompting . IEEE Trans. Geosci. Remote Sens. (2025). Li, C. et al. VIDHALLUC: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding. CVPR (2025). Huang, X. Geo-hallucination in urban analytics: What it is and why it matters. Environment and Planning B-Urban Analytics and City Science (2025). Chu & Shiyong༆ Yuwei Chen.. Weaponizing cognitive bias in autonomous systems: a framework for black-box inference attacks . Front. Artif. Intell. (2025). Hu, T. et al. 3DBench: A scalable benchmark for object and scene-level instruction-tuning of 3D large language models. Neural Networks (2025). Wang, Z. et al. Instruction-guided path planning with 3D semantic maps for vision-language navigation. Neurocomputing (2025). Faria, F. T. et al. Towards Robust Chain-of-Thought Prompting with Self-Consistency for Remote Sensing VQA: An Empirical Study Across Large Multimodal Models. Mathematics (2025). Yang, A. et al. Evaluating and enhancing spatial cognition abilities of large language models . Int. J. Geogr. Inf. Sci. (2025). Liu, F. et al. Visual Spat. Reasoning Trans. Association Comput. Linguistics (2023). Li, C. M. & Zhang A dual-dimension collaborative enhancement framework to boost language model spatial semantic understanding ( Annals of the New York Academy of Sciences, 2025). Ingale, A. A. Spatial Sequence Reasoning in Large Language Models… Dissertation/Thesis (2025). Thai, A. Mutual Exclusivity Bias and Spatial Reasoning in Vision-language Models. Dissertation/Thesis (2025). Zhang, Z. et al. Dual-level dynamic heterogeneous graph network for video question answering . Neural Networks (2026). Yusuf, A. et al. Graph-enhanced visual representations and question-guided dual attention for visual question answering . Neurocomputing (2025). Lai, C. et al. Object-Centric Cross-Modal Knowledge Reasoning for Future Event Prediction in Videos . IEEE Trans. Circuits Syst. Video Technol. (2024). Honerkamp, D. et al. Language-Grounded Dynamic Scene Graphs for Interactive Object Search With Mobile Manipulation . IEEE Rob. Autom. Lett. (2024). Huang, Z. et al. AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding . IEEE Trans. Multimedia (2025). Padhan, S. Probabilistic Metric-Semantic Grounding for Embodied AI Composing Bayesian Spatial Kernels, VLMs, and 3D Scene Graphs . Dissertation/Thesis (2025). Geng, L. et al. Pseudo-EV: Enhancing 3D Visual Grounding With Pseudo Embodied Viewpoint . IEEE Trans. Circuits Syst. Video Technol. (2025). Xue, Y. et al. Spatial Knowledge Graph-Guided Multimodal Synthesis . IEEE Trans. Audio Speech Lang. Process. (2025). Hossain, M. et al. CM-SC: Cross-modal spatial-channel attention network for image captioning . Displays (2025). Park, K. et al. Learning Compositionality from Multifaceted Synthetic Data for Language-based Object Detection . Int. J. Comput. Vision (2025). Papadopoulos, G. et al. SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks . Computers (2025). Sterz, H. et al. DARE: Diverse Visual Question Answering with Robustness Evaluation ( Transactions of the Association for Computational Linguistics, 2025). Chen, C. et al. Towards bias-aware visual question answering: Rectifying and mitigating comprehension biases (Expert Systems with Applications, 2025). Gao, Y. et al. Adaptive Conditional Reasoning for Remote Sensing Visual Question Answering. Remote Sensing (2025). Pei, B. et al. Guiding Audio-Visual Question Answering with Collective Question Reasoning . Int. J. Comput. Vision (2025). Kim, W. et al. An Image Grid Can Be Worth a Video: Zero-Shot Video Question Answering Using a VLM . IEEE Access. (2024). Zargarzadeh, S. et al. From Decision to Action in Surgical Autonomy: Multi-Modal Large Language Models for Robot-Assisted Blood Suction . IEEE Rob. Autom. Lett. (2025). Mehta, V. et al. Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy . Sensors (2025). Hu, S. et al. TARAD: Task-Aware Robot Affordance-Centric Diffusion Policy Learned From LLM-Generated Demonstrations . IEEE Rob. Autom. Lett. (2025). Nan, Y. et al. Beyond the Hype: A Dispassionate Look at Vision-Language Models in Medical Scenario . IEEE Trans. Neural Networks Learn. Syst. (2025). Lin, B. et al. NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning . IEEE Trans. Pattern Anal. Mach. Intell. (2025). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 06 Apr, 2026 Reviewers agreed at journal 15 Mar, 2026 Reviewers invited by journal 09 Mar, 2026 Editor assigned by journal 02 Feb, 2026 Editor invited by journal 02 Feb, 2026 Submission checks completed at journal 30 Jan, 2026 First submitted to journal 30 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8634056","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":604809301,"identity":"3cdf9819-dcac-44ea-9ed7-7eef0dbbfbd9","order_by":0,"name":"Wang Zhenxing","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuUlEQVRIiWNgGAWjYDCCA2Bkw0OyljQStQDBYRJ08B1vf3jw647zMvzszQcYPu6pJaxF8swZg8OyZ27zSPYcS2Cc8ew4YS0GN3IYDku23eYBMgyYeQ4cI0LL/ecPgFrO8djff/+BSC03GAwOfmw7wGMgwcMA1FJDjF9yDA4ztiXzSJxJMzg448ABwlr4jh9//PFnm509f/vhhw8+HKgjrAUEmGHxeIDoCGL8gWATacsoGAWjYBSMKAAAFoVCm01z/IgAAAAASUVORK5CYII=","orcid":"","institution":"上海第二工业大学","correspondingAuthor":true,"prefix":"","firstName":"Wang","middleName":"","lastName":"Zhenxing","suffix":""},{"id":604809302,"identity":"5927a878-0a07-4399-8844-4ff428bd9174","order_by":1,"name":"Ruidi Qi","email":"","orcid":"","institution":"上海第二工业大学","correspondingAuthor":false,"prefix":"","firstName":"Ruidi","middleName":"","lastName":"Qi","suffix":""},{"id":604809303,"identity":"c425fa07-4303-4e5b-a33b-ee998659a841","order_by":2,"name":"Ziyan Wu","email":"","orcid":"","institution":"上海第二工业大学","correspondingAuthor":false,"prefix":"","firstName":"Ziyan","middleName":"","lastName":"Wu","suffix":""},{"id":604809304,"identity":"90e2eafc-c5c3-46a7-bf2e-cc76686dda3e","order_by":3,"name":"Xuan Dou","email":"","orcid":"","institution":"上海第二工业大学","correspondingAuthor":false,"prefix":"","firstName":"Xuan","middleName":"","lastName":"Dou","suffix":""},{"id":604809305,"identity":"9cf848ee-f8c1-4be2-ab5a-54b89553c5a7","order_by":4,"name":"Dehu Du","email":"","orcid":"","institution":"上海第二工业大学","correspondingAuthor":false,"prefix":"","firstName":"Dehu","middleName":"","lastName":"Du","suffix":""}],"badges":[],"createdAt":"2026-01-19 01:38:04","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8634056/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8634056/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104593585,"identity":"7a3a5ee5-ef55-4fc6-af19-ecf213053d99","added_by":"auto","created_at":"2026-03-13 17:42:10","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":180112,"visible":true,"origin":"","legend":"\u003cp\u003eLegend not included with this version\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/637cfd1deb1e34f6a8d3b5fa.png"},{"id":104781591,"identity":"332e23c5-9d14-49d9-b64d-1682bff026b5","added_by":"auto","created_at":"2026-03-17 07:55:58","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":123124,"visible":true,"origin":"","legend":"\u003cp\u003eillustrates our foundational train- ing process. For a given \u0026nbsp;standard RGB image, we first employ the DepthAnything tool to extract its spa- tial information, generating a corresponding grayscale depth map where lighter shades indicate proximity to the camera, and darker shades suggest greater distance. Concurrently, we extract point cloud data[22,23]. Due to the extensive character length of point cloud data when stored as pure text—exceeding typical input constraints for large language models (LLMs)—we retain approx- imately 10% of this information. The original image and its depth map are fed into the multimodal language model’s (MLLM) visual embedding module[24-25], while the textual point cloud data and an extended image descrip- tion (exceeding 300 characters) are inputted into the text embedding module. We thenperform instruction tuning on the llava1.5-7b model using the instructions depicted.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/111f413e3f6e333f7b9a91d1.png"},{"id":104593587,"identity":"bac727d0-ee1f-41ce-9b95-1d15bc5ae66d","added_by":"auto","created_at":"2026-03-13 17:42:10","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":220668,"visible":true,"origin":"","legend":"\u003cp\u003eIt is evident that the inclusion of supplementary instructions in our experimental design enables the model to appropriately respond to conventional inputs (i.e., those lacking additional information), as demonstrated in the second example. In contrast, when such instructions are omitted, model performance significantly declines, as observed in the first sentence. This comparison highlights the critical role of instruction design in modulating the model's ability to generalize beyond training conditions.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/2656339a47bd8d46c4a2a134.png"},{"id":104782012,"identity":"193023a9-fe6b-49b6-9f19-6601be085a9a","added_by":"auto","created_at":"2026-03-17 07:56:42","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":193081,"visible":true,"origin":"","legend":"\u003cp\u003eWe detail the process by which we obtained\u003c/p\u003e\n\u003cp\u003ethree distinct datasets: original images, depth images, and sampled point cloud data. During the instruc- tion fine-tuning phase, the directives we employed are illustrated on the right, where the red segments indicate placeholders for image and point cloud data. At the time of actual input, the corresponding information is directly conveyed to the multi-modal large language model (MLLM). Upon completion of this setup, we standardly appended the directive \"We will not provide depth maps and point cloud information in the future\" to reflect the common practice in everyday interactions of not providing precise depth and point cloud data.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/9775a5991bb8a51d8a31f166.png"},{"id":104593590,"identity":"6c7395d6-6197-496d-80b7-853531c9cf3c","added_by":"auto","created_at":"2026-03-13 17:42:11","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":170744,"visible":true,"origin":"","legend":"\u003cp\u003eIn our processing methodology, both the orig- inal image Io and depth image Id are fed into LLaVA’s visual encoder, while the synthesized language instruction T is input into the text encoder. Following this, the combined inputs are subjected to LoRA fine-tuning on LLaVA, rather than a full-scale fine-tuning, to optimize model performance.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/7eca94f51e5a5ccf1fec03de.png"},{"id":104593588,"identity":"736db8cc-bf90-4235-a299-fe360e093358","added_by":"auto","created_at":"2026-03-13 17:42:10","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":308962,"visible":true,"origin":"","legend":"\u003cp\u003eWe can see that while both humans and GPT4v consider the effect to be satisfactory, the CIDEr indica-tor scores 0.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/1f9f398158afe50da13f0e64.png"},{"id":104784760,"identity":"70fd35f7-7564-4eb2-b474-10875336b788","added_by":"auto","created_at":"2026-03-17 08:08:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1553897,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8634056/v1/edd64460-227e-417c-9c2e-89f381f47620.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Enhancing Spatial Cognition in MLLMs with Depth Maps and Point Cloud Data","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn recent studies on multimodal large-scale language models, Vision-Language Models (VLMs) have demonstrated significant capabilities in visual comprehension and text representation across various tasks, such as image captioning and Visual Question Answering (VQA).[1-3] These studies predominantly utilize standard datasets comprising images of everyday human activities alongside concise descriptions, such as COCO, Flickr8K, and VQAV2. However, two main issues persist with these datasets. Firstly, the images are\u0026nbsp;conventional 2D RGB images\u0026nbsp;lacking spatial information .[4-6]\u0026nbsp;Secondly, the image descriptions are often simplistic, failing to capture intricate and contextual details beyond the image itself . Consequently, models must independently infer spatial information, which leads to limitations in spatial reasoning[7]\u0026nbsp;and detailed textual output, with many current models producing relatively brief text of about 10-30 words. Extending the text length without sufficient training data can induce hallucinations[8,9], potentially explaining some of these phenomena.\u003c/p\u003e\n\u003cp\u003eRecently, instruction tuning[10] \u0026mdash; a supervised fine-tuning method that trains models to follow human-provided natural language instructions[11] \u0026mdash; has emerged as a highly effective technique for enhancing multimodal large-scale language models (MLLMs).[12] By fine-tuning MLLMs with natural language instructions across diverse tasks, this approach improves the models\u0026rsquo; visual input comprehension and precise image information representation. Instruction tuning\u0026rsquo;s strength lies in exposing models to a wide semantic context during training, thus enhancing their ability to process visual data without predefined contextual knowledge. The application of Low-Rank Adaptation (LoRA) \u0026nbsp;fine-tuning has also gained traction in this domain. LoRA facilitates parameter-efficient fine-tuning by integrating low-rank matrices into large pre-trained models, thereby significantly reducing computational overhead.\u003c/p\u003e\n\u003cp\u003eMoreover, monocular depth estimation is a fundamental problem with extensive applications in robotics and autonomous driving, aiming to infer depth information from flat images. Although specific visual processing modules are typically required, most current MLLMs lack such components. Nevertheless, given the robust comprehension abilities of MLLMs,\u0026nbsp;they can extract requisite information, such as medical imagery. However, research on actively providing MLLMs with specialized image data to enrich their understanding of the everyday world remains limited.\u003c/p\u003e\n\u003cp\u003eSimultaneously, long-form image description presents a challenging research area. The prevalent image-text contrastive learning pre-training model, CLIP, exhibits exceptional image-text matching capabilities across a wide array of tasks. However, due to architectural constraints, CLIP\u0026rsquo;s token length is limited to 77 tokens. LongCLIP represents a significant improvement, extending the token input limit beyond 300 while retaining CLIP\u0026rsquo;s original strengths. For image description tasks, the CIDEr metric evaluates consistency by computing cosine similarity between generated and reference texts. However, CIDEr faces challenges with evaluating extremely lengthy texts, as it relies on multiple reference texts whose availability and quality may diminish with increasing text length, potentially impacting evaluation accuracy and reliability.\u003c/p\u003e\n\u003cp\u003eIn summary, our contributions are as follows: By synergizing instruction tuning with LoRA fine-tuning, integrating depth maps and point cloud information extracted using Depth anything into instructions and input images during training, we successfully fine-tuned a model, WeCanSee-llava, based on llava1.5, with comprehensive experimental validation.\u003c/p\u003e\n\u003cp\u003eIn Image Textualization, we trained the model to enhance image information acquisition and foster the ability to generate extended text descriptions. Utilizing LongCLIP as an evaluation tool, we assessed the model\u0026rsquo;s outputs by calculating similarity scores with generated texts and images, employing the RefLongCLIP_score for intuitive demonstration of this similarity.\u003c/p\u003e"},{"header":"Related Work ","content":"\u003cp\u003e\u003cstrong\u003e2.1\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eZero-shot depth estimation \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the realm of spatial cognition, a novel research endeavor known as Visual Spatial Description (VSD) has emerged[13-17], with several proposed solutions. These include the incorporation of external 3D spatial understanding modules[18-21] and the introduction of visual spatial understanding instructions during fine-tuning. In our work, we leverage additional depth estimation images, employing methods such as weakly supervised training through the collection of extensive training images, and recently developed techniques utilizing affine-invariant loss to disregard variances in depth scales and inter-dataset variations, thereby deriving relative depth information. Moreover, robust relative depth estimation models can be effectively adapted for generalized metric depth estimation by fine-tuning on metric depth data. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.2\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eRef_LongCLIP Score\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRef_LongCLIP is an evaluation metric based on LongCLIP that measures image-text similarity. Its output scores are normalized to the range [0, 100], where a higher score indicates greater semantic closeness between the generated text and the image.\u003c/p\u003e\n\u003cp\u003eThe domain of long-form image description remains a formidable research challenge. When generating extended text, the intrinsic limitations of the CIDEr metric can result in low scores even when the semantic similarity between texts is high. To address this, CLIPScore is introduced as an evaluation metric for long text generation. It operates without requiring reference captions by calculating the cosine similarity between embeddings of images and machine-generated captions within the CLIP model. Despite the widespread use of CLIP for its powerful image-text matching capabilities, its input token length is constrained to 77 due to architectural limitations. LongCLIP significantly enhances this by extending the input token length to over 300 while maintaining the original strengths of CLIP.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.3\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eInstruction Fine-tuning \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eInstruction tuning has recently captivated attention for its potential to enhance the understanding of large language models (LLMs) regarding specific task requirements through meticulously crafted input instructions. In the multimodal domain, researchers have applied instruction tuning by utilizing small but high-quality instruction datasets to refine image-to-text multimodal models. This is exemplified by InstructPix2Pix, which enables rapid image editing in response to user instructions within seconds.\u003c/p\u003e"},{"header":"Method","content":"\u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Visual Depth Information Extraction\u003c/h2\u003e \u003cp\u003eTo effectively capture the depth information of im- ages, we utilized the Depth-Anything model as our primary image depth extraction tool. This model comprises two components: DepthNet for depth prediction and the PosNet for estimating the pose between adjacent monocular views. However, given that our focus is primarily on static images, we employed only the DepthNet module. The DepthNet encoder consists of four stages, leveraging the Cascaded Dilated Convolu- tion (CDC) and Local-Global Feature Interaction\u003c/p\u003e \u003cp\u003e(LGFI) modules to extract rich hierarchical fea- tures. Specifically, let us consider an original im- age denoted as I.Upon Fd processing with Depth- Anything, represented as I\u003csub\u003ed\u003c/sub\u003e, we obtain results Idand T\u003csub\u003ed\u003c/sub\u003e, corresponding to the depth map and point cloud data, respectively. These can be formu- lated as follows:\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e(Id, Td)\u0026thinsp;=\u0026thinsp;Fd(I)\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn this context,\u0026nbsp;Id\u0026nbsp;represents the relative depth data of objects in the image with respect to the camera, while\u0026nbsp;T\u003csub\u003ed\u003c/sub\u003e provides detailed spatial and morphological information. By processing all utilized data in this manner, we derive the corresponding depth map and point cloud data for the original image. However, as the point cloud data is stored in text format, even a relatively simple representation of an image typically comprises several hundred points. Each individual point requires approximately 45 characters for storage, resulting in a textual length of about 13,000 characters, which often significantly exceeds the input capacity of most large language models. To address this, we employed the technique of Statistical Outlier Removal using Kernel Density Estimation [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] to filter the data, ultimately retaining approximately 10% of the point cloud data. Despite the retained text still being lengthy, it is within the processable range for large language models.\u003c/p\u003e \u003cp\u003eThe filtering procedure is as follows: For a given point cloud dataset P, we calculate the average distance d\u003csub\u003ei\u003c/sub\u003e for each point p\u003csub\u003ei\u003c/sub\u003e to its k-nearest neighbors. Let N\u003csub\u003ei\u003c/sub\u003e denote the neighborhood of point p\u003csub\u003ei\u003c/sub\u003e,The average distance is calculated as:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:{d}_{i}=\\frac{1}{k}\\sum\\:_{j=1}^{k}||{p}_{i}-{p}_{ij}||$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eHere, ||p\u003csub\u003ei\u003c/sub\u003e-p\u003csub\u003eij\u003c/sub\u003e|| represents the Euclidean distance. Subsequently, we compute the mean η and standard deviation σ of these average distances across the entire dataset. A threshold T is established to eliminate outliers, defined as:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:T={\\eta\\:}+{\\alpha\\:}\\bullet\\:{\\sigma\\:}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\alpha\\:}\\)\u003c/span\u003e\u003c/span\u003e is a user-defined multiplier. Points satisfying d\u003csub\u003ei\u003c/sub\u003e \u0026le; T are retained. Given the variability among images, the optimal k and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\alpha\\:}\\)\u003c/span\u003e\u003c/span\u003e may vary; however, we generally preserve approximately 10% of the point cloud data. Hence, the refined point cloud data \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{T}_{d}^{{\\prime\\:}}\\)\u003c/span\u003e\u003c/span\u003e is obtained from the original T\u003csub\u003ed\u003c/sub\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Instruction Fine-tuning\u003c/h2\u003e \u003cp\u003eOur objective is to enable the multi-modal large language model (MLLM) to acquire spatial information that is challenging to discern from RGB images alone. Consequently, the design of our instructions is tailored to facilitate the learning of such spatial data. Our instructions typically prompt the model to focus on spatial comprehension, exemplified by directives such as: \"Please refer to the original image and depth map, along with the provided point cloud information...\" Furthermore, we conclude these instructions with the emphatic statement, \"We will not provide depth maps and point cloud information subsequently,\" to ensure the model does not become dependent on this type of input. Subsequent experiments have demonstrated the significance of this latter statement in enhancing the model\u0026rsquo;s learning capabilities.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe utilized LLaVA 1.5 as our baseline model. Building upon its instruction fine-tuning guided by GPT-4\u0026rsquo;s knowledge, we incorporated additional instructions processed through the aforementioned methodology for further refinement, employing low-rank adaptation (LoRA) techniques.\u003c/p\u003e \u003cp\u003eIn summary, our approach leverages the spatial information extraction model, Depth Anything, to obtain depth maps and point cloud data from images. After sampling the point cloud data, we integrate these components with the original image to create a cohesive visual-language instruction set. This set is then used for fine-tuning LLaVA 1.5 through visual instruction tuning, employing low-rank adaptation (LoRA) techniques. The resultant model, which we have named WCSLLaVA, serves as a comprehensive visual-language assistant, embodying the concept of \"We can see everything.\"\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ethree distinct datasets: original images, depth images, and sampled point cloud data. During the instruc- tion fine-tuning phase, the directives we employed are illustrated on the right, where the red segments indicate placeholders for image and point cloud data. At the time of actual input, the corresponding information is directly conveyed to the multi-modal large language model (MLLM). Upon completion of this setup, we standardly appended the directive \"We will not provide depth maps and point cloud information in the future\" to reflect the common practice in everyday interactions of not providing precise depth and point cloud data.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMETEOR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eROUGE_L\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRef_longCLIP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSemantic Similarity\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWCSLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e18.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9997\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLLaVA1.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e16.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9996\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBLIP2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e17.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e-11.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9993\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"6\"\u003e\u003cb\u003eNote\u003c/b\u003e: \u003cem\u003eThe Ref_LongCLIP score ranges from 0 to 100, with higher values indicating greater semantic similarity between the generated text and the image.\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;1: Performance Metrics for Different Models on COCO-long and SA1B Datasets\u003c/p\u003e \u003c/div\u003e"},{"header":"Experiment","content":"\u003cp\u003eWe employed Depth-Anything-V2-Large as our spatial information extraction model, with LLaVA-1.5-7B-hf serving as the baseline model. Although the baseline model's maximum input length is sufficient to accommodate our extended text-based point cloud data, we applied the aforementioned sampling procedure to expedite the training process.\u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Dataset and Training Settings\u003c/h2\u003e \u003cp\u003eOur dataset consisted of the COCO-long subset, which was synthetically generated with GPT-4V, and the refined SA1B dataset, enhanced through QwenVL. This compilation, drawn from the Image-Textualization Dataset, comprised approximately 73,000 image-long text pairs, with an average text length of 328 tokens. Of these, 90% were allocated for training and 10% for evaluation.\u003c/p\u003e \u003cp\u003eThe training was conducted over approximately seven days on three 4090Ti 24GB nodes. It spanned 100 epochs with a batch size of 32. The initial learning rate was set at 1e-5 and was dynamically adjusted, tapering down to approximately 1e-8 during the final training phases.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Evaluation indicators\u003c/h2\u003e \u003cp\u003eTo comprehensively assess caption quality across varying text lengths, we employ a combination of metrics, each with distinct characteristics. Traditional metrics\u0026mdash;BLEU-4, METEOR, and ROUGE-L\u0026mdash;primarily measure n-gram overlap and are most reliable for evaluating short captions (typically\u0026thinsp;\u0026le;\u0026thinsp;30 words); they may underestimate longer, lexically diverse descriptions. CIDEr, though designed for image captioning, is highly sensitive to text length and lexical alignment, often penalizing semantically correct but phrasally divergent long descriptions. For robust long-form evaluation, we prioritize Ref_LongCLIP, a LongCLIP-based metric that computes image-text similarity and outputs scores normalized to the range [0, 100]. Additionally, we use Semantic Similarity to measure embedding-level correspondence between texts, complementing the lexical-based metrics. In our analysis, Ref_LongCLIP and Semantic Similarity serve as primary metrics for long-text assessment, while traditional metrics provide supplementary reference.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Experimental Result\u003c/h2\u003e \u003cp\u003eTo assess whether the trained model successfully learned spatial information from images, we initially employed conventional image description evaluation metrics such as BLEU-4 (B4), METEOR, ROUGE-L, SPICE, and CIDEr. However, we observed inconsistencies, particularly with CIDEr, which occasionally scored zero despite semantic similarity in some human-acceptable text descriptions (e.g., the image used initially received a CIDEr score of zero). Given CIDEr's sensitivity to long texts, we decided to exclude certain metrics. Instead, we found that the dot product similarity between two long texts and LongCLIP's computation of long text-image similarity were more suitable for our evaluation objectives. Consequently, we adopted the LongBERT-based dot product similarity scoring and Ref_LongCLIP as primary evaluation metrics, with traditional metrics serving as supplementary.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eUnderstanding CIDEr Failures in Long-Text Evaluation.The observed zero CIDEr scores for semantically acceptable descriptions stem from its design for short captions (typically\u0026thinsp;\u0026lt;\u0026thinsp;20 words). CIDEr relies on n-gram overlap between generated and reference texts, which diminishes as text length increases due to lexical variation and structural divergence. For instance, in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, the two descriptions differ in phrasing and detail yet convey equivalent semantics. To further illustrate, we provide additional examples in Appendix A, where CIDEr similarly fails to reflect human judgment. This limitation underscores the need for metrics better suited to long-form generation, such as Ref_LongCLIP.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eZero-shot Evaluation Accuracy on VQAV2 and VIZWIZ Datasets\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWCSLLaVA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaVA1.5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBLIP2\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVQAV2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e40.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e39.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e36.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVIZWIZ\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e32.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e24.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e14.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTo qualitatively illustrate the improvements of WCSLLaVA, we compare generated captions for the image in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e across baseline models. BLIP-2 produces a concise but limited description: \u0026ldquo;Two children using a computer.\u0026rdquo; LLaVA-1.5 offers more detail: \u0026ldquo;Two children sit in front of an old computer in a room with books and other devices.\u0026rdquo; In contrast, our WCSLLaVA generates a longer, more spatially and contextually rich narrative: \u0026ldquo;In a softly decorated room, two children sit before an old-style computer... The scene evokes a nostalgic, timeless atmosphere.\u0026rdquo;[\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] This comparison demonstrates that our model, enhanced with spatial information, produces more descriptive and spatially-aware captions.\u003c/p\u003e \u003cp\u003eThe BLIP2 model, which generates relatively short texts (approximately 20 characters), is penalized on Ref_LongCLIP due to the preference for longer reference texts, highlighting its limitations in producing ultra-long image descriptions. In contrast, the fine-tuned WCSLLaVA demonstrates superior performance in generating ultra-long text descriptions from images without sacrificing prior knowledge.\u003c/p\u003e \u003cp\u003eTo further validate the model's acquisition of spatial knowledge, we conducted zero-shot evaluations on the VQAV2 and VIZWIZ datasets[\u003cspan additionalcitationids=\"CR30 CR31 CR32\" citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], which revealed notable improvements. Despite using a straightforward accuracy metric without semantic scoring, our model outperforms the baseline under consistent conditions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAblation Study Results: The Effect of Spatial Data on Performance\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConfiguration\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMETEOR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eROUGE_L\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRef_longCLIP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSemantic Similarity\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e16.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9996\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImage and Depth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e18.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9997\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImage and Pointcloud\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e17.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9996\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWCSLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e18.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e57.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.9997\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e4.4 Ablation Study\u003c/h2\u003e \u003cp\u003eOur decision to incorporate both depth maps and point cloud data as additional inputs is substantiated through ablation studies. We evaluated three configurations\u0026mdash;image-only, image-depth, and image-pointcloud\u0026mdash;each trained with the same parameters and dataset. The results affirm that combining depth maps with point cloud data yields superior outcomes.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study effectively demonstrates the potential to enhance the spatial cognition capabilities of multimodal large language models (MLLMs) through the integration of depth maps and point cloud data. By crafting sophisticated instructions and employing Low-Rank Adaptation (LoRA) fine-tuning techniques, we developed a model named WeCanSee-LLaVA. This model exhibits superior performance in tasks such as image captioning and visual question answering, particularly excelling in handling intricate spatial information. Our experimental results reveal that the integration of spatial data not only addresses the limitations inherent in traditional RGB images and succinct text descriptions but also significantly enhances the quality and detail of image-to-text generation, especially in producing extended textual narratives. This research offers novel insights and methodologies for advancing MLLMs\u0026rsquo; comprehension and articulation of complex spatial relationships in multimodal tasks, highlighting a promising avenue for broad applications.[\u003cspan additionalcitationids=\"CR35 CR36 CR37\" citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/p\u003e"},{"header":"Limitations","content":"\u003cp\u003eAlthough our approach has demonstrated the capability to enable MLLMs to learn spatial information from planar images, this learning process does not entirely eliminate the issue of hallucinations but mitigates the issue to a certain extent rather than eliminating it entirely. Furthermore, when generating exceptionally long texts, conventional metrics are inadequate for effectively quantifying and evaluating the generated content. The metrics currently employed are limited to basic testing, underscoring the necessity for the development of evaluation methods tailored to these specific requirements. Addressing these challenges mandates further investigation and research.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003e7 Funding\u003c/h2\u003e \u003cp\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eZ.Wang and R.Q. conceived the experiments, conducted the model training and analysis, and wrote the main manuscript text. Z.Wu contributed to the data preprocessing and collection. D. Du contributed to the methodology discussion and provided technical support throughout the research. X.D. prepared Figures 1\u0026ndash;6. All authors reviewed the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and/or analysed during the current study are not publicly available due to the large volume and proprietary format of the raw depth maps and point cloud data, but are available from the corresponding author on reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eZhu, Y. et al. \u003cem\u003eChatNav: Leveraging LLM to Zero-Shot Semantic Reasoning in Object Navigation\u003c/em\u003e. \u003cem\u003eIEEE Trans. Circuits Syst. Video Technol.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, R. et al. \u003cem\u003eSce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning\u003c/em\u003e. \u003cem\u003eIEEE Rob. Autom. Lett.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, D. et al. \u003cem\u003eEmpowering natural human-robot collaboration through multimodal language models and spatial intelligence: Pathways and perspectives\u003c/em\u003e ( Robotics and Computer-Integrated Manufacturing, 2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiong, H. et al. \u003cem\u003e3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding\u003c/em\u003e. \u003cem\u003eIEEE Trans. Multimedia\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeng Liang ༆ Jianqin Yin. \u003cem\u003eViewInfer3D: 3D Visual Grounding Based on Embodied Viewpoint Inference\u003c/em\u003e. \u003cem\u003eIEEE Rob. Autom. Lett.\u003c/em\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, W. et al. \u003cem\u003eEarthGPT-X: A Spatial MLLM for Multilevel Multisource Remote Sensing Imagery Understanding With Visual Prompting\u003c/em\u003e. \u003cem\u003eIEEE Trans. Geosci. Remote Sens.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, C. et al. \u003cem\u003eVIDHALLUC: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding.\u003c/em\u003e CVPR (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, X. \u003cem\u003eGeo-hallucination in urban analytics: What it is and why it matters.\u003c/em\u003e Environment and Planning B-Urban Analytics and City Science (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChu \u0026amp; Shiyong༆ Yuwei Chen.. \u003cem\u003eWeaponizing cognitive bias in autonomous systems: a framework for black-box inference attacks\u003c/em\u003e. \u003cem\u003eFront. Artif. Intell.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu, T. et al. \u003cem\u003e3DBench: A scalable benchmark for object and scene-level instruction-tuning of 3D large language models.\u003c/em\u003e Neural Networks (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Z. et al. \u003cem\u003eInstruction-guided path planning with 3D semantic maps for vision-language navigation.\u003c/em\u003e Neurocomputing (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFaria, F. T. et al. \u003cem\u003eTowards Robust Chain-of-Thought Prompting with Self-Consistency for Remote Sensing VQA: An Empirical Study Across Large Multimodal Models.\u003c/em\u003e Mathematics (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, A. et al. \u003cem\u003eEvaluating and enhancing spatial cognition abilities of large language models\u003c/em\u003e. \u003cem\u003eInt. J. Geogr. Inf. Sci.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, F. et al. \u003cem\u003eVisual Spat. Reasoning Trans. Association Comput. Linguistics\u003c/em\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, C. M. \u0026amp; Zhang \u003cem\u003eA dual-dimension collaborative enhancement framework to boost language model spatial semantic understanding\u003c/em\u003e ( Annals of the New York Academy of Sciences, 2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIngale, A. A. \u003cem\u003eSpatial Sequence Reasoning in Large Language Models\u0026hellip;\u003c/em\u003e Dissertation/Thesis (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThai, A. \u003cem\u003eMutual Exclusivity Bias and Spatial Reasoning in Vision-language Models.\u003c/em\u003e Dissertation/Thesis (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Z. et al. \u003cem\u003eDual-level dynamic heterogeneous graph network for video question answering\u003c/em\u003e. Neural Networks (2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYusuf, A. et al. \u003cem\u003eGraph-enhanced visual representations and question-guided dual attention for visual question answering\u003c/em\u003e. Neurocomputing (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLai, C. et al. \u003cem\u003eObject-Centric Cross-Modal Knowledge Reasoning for Future Event Prediction in Videos\u003c/em\u003e. \u003cem\u003eIEEE Trans. Circuits Syst. Video Technol.\u003c/em\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHonerkamp, D. et al. \u003cem\u003eLanguage-Grounded Dynamic Scene Graphs for Interactive Object Search With Mobile Manipulation\u003c/em\u003e. \u003cem\u003eIEEE Rob. Autom. Lett.\u003c/em\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, Z. et al. \u003cem\u003eAutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding\u003c/em\u003e. \u003cem\u003eIEEE Trans. Multimedia\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePadhan, S. \u003cem\u003eProbabilistic Metric-Semantic Grounding for Embodied AI Composing Bayesian Spatial Kernels, VLMs, and 3D Scene Graphs\u003c/em\u003e. Dissertation/Thesis (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeng, L. et al. \u003cem\u003ePseudo-EV: Enhancing 3D Visual Grounding With Pseudo Embodied Viewpoint\u003c/em\u003e. \u003cem\u003eIEEE Trans. Circuits Syst. Video Technol.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXue, Y. et al. \u003cem\u003eSpatial Knowledge Graph-Guided Multimodal Synthesis\u003c/em\u003e. \u003cem\u003eIEEE Trans. Audio Speech Lang. Process.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHossain, M. et al. \u003cem\u003eCM-SC: Cross-modal spatial-channel attention network for image captioning\u003c/em\u003e. Displays (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark, K. et al. \u003cem\u003eLearning Compositionality from Multifaceted Synthetic Data for Language-based Object Detection\u003c/em\u003e. \u003cem\u003eInt. J. Comput. Vision\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePapadopoulos, G. et al. \u003cem\u003eSADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks\u003c/em\u003e. Computers (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSterz, H. et al. \u003cem\u003eDARE: Diverse Visual Question Answering with Robustness Evaluation\u003c/em\u003e ( Transactions of the Association for Computational Linguistics, 2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, C. et al. \u003cem\u003eTowards bias-aware visual question answering: Rectifying and mitigating comprehension biases\u003c/em\u003e (Expert Systems with Applications, 2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao, Y. et al. \u003cem\u003eAdaptive Conditional Reasoning for Remote Sensing Visual Question Answering.\u003c/em\u003e Remote Sensing (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePei, B. et al. \u003cem\u003eGuiding Audio-Visual Question Answering with Collective Question Reasoning\u003c/em\u003e. \u003cem\u003eInt. J. Comput. Vision\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, W. et al. \u003cem\u003eAn Image Grid Can Be Worth a Video: Zero-Shot Video Question Answering Using a VLM\u003c/em\u003e. \u003cem\u003eIEEE Access.\u003c/em\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZargarzadeh, S. et al. \u003cem\u003eFrom Decision to Action in Surgical Autonomy: Multi-Modal Large Language Models for Robot-Assisted Blood Suction\u003c/em\u003e. \u003cem\u003eIEEE Rob. Autom. Lett.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMehta, V. et al. \u003cem\u003eLarge Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy\u003c/em\u003e. \u003cem\u003eSensors\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu, S. et al. \u003cem\u003eTARAD: Task-Aware Robot Affordance-Centric Diffusion Policy Learned From LLM-Generated Demonstrations\u003c/em\u003e. \u003cem\u003eIEEE Rob. Autom. Lett.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNan, Y. et al. \u003cem\u003eBeyond the Hype: A Dispassionate Look at Vision-Language Models in Medical Scenario\u003c/em\u003e. \u003cem\u003eIEEE Trans. Neural Networks Learn. Syst.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, B. et al. \u003cem\u003eNavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning\u003c/em\u003e. \u003cem\u003eIEEE Trans. Pattern Anal. Mach. Intell.\u003c/em\u003e (2025).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"multimodal large language models (MLLMs), spatial cognition, depth maps, point cloud data, visual question answering (VQA), Ref_LongCLIPScore","lastPublishedDoi":"10.21203/rs.3.rs-8634056/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8634056/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eContemporary multimodal large language models (MLLMs), particularly those integrating visual and textual modalities, have demonstrated remarkable capabilities in both image comprehension and text generation. However, current multimodal learning paradigms predominantly focus on RGB images and traditional text, often exhibiting limitations in spatial cognition. In this study, we enhance the spatial understanding of MLLMs by preprocessing raw images to extract spatial information such as depth maps and point cloud data, subsequently incorporating these into the learning process. Additionally, we employ instruction tuning techniques with comprehensive and detailed textual descriptions to enrich the model’s spatial awareness. Our experiments reveal that models trained with these enhancements surpass baseline models in tasks such as image captioning and visual question answering (VQA). Although traditional metrics such as CIDEr and ROUGE-L show improvement, they fail to capture the model's enhanced spatial reasoning abilities, necessitating complementary evaluation methods like Ref_LongCLIPScore. Empirically, we observed statistically significant improvements: a 1.5% absolute increase in Ref_LongCLIPScore (p \u0026lt; 0.05) and a 1.2% boost in the average accuracy of VQA tasks (p \u0026lt; 0.05). These gains underscore the model’s superior performance in describing spatial relationships within images. Our model weights are publicly available at huggingface.co/fisheries/wcsllava/tree/main.\u003c/p\u003e","manuscriptTitle":"Enhancing Spatial Cognition in MLLMs with Depth Maps and Point Cloud Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-13 17:42:06","doi":"10.21203/rs.3.rs-8634056/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-04-06T15:00:05+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"181711710079502330833112911426522844913","date":"2026-03-15T07:58:58+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-10T02:00:58+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-03T02:53:07+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-02T12:22:57+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-30T07:56:18+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-01-30T07:40:11+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ecc96e63-74ba-4c70-b67b-a3a9118f0802","owner":[],"postedDate":"March 13th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":64363925,"name":"Biological sciences/Neuroscience"},{"id":64363926,"name":"Biological sciences/Psychology"},{"id":64363927,"name":"Social science/Psychology"}],"tags":[],"updatedAt":"2026-03-13T17:42:06+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-13 17:42:06","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8634056","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8634056","identity":"rs-8634056","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.