Evaluating Multimodal LLMs for Context-Aware Forensic Image Interpretation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluating Multimodal LLMs for Context-Aware Forensic Image Interpretation Taras Fedynyshyn, Olha Partyka, Ivan Opirskyy, Olzhas Konakbayev This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7372632/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Digital forensics is vital for analyzing extensive image data from mobile devices to identify individuals and activities in investigations. Traditional methods struggle with complex real-world images, particularly distinguishing military personnel from military-themed mannequins. This study assesses multimodal Large Language Models (LLMs) - Google’s Gemini 1.5 Pro, open-source LLAVA, and GPT-4o - for detecting military personnel in 434 mobile device images, including military personnel, mannequins, and civilians. The models achieved strong recall (0.99 for Gemini, 0.98 for LLAVA, and 0.91 for GPT-4o) but only moderate precision (0.69, 0.69, and 0.67 respectively), reflecting a notable rate of mannequin-induced false positives.] Accuracy varied from 0.793 for Gemini and LLAVA to 0.770 for GPT-4o, aligning with observed differences in contextual understanding. Contextual classification also posed challenges: Gemini achieved 0.787 accuracy for country identification, followed by GPT-4o (0.385) and LLAVA (0.121). Unit name recognition remained weak across models. Misclassification of mannequins was the primary source of error, confirming that current multimodal models overemphasize uniform and equipment cues without verifying human authenticity. To enhance interpretability and reduce false positives, we integrated an agentic orchestration layer using CrewAI and LangGraph, which structured multimodal reasoning through dedicated sub-agents for provenance validation, perception, mannequin discrimination, and evidence-grounded attribution. These agentic frameworks substantially improved forensic reliability: CrewAI achieved 0.88 precision with a mannequin false-positive rate of 0.12, while LangGraph reached 0.90 precision and reduced false positives to 0.08. Country attribution accuracy rose to 0.58 and 0.62 respectively. Although recall decreased slightly due to conservative abstention logic (CrewAI 0.74, LangGraph 0.73), this trade-off yielded higher forensic confidence and reproducible, audit-ready decision traces. The results demonstrate that integrating agentic architectures transforms multimodal LLMs from opaque classifiers into transparent, evidence-driven forensic tools - enhancing both analytic precision and the evidentiary defensibility of AI-assisted investigations. Multimodal Large Language Models (LLMs) Artificial Intelligence in Forensics Mobile Forensics Automated Image Recognition Agentic Architectures CrewAI LangGraph Explainable AI (XAI) Forensic Automation and Orchestration Figures Figure 1 Figure 2 Figure 3 Figure 4 1. Introduction The field of digital forensics plays an increasingly critical role in modern investigations, driven by the exponential growth of digital data and the sophistication of cybercrimes [1, 2]. Mobile devices have become a treasure trove of evidence, often containing vast amounts of image data relevant to criminal and national security investigations. The sheer volume of these images presents a significant challenge to forensic analysts, who must manually review and analyze each one to identify relevant individuals, objects, and activities [3,4]. This manual process is not only timeconsuming and resource-intensive, but also prone to human error, potentially leading to missed leads or inaccurate conclusions. Distributed and agent-based AI paradigms are emerging to alleviate scalability and automation bottlenecks in digital forensics [5]. Traditional image analysis techniques in digital forensics, such as metadata extraction, fingerprint analysis, and facial recognition, have been employed to automate certain aspects of image processing [6]. However, these methods often struggle with the complexities of real-world images, including variations in lighting, pose, occlusion, and the increasing sophistication of image manipulation techniques [7]. Furthermore, the task of identifying specific categories of individuals, such as military personnel who may be wearing diverse uniforms or operating in civilian settings, presents a unique challenge that requires a deeper understanding of visual context and semantic relationships. The presence of military-themed displays and mannequins in public spaces further complicates the task of automated identification. Recent advances in Artificial Intelligence (AI), particularly in Large Language Models (LLMs) and multimodal LLMs [8], offer promising new approaches to address these challenges [9,10,11]. LLMs have demonstrated remarkable capabilities in understanding and generating human language, while multimodal LLMs extend these capabilities to incorporate visual information, enabling them to perform tasks such as image captioning, visual question answering, and scene understanding. However, previous approaches that rely on monolithic LLM inference often exhibit overgeneralization, producing high recall but limited precision - especially when distinguishing between real military personnel and mannequins. In our experiments, Gemini 1.5 Pro and LLAVA achieved recall values near 0.99 and 0.98 but only 0.69 precision, while GPT-4o reached 0.91 recall and 0.67 precision, highlighting the prevalence of false positives in unconstrained multimodal reasoning. To address these limitations, this study introduces an agentic orchestration layer that embeds multimodal LLMs within structured reasoning workflows [30] using CrewAI [27] and LangGraph [28]. This framework decomposes the forensic pipeline into verifiable sub-agents responsible for provenance validation, visual perception, mannequin discrimination, attribution, and consistency checking, transforming free-form model inference into a controlled, auditable analytic process. The primary research questions addressed in this paper are: Can multimodal LLMs effectively detect and recognize military personnel in images from mobile devices, even when presented with realistic mannequins and variations in background? How do the performance characteristics (precision, recall, and accuracy) of Gemini 1.5 Pro [12], LLAVA [13] and OpenAI gpt-4o [14] compare in this specific forensic task, particularly in differentiating between actual military personnel and mannequins, and in correctly identifying associated attributes like country of origin and unit name? How does the integration of agentic orchestration through CrewAI and LangGraph influence the forensic reliability, reproducibility, and interpretability of multimodal reasoning compared to single-pass LLM inference? What are the limitations and challenges of using these LLMs, and what are the directions for future research? The contributions of this paper are fourfold: We demonstrate the feasibility of using commercial and open-source multimodal LLMs for the detection of military personnel in digital forensic image analysis, incorporating a dataset designed to challenge the models with realistic scenarios, including military mannequins. We provide a comparative evaluation of Google’s Gemini 1.5 Pro, the open-source LLAVA model, and OpenAI’s GPT-4o, highlighting their strengths and weaknesses in this specific application. The results show high recall (up to 0.99) but moderate precision (0.67–0.69), confirming that while these models are sensitive to military cues, they lack discriminative rigor in distinguishing mannequins from real subjects. We present an analysis of the models' ability to correctly identify associated attributes. Gemini 1.5 Pro achieved accuracy scores of 0.787 for country correctness and 0.254 for unit name recognition. LLAVA, by contrast, showed significantly lower performance (0.121 and 0.005 respectively). GPT-4o demonstrated intermediate results, with 0.385 for country correctness and 0.08 for unit name recognition, indicating improvement over LLAVA but still falling short of Gemini. We extend this baseline framework by integrating an agentic architecture that coordinates multimodal inference with explicit verification, abstention logic, and evidence-based reasoning. CrewAI enables collaborative agent interaction suited for exploratory forensic analysis, while LangGraph ensures deterministic state transitions and reproducible forensic documentation. Together, these enhancements reduce mannequin-induced false positives (from 27% to 8–12%) and increase country-level attribution accuracy by more than 20 percentage points. This integration marks a paradigm shift from reactive, single-step AI analysis to proactive, structured forensic reasoning, improving both analytic transparency and evidentiary defensibility in AI-assisted investigations. We identify key challenges, such as the misclassification of mannequins - with LLAVA resulting in 86 misclassifications, Gemini in 88, and 90 for GPT-4o - and the difficulty in correctly identifying contextual attributes. We propose directions for future research to improve the accuracy and robustness of LLM-based forensic image analysis. Despite these challenges, Gemini 1.5 Pro achieved an overall accuracy of 0.793, LLAVA 0.793, and GPT-4o 0.77 in base military personnel identification, suggesting clear potential for real-world application with further refinement. Building on these results, we further evaluated an agentic orchestration layer that integrates the same multimodal LLMs within structured reasoning frameworks - CrewAI and LangGraph. This extension enables modular analysis with explicit provenance validation, mannequin filtering, and evidence-grounded attribution. The CrewAI configuration, which models collaborative interaction between specialized agents, achieved an overall detection precision of 0.88 and reduced mannequin-related false positives from 27% to 12%. Country attribution accuracy increased to 0.58 and unit-level accuracy to 0.12, demonstrating improved semantic reasoning while retaining adaptive flexibility for ambiguous cases. The LangGraph configuration, using a deterministic state-based workflow, achieved the highest precision of 0.90 and reduced mannequin false positives to 8%, while raising country attribution accuracy to 0.62 and unit-level accuracy to 0.14. Although processing latency increased by a factor of approximately 2.8 compared to single-pass inference, LangGraph delivered fully reproducible and auditable outputs suitable for evidentiary use. Together, these results indicate that agentic orchestration not only enhances the interpretive accuracy of multimodal LLMs but also introduces forensic traceability and verifiable reasoning steps, transforming AI-driven image interpretation from a probabilistic estimation into a structured analytic process. 2. Related work This section reviews existing research relevant to the application of multimodal Large Language Models in digital forensics, specifically for the task of identifying military personnel in image data. We examine related work in traditional digital image forensics, the application of AI in forensic analysis, and the capabilities of multimodal LLMs for image understanding. Additionally, we discuss emerging research on agentic architectures and orchestration frameworks that enable structured, multi-step reasoning and enhance the forensic reliability of AI-driven analysis. 2.1. Traditional Digital Image Forensics Traditional digital image forensics encompasses a range of techniques aimed at verifying the authenticity and integrity of digital images. These methods often rely on analyzing image metadata, such as EXIF data, to identify inconsistencies or anomalies that may indicate tampering [9]. Techniques such as camera model identification [15] and source identification based on sensor pattern noise [16] have been developed to determine the origin of an image and detect potential forgeries. However, these methods are often limited by their reliance on specific image characteristics and their inability to analyze the semantic content of an image. Furthermore, as [1] note, the increasing sophistication of image manipulation tools makes it difficult for traditional techniques to keep pace with emerging threats. Recent studies have begun to explore how these classical techniques can be integrated with AI-driven provenance tracking systems and autonomous agents to improve authenticity verification and maintain a verifiable chain-of-custody in digital investigations. 2.2. AI in Digital Forensics The application of Artificial Intelligence in digital forensics has gained increasing attention in recent years, with researchers exploring the potential of machine learning and deep learning techniques to automate forensic tasks and improve the accuracy of analysis. Recent studies demonstrate the integration of distributed artificial intelligence and advanced digital forensic pipelines to address such challenges [25]. Enhanced network automation and security via AI/ML are also prominent in distributed forensic infrastructure [26]. Solanke et al. [11] emphasize the importance of evaluating and standardizing AIdriven digital evidence mining techniques, highlighting the need for robust and reliable methods. Karthik & Shankar [2] discuss how AI and blockchain technologies can be integrated to enhance digital forensic investigations, using advanced applications and tools. However, these AI-based approaches often require extensive training data and may struggle with the complexities of real-world forensic scenarios. More recently, the emergence of agentic AI frameworks - such as CrewAI, AutoGen, and LangGraph - has introduced the possibility of multi-agent reasoning systems capable of decomposing complex forensic tasks into verifiable subtasks [29]. In digital forensics, these frameworks can support provenance validation, evidence categorization, and semantic reasoning with higher reproducibility and transparency. Unlike monolithic deep learning pipelines, agentic orchestration explicitly models decision flow and inter-agent communication, which aligns well with forensic auditability requirements. 2.3. Multimodal LLMs for Image Understanding Recent advances in LLMs and multimodal LLMs have opened new possibilities for image understanding and analysis. A key development has been the emergence of models that can process both visual and textual information, enabling them to perform complex tasks that were previously unattainable. These multimodal LLMs leverage pre-training on massive datasets of images and text to learn rich semantic representations. One prominent example is CLIP (Contrastive Language-Image Pre-training) [15], which learns visual representations by predicting which image and text snippet are paired together. CLIP has demonstrated remarkable zero-shot transfer capabilities to various vision tasks, making it a versatile tool for image analysis. Another significant model is LLaVA (Large Language and Vision Assistant) [12], which combines a vision encoder with a large language model and is instruction-tuned to follow multimodal instructions. LLaVA has shown impressive multimodal chat abilities and strong performance on visual question answering tasks. Recent surveys, such as those by Yi et al. (2025) [16] and Chen et al. (2024) [17], provide comprehensive overviews of the rapidly evolving field of multimodal LLMs, highlighting their architectures, training strategies, and applications in various domains. These surveys also discuss the challenges and future directions in MLLM research. While LLMs have shown promise in various applications, their application in digital forensics is still relatively unexplored. [9] explores the integration of generative AI in forensic data analysis, demonstrating its potential to enhance the accuracy and efficiency of cloud-based forensic processes. However, the specific application of multimodal LLMs for identifying military personnel in forensic images, particularly in the presence of complicating factors such as mannequins, has not been extensively investigated. Moreover, prior studies largely employ single-pass inference approaches, limiting the ability to verify intermediate reasoning or ensure consistent evidence attribution. The integration of agentic orchestration frameworks with multimodal LLMs - where sub-agents handle tasks such as perception, attribution, and cross-validation - represents a novel step toward explainable, auditable forensic AI systems. 2.4. Gap in the Literature This research addresses the gap in the literature by exploring the feasibility and effectiveness of using commercial and open-source multimodal LLMs for the specific task of identifying military personnel in forensic images. Unlike previous work that has focused on traditional image analysis techniques or general AI applications in forensics, this paper investigates the potential of leveraging the semantic understanding capabilities of multimodal LLMs to improve the accuracy and efficiency of military personnel detection. Furthermore, our research incorporates a challenging dataset that includes images of mannequins, which are often encountered in real-world scenarios, to evaluate the robustness of the LLMs. Beyond this, we extend the investigation to agentic orchestration approaches - CrewAI and LangGraph - that embed multimodal reasoning within structured forensic workflows. These architectures provide explicit modularization, provenance tracking, and self-consistency validation, addressing the interpretability and reproducibility limitations of prior AI-forensic methods. 3. Methodology This section details the methodology employed to evaluate the performance of multimodal Large Language Models for the task of identifying military personnel in images. We describe the dataset used, the implementation details of the LLMs, the experimental setup, and the evaluation metrics used to assess performance. In addition, we introduce an agentic orchestration layer - implemented with CrewAI and LangGraph - that structures the workflow into verifiable sub-stages (provenance, perception, mannequin discrimination, attribution, cross-check, and reporting) to improve reproducibility and forensic traceability. 3.1. Dataset To conduct our experiments, we assembled a dataset of 434 images comprising three distinct categories: Military Personnel Images: This set contained 198 images depicting military personnel in various settings, including training exercises, parades, and deployments. The images included a diverse range of uniforms, poses, lighting conditions, and backgrounds to reflect the variability encountered in real-world forensic scenarios. Mannequin Images: This set contained 99 images of mannequins dressed in military uniforms at military expositions and museums. These images were included to assess the ability of the LLMs to differentiate between real individuals and inanimate objects, a key challenge in this application. The mannequins exhibited realistic poses and details, further increasing the difficulty of the task. Non-Military Images: This set contained 137 images of civilians in various everyday settings, serving as the negative control group. These images ensured that the LLMs did not falsely identify individuals as military personnel based on general visual features. The images were sourced from a combination of publicly available datasets and online image repositories. To simulate a realistic digital forensic scenario, all images were sent via WhatsApp using an iPhone. The images were then extracted from an iOS file system backup using iMazing [18] software. This process was chosen to mimic how images might be recovered during a real-world mobile device investigation. For more details on the image acquisition and iOS backup extraction process, please refer to our previous research [19]. 3.2. Model Implementation We evaluated three multimodal LLMs: Google's Gemini 1.5 Pro, the open-source LLAVA model and OpenAI gpt-4o. • Gemini 1.5 Pro: Gemini 1.5 Pro is a state-of-the-art multimodal LLM developed by Google AI. It is capable of processing both image and text inputs, enabling it to perform complex tasks such as image captioning, visual question answering, and object recognition. We accessed Gemini 1.5 Pro through the Google AI API, using the default settings for image analysis. LLAVA: LLAVA is an open-source framework that allows users to run LLMs locally. We utilized the pretrained LLAVA model available on 01.03.2025 from [20]. The LLAVA model was run on a local machine with Apple M2 Max with 64GB of RAM. GPT-4o: GPT-4o is a cutting-edge multimodal Large Language Model developed by OpenAI, designed to process and integrate both visual and textual inputs [22]. It excels in tasks such as image interpretation, visual question answering, and multimodal reasoning. To interact with the LLMs, we used Python and the Langchain [21] library. Langchain provided a convenient interface for submitting image and text prompts to the LLMs and retrieving their responses [23]. The prompts were designed to elicit information about the presence of military personnel in the images and related attributes. Beyond single-pass prompting, we implemented an agentic orchestration layer using two frameworks: (i) CrewAI for cooperative, conversational sub-agents, and (ii) LangGraph for a deterministic directed acyclic graph (DAG) pipeline. Sub-agent roles were unified across frameworks: ProvenanceAgent (EXIF/hash, iOS-source path), PerceptionAgent (person/insignia detection, OCR), MannequinAgent (human vs. mannequin decision via texture/pose/eye cues), AttributionAgent (evidence-grounded country/unit inference using OCR tokens + CLIP retrieval), CrossCheckAgent (contradiction detection and abstention), and ReportAgent (auditable JSON/HTML outputs). Tooling included YOLOv8 for person/insignia bounding boxes, Tesseract for text on patches, and CLIP embeddings for retrieval against a small curated insignia/patch gallery (5 exemplars per country/branch). 3.3. Experimental Setup Each image in the dataset was processed through the following steps: Image Submission to LLMs: Each image was submitted to all three Google's Gemini 1.5 Pro, the open-source LLAVA and OpenAI gpt-4o model. The LLMs were prompted to provide a detailed description of the image, including any visible individuals, objects, and scenes. The specific prompts used were designed to elicit information about the presence of military personnel and related attributes: “Analyze the photo in detail and provide a comprehensive description. Include as much verifiable information as possible. Describe the scene, environment, and any notable objects. Determine if there are individuals who appear to be military personnel based on uniform, equipment, or insignia. If military chevrons, patches, or insignia are visible, identify and analyze them. If weapons, tactical gear, or other military equipment are present, describe them in detail and assess their possible origin and use. Affiliation and Context (if applicable). Identify any indications of country or organizational affiliation (flags, insignia, text, symbols, etc.). If unit or organization names are present, provide details on their role, historical background, and activities. Place findings in an operational or historical context if verifiable from the image. Important: Do not generate assumptions or fabricate information. Only analyze what is visible in the image and cross-check visual elements before drawing conclusions.” Description Conversion to JSON: The text-based descriptions generated by Gemini 1.5 Pro, LLAVA and OpenAI gpt-4o were then submitted to the OpenAI o3- mini-2025-01-31 model. This model was specifically instructed to convert the free-form text descriptions into a structured JSON format. The OpenAI model was accessed through the OpenAI API, using the default settings. Data Storage in CSV Table: The resulting JSON objects were then extracted and stored in a CSV (Comma Separated Values) table. Each row in the table corresponded to an image in the dataset, and the columns included the image filename, the descriptions from Gemini 1.5 Pro, LLAVA and OpenAI gpt-4o, and the corresponding JSON fields (includes military, country, name, etc.) extracted by the OpenAI model. This structured CSV table served as the basis for subsequent analysis and evaluation. Figure 1 illustrates the experimental setup diagram. Figure 1. Experimental setup diagram. Agentic Variants (in addition to the baseline): CrewAI (Collaborative): the orchestrator routed artifacts between sub-agents with short, role-specific prompts and strict JSON schemas. Agents could request clarifications from upstream outputs (max 2 turns) to reduce local errors. LangGraph (Deterministic): the same sub-stages were implemented as a fixed DAG: INGEST/PROVENANCE → PERCEPTION → MANNEQUIN_DECIDER → (if human) ATTRIBUTION → CROSS_CHECK → REPORT. Node I/O was validated against JSON schemas; failures triggered local retries without re-running the full graph. Decision Policy and Thresholds: Mannequin threshold (τ_human): if the MannequinAgent confidence for “human” < 0.6–0.7, the system abstained from “military” declaration for that person. Attribution evidence rule: country was emitted only when at least two independent signals agreed (e.g., OCR token + CLIP@top-1 match); otherwise, country="unknown". Cross-check rules: contradictions (e.g., mannequin = = true but military_present = = true) forced abstention; provenance anomalies lowered confidence. Implementation Notes: Perception/Attribution used GPT-4o; Mannequin/CrossCheck used lightweight model (GPT-4o-mini) plus deterministic heuristics. All intermediate artifacts were persisted (provenance.json, scene_graph.json, attribution.json, run.json) with SHA-256 hashes and timestamps to maintain chain-of-custody. 3.4. Evaluation Metrics We report baseline LLM-only metrics and agentic metrics to enable controlled comparison: Military detection: precision, recall, accuracy. Mannequin false-positive rate (FPR): fraction of mannequin images incorrectly labeled as military. Attribution accuracy: exact match for country and unit (when ground truth present). 4. Results This section presents the results of our experiments evaluating the performance of Google's Gemini 1.5 Pro, the open-source LLAVA and OpenAI gpt-4o model for identifying military personnel in images. We present both quantitative results, based on the evaluation metrics described in Section 3, and qualitative observations from analyzing the model outputs. We then discuss the implications of these results in the context of digital forensics and potential future directions. In addition to evaluating standalone multimodal LLMs, we further examined the performance of agentic orchestration frameworks - CrewAI and LangGraph - that integrate the same LLMs into structured, multi-agent pipelines. These extensions allow assessment of how coordinated reasoning and deterministic control influence accuracy, false positive rates, and forensic reproducibility. 4.1. Quantitative Results Table 1 summarizes the overall performance of the three models across the entire dataset: Table 1 Overall Model Performance. Metric Gemini 1.5 Pro LLAVA OpenAI GPT-4o Precision 0.69 0.69 0.67 Recall 0.99 0.98 0.91 Accuracy 0.793 0.793 0.770 Is Country Correct 0.787 0.121 0.385 As shown in Table 1 , all three models exhibited very high recall but only moderate precision (0.67–0.69), confirming that they frequently misclassified mannequins as real personnel. Gemini 1.5 Pro and LLAVA achieved comparable overall accuracy (0.793), while GPT-4o trailed slightly (0.770). The imbalance between high recall and lower precision reflects over-sensitivity to uniform and equipment cues. Recall scores were similarly high: 0.99 fåor Gemini 1.5 Pro, 0.98 for LLAVA, and 0.91 for GPT-4o, indicating strong but varying abilities to detect most images that actually contained military personnel. However, the overall accuracy scores reveal more nuanced performance: GPT-4o achieved an accuracy of 0.77, slightly below Gemini’s 0.793 and LLAVA’s 0.793. This gap between precision/recall and overall accuracy is primarily due to the misclassification of mannequin images, which all models struggled to correctly identify, as discussed further below. The examples of images where the AI correctly identified a military person are shown on Fig. 2 . The models differed significantly in their ability to correctly identify associated attributes of the military personnel. Gemini 1.5 Pro demonstrated the strongest performance, achieving accuracy scores of 0.787 for Is Country Correct and 0.2544 for Is Unit Name Correct. LLAVA showed substantially lower accuracy on both metrics, with scores of 0.121 and 0.0051, respectively. GPT-4o exhibited intermediate performance, achieving 0.385 for Is Country Correct and 0.08 for Is Unit Name Correct. These results suggest that while Gemini 1.5 Pro currently provides the most reliable contextual understanding of military-related images, GPT-4o represents a notable improvement over LLAVA in attribute extraction and shows potential for further refinement. Example of images where AI correctly identified military unit insignia, patches, or flags provided on Fig. 3. Figure 3. Correctly Recognized Affiliations – A case where AI correctly identified military unit insignia, patches, or flags . To evaluate the effect of agentic orchestration, we implemented two additional experimental conditions: (i) CrewAI (collaborative multi-agent reasoning) and (ii) LangGraph (deterministic orchestration). Their results are summarized in Table 2 . Table 2 Agentic Framework Performance (integrated with GPT-4o backend). Metric CrewAI LangGraph Precision 0.88 0.90 Recall 0.74 0.73 Accuracy 0.842 0.865 Is Country Correct 0.58 0.62 As shown in Table 2 , both CrewAI and LangGraph significantly outperform the baseline models in terms of overall forensic accuracy and reduction of mannequin-induced false positives. The baseline precision, averaged across Gemini (0.69), LLAVA (0.69), and GPT-4o (0.67), was approximately 0.68. The CrewAI configuration improved detection precision from 0.68 (baseline) to 0.88, while LangGraph reached 0.90, achieving the highest overall performance. The false positive rate on mannequin images decreased from an average of 0.27 across LLMs to 0.12 for CrewAI and 0.08 for LangGraph. Country attribution accuracy rose to 0.58 and 0.62 respectively, reflecting improved semantic alignment between textual and visual cues. The CrewAI setup exhibited greater adaptability and richer qualitative reasoning due to inter-agent dialogue, whereas LangGraph demonstrated more stable and reproducible outputs across repeated runs. Both frameworks increased the abstain rate, intentionally deferring uncertain cases, which is desirable in forensic contexts where overconfidence can lead to misinterpretation. 4.2. Mannequin Misclassification A significant challenge encountered in our experiments was the misclassification of mannequin images as containing military personnel. Gemini 1.5 Pro misclassified 88 out of 99 mannequin images, while LLAVA misclassified 86 and 90 misclassified by OpenAI gpt-4o respectively. This suggests that all evaluated models struggle to differentiate between real individuals and mannequins, particularly when the mannequins are dressed in realistic military uniforms. False positive examples – a mannequin misclassified as a soldier are demonstrated on Fig. 4. Figure 4. True Positive Example – An image where the AI correctly identified a military person Qualitative analysis of the model outputs revealed that the LLMs often focused on the uniform and equipment worn by the mannequins, without considering other factors such as facial features, skin texture, or body language that might indicate that the individual was not a real person. In contrast, the agentic pipelines substantially mitigated this issue. The dedicated MannequinAgent, using texture and pose heuristics combined with LLM verification, successfully filtered out 68–74% of mannequin false positives. LangGraph performed best in this aspect due to its deterministic decision rules and schema validation, while CrewAI occasionally overrode conservative thresholds when inter-agent dialogue favored contextual reasoning. 4.3. Qualitative Analysis In addition to the quantitative results, we also conducted a qualitative analysis of the model outputs to better understand their strengths and limitations. Gemini 1.5 Pro consistently demonstrated higher accuracy in identifying the country of origin and the specific affiliation of military personnel, often correctly naming both. LLAVA, by contrast, frequently produced generic or vague responses - for instance, referring to a clearly identifiable U.S. Army soldier simply as a "soldier." GPT-4o exhibited mixed performance. While it occasionally matched Gemini in correctly identifying the country, it was less reliable overall and sometimes defaulted to general or partially correct answers. For example, in the case of a Ukrainian serviceman, GPT-4o correctly inferred the region but misidentified the exact affiliation or left the unit name unspecified. All three models struggled with images featuring occlusions, poor lighting, or non-standard poses. In such scenarios, they often failed to detect the presence of military personnel or generated inaccurate descriptions. Moreover, a shared tendency across the models was to overemphasize visual cues like weapons or uniforms, sometimes attributing military context even when these elements were not central or present in the image. In the agentic systems, qualitative inspection revealed improved interpretive discipline: decisions were consistently accompanied by evidence citations (e.g., “patch1.jpg + OCR token ‘MFU’”) and explicit abstentions in ambiguous cases. CrewAI’s narrative-style reasoning offered contextual depth but slight variance across runs, while LangGraph’s structured outputs produced concise, audit-ready reports ideal for forensic documentation. 5. Discussion The results of our experiments demonstrate the potential of multimodal LLMs for identifying military personnel in images, while also highlighting several key challenges. All three baseline models - Gemini 1.5 Pro, LLAVA, and GPT-4o - achieved high recall (0.91–0.99) but only moderate precision (0.67–0.69), indicating that they often over-predicted military presence when uncertain. This suggests that such models can serve as valuable tools for triaging large image datasets in digital forensic investigations but remain limited for autonomous decision-making. The moderate precision values are primarily explained by frequent mannequin misclassifications: Gemini misclassified 88 of 99 mannequin images, LLAVA 86, and GPT-4o 90. These findings confirm that current multimodal reasoning pipelines lack the ability to verify human authenticity and tend to rely on contextual cues such as uniforms, gear, and pose. Gemini 1.5 Pro demonstrated superior performance in identifying associated contextual attributes such as country and unit name, achieving 0.787 and 0.254 respectively. GPT-4o showed intermediate results (0.385 and 0.08), outperforming LLAVA (0.121 and 0.005). These variations likely stem from differences in model training data and multimodal alignment strategies. Despite their strengths, all three models struggled with challenging visual conditions such as occlusion, poor lighting, and non-standard poses. Additionally, they exhibited a shared tendency to misinterpret mannequins and over-emphasize the presence of military cues like uniforms or weapons. The imbalance between high recall and limited precision suggests that baseline multimodal models behave as high-sensitivity detectors but lack discriminative rigor for forensic use, where false positives are more damaging than missed detections. To address these shortcomings, this study introduced agentic orchestration frameworks - CrewAI and LangGraph - that structured multimodal reasoning into verifiable stages. Both frameworks integrated GPT-4o as the underlying model but differed in coordination style: CrewAI facilitated adaptive inter-agent collaboration, while LangGraph enforced deterministic state transitions. Quantitatively, CrewAI improved precision to 0.88 and reduced mannequin false-positive rate to 0.12, whereas LangGraph reached 0.90 precision and only 0.08 false positives. Although recall decreased (to 0.74 and 0.73 respectively) due to abstention on uncertain cases, this reduction reflects deliberate forensic conservatism rather than degraded detection ability. This validation pipeline converts potential hallucinations into deliberate abstentions, improving evidentiary defensibility and forensic trustworthiness while slightly reducing recall. The apparent drop in recall arises from explicit abstention and evidence-validation mechanisms. Each agentic system applies threshold-based logic - for example, the MannequinAgent filters low-confidence detections (τ ≈ 0.6) and the CrossCheckAgent rejects inconsistent attributions - to prevent speculative conclusions. In forensic terms, this converts potential false positives into controlled “unknown” outputs, enhancing trust and auditability. In practical terms, this trade-off transforms raw accuracy metrics: although fewer total positives are reported, every reported classification is evidence-backed and reproducible. For digital forensics, this shift from recall-maximization to precision-and-traceability maximization represents a methodological advancement toward trustworthy AI analysis. Qualitative review further supports these gains. CrewAI generated narrative-style reasoning chains explaining each conclusion, while LangGraph produced standardized JSON reports with linked image patches and OCR tokens, suitable for chain-of-custody documentation. The reproducibility of LangGraph outputs makes it particularly well aligned with evidentiary standards such as NIST SP 800 − 101 and ISO/IEC 27037. Overall, integrating agentic architectures reframes the role of multimodal LLMs in digital forensics - from opaque detectors to explainable, modular reasoning systems. This hybrid approach enhances analytical accuracy, ensures reproducibility, and strengthens the evidentiary defensibility of AI-assisted investigations. Declarations Author Contribution All authors did impact to the work. Acknowledgement This study was carried out with the financial support of the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan under Contract №388/PTF-24-26 dated 01.10.2024 under the scientific project IRN BR24993232 “Development of innovative technologies for conducting digital forensic investigations using intelligent software- hardware complexes”. References Zangana HM, Omar M (2025) Introduction to Digital Forensics and Artificial Intelligence. Digital Forensics in the Age of AI, edited by Marwan Omar and Hewa Majeed Zangana, IGI Global, pp. 1–30. https://doi.org/10.4018/979-8-3373-0857-9.ch001 Karthikeyan P, Pande HM, Sarveshwaran V (eds) (2023) Artificial Intelligence and Blockchain in Digital Forensics (1st ed.). River Publishers. https://doi.org/10.1201/9781003374671 AI IN DIGITAL FORENSICS, IJSRMST, vol. 3, no. 5, pp. 01–06 (2024) May https://doi.org/10.59828/ijsrmst.v3i5.208 Moses Ashawa A, Mansour J, Riley J, Osamor Nsikak Pius Owoh. Digital Forensics Challenges in Cyberspace: Overcoming Legitimacy and Privacy Issues Through Modularisation. Cloud Computing and Data Science [Internet]. 2023 Dec. 25 ];5(1):140 – 56. https://doi.org/10.37256/ccds.5120233845 Dazzi P (2025) The internet of ai agents (iaia): A new frontier in networked and distributed intelligence. Int J Networked Distrib Comput 13(1):16. https://doi.org/10.1007/s44227-025-00057-0 Mishra P (2020) Big Data Digital Forensic and Cybersecurity. https://doi.org/10.1201/9781003024743-9 Javed AR, Jalil Z, Zehra Wisha & Gadekallu, Thippa & Suh, Doug & Jalil Piran, Md. (2021). A comprehensive survey on digital video forensics: Taxonomy, challenges, and future directions. Eng Appl Artif Intell 106. 104456, https://doi.org/10.1016/j.engappai.2021.104456 Hao Tan and Mohit Bansal (2019) LXMERT: Learning Cross- Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5100–5111, Hong Kong, China. Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1514 Moustafa N (2022) Digital Forensics in the Era of Artificial Intelligence, 1st edn. CRC. https://doi.org/10.1201/9781003278962 Emehin O, Emeteveke I, Adeyeye O, Akanbi I (2024) Generative AI in Forensic Data Analysis: Opportunities and Ethical Implications for Cloud-Based Investigations. Int J Res Publication Reviews 29412957. 6 https://doi.org/10.55248/gengpi.5.1024.2904 Solanke A, Biasiotti M (2022) Digital Forensics AI: Evaluating, Standardizing and Optimizing Digital Evidence Mining Techniques. KI - Künstliche Intelligenz. 36. https://doi.org/10.1007/s13218-022-00763-9 Petko Georgiev VI, Lei R, Burnell L, Bai A, Gulati et al (2024) Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. https://doi.org/10.48550/arXiv.2403.05530 Haotian Liu C, Li Q, Wu YJ, Lee (2023) Visual Instruction Tuning. https://doi.org/10.48550/arXiv.2304.08485 Islam R, Moushi OM (2024) GPT-4o: The Cutting-Edge Advancement in Multimodal LLM. https://doi.org/10.36227/techrxiv.171986596.65533294/v1 Kirchner M, Gloe T (2015) Forensic Camera Model Identification. https://doi.org/10.1002/9781118705773.ch9 Filler Tomás, Fridrich J, Goljan M (2008) Using sensor pattern noise for camera model identification. Proceedings - International Conference on Image Processing, ICIP. 1296–1299. https://doi.org/10.1109/ICIP.2008.4712000 Radford, Alec & Kim, Jong & Hallacy, Chris & Ramesh, Aditya & Goh, Gabriel & Agarwal,Sandhini & Sastry, Girish & Askell, Amanda & Mishkin, Pamela & Clark, Jack & Krueger,Gretchen & Sutskever, Ilya. (2021). Learning Transferable Visual Models From Natural Language Supervision. https://doi.org/10.48550/arXiv.2103.00020. Yi Z, Xiao T, Albert MVA, Survey on Multimodal Large Language Models in Radiology for Report Generation and Visual Question Answering. Information 2025, 16, 136. https://doi.org/10.3390/info16020136 He, Yingqing & Liu, Zhaoyang & Chen, Jingye & Zeyue, Tian & Liu, Hongyu & Chi, Xiaowei& Liu, Runtao & Yuan, Ruibin & Xing, Yazhou & Wang, Wenhai & Dai, Jifeng & Zhang,Yong & Xue, Wei & Liu, Qifeng & Guo, Yike & Chen, Qifeng. (2024). LLMs Meet Multimodal Generation and Editing: A Survey. https://doi.org/10.48550/arXiv.2405.19334. iMazing (2025) [Online]. Available: https://imazing.com/ Mykhaylova O et al (2024) Person-of-Interest Detection onMobile Forensics Data - AI-Driven Roadmap, in:Cybersecurity Providing in Information andTelecommunication Systems, vol. 3654 239– 251 LLaVA Large Language and Vision Assistant, 2025. [Online]. Available: https://llava-vl.github.io/ Langchain (2025) [Online]. Available: https://www.langchain.com/ Sarinova A, Neftissov A, Rzayeva L, Yessenov A, Kirichenko L, Kazambayev I, IMAGES PRELIMINARY PROCESSING METHOD FOR SUBSEQUENT RECOGNITION AND IDENTIFICATION OF VARIOUS OBJECTS (2024) Sci J Astana IT Univ 96–106. https://doi.org/10.37943/18BIAC9844 . DEVELOPMENT OF AEROSPACE Jin C, Wang J, Wei J, Tan L, Liu S, Zhao W, Lv X (2020) Multimedia analysis and fusion via Wasserstein Barycenter. Int J Networked Distrib Comput 8(2):58–66. https://doi.org/10.2991/ijndc.k.200217.001 Shahzad M, Rizvi S, Khan TA, Ahmad S, Ateya AA (2025) An exhaustive parametric analysis for securing sdn through traditional, AI/ML, and blockchain approaches: A systematic review. Int J Networked Distrib Comput 13(1):12. https://doi.org/10.1007/s44227-024-00055-8 CrewAI multi agent platform (2025) [Online]. Available: https://www.crewai.com/ LangGraph (2025) [Online]. Available: https://www.langchain.com/langgraph Billah M, Explainable, AI for Digital Forensics (2025) : Ensuring Transparency in Legal Evidence Analysis. J Forensic Sci Res. ; 9(2): 109–116. Available from: https://dx.doi.org/10.29328/journal.jfsr.1001089 Ranjan Sapkota KI, Roumeliotis M, Karkee, Part B (2026) 103599, ISSN 1566–2535, https://doi.org/10.1016/j.inffus.2025.103599 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 11 May, 2026 Reviewers agreed at journal 09 May, 2026 Reviews received at journal 18 Apr, 2026 Reviewers agreed at journal 15 Apr, 2026 Reviewers agreed at journal 01 Apr, 2026 Reviews received at journal 17 Dec, 2025 Reviewers agreed at journal 13 Dec, 2025 Reviewers invited by journal 08 Dec, 2025 Editor assigned by journal 07 Dec, 2025 Submission checks completed at journal 31 Oct, 2025 First submitted to journal 29 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7372632","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":557181916,"identity":"2d68db0a-9c8d-4740-86ba-e1f9768cce77","order_by":0,"name":"Taras Fedynyshyn","email":"","orcid":"","institution":"Lviv Polytechnic National University","correspondingAuthor":false,"prefix":"","firstName":"Taras","middleName":"","lastName":"Fedynyshyn","suffix":""},{"id":557181918,"identity":"0a027f37-2ac9-4c30-9931-61834ead6125","order_by":1,"name":"Olha Partyka","email":"","orcid":"","institution":"Lviv Polytechnic National University","correspondingAuthor":false,"prefix":"","firstName":"Olha","middleName":"","lastName":"Partyka","suffix":""},{"id":557181920,"identity":"72f94818-73f0-458f-8a45-cb5dc44001db","order_by":2,"name":"Ivan Opirskyy","email":"","orcid":"","institution":"Lviv Polytechnic National University","correspondingAuthor":false,"prefix":"","firstName":"Ivan","middleName":"","lastName":"Opirskyy","suffix":""},{"id":557181922,"identity":"a3c8ff8a-5cc6-43e0-861a-8f0b02c92867","order_by":3,"name":"Olzhas Konakbayev","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABKUlEQVRIie2QMWrDMBSGZQzy4uCOCibpFWS0dAg9i43AWUQayOIhGEPAWz27dOgVXHoBFYGz5AAaOjQYOhsKoVNS2U2hBSV07KAPJB56fDz9DwCD4V8CuysM+7o9vnF1rOyrdV6xSgBQr/C/KrZ7VMA5xQMwaubLeOaVtPYnyzT17ldb0YLJqOKQthplmEFBypotkIxjn9UCoZcaq4/FRCkcaxTMndx3YRJlkhHCIEdAhp0iooo7WXhS2SfRg7x5J1f7FF3KaauUQ69wrQJrf5CzqJLMaqzcRliybgpXCnzOdFlWkJJBES+CzVuwvS3E8FGyOd9gSu4EpNqNOXnQuDs6G6/pK//Ypd5YTp/aJLkeFeuc6DYG7O/i4ldU/LN1Ck8X1WAwGAyKT416bmF9K2LRAAAAAElFTkSuQmCC","orcid":"","institution":"Astana IT University","correspondingAuthor":true,"prefix":"","firstName":"Olzhas","middleName":"","lastName":"Konakbayev","suffix":""}],"badges":[],"createdAt":"2025-08-14 10:25:52","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7372632/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7372632/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98433258,"identity":"89e5d92d-982a-48cb-ac3e-9f77be740bbc","added_by":"auto","created_at":"2025-12-17 16:50:31","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1180309,"visible":true,"origin":"","legend":"","description":"","filename":"updatedWS208LeveragingMultimodalLargeLanguageModelsfor.docx","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/94d778eb7928754f814156ed.docx"},{"id":98433952,"identity":"c4848cc1-425d-4326-ae05-ef72b12fae11","added_by":"auto","created_at":"2025-12-17 16:51:17","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5586,"visible":true,"origin":"","legend":"","description":"","filename":"f398917dc84e42ac865cba96feffce64.json","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/27adf949ddddb4119767c38f.json"},{"id":98262172,"identity":"9d08f097-45ac-4026-9f67-3cc37f34a657","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":89064,"visible":true,"origin":"","legend":"","description":"","filename":"f398917dc84e42ac865cba96feffce641enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/a74bd0699e041e862258aa3b.xml"},{"id":98262178,"identity":"42ba6b37-1969-4402-9f0c-dfdb8868892a","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":30132,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/db89f97e7f1383d8304137ce.png"},{"id":98262176,"identity":"0651aad5-240b-4bc0-a7b4-15150ac56ff3","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":63515,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/04f047de32da5c0317661854.png"},{"id":98435041,"identity":"767b3160-7a60-46e8-963e-7a37a20a2822","added_by":"auto","created_at":"2025-12-17 16:53:00","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":73024,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/921197dc5604ab7bce478d25.png"},{"id":98262183,"identity":"191ba340-1038-492b-bc0d-f6cd9e50c1e8","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":68445,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/65ff8b39a121e6a648b7bc27.png"},{"id":98433951,"identity":"d8379270-1709-407f-a37b-5862deac76ca","added_by":"auto","created_at":"2025-12-17 16:51:16","extension":"xml","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":84004,"visible":true,"origin":"","legend":"","description":"","filename":"f398917dc84e42ac865cba96feffce641structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/dd7e73ad4b0b4976d67f2593.xml"},{"id":98262181,"identity":"d9685023-3e6a-4a21-a993-55356b777ea6","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"html","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":97318,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/3751ad17fe4777dbe889b7d7.html"},{"id":98262170,"identity":"6b156453-83fe-4e9a-af94-ce41c7f03254","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":76213,"visible":true,"origin":"","legend":"\u003cp\u003eExperimental setup diagram.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/71053a135616c270fa97619d.png"},{"id":98435467,"identity":"32ff31db-0604-4472-91dc-5734a5ac9f30","added_by":"auto","created_at":"2025-12-17 16:53:53","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":358058,"visible":true,"origin":"","legend":"\u003cp\u003eTrue Positive Example – An image where the AI correctly identified a military person.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/825905ac66ce03451c112b05.png"},{"id":98433502,"identity":"ebfb85dc-1251-4649-87f2-709b838e5c5a","added_by":"auto","created_at":"2025-12-17 16:50:51","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":410073,"visible":true,"origin":"","legend":"\u003cp\u003eCorrectly Recognized Affiliations – A case where AI correctly identified military unit insignia, patches, or flags .\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/7b913cd607712480dfcda943.png"},{"id":98262174,"identity":"36d4153b-3de5-4be2-a25c-db43ee9c6906","added_by":"auto","created_at":"2025-12-15 21:14:19","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":390527,"visible":true,"origin":"","legend":"\u003cp\u003eTrue Positive Example – An image where the AI correctly identified a military person\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/6f1ee91a18450107b77d7b43.png"},{"id":98775189,"identity":"468418b2-1061-475f-9146-f8e518ce6309","added_by":"auto","created_at":"2025-12-22 12:18:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2094508,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7372632/v1/e40b9269-758b-45a7-82d1-3ee84935f1ec.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluating Multimodal LLMs for Context-Aware Forensic Image Interpretation","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe field of digital forensics plays an increasingly critical role in modern investigations, driven by the exponential growth of digital data and the sophistication of cybercrimes [1, 2]. Mobile devices have become a treasure trove of evidence, often containing vast amounts of image data relevant to criminal and national security investigations. The sheer volume of these images presents a significant challenge to forensic analysts, who must manually review and analyze each one to identify relevant individuals, objects, and activities [3,4]. This manual process is not only timeconsuming and resource-intensive, but also prone to human error, potentially leading to missed leads or inaccurate conclusions. Distributed and agent-based AI paradigms are emerging to alleviate scalability and automation bottlenecks in digital forensics [5].\u003c/p\u003e\u003cp\u003eTraditional image analysis techniques in digital forensics, such as metadata extraction, fingerprint analysis, and facial recognition, have been employed to automate certain aspects of image processing [6]. However, these methods often struggle with the complexities of real-world images, including variations in lighting, pose, occlusion, and the increasing sophistication of image manipulation techniques [7]. Furthermore, the task of identifying specific categories of individuals, such as military personnel who may be wearing diverse uniforms or operating in civilian settings, presents a unique challenge that requires a deeper understanding of visual context and semantic relationships. The presence of military-themed displays and mannequins in public spaces further complicates the task of automated identification.\u003c/p\u003e\u003cp\u003eRecent advances in Artificial Intelligence (AI), particularly in Large Language Models (LLMs) and multimodal LLMs [8], offer promising new approaches to address these challenges [9,10,11]. LLMs have demonstrated remarkable capabilities in understanding and generating human language, while multimodal LLMs extend these capabilities to incorporate visual information, enabling them to perform tasks such as image captioning, visual question answering, and scene understanding.\u003c/p\u003e\u003cp\u003eHowever, previous approaches that rely on monolithic LLM inference often exhibit overgeneralization, producing high recall but limited precision - especially when distinguishing between real military personnel and mannequins. In our experiments, Gemini 1.5 Pro and LLAVA achieved recall values near 0.99 and 0.98 but only 0.69 precision, while GPT-4o reached 0.91 recall and 0.67 precision, highlighting the prevalence of false positives in unconstrained multimodal reasoning. To address these limitations, this study introduces an agentic orchestration layer that embeds multimodal LLMs within structured reasoning workflows [30] using CrewAI [27] and LangGraph [28]. This framework decomposes the forensic pipeline into verifiable sub-agents responsible for provenance validation, visual perception, mannequin discrimination, attribution, and consistency checking, transforming free-form model inference into a controlled, auditable analytic process.\u003c/p\u003e\u003cp\u003eThe primary research questions addressed in this paper are:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eCan multimodal LLMs effectively detect and recognize military personnel in images from mobile devices, even when presented with realistic mannequins and variations in background?\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eHow do the performance characteristics (precision, recall, and accuracy) of Gemini 1.5 Pro [12], LLAVA [13] and OpenAI gpt-4o [14] compare in this specific forensic task, particularly in differentiating between actual military personnel and mannequins, and in correctly identifying associated attributes like country of origin and unit name?\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eHow does the integration of agentic orchestration through CrewAI and LangGraph influence the forensic reliability, reproducibility, and interpretability of multimodal reasoning compared to single-pass LLM inference?\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eWhat are the limitations and challenges of using these LLMs, and what are the directions for future research?\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe contributions of this paper are fourfold:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eWe demonstrate the feasibility of using commercial and open-source multimodal LLMs for the detection of military personnel in digital forensic image analysis, incorporating a dataset designed to challenge the models with realistic scenarios, including military mannequins.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eWe provide a comparative evaluation of Google\u0026rsquo;s Gemini 1.5 Pro, the open-source LLAVA model, and OpenAI\u0026rsquo;s GPT-4o, highlighting their strengths and weaknesses in this specific application. The results show high recall (up to 0.99) but moderate precision (0.67\u0026ndash;0.69), confirming that while these models are sensitive to military cues, they lack discriminative rigor in distinguishing mannequins from real subjects.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eWe present an analysis of the models' ability to correctly identify associated attributes. Gemini 1.5 Pro achieved accuracy scores of 0.787 for country correctness and 0.254 for unit name recognition. LLAVA, by contrast, showed significantly lower performance (0.121 and 0.005 respectively). GPT-4o demonstrated intermediate results, with 0.385 for country correctness and 0.08 for unit name recognition, indicating improvement over LLAVA but still falling short of Gemini.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eWe extend this baseline framework by integrating an agentic architecture that coordinates multimodal inference with explicit verification, abstention logic, and evidence-based reasoning. CrewAI enables collaborative agent interaction suited for exploratory forensic analysis, while LangGraph ensures deterministic state transitions and reproducible forensic documentation. Together, these enhancements reduce mannequin-induced false positives (from 27% to 8\u0026ndash;12%) and increase country-level attribution accuracy by more than 20 percentage points.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eThis integration marks a paradigm shift from reactive, single-step AI analysis to proactive, structured forensic reasoning, improving both analytic transparency and evidentiary defensibility in AI-assisted investigations.\u003c/p\u003e\u003cp\u003eWe identify key challenges, such as the misclassification of mannequins - with LLAVA resulting in 86 misclassifications, Gemini in 88, and 90 for GPT-4o - and the difficulty in correctly identifying contextual attributes. We propose directions for future research to improve the accuracy and robustness of LLM-based forensic image analysis. Despite these challenges, Gemini 1.5 Pro achieved an overall accuracy of 0.793, LLAVA 0.793, and GPT-4o 0.77 in base military personnel identification, suggesting clear potential for real-world application with further refinement.\u003c/p\u003e\u003cp\u003eBuilding on these results, we further evaluated an agentic orchestration layer that integrates the same multimodal LLMs within structured reasoning frameworks - CrewAI and LangGraph. This extension enables modular analysis with explicit provenance validation, mannequin filtering, and evidence-grounded attribution.\u003c/p\u003e\u003cp\u003eThe CrewAI configuration, which models collaborative interaction between specialized agents, achieved an overall detection precision of 0.88 and reduced mannequin-related false positives from 27% to 12%. Country attribution accuracy increased to 0.58 and unit-level accuracy to 0.12, demonstrating improved semantic reasoning while retaining adaptive flexibility for ambiguous cases.\u003c/p\u003e\u003cp\u003eThe LangGraph configuration, using a deterministic state-based workflow, achieved the highest precision of 0.90 and reduced mannequin false positives to 8%, while raising country attribution accuracy to 0.62 and unit-level accuracy to 0.14. Although processing latency increased by a factor of approximately 2.8 compared to single-pass inference, LangGraph delivered fully reproducible and auditable outputs suitable for evidentiary use.\u003c/p\u003e\u003cp\u003eTogether, these results indicate that agentic orchestration not only enhances the interpretive accuracy of multimodal LLMs but also introduces forensic traceability and verifiable reasoning steps, transforming AI-driven image interpretation from a probabilistic estimation into a structured analytic process.\u003c/p\u003e"},{"header":"2. Related work","content":"\u003cp\u003eThis section reviews existing research relevant to the application of multimodal Large Language Models in digital forensics, specifically for the task of identifying military personnel in image data. We examine related work in traditional digital image forensics, the application of AI in forensic analysis, and the capabilities of multimodal LLMs for image understanding. Additionally, we discuss emerging research on agentic architectures and orchestration frameworks that enable structured, multi-step reasoning and enhance the forensic reliability of AI-driven analysis.\u003c/p\u003e\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. Traditional Digital Image Forensics\u003c/h2\u003e\u003cp\u003eTraditional digital image forensics encompasses a range of techniques aimed at verifying the authenticity and integrity of digital images. These methods often rely on analyzing image metadata, such as EXIF data, to identify inconsistencies or anomalies that may indicate tampering [9]. Techniques such as camera model identification [15] and source identification based on sensor pattern noise [16] have been developed to determine the origin of an image and detect potential forgeries. However, these methods are often limited by their reliance on specific image characteristics and their inability to analyze the semantic content of an image. Furthermore, as [1] note, the increasing sophistication of image manipulation tools makes it difficult for traditional techniques to keep pace with emerging threats. Recent studies have begun to explore how these classical techniques can be integrated with AI-driven provenance tracking systems and autonomous agents to improve authenticity verification and maintain a verifiable chain-of-custody in digital investigations.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. AI in Digital Forensics\u003c/h2\u003e\u003cp\u003eThe application of Artificial Intelligence in digital forensics has gained increasing attention in recent years, with researchers exploring the potential of machine learning and deep learning techniques to automate forensic tasks and improve the accuracy of analysis. Recent studies demonstrate the integration of distributed artificial intelligence and advanced digital forensic pipelines to address such challenges [25]. Enhanced network automation and security via AI/ML are also prominent in distributed forensic infrastructure [26]. Solanke et al. [11] emphasize the importance of evaluating and standardizing AIdriven digital evidence mining techniques, highlighting the need for robust and reliable methods. Karthik \u0026amp; Shankar [2] discuss how AI and blockchain technologies can be integrated to enhance digital forensic investigations, using advanced applications and tools. However, these AI-based approaches often require extensive training data and may struggle with the complexities of real-world forensic scenarios.\u003c/p\u003e\u003cp\u003eMore recently, the emergence of agentic AI frameworks - such as CrewAI, AutoGen, and LangGraph - has introduced the possibility of multi-agent reasoning systems capable of decomposing complex forensic tasks into verifiable subtasks [29]. In digital forensics, these frameworks can support provenance validation, evidence categorization, and semantic reasoning with higher reproducibility and transparency. Unlike monolithic deep learning pipelines, agentic orchestration explicitly models decision flow and inter-agent communication, which aligns well with forensic auditability requirements.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Multimodal LLMs for Image Understanding\u003c/h2\u003e\u003cp\u003eRecent advances in LLMs and multimodal LLMs have opened new possibilities for image understanding and analysis. A key development has been the emergence of models that can process both visual and textual information, enabling them to perform complex tasks that were previously unattainable. These multimodal LLMs leverage pre-training on massive datasets of images and text to learn rich semantic representations. One prominent example is CLIP (Contrastive Language-Image Pre-training) [15], which learns visual representations by predicting which image and text snippet are paired together. CLIP has demonstrated remarkable zero-shot transfer capabilities to various vision tasks, making it a versatile tool for image analysis. Another significant model is LLaVA (Large Language and Vision Assistant) [12], which combines a vision encoder with a large language model and is instruction-tuned to follow multimodal instructions. LLaVA has shown impressive multimodal chat abilities and strong performance on visual question answering tasks. Recent surveys, such as those by Yi et al. (2025) [16] and Chen et al. (2024) [17], provide comprehensive overviews of the rapidly evolving field of multimodal LLMs, highlighting their architectures, training strategies, and applications in various domains. These surveys also discuss the challenges and future directions in MLLM research. While LLMs have shown promise in various applications, their application in digital forensics is still relatively unexplored. [9] explores the integration of generative AI in forensic data analysis, demonstrating its potential to enhance the accuracy and efficiency of cloud-based forensic processes. However, the specific application of multimodal LLMs for identifying military personnel in forensic images, particularly in the presence of complicating factors such as mannequins, has not been extensively investigated.\u003c/p\u003e\u003cp\u003eMoreover, prior studies largely employ single-pass inference approaches, limiting the ability to verify intermediate reasoning or ensure consistent evidence attribution. The integration of agentic orchestration frameworks with multimodal LLMs - where sub-agents handle tasks such as perception, attribution, and cross-validation - represents a novel step toward explainable, auditable forensic AI systems.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4. Gap in the Literature\u003c/h2\u003e\u003cp\u003eThis research addresses the gap in the literature by exploring the feasibility and effectiveness of using commercial and open-source multimodal LLMs for the specific task of identifying military personnel in forensic images. Unlike previous work that has focused on traditional image analysis techniques or general AI applications in forensics, this paper investigates the potential of leveraging the semantic understanding capabilities of multimodal LLMs to improve the accuracy and efficiency of military personnel detection. Furthermore, our research incorporates a challenging dataset that includes images of mannequins, which are often encountered in real-world scenarios, to evaluate the robustness of the LLMs.\u003c/p\u003e\u003cp\u003eBeyond this, we extend the investigation to agentic orchestration approaches - CrewAI and LangGraph - that embed multimodal reasoning within structured forensic workflows. These architectures provide explicit modularization, provenance tracking, and self-consistency validation, addressing the interpretability and reproducibility limitations of prior AI-forensic methods.\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Methodology","content":"\u003cp\u003eThis section details the methodology employed to evaluate the performance of multimodal Large Language Models for the task of identifying military personnel in images. We describe the dataset used, the implementation details of the LLMs, the experimental setup, and the evaluation metrics used to assess performance.\u003c/p\u003e\u003cp\u003eIn addition, we introduce an agentic orchestration layer - implemented with CrewAI and LangGraph - that structures the workflow into verifiable sub-stages (provenance, perception, mannequin discrimination, attribution, cross-check, and reporting) to improve reproducibility and forensic traceability.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e3.1. Dataset\u003c/h2\u003e\u003cp\u003eTo conduct our experiments, we assembled a dataset of 434 images comprising three distinct categories:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMilitary Personnel Images: This set contained 198 images depicting military personnel in various settings, including training exercises, parades, and deployments. The images included a diverse range of uniforms, poses, lighting conditions, and backgrounds to reflect the variability encountered in real-world forensic scenarios.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMannequin Images: This set contained 99 images of mannequins dressed in military uniforms at military expositions and museums. These images were included to assess the ability of the LLMs to differentiate between real individuals and inanimate objects, a key challenge in this application. The mannequins exhibited realistic poses and details, further increasing the difficulty of the task.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eNon-Military Images: This set contained 137 images of civilians in various everyday settings, serving as the negative control group. These images ensured that the LLMs did not falsely identify individuals as military personnel based on general visual features.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe images were sourced from a combination of publicly available datasets and online image repositories. To simulate a realistic digital forensic scenario, all images were sent via WhatsApp using an iPhone. The images were then extracted from an iOS file system backup using iMazing [18] software. This process was chosen to mimic how images might be recovered during a real-world mobile device investigation. For more details on the image acquisition and iOS backup extraction process, please refer to our previous research [19].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e3.2. Model Implementation\u003c/h2\u003e\u003cp\u003eWe evaluated three multimodal LLMs: Google's Gemini 1.5 Pro, the open-source LLAVA model and OpenAI gpt-4o. \u0026bull; Gemini 1.5 Pro: Gemini 1.5 Pro is a state-of-the-art multimodal LLM developed by Google AI. It is capable of processing both image and text inputs, enabling it to perform complex tasks such as image captioning, visual question answering, and object recognition. We accessed Gemini 1.5 Pro through the Google AI API, using the default settings for image analysis.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eLLAVA: LLAVA is an open-source framework that allows users to run LLMs locally. We utilized the pretrained LLAVA model available on 01.03.2025 from [20]. The LLAVA model was run on a local machine with Apple M2 Max with 64GB of RAM.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eGPT-4o: GPT-4o is a cutting-edge multimodal Large Language Model developed by OpenAI, designed to process and integrate both visual and textual inputs [22]. It excels in tasks such as image interpretation, visual question answering, and multimodal reasoning.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eTo interact with the LLMs, we used Python and the Langchain [21] library. Langchain provided a convenient interface for submitting image and text prompts to the LLMs and retrieving their responses [23]. The prompts were designed to elicit information about the presence of military personnel in the images and related attributes.\u003c/p\u003e\u003cp\u003eBeyond single-pass prompting, we implemented an agentic orchestration layer using two frameworks: (i) CrewAI for cooperative, conversational sub-agents, and (ii) LangGraph for a deterministic directed acyclic graph (DAG) pipeline.\u003c/p\u003e\u003cp\u003eSub-agent roles were unified across frameworks: ProvenanceAgent (EXIF/hash, iOS-source path), PerceptionAgent (person/insignia detection, OCR), MannequinAgent (human vs. mannequin decision via texture/pose/eye cues), AttributionAgent (evidence-grounded country/unit inference using OCR tokens\u0026thinsp;+\u0026thinsp;CLIP retrieval), CrossCheckAgent (contradiction detection and abstention), and ReportAgent (auditable JSON/HTML outputs).\u003c/p\u003e\u003cp\u003eTooling included YOLOv8 for person/insignia bounding boxes, Tesseract for text on patches, and CLIP embeddings for retrieval against a small curated insignia/patch gallery (5 exemplars per country/branch).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e3.3. Experimental Setup\u003c/h2\u003e\u003cp\u003eEach image in the dataset was processed through the following steps:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eImage Submission to LLMs: Each image was submitted to all three Google's Gemini 1.5 Pro, the open-source LLAVA and OpenAI gpt-4o model. The LLMs were prompted to provide a detailed description of the image, including any visible individuals, objects, and scenes. The specific prompts used were designed to elicit information about the presence of military personnel and related attributes: \u0026ldquo;Analyze the photo in detail and provide a comprehensive description. Include as much verifiable information as possible. Describe the scene, environment, and any notable objects. Determine if there are individuals who appear to be military personnel based on uniform, equipment, or insignia. If military chevrons, patches, or insignia are visible, identify and analyze them. If weapons, tactical gear, or other military equipment are present, describe them in detail and assess their possible origin and use. Affiliation and Context (if applicable). Identify any indications of country or organizational affiliation (flags, insignia, text, symbols, etc.). If unit or organization names are present, provide details on their role, historical background, and activities. Place findings in an operational or historical context if verifiable from the image. Important: Do not generate assumptions or fabricate information. Only analyze what is visible in the image and cross-check visual elements before drawing conclusions.\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eDescription Conversion to JSON: The text-based descriptions generated by Gemini 1.5 Pro, LLAVA and OpenAI gpt-4o were then submitted to the OpenAI o3- mini-2025-01-31 model. This model was specifically instructed to convert the free-form text descriptions into a structured JSON format. The OpenAI model was accessed through the OpenAI API, using the default settings.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eData Storage in CSV Table: The resulting JSON objects were then extracted and stored in a CSV (Comma Separated Values) table. Each row in the table corresponded to an image in the dataset, and the columns included the image filename, the descriptions from Gemini 1.5 Pro, LLAVA and OpenAI gpt-4o, and the corresponding JSON fields (includes military, country, name, etc.) extracted by the OpenAI model. This structured CSV table served as the basis for subsequent analysis and evaluation.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eFigure 1 illustrates the experimental setup diagram.\u003c/p\u003e\u003cp\u003eFigure\u0026nbsp;1. Experimental setup diagram.\u003c/p\u003e\u003cp\u003eAgentic Variants (in addition to the baseline):\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eCrewAI (Collaborative): the orchestrator routed artifacts between sub-agents with short, role-specific prompts and strict JSON schemas. Agents could request clarifications from upstream outputs (max 2 turns) to reduce local errors.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLangGraph (Deterministic): the same sub-stages were implemented as a fixed DAG: INGEST/PROVENANCE \u0026rarr; PERCEPTION \u0026rarr; MANNEQUIN_DECIDER \u0026rarr; (if human) ATTRIBUTION \u0026rarr; CROSS_CHECK \u0026rarr; REPORT. Node I/O was validated against JSON schemas; failures triggered local retries without re-running the full graph.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eDecision Policy and Thresholds:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMannequin threshold (τ_human): if the MannequinAgent confidence for \u0026ldquo;human\u0026rdquo; \u0026lt; 0.6\u0026ndash;0.7, the system abstained from \u0026ldquo;military\u0026rdquo; declaration for that person.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAttribution evidence rule: country was emitted only when at least two independent signals agreed (e.g., OCR token\u0026thinsp;+\u0026thinsp;CLIP@top-1 match); otherwise, country=\"unknown\".\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eCross-check rules: contradictions (e.g., mannequin\u0026thinsp;=\u0026thinsp;=\u0026thinsp;true but military_present\u0026thinsp;=\u0026thinsp;=\u0026thinsp;true) forced abstention; provenance anomalies lowered confidence.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eImplementation Notes:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003ePerception/Attribution used GPT-4o; Mannequin/CrossCheck used lightweight model (GPT-4o-mini) plus deterministic heuristics.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAll intermediate artifacts were persisted (provenance.json, scene_graph.json, attribution.json, run.json) with SHA-256 hashes and timestamps to maintain chain-of-custody.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.4. Evaluation Metrics\u003c/h2\u003e\u003cp\u003eWe report baseline LLM-only metrics and agentic metrics to enable controlled comparison:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMilitary detection: precision, recall, accuracy.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMannequin false-positive rate (FPR): fraction of mannequin images incorrectly labeled as military.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAttribution accuracy: exact match for country and unit (when ground truth present).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Results","content":"\u003cp\u003eThis section presents the results of our experiments evaluating the performance of Google's Gemini 1.5 Pro, the open-source LLAVA and OpenAI gpt-4o model for identifying military personnel in images. We present both quantitative results, based on the evaluation metrics described in Section 3, and qualitative observations from analyzing the model outputs. We then discuss the implications of these results in the context of digital forensics and potential future directions.\u003c/p\u003e\u003cp\u003eIn addition to evaluating standalone multimodal LLMs, we further examined the performance of agentic orchestration frameworks - CrewAI and LangGraph - that integrate the same LLMs into structured, multi-agent pipelines. These extensions allow assessment of how coordinated reasoning and deterministic control influence accuracy, false positive rates, and forensic reproducibility.\u003c/p\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e4.1. Quantitative Results\u003c/h2\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the overall performance of the three models across the entire dataset:\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eOverall Model Performance.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMetric\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGemini 1.5 Pro\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLAVA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eOpenAI GPT-4o\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.69\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.69\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.99\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.91\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.793\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.793\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.770\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eIs Country Correct\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.787\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.121\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.385\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, all three models exhibited very high recall but only moderate precision (0.67\u0026ndash;0.69), confirming that they frequently misclassified mannequins as real personnel. Gemini 1.5 Pro and LLAVA achieved comparable overall accuracy (0.793), while GPT-4o trailed slightly (0.770). The imbalance between high recall and lower precision reflects over-sensitivity to uniform and equipment cues. Recall scores were similarly high: 0.99 f\u0026aring;or Gemini 1.5 Pro, 0.98 for LLAVA, and 0.91 for GPT-4o, indicating strong but varying abilities to detect most images that actually contained military personnel. However, the overall accuracy scores reveal more nuanced performance: GPT-4o achieved an accuracy of 0.77, slightly below Gemini\u0026rsquo;s 0.793 and LLAVA\u0026rsquo;s 0.793. This gap between precision/recall and overall accuracy is primarily due to the misclassification of mannequin images, which all models struggled to correctly identify, as discussed further below. The examples of images where the AI correctly identified a military person are shown on Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe models differed significantly in their ability to correctly identify associated attributes of the military personnel. Gemini 1.5 Pro demonstrated the strongest performance, achieving accuracy scores of 0.787 for Is Country Correct and 0.2544 for Is Unit Name Correct. LLAVA showed substantially lower accuracy on both metrics, with scores of 0.121 and 0.0051, respectively. GPT-4o exhibited intermediate performance, achieving 0.385 for Is Country Correct and 0.08 for Is Unit Name Correct. These results suggest that while Gemini 1.5 Pro currently provides the most reliable contextual understanding of military-related images, GPT-4o represents a notable improvement over LLAVA in attribute extraction and shows potential for further refinement. Example of images where AI correctly identified military unit insignia, patches, or flags provided on Fig.\u0026nbsp;3.\u003c/p\u003e\u003cp\u003eFigure\u0026nbsp;3. Correctly Recognized Affiliations \u0026ndash; A case where AI correctly identified military unit insignia, patches, or flags\u003c/p\u003e\u003cp\u003e.\u003c/p\u003e\u003cp\u003eTo evaluate the effect of agentic orchestration, we implemented two additional experimental conditions: (i) CrewAI (collaborative multi-agent reasoning) and (ii) LangGraph (deterministic orchestration). Their results are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eAgentic Framework Performance (integrated with GPT-4o backend).\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMetric\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eCrewAI\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLangGraph\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.90\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.74\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.73\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.842\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.865\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eIs Country Correct\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.62\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, both CrewAI and LangGraph significantly outperform the baseline models in terms of overall forensic accuracy and reduction of mannequin-induced false positives. The baseline precision, averaged across Gemini (0.69), LLAVA (0.69), and GPT-4o (0.67), was approximately 0.68. The CrewAI configuration improved detection precision from 0.68 (baseline) to 0.88, while LangGraph reached 0.90, achieving the highest overall performance. The false positive rate on mannequin images decreased from an average of 0.27 across LLMs to 0.12 for CrewAI and 0.08 for LangGraph. Country attribution accuracy rose to 0.58 and 0.62 respectively, reflecting improved semantic alignment between textual and visual cues.\u003c/p\u003e\u003cp\u003eThe CrewAI setup exhibited greater adaptability and richer qualitative reasoning due to inter-agent dialogue, whereas LangGraph demonstrated more stable and reproducible outputs across repeated runs. Both frameworks increased the abstain rate, intentionally deferring uncertain cases, which is desirable in forensic contexts where overconfidence can lead to misinterpretation.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003e4.2. Mannequin Misclassification\u003c/h2\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eA significant challenge encountered in our experiments was the misclassification of mannequin images as containing military personnel. Gemini 1.5 Pro misclassified 88 out of 99 mannequin images, while LLAVA misclassified 86 and 90 misclassified by OpenAI gpt-4o respectively. This suggests that all evaluated models struggle to differentiate between real individuals and mannequins, particularly when the mannequins are dressed in realistic military uniforms. False positive examples \u0026ndash; a mannequin misclassified as a soldier are demonstrated on Fig.\u0026nbsp;4.\u003c/p\u003e\u003cp\u003eFigure\u0026nbsp;4. True Positive Example \u0026ndash; An image where the AI correctly identified a military person\u003c/p\u003e\u003cp\u003eQualitative analysis of the model outputs revealed that the LLMs often focused on the uniform and equipment worn by the mannequins, without considering other factors such as facial features, skin texture, or body language that might indicate that the individual was not a real person.\u003c/p\u003e\u003cp\u003eIn contrast, the agentic pipelines substantially mitigated this issue. The dedicated MannequinAgent, using texture and pose heuristics combined with LLM verification, successfully filtered out 68\u0026ndash;74% of mannequin false positives. LangGraph performed best in this aspect due to its deterministic decision rules and schema validation, while CrewAI occasionally overrode conservative thresholds when inter-agent dialogue favored contextual reasoning.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e4.3. Qualitative Analysis\u003c/h2\u003e\u003cp\u003eIn addition to the quantitative results, we also conducted a qualitative analysis of the model outputs to better understand their strengths and limitations. Gemini 1.5 Pro consistently demonstrated higher accuracy in identifying the country of origin and the specific affiliation of military personnel, often correctly naming both. LLAVA, by contrast, frequently produced generic or vague responses - for instance, referring to a clearly identifiable U.S. Army soldier simply as a \"soldier.\"\u003c/p\u003e\u003cp\u003eGPT-4o exhibited mixed performance. While it occasionally matched Gemini in correctly identifying the country, it was less reliable overall and sometimes defaulted to general or partially correct answers. For example, in the case of a Ukrainian serviceman, GPT-4o correctly inferred the region but misidentified the exact affiliation or left the unit name unspecified.\u003c/p\u003e\u003cp\u003eAll three models struggled with images featuring occlusions, poor lighting, or non-standard poses. In such scenarios, they often failed to detect the presence of military personnel or generated inaccurate descriptions. Moreover, a shared tendency across the models was to overemphasize visual cues like weapons or uniforms, sometimes attributing military context even when these elements were not central or present in the image.\u003c/p\u003e\u003cp\u003eIn the agentic systems, qualitative inspection revealed improved interpretive discipline: decisions were consistently accompanied by evidence citations (e.g., \u0026ldquo;patch1.jpg\u0026thinsp;+\u0026thinsp;OCR token \u0026lsquo;MFU\u0026rsquo;\u0026rdquo;) and explicit abstentions in ambiguous cases. CrewAI\u0026rsquo;s narrative-style reasoning offered contextual depth but slight variance across runs, while LangGraph\u0026rsquo;s structured outputs produced concise, audit-ready reports ideal for forensic documentation.\u003c/p\u003e\u003c/div\u003e"},{"header":"5. Discussion","content":"\u003cp\u003eThe results of our experiments demonstrate the potential of multimodal LLMs for identifying military personnel in images, while also highlighting several key challenges. All three baseline models - Gemini 1.5 Pro, LLAVA, and GPT-4o - achieved high recall (0.91\u0026ndash;0.99) but only moderate precision (0.67\u0026ndash;0.69), indicating that they often over-predicted military presence when uncertain. This suggests that such models can serve as valuable tools for triaging large image datasets in digital forensic investigations but remain limited for autonomous decision-making.\u003c/p\u003e\u003cp\u003eThe moderate precision values are primarily explained by frequent mannequin misclassifications: Gemini misclassified 88 of 99 mannequin images, LLAVA 86, and GPT-4o 90. These findings confirm that current multimodal reasoning pipelines lack the ability to verify human authenticity and tend to rely on contextual cues such as uniforms, gear, and pose. Gemini 1.5 Pro demonstrated superior performance in identifying associated contextual attributes such as country and unit name, achieving 0.787 and 0.254 respectively. GPT-4o showed intermediate results (0.385 and 0.08), outperforming LLAVA (0.121 and 0.005). These variations likely stem from differences in model training data and multimodal alignment strategies.\u003c/p\u003e\u003cp\u003eDespite their strengths, all three models struggled with challenging visual conditions such as occlusion, poor lighting, and non-standard poses. Additionally, they exhibited a shared tendency to misinterpret mannequins and over-emphasize the presence of military cues like uniforms or weapons. The imbalance between high recall and limited precision suggests that baseline multimodal models behave as high-sensitivity detectors but lack discriminative rigor for forensic use, where false positives are more damaging than missed detections.\u003c/p\u003e\u003cp\u003eTo address these shortcomings, this study introduced agentic orchestration frameworks - CrewAI and LangGraph - that structured multimodal reasoning into verifiable stages. Both frameworks integrated GPT-4o as the underlying model but differed in coordination style: CrewAI facilitated adaptive inter-agent collaboration, while LangGraph enforced deterministic state transitions.\u003c/p\u003e\u003cp\u003eQuantitatively, CrewAI improved precision to 0.88 and reduced mannequin false-positive rate to 0.12, whereas LangGraph reached 0.90 precision and only 0.08 false positives. Although recall decreased (to 0.74 and 0.73 respectively) due to abstention on uncertain cases, this reduction reflects deliberate forensic conservatism rather than degraded detection ability. This validation pipeline converts potential hallucinations into deliberate abstentions, improving evidentiary defensibility and forensic trustworthiness while slightly reducing recall.\u003c/p\u003e\u003cp\u003eThe apparent drop in recall arises from explicit abstention and evidence-validation mechanisms. Each agentic system applies threshold-based logic - for example, the MannequinAgent filters low-confidence detections (τ\u0026thinsp;\u0026asymp;\u0026thinsp;0.6) and the CrossCheckAgent rejects inconsistent attributions - to prevent speculative conclusions. In forensic terms, this converts potential false positives into controlled \u0026ldquo;unknown\u0026rdquo; outputs, enhancing trust and auditability.\u003c/p\u003e\u003cp\u003eIn practical terms, this trade-off transforms raw accuracy metrics: although fewer total positives are reported, every reported classification is evidence-backed and reproducible. For digital forensics, this shift from recall-maximization to precision-and-traceability maximization represents a methodological advancement toward trustworthy AI analysis.\u003c/p\u003e\u003cp\u003eQualitative review further supports these gains. CrewAI generated narrative-style reasoning chains explaining each conclusion, while LangGraph produced standardized JSON reports with linked image patches and OCR tokens, suitable for chain-of-custody documentation. The reproducibility of LangGraph outputs makes it particularly well aligned with evidentiary standards such as NIST SP 800\u0026thinsp;\u0026minus;\u0026thinsp;101 and ISO/IEC 27037.\u003c/p\u003e\u003cp\u003eOverall, integrating agentic architectures reframes the role of multimodal LLMs in digital forensics - from opaque detectors to explainable, modular reasoning systems. This hybrid approach enhances analytical accuracy, ensures reproducibility, and strengthens the evidentiary defensibility of AI-assisted investigations.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors did impact to the work.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThis study was carried out with the financial support of the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan under Contract №388/PTF-24-26 dated 01.10.2024 under the scientific project IRN BR24993232 \u0026ldquo;Development of innovative technologies for conducting digital forensic investigations using intelligent software- hardware complexes\u0026rdquo;.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eZangana HM, Omar M (2025) Introduction to Digital Forensics and Artificial Intelligence. Digital Forensics in the Age of AI, edited by Marwan Omar and Hewa Majeed Zangana, IGI Global, pp. 1\u0026ndash;30. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.4018/979-8-3373-0857-9.ch001\u003c/span\u003e\u003cspan address=\"10.4018/979-8-3373-0857-9.ch001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKarthikeyan P, Pande HM, Sarveshwaran V (eds) (2023) Artificial Intelligence and Blockchain in Digital Forensics (1st ed.). River Publishers. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1201/9781003374671\u003c/span\u003e\u003cspan address=\"10.1201/9781003374671\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAI IN DIGITAL FORENSICS, IJSRMST, vol. 3, no. 5, pp. 01\u0026ndash;06 (2024) May \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.59828/ijsrmst.v3i5.208\u003c/span\u003e\u003cspan address=\"10.59828/ijsrmst.v3i5.208\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoses Ashawa A, Mansour J, Riley J, Osamor Nsikak Pius Owoh. Digital Forensics Challenges in Cyberspace: Overcoming Legitimacy and Privacy Issues Through Modularisation. Cloud Computing and Data Science [Internet]. 2023 Dec. 25 ];5(1):140\u0026thinsp;\u0026ndash;\u0026thinsp;56. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.37256/ccds.5120233845\u003c/span\u003e\u003cspan address=\"10.37256/ccds.5120233845\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDazzi P (2025) The internet of ai agents (iaia): A new frontier in networked and distributed intelligence. Int J Networked Distrib Comput 13(1):16. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s44227-025-00057-0\u003c/span\u003e\u003cspan address=\"10.1007/s44227-025-00057-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMishra P (2020) Big Data Digital Forensic and Cybersecurity. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1201/9781003024743-9\u003c/span\u003e\u003cspan address=\"10.1201/9781003024743-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJaved AR, Jalil Z, Zehra Wisha \u0026amp; Gadekallu, Thippa \u0026amp; Suh, Doug \u0026amp; Jalil Piran, Md. (2021). A comprehensive survey on digital video forensics: Taxonomy, challenges, and future directions. Eng Appl Artif Intell 106. 104456, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.engappai.2021.104456\u003c/span\u003e\u003cspan address=\"10.1016/j.engappai.2021.104456\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHao Tan and Mohit Bansal (2019) LXMERT: Learning Cross- Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5100\u0026ndash;5111, Hong Kong, China. Association for Computational Linguistics. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18653/v1/D19-1514\u003c/span\u003e\u003cspan address=\"10.18653/v1/D19-1514\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoustafa N (2022) Digital Forensics in the Era of Artificial Intelligence, 1st edn. CRC. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1201/9781003278962\u003c/span\u003e\u003cspan address=\"10.1201/9781003278962\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEmehin O, Emeteveke I, Adeyeye O, Akanbi I (2024) Generative AI in Forensic Data Analysis: Opportunities and Ethical Implications for Cloud-Based Investigations. Int J Res Publication Reviews 29412957. 6\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.55248/gengpi.5.1024.2904\u003c/span\u003e\u003cspan address=\"10.55248/gengpi.5.1024.2904\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSolanke A, Biasiotti M (2022) Digital Forensics AI: Evaluating, Standardizing and Optimizing Digital Evidence Mining Techniques. KI - K\u0026uuml;nstliche Intelligenz. 36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s13218-022-00763-9\u003c/span\u003e\u003cspan address=\"10.1007/s13218-022-00763-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePetko Georgiev VI, Lei R, Burnell L, Bai A, Gulati et al (2024) Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2403.05530\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2403.05530\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHaotian Liu C, Li Q, Wu YJ, Lee (2023) Visual Instruction Tuning. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2304.08485\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2304.08485\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIslam R, Moushi OM (2024) GPT-4o: The Cutting-Edge Advancement in Multimodal LLM. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.36227/techrxiv.171986596.65533294/v1\u003c/span\u003e\u003cspan address=\"10.36227/techrxiv.171986596.65533294/v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKirchner M, Gloe T (2015) Forensic Camera Model Identification. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/9781118705773.ch9\u003c/span\u003e\u003cspan address=\"10.1002/9781118705773.ch9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFiller Tom\u0026aacute;s, Fridrich J, Goljan M (2008) Using sensor pattern noise for camera model identification. Proceedings - International Conference on Image Processing, ICIP. 1296\u0026ndash;1299. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICIP.2008.4712000\u003c/span\u003e\u003cspan address=\"10.1109/ICIP.2008.4712000\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRadford, Alec \u0026amp; Kim, Jong \u0026amp; Hallacy, Chris \u0026amp; Ramesh, Aditya \u0026amp; Goh, Gabriel \u0026amp; Agarwal,Sandhini \u0026amp; Sastry, Girish \u0026amp; Askell, Amanda \u0026amp; Mishkin, Pamela \u0026amp; Clark, Jack \u0026amp; Krueger,Gretchen \u0026amp; Sutskever, Ilya. (2021). Learning Transferable Visual Models From Natural Language Supervision. https://doi.org/10.48550/arXiv.2103.00020.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYi Z, Xiao T, Albert MVA, Survey on Multimodal Large Language Models in Radiology for Report Generation and Visual Question Answering. Information 2025, 16, 136. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/info16020136\u003c/span\u003e\u003cspan address=\"10.3390/info16020136\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHe, Yingqing \u0026amp; Liu, Zhaoyang \u0026amp; Chen, Jingye \u0026amp; Zeyue, Tian \u0026amp; Liu, Hongyu \u0026amp; Chi, Xiaowei\u0026amp; Liu, Runtao \u0026amp; Yuan, Ruibin \u0026amp; Xing, Yazhou \u0026amp; Wang, Wenhai \u0026amp; Dai, Jifeng \u0026amp; Zhang,Yong \u0026amp; Xue, Wei \u0026amp; Liu, Qifeng \u0026amp; Guo, Yike \u0026amp; Chen, Qifeng. (2024). LLMs Meet Multimodal Generation and Editing: A Survey. https://doi.org/10.48550/arXiv.2405.19334.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eiMazing (2025) [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://imazing.com/\u003c/span\u003e\u003cspan address=\"https://imazing.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMykhaylova O et al (2024) Person-of-Interest Detection onMobile Forensics Data - AI-Driven Roadmap, in:Cybersecurity Providing in Information andTelecommunication Systems, vol. 3654 239\u0026ndash; 251\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLLaVA Large Language and Vision Assistant, 2025. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://llava-vl.github.io/\u003c/span\u003e\u003cspan address=\"https://llava-vl.github.io/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLangchain (2025) [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.langchain.com/\u003c/span\u003e\u003cspan address=\"https://www.langchain.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSarinova A, Neftissov A, Rzayeva L, Yessenov A, Kirichenko L, Kazambayev I, IMAGES PRELIMINARY PROCESSING METHOD FOR SUBSEQUENT RECOGNITION AND IDENTIFICATION OF VARIOUS OBJECTS (2024) Sci J Astana IT Univ 96\u0026ndash;106. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.37943/18BIAC9844\u003c/span\u003e\u003cspan address=\"10.37943/18BIAC9844\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. DEVELOPMENT OF AEROSPACE\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJin C, Wang J, Wei J, Tan L, Liu S, Zhao W, Lv X (2020) Multimedia analysis and fusion via Wasserstein Barycenter. Int J Networked Distrib Comput 8(2):58\u0026ndash;66. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2991/ijndc.k.200217.001\u003c/span\u003e\u003cspan address=\"10.2991/ijndc.k.200217.001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShahzad M, Rizvi S, Khan TA, Ahmad S, Ateya AA (2025) An exhaustive parametric analysis for securing sdn through traditional, AI/ML, and blockchain approaches: A systematic review. Int J Networked Distrib Comput 13(1):12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s44227-024-00055-8\u003c/span\u003e\u003cspan address=\"10.1007/s44227-024-00055-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCrewAI multi agent platform (2025) [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.crewai.com/\u003c/span\u003e\u003cspan address=\"https://www.crewai.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLangGraph (2025) [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.langchain.com/langgraph\u003c/span\u003e\u003cspan address=\"https://www.langchain.com/langgraph\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBillah M, Explainable, AI for Digital Forensics (2025) : Ensuring Transparency in Legal Evidence Analysis. J Forensic Sci Res. ; 9(2): 109\u0026ndash;116. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dx.doi.org/10.29328/journal.jfsr.1001089\u003c/span\u003e\u003cspan address=\"10.29328/journal.jfsr.1001089\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRanjan Sapkota KI, Roumeliotis M, Karkee, Part B (2026) 103599, ISSN 1566\u0026ndash;2535, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.inffus.2025.103599\u003c/span\u003e\u003cspan address=\"10.1016/j.inffus.2025.103599\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-networked-and-distributed-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [International Journal of Networked and Distributed Computing](https://link.springer.com/journal/44227)","snPcode":"44227","submissionUrl":"https://submission.springernature.com/new-submission/44227/3","title":"International Journal of Networked and Distributed Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Open","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Multimodal Large Language Models (LLMs), Artificial Intelligence in Forensics, Mobile Forensics, Automated Image Recognition, Agentic Architectures, CrewAI, LangGraph, Explainable AI (XAI), Forensic Automation and Orchestration","lastPublishedDoi":"10.21203/rs.3.rs-7372632/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7372632/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDigital forensics is vital for analyzing extensive image data from mobile devices to identify individuals and activities in investigations. Traditional methods struggle with complex real-world images, particularly distinguishing military personnel from military-themed mannequins. This study assesses multimodal Large Language Models (LLMs) - Google\u0026rsquo;s Gemini 1.5 Pro, open-source LLAVA, and GPT-4o - for detecting military personnel in 434 mobile device images, including military personnel, mannequins, and civilians. The models achieved strong recall (0.99 for Gemini, 0.98 for LLAVA, and 0.91 for GPT-4o) but only moderate precision (0.69, 0.69, and 0.67 respectively), reflecting a notable rate of mannequin-induced false positives.] Accuracy varied from 0.793 for Gemini and LLAVA to 0.770 for GPT-4o, aligning with observed differences in contextual understanding. Contextual classification also posed challenges: Gemini achieved 0.787 accuracy for country identification, followed by GPT-4o (0.385) and LLAVA (0.121). Unit name recognition remained weak across models. Misclassification of mannequins was the primary source of error, confirming that current multimodal models overemphasize uniform and equipment cues without verifying human authenticity. To enhance interpretability and reduce false positives, we integrated an agentic orchestration layer using CrewAI and LangGraph, which structured multimodal reasoning through dedicated sub-agents for provenance validation, perception, mannequin discrimination, and evidence-grounded attribution. These agentic frameworks substantially improved forensic reliability: CrewAI achieved 0.88 precision with a mannequin false-positive rate of 0.12, while LangGraph reached 0.90 precision and reduced false positives to 0.08. Country attribution accuracy rose to 0.58 and 0.62 respectively. Although recall decreased slightly due to conservative abstention logic (CrewAI 0.74, LangGraph 0.73), this trade-off yielded higher forensic confidence and reproducible, audit-ready decision traces. The results demonstrate that integrating agentic architectures transforms multimodal LLMs from opaque classifiers into transparent, evidence-driven forensic tools - enhancing both analytic precision and the evidentiary defensibility of AI-assisted investigations.\u003c/p\u003e","manuscriptTitle":"Evaluating Multimodal LLMs for Context-Aware Forensic Image Interpretation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-15 21:14:14","doi":"10.21203/rs.3.rs-7372632/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-11T13:54:53+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"107460267869932911622145920082116366907","date":"2026-05-09T14:17:05+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-18T17:13:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"309418418836503663272484909231268975322","date":"2026-04-15T10:22:36+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"102741948278723272058090352225074716372","date":"2026-04-01T11:31:06+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-17T17:14:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"273156789410990019524111325293237152004","date":"2025-12-13T18:39:53+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-12-08T18:26:33+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-12-08T03:16:31+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-10-31T05:46:43+00:00","index":"","fulltext":""},{"type":"submitted","content":"International Journal of Networked and Distributed Computing","date":"2025-10-29T19:12:17+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-networked-and-distributed-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [International Journal of Networked and Distributed Computing](https://link.springer.com/journal/44227)","snPcode":"44227","submissionUrl":"https://submission.springernature.com/new-submission/44227/3","title":"International Journal of Networked and Distributed Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Open","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ed02bbf1-c9a6-4cbf-b2ef-5eb9210e0c48","owner":[],"postedDate":"December 15th, 2025","published":true,"recentEditorialEvents":[{"type":"decision","content":"Accepted","date":"2026-05-11T20:39:54+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-11T13:54:53+00:00","index":50,"fulltext":""},{"type":"reviewerAgreed","content":"107460267869932911622145920082116366907","date":"2026-05-09T14:17:05+00:00","index":49,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-11T20:55:12+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-15 21:14:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7372632","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7372632","identity":"rs-7372632","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.