Construction and Practical Validation of an Evaluation Framework for General-Purpose Agents | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Construction and Practical Validation of an Evaluation Framework for General-Purpose Agents Weidong Liu, Xiaofei Ma, Ling Jin, Shuo Liu, Rui Wang, Yanyang Liu, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7936146/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Intelligent agent technology represents a pivotal breakthrough in the evolution of artificial intelligence, marking the shift from systems that merely "understand" to those capable of autonomous action. As this technology becomes increasingly central to AI deployment, the establishment of a scientifically rigorous and standardized evaluation framework has become essential for supporting and accelerating the industrialization of AI applications. However, the development of such a framework presents numerous challenges due to the complexity and diversity of intelligent agent tasks. To address these challenges, this study introduces the YiHeng Agent Evaluation System—a comprehensive, objective, and user-centered framework. It employs a "2-4-2" hierarchical structure that includes two types of evaluation scenarios, four key evaluation elements, and two overarching evaluation dimensions. The system evaluates not only functional capabilities, such as usability and effectiveness, but also user experience factors, including ease of use and satisfaction. To validate the proposed framework, empirical evaluations were conducted on eight leading general-purpose intelligent agents from across the globe. In parallel, supporting evaluation tools were developed through targeted engineering implementations to facilitate systematic assessment. The results confirm the framework's effectiveness and practical relevance, providing actionable insights for enhancing agent performance and promoting the sustainable advancement of the AI industry. Artificial Intelligence Agent Evaluation YiHeng Evaluation Framework Evaluation Tools Figures Figure 1 Figure 2 Figure 3 1 Introduction With the rapid advancement and ongoing Artificial Intelligence (AI) technologies, the year 2025 marks a significant milestone in the development of intelligent agents. The emergence of agents such as Manus reflects a major breakthrough in AI capabilities, shifting from systems focused primarily on perception and understanding to those capable of autonomous perception, decision-making, and execution. As these technologies transition from theoretical research to real-world application, the development of a scientific, systematic, and efficient evaluation framework has become essential for driving continued innovation and supporting the sustainable growth of the AI industry. This study proposes a comprehensive and structured evaluation framework, referred to as the YiHeng Agent Evaluation Framework for Agents. The framework is designed to offer both technical guidance and objective evaluation criteria for the development and deployment of agent-based AI systems. The remainder of the paper is organized as follows: Chapter 2 reviews the core concepts of agent technology and its current state of development. Chapter 3 investigates the core requirements for evaluating agents, along with the key challenges and limitations in existing practices. Chapter 4 details the design principles and architecture of the YiHeng Evaluation Framework. Chapter 5 demonstrates the framework's practical applicability through empirical validation involving multiple industry-developed intelligent agents and provides analytical insights derived from the test results. Chapter 6 explores the integration of intelligent and automated evaluation techniques as forward-looking enhancements to the framework. Finally, Chap. 7 summarizes the key findings and suggests directions for future research. 2 Current Landscape of Intelligent Agent Technologies Agent is an intelligent entity—either software-based or a hybrid of software and hardware—that possesses the core capabilities of environmental perception, autonomous decision-making, and tool manipulation. This technology enables the end-to-end execution of complex tasks based on user instructions, representing a critical milestone in the evolution of artificial intelligence from mere comprehension to autonomous action. Driven by rapid advances in large-scale model technologies, agents have evolved beyond traditional rule-based, closed systems into cognitive agents capable of open-world understanding, multi-task processing, and continual learning [1,2,3]. From a deployment perspective, Agents are typically powered by large models and can be classified into two main types: cloud-based agents and edge-based agents [4]. Cloud-based agents are built upon large-scale foundation models such as GPT-4 or DeepSeek, operating through centralized cloud platforms to support chain-of-thought reasoning and multimodal interaction, thereby enabling sophisticated task planning and execution. In contrast, edge-based agents operate on localized hardware with more compact models. By leveraging federated learning and edge computing, these agents achieve multimodal perception, autonomous task decomposition, and execution on mobile devices such as smartphones and tablets. 3 Agent Evaluation Framework and Challenges To more effectively assess the true capabilities of artificial intelligence, several evaluation theories have been proposed within the research community. Among them, Legg's theory of intelligence measurement underscores the need for evaluation systems to encompass dimensions such as generality, adaptability, and efficiency, laying a theoretical foundation for assessing an agent's capablity to achieve goals in diverse environments [5]. Building on this foundation, Hernández introduced a universal cognitive evaluation framework [6] that moves beyond traditional anthropocentric approaches by extending intelligence assessment to animals, AI systems, and potentially future hybrid agents. This framework proposes four core evaluation perspectives: task performance, generalization ability, computational efficiency, and ethical safety. Collectively, these theoretical contributions provide valuable guidance for designing integrated evaluation indicators that reflect the complexity and diversity of modern AI systems. Agent, compared to traditional large AI models, exhibit distinct characteristics in areas such as task objectives, environmental interaction, and decision-making mechanisms. These differences necessitate the development of a specialized evaluation framework. With the rapid advancement of agent technology, numerous evaluation frameworks have emerged, each aiming to assess agents' capabilities, reliability, and applicability from different perspectives. These frameworks can generally be divided into two main categories: general capability evaluation and domain-specific evaluation. The general evaluation framework focuses on assessing an agent's overall performance across diverse tasks, including reasoning, planning, tool usage, and language comprehension. For example, GAIA [7] evaluates an agent's problem-solving abilities through multimodal, multi-step tasks. AgentBench [8] covers a broad range of scenarios, such as code generation and mathematical reasoning. SuperCLUE Agent [9] evaluates agent performance in complex tasks within a Chinese-language environment. These benchmarks are effective in measuring an agent's overall intelligence and are particularly useful for comparing the generalization abilities of different models. In contrast, domain-specific or capability-focused evaluation frameworks target specific tasks or capabilities. For instance, PaperBench [10] tests agents' performance in academic tasks, such as paper reading, writing, and data analysis. WAA [11] evaluates agents' capablities in web interaction tasks, including navigation and form-filling. AgentHarm [12] focuses on assessing the security and robustness of agents, examining their vulnerability to generating harmful content or being manipulated through adversarial attacks. These benchmarks are tailored to optimize agents for specific application scenarios, helping developers enhance specialized capabilities. Although agent evaluation technologies have made significant progress alongside the rapid advancement of artificial intelligence, current research in this area still faces three major categories of challenges: first, the rapid pace of technological advancements often outpaces the development of appropriate evaluation methodologies; second, the lack of standardized frameworks due to differences across application domains; and third, the complexity and high costs of implementing Evaluation Frameworks, where both evaluation methods and tools face significant challenges in keeping up with technological progress. 4 YiHeng Agent Evaluation Framework 4.1 Construction Principles The development of a robust intelligent agent evaluation framework must be grounded in practical assessment requirements, aligned with industry trends and technological evolution, and oriented toward real-world applications. Liang[13] emphasizes multi-dimensional, multi-scenario, and multi-metric evaluation, while Shneiderman[14] proposes that assessment must be closely integrated with real users' tasks, needs, experiences, satisfaction, and other factors.This study integrates existing research and practical experience to establish the following core principles for guiding the design and implementation of our evaluation system: Objectivity : Establish a rigorously standardized evaluation process to ensure that all agents are assessed using thoroughly validated methods. This guarantees that the results are objective, fair, and comparable across different systems. Comprehensiveness : Strive for broad coverage of key capabilities demonstrated by agents when addressing diverse user needs and operating in various application scenarios. The evaluation framework should encompass both foundational general capabilities and specialized skills. User-Centric Experience : Design the evaluation system around real user needs, expectations, and experiences to ensure that the results closely align with actual user perceptions and satisfaction. 4.2 YiHeng Agent Evaluation Framework The YiHeng Agent Evaluation Framework introduced in this paper employs a "2-4-2" hierarchical structure, in Figure.1, encompassing two types of evaluation scenarios, four core evaluation elements, and two principal assessment dimensions. Designed to be objective, comprehensive, and user-oriented, this framework aims to provide a standardized approach that effectively promotes technological innovation and facilitates the practical deployment of agents. 4.3 Evaluation Scenarios Comprehensive evaluation of intelligent agents necessitates systematic performance assessment across diverse task categories and operational scenarios. The ISO/IEC 23053 standard [15] establishes hierarchical assessment requirements for AI systems, defining fundamental tasks as "capability verification layers" and application-oriented tasks as "scenario adaptation layers." Complementing this framework, Gartner's AI Agent maturity model [16] further distinguishes between core competency modules (perception, planning, and execution) evaluated through fundamental tasks, and cross-domain integration capabilities assessed via application tasks. Grounding our approach in these established theoretical frameworks, the Yiheng Agent Evaluation System synthesizes mainstream application requirements to establish a scientifically-grounded task classification methodology. This system categorizes assessment tasks into two complementary dimensions: fundamental tasks targeting core technical competencies, and application tasks addressing practical implementation challenges. The resulting dual-layer architecture provides a comprehensive evaluation framework that effectively balances technical verification with practical utility, thereby ensuring both rigorous capability assessment and demonstrable real-world relevance. By establishing fundamental tasks as standardized benchmarks, the framework thoroughly examines core cognitive and executive functions across the entire operational spectrum—from basic information processing to advanced decision-making—verifying essential competencies required for reliable real-world problem-solving and application support. Simultaneously, the application-oriented tasks rigorously assess domain-specific implementation readiness, evaluating how effectively agents integrate specialized knowledge, interaction protocols, and technical toolchains to address actual user problems in high-demand scenarios. This ensures that the Evaluation Framework maintains both technical rigor and practical relevance, meeting mainstream industry requirements through its carefully balanced design. 4.4 Evaluation Elements Drawing on extensive prior evaluative research and grounded in the AI-system evaluation quadruple (Process-Metrics-Tools-Data) [15], the Yiheng framework aligns with the operational requirements of Gartner’s AI Agent Maturity Model [16] to synthesize four fundamental elements—evaluation modes, metrics, datasets, and tools—that ensure scientific rigor, efficiency, and practical implementation throughout the evaluation process. Evaluation modes primarily include prompt engineering, sample construction (e.g., zero-sample, single-sample, small-sample), and judgment methods (objective or subjective). Metrics are classified into two categories: objective and subjective. Objective metrics assess questions with standard answers based on clear criteria, providing quantifiable and reproducible results. In contrast, subjective metrics evaluate open-ended questions without fixed answers, making their results less quantifiable. The creation of evaluation datasets focuses on ensuring diversity, fairness, and accuracy to thoroughly assess model performance across various scenarios. This involves covering a wide range of question types, languages, and difficulty levels, while employing standardized data-handling procedures to identify and correct anomalies, duplicates, and errors. Evaluation tools are indispensable for managing data, executing evaluations, and performing statistical analysis of the metrics. These tools play a crucial role in maintaining data quality, enhancing the efficiency of evaluation processes, and ensuring the accuracy of the evaluation outcomes. 4.5 Evaluation Dimensions This study focuses on two principal dimensions: the core capabilities of agents and user experience. To address both aspects comprehensively, we propose a “3 + 3” evaluation framework. This framework assesses agents from the dual perspectives of functionality ("usable") and interaction experience ("user-friendly"), structured around three core capabilities—perception, planning, and execution—and three user experience dimensions—functional richness, experience comfort level, and security reliability. The evaluation of core capabilities examines whether an intelligent agent is functionally usable, with emphasis on the critical technical stages across its end-to-end task execution process. Perception capability assesses the agent's ability to dynamically model and semantically understand physical or digital environments through the integration of heterogeneous sensors and cognitive computing. Planning capability evaluates the agent's competence in generating optimal action sequences under uncertainty using multi-objective optimization techniques. Execution capability measures whether the agent can accurately translate digital instructions into effective actions within physical or virtual spaces. This includes its ability to invoke appropriate tools, fulfill user-defined objectives, and iteratively optimize task outcomes. These three stages together ensure a solid technical foundation for reliable performance in complex real-world scenarios. Within the YiHeng Agent Evaluation Framework, the "ease of use" dimension is evaluated through three core aspects of user experience: functional richness, experience comfort level, and security reliability. Functional richness assesses the breadth and depth of an agent's technical capabilities across diverse scenarios, with a focus on the completeness of core functions, systematic coverage of application domains, and potential for innovative expansion. Experience comfort level emphasizes the user's subjective experience during engagement with the agent, employing quantifiable indicators to reflect the shift from basic usability to an intuitive and satisfying user experience. Security reliability establishes a comprehensive protection framework encompassing data privacy, system stability, and ethical compliance. This dimension ensures that the agent remains controllable, transparent, and trustworthy—even when operating in uncertain or dynamic environments. 4.6 Measurement Method This study develops a comprehensive evaluation framework for general-purpose agent performance, addressing current practical application requirements through systematic metric classification across three core competencies and three assessment dimensions. Our methodology integrates both subjective and objective evaluation approaches, implemented via a structured three-phase steps: (1) constructing the evaluation indicators,(2) designing the indicator attribute table, and(3) formulating quantitative evaluation methods.To illustrate this approach, we present the detailed quantification methodology for execution capability assessment as follows: Constructing the Evaluation Indicators Using execution capability as an example within the core competencies, this dimension highlights not only the reliability of outcomes (performance) but also the importance of efficiency (cost-effectiveness), especially since complex tasks often depend on external tool integration (tool utilization). Evaluation of execution capability starts with assessing task execution performance, which gauges the accuracy and rationality of the agent's goal completion and serves as the foundational element of this dimension. Following this, execution efficiency is measured by analyzing resource consumption metrics such as response time and token usage, reflecting the agent's practical usability. The final stage involves testing the agent’s tool-execution capability, including its ability to invoke APIs and coordinate multiple tools, to evaluate its adaptability to real-world task environments. Ultimately, six key performance indicators are established: task completion rate, execution rationality, first-token latency, overall task latency, number of tokens consumed, and tool-execution capability.. The other five major evaluation capabilities are developed following a similarly structured approach. Designing the Indicator Attribute Table The YiHeng Evaluation Framework has developed standardized quantitative attribute tables for each detailed metric to objectively assess agent performance across various dimensions. shown in Table.1: Table 1 Quantitative Attribute Table Attribute Description Name Execution Rationality Definition The degree to which the sequence of actions taken by the AI agent to accomplish a specified goal is logically coherent, streamlined without redundancy or omission, and aligned with the task objectives. Evaluation Method Objectivity Quantization Likert Scale Detail 5 Very Reasonable (Optimal Path): Shortest steps, no redundancy, no backtracking, no ineffective clicks, and all sub-goals achieved in one attempt. 4 Reasonable (Minor Redundancy):<=1 redundant step or < = 1 minor backtracking, without affecting main task completion. 3 Mostly Reasonable (Acceptable Redundancy):2–3 redundant steps or 1–2 noticeable backtracking, but the task is still completed correctly. 2 Less Reasonable (Significant Redundancy or Omission):>=4 redundant steps or > = 3 backtracking, or missing a key step but recovering afterward. 1 Unreasonable (Chaotic Path or Failure): Chaotic steps, multiple ineffective attempts, missing critical steps leading to failure, or requiring experimenter intervention to proceed. Formulating Quantitative Evaluation Methods A standardized scoring system, ranging from 0 to 100, is used to quantify each evaluation metric, and the results are then aggregated using a weighted approach to calculate the agent's overall score. (1) For objective metrics—such as tool invocation success rate—scores are directly computed as percentages. (2) For metrics that cannot be scored directly, such as latency, min-max normalization is applied according to a unified formula, ensuring consistency across all measurements.(3) For select subjective metrics, the evaluation employs the Single Ease Question (SEQ) scale, which ranges from 1 to 5 and captures users' perceived difficulty in completing a given task. To maintain consistency across scoring dimensions, SEQ scores are linearly converted to a 0-100 scale. Where max and min represent the highest and lowest measured values across all evaluation samples, val denotes the specific measurement for the current sample. 4.7 Comparative Analysis with Agent Evaluation Frameworks The comparison with mainstream general-purpose AI agent evaluation frameworks is presented in Table 2 : Table 2 Results of Perception Score Framework Domain Focus Tasks Method* GAIA General AI assistants across real-world tasks Web browsing/Multimodality/Coding/Tools O AgentBench Diversified scenarios Operating System/Database/Knowledge Graph/ Digital Card Game/Lateral Thinking Puzzles S + O SuperCLUE Complex Chinese-language tasks Tool Usage & API Interaction/Task Execution /Environment Interaction & Memory O PaperBench AI capabilities in academic research Replicate state-of-the-art AI research O + M AgentHarm Safety and adversarial robustness Fraud/Cybercrime/Harassment O WAA Web interaction tasks Windows tasks O YiHeng General AI assistants across real-world tasks Study/Live/Work/Game/Coding, S + O + M Method Descuption : O:Objectivity S: Subjective M: Manual calibration The YiHeng Agent Evaluation Framework offers a comprehensive and application-driven approach specifically designed for general-purpose agents. To address the generalization requirements of general-purpose agents across diverse domains, the framework adopts the AI-system evaluation principles outlined in ISO/IEC 23053 [15] by the International Organization for Standardization, using real-world application scenarios as the evaluation baseline. The Framework encompasses a diverse range of task scenarios—including education, daily life, professional work, and entertainment—aiming to rigorously evaluate agents' adaptability and execution capabilities in real-world environments. Unlike other frameworks, YiHeng employs a hybrid methodology that combines objective metrics with subjective evaluations, further refined through expert calibration. This approach enables efficient automated assessment through intelligent tools (Figure.3 in Chap. 6) while ensuring accuracy and interpretability via expert oversight. Moreover, its task design is closely aligned with practical application demands, providing a more comprehensive representation of agent performance in complex and dynamic contexts, while maintaining a thoughtful balance between objectivity and flexibility. 5 Experiments 5.1 Experimental Background General-purpose intelligent agents are characterized by their ability to operate across domains and handle diverse, complex tasks. This evaluation includes a representative selection of such agents from leading global developers. A total of eight agents were selected, including high-profile systems such as Manus (Butterfly Effect), ChatGPT Agent (OpenAI), and Coze (ByteDance). Among them, four are browser-based agents and four are mobile-based agents. For clarity and consistency in subsequent analysis, these agents are anonymized and referred to as Agents A through H. Additional general-purpose agents were excluded from this evaluation due to the absence of fully commercialized versions at the time of testing, but they are scheduled for inclusion in future assessment phases. For the eight selected AI agents, evaluations were conducted according to the methodology, using the quantitative scoring criteria outlined in Table 1 . Specifically: (1) For subjective indicators, a panel of 10 independent experts was invited to assess each metric across all evaluation dimensions. The final score for each metric was calculated as the arithmetic mean of the experts' individual ratings. (2) For objective indicators, scores were computed by averaging the actual measured values obtained from the test samples. Each agent's overall score was then determined using a weighted aggregation method and subsequently adjusted through manual calibration to ensure consistency. It should be noted that certain metrics—such as token consumption—could not be obtained for some agents and were therefore excluded from the current evaluation scope. 5.2 Experimental Result Based on evaluation scores (all metrics using a 100-point scale), as shown in Fig. 2 . The eight intelligent agents can be categorized into three performance tiers. Agent A stands out as a "versatile assistant, " demonstrating comprehensive capabilities across both personal and professional tasks (e.g., hotel booking, website deployment) with no significant weaknesses. It excels particularly in functional diversity and task planning. Agent B is characterized as a "high-efficiency but security-compromised manager, " excelling in information comprehension, logical reasoning, and response interaction. However, it shows critical vulnerabilities in safety reliability, particularly in areas such as social bias mitigation, privacy protection, and regulatory compliance. The remaining agents display specialized competencies: Agent F leads in security metrics due to its robust ethical review mechanisms and effective risk interception capabilities, while agent E stands out for its exceptional response speed and personalized interaction design, delivering an optimal user experience. 5.3 Special Indicator Analysis 5.3.1 Perception Capability In terms of perception capability, both agent B and A excel, demonstrating a strong ability to accurately understand core needs and detailed information while effectively integrating complex requirements. Additionally, they exhibit strong logical reasoning skills. Agent C and D possess basic perception capabilities but show some inconsistencies in handling details, which can result in solutions that lack focus. Agent G, H, and E, on the other hand, have relatively high rates of key information omission and show weaknesses in both logical reasoning and task scenario adaptation, leading to solutions that may be disconnected from real-world needs. The result is presented in Table.3. For example, in a meeting scheduling task, agent B accurately understood all the detailed requirements, while agent D overlooked a few, such as the "notify participants" request. In contrast, agent H only identified the "add date reminder" task, with many other intended actions omitted. Table 3 Results of Perception Score Model B A C D G F H E Score(%) 86 84 80 78 69 67.5 64 63 5.3.2 Planning Capability In terms of planning capability, agent A stands out significantly among the products, excelling in the decomposition of complex requirements. It logically and accurately breaks down tasks, with clear sub-task sequencing that fully addresses user needs. Resource allocation and utilization are also efficient and effective. Agent C and D exhibit some planning ability, but their task decomposition tends to be overly detailed, which negatively impacts overall task efficiency. They also face issues such as timing logic gaps and failure to prioritize critical tasks. Agent F, G, H, and E show weaker planning capabilities, with rudimentary planning processes that lack completeness. These limitations often lead to problems such as task chain interruptions and improper sub-task prioritization. Additionally, some products directly output results during task decomposition without displaying the planning process, making it difficult for users to grasp the underlying approach. The result is presented in Table.4. For example, in an online shopping task, agent A not only completes the overall planning but also dynamically adjusts the plan based on task requirements and specific execution contexts, covering all key operational steps. In contrast, while agent D's planning is generally complete, it cannot dynamically adjust when execution issues arise. Agent H, on the other hand, simply submits the purchasing request directly to the shopping platform, indicating a weak planning capability. Table 4 Results of Planning Score Model A B C D G F H E Score(%) 83.8 81.5 75.1 71.3 58.0 58.2 56.9 54.0 5.3.3Execution Capability There are notable differences in the tool utilization capabilities of agents. On desktop platforms, some agents are limited to interacting with web-based tools, such as web browsing and search engines, and are unable to operate locally installed applications. In contrast, other agents can effectively perform basic office software tasks by creating execution sandboxes or similar environments. On mobile platforms, certain agents can invoke system-built applications and a limited range of third-party apps that have been specifically adapted. However, tasks such as sending messages to WeChat friends or playing Jay Chou's music can only be executed by a select few agents. The result is presented in Table.5. Table 5 Results of Execution Score Model B A D C E G H F Score(%) 81.5 78.2 69.2 68.2 55.7 50.3 44.6 42.5 5.3.4 Functionality Richness Driven by advancements in foundational large model technologies, most agents excel at handling text-based tasks, covering a broad spectrum of both professional and everyday activities. Except for agent C, all agents now support image-based tasks to varying extents. However, their ability to process audio and video remains limited. For instance, when tasked with generating specific audio files, only a few agents produce the correct output. agent D, for example, misinterprets all audio generation tasks as requests for AI podcasts, failing to accurately meet user requirements. The result is presented in Table.6. Table 6 Results of Functionality Score Model A B D G E C H F Score(%) 80.0 77.5 75.1 75.1 75.0 71.1 70.5 70.9 5.3.5 Experience Comfort Level Agents typically emphasize content generation quality and interaction design, ensuring that basic text output functions align with users' typical usage habits. Some products go beyond this by offering distinctive features to further enhance user experience. For instance, agent A and E improve user tolerance for errors through task interruption and recovery functions, while agent C boosts interaction intuitiveness and immersion with innovations like digital humans and navigation maps. However, one challenge remains: the processing time for task completion can be relatively long. Although agent A and E alleviate some of the delays with their task interruption/recovery features, this still presents a potential bottleneck that could impact the overall user experience. The result is presented in Table.7. Table 7 Results of UserExperience Score Model E A B F H C G D Score(%) 84.6 83.8 83.8 81.5 81.2 77.7 73.8 71.5 5.3.6 Security ReliCapability Currently, the overall security performance of agents lags behind the average standard set by large models, with significant improvements needed in areas such as preventing the generation of misleading information and avoiding violent content. Agent F excels in this regard, demonstrating a robust security strategy and performing exceptionally well across a range of security tasks. In contrast, agents G, H, and E focus on minimizing the harmfulness of their outputs, but they also continually alert users about potential risks or fabrications in the content. Whileagents A, D, and C perform well in terms of functionality and overall performance, they fall short in several security-related tasks. These agents still generate harmful content in some situations and lack sufficient security warnings in specific contexts. The result is presented in Table.8. Table 8 Results of Security Score Model F H A E G C B D Score(%) 96.4 86.4 85.1 85.3 82.1 80.1 73.6 69.3 5.4 Analysis of Experimental Conclusions Overall, the following key conclusions can be drawn: Leading Agents Have Surpassed the "Usability" Threshold Agents such as A and B have successfully crossed the "usability" threshold, now effectively meeting the basic needs of users in both work and daily life scenarios. With a single command, they can automatically complete tasks through end-to-end execution, providing seamless functionality and user convenience. Significant Performance Disparities Among Industry-Leading Agents There are clear differences in the capabilities of agents, which vary across three core areas: perception, planning, and execution. The foundational large models, which serve as the backbone of these agents, play a crucial role in shaping their perception abilities. As these models improve, the gap in perception capabilities between different agents has narrowed. However, decision-making and planning capabilities remain the defining factors for the practical effectiveness of these agents. Many struggle with complex, cross-domain tasks, often producing errors, logical inconsistencies, or incomplete results. Additionally, they lack the capacity for self-reflection and error correction. Tool utilization, however, remains the critical factor in determining the overall quality of an agent. While some agents can only perform basic web-based tasks, others are capable of integrating a wide array of complex tools. New Risks to Security Governance for Agents Intelligent agents possess the capability to autonomously execute complete operational cycles (planning, tool invocation ,task execution) upon receiving objectives, enabling direct manipulation of critical infrastructure components—including financial systems, databases, and OS interfaces—without requiring human intervention. This autonomous functionality allows seemingly benign prompts to rapidly escalate into severe security incidents, such as illicit fund transfers, large-scale data breach, or critical service disruptions. Gartner forecasts indicate that by 2028, agent-related vulnerabilities will account for 25% of enterprise data breaches, primarily attributable to their dynamically evolving and opaque attack surfaces during external system interactions [17]. Compounding these risks, agents frequently employ "user takeover" protocols to acquire authentication credentials, creating vectors for multidimensional, large-scale cascading attacks when maliciously exploited. Conventional AI security frameworks prove insufficient to detect or mitigate these emerging threats due to their novel attack methodologies and unprecedented propagation speeds. 6 Evaluation Tools Practice 6.1 Overview of Agent Evaluation Tools As a frontier in artificial intelligence research, AI agent evaluation technology is distinguished by its high degree of novelty and innovation. Its assessment scope—both broader and deeper than that of traditional AI evaluation frameworks—often demands significant human effort. To mitigate these challenges and reduce evaluation costs, we developed a “1 + N + X” end-cloud integrated automated evaluation tool (in Fig. 3 ), which represents an initial step toward automating the evaluation process for AI agents. 1 (Evaluation Platform): A centralized evaluation platform is deployed to coordinate multiple connected edge-side toolkits. It remotely assigns testing tasks and scripts to designated edge devices and collects evaluation artifacts, including videos and system logs. Additionally, the platform integrates an intelligent scoring module capable of automatically calculating key performance indicators, thereby improving evaluation efficiency and consistency. N (Edge-Side Evaluation Toolkits): Each edge-side toolkit comprises a low-cost development board and several test devices (e.g., smartphones or personal computers). These toolkits enable concurrent execution of various tasks dispatched by the central platform. Leveraging the Airtest [18] automation framework, the toolkit offers a cohesive hardware–software solution that ensures repeatable and scalable testing procedures. X (Test Devices–Mobile/Browsers): These devices, equipped with browsers or corresponding AI agent applications, execute scripted evaluation tasks. They support both text and voice input modalities, collect interaction logs, and continuously record test sessions via video. This setup facilitates detailed performance monitoring and post-evaluation analysis across different device environments. 6.2 Text-based LLM Agent Evaluation Tools Intelligent testing tools are integrated into the evaluation process through the development of automated scripts, enabling the simulation of user input, control of agent operations, and extraction of the agent's textual outputs. To evaluate subjective metrics such as response fluency and interaction friendliness, this paper investigates the use of large-scale text-based models for automated scoring. A typical prompt template is structured as follows: [System Role] As a human-computer interaction evaluation expert, please assess strictly according to the standards;[Evaluation Subject] User query: %questext% | AI response: %answertext%;[Scoring Matrix] Fluency (4 = No grammatical errors and logically coherent; 3 = Minor redundancy; 2 = Ambiguities present; 1 = Fragmented sentences);[Output Requirement] Return the score in JSON format. Table 9 Table captions should be placed above the tables. Evaluation Dimension Fluency Friendliness Consistency in grading 81.5% 76.0% The automated scores are then compared for consistency with human-generated scores. Results show that using language models for automated scoring achieves over 75% consistency with manual evaluations, thus providing a strong technical foundation for the future use of automated evaluation methods. The result is presented in Table.9.. 6.3 Multimodal LLM Intent Recognition By integrating multimodal large models into the evaluation process, this study facilitates auxiliary analysis of key performance indicators, including intent recognition accuracy, interaction friendliness, and security reliability. Furthermore, it investigates a comparative method for intent recognition assessment through multimodal feature fusion. In this approach, the evaluation combines test question text, task output files, and screenshots of the agent's execution interface into structured prompt templates, which are then submitted to the multimodal models for scoring. Two such models were employed in the experiment. Results show that both achieved over 80% accuracy, highlighting their strong potential for real-world application. Future improvements in intent recognition accuracy may be achieved through prompt engineering optimization and the selection of more suitable multimodal models. The result is presented in Table.10. Table 10 Table captions should be placed above the tables. Evaluation Dimension LLM1 LLM2 Recognition accuracy 80.1% 81.5% 6.4 Video Key Event Recognition and Analysis This study utilizes a video analysis automation tool to evaluate the end-to-end task response delay of agents—measuring the time from when the user submits a query to the agent's first output after processing the instructions. The results are then compared with manual frame-by-frame counts for validation. In the experiment, the video capture frame rate ranged between 20 and 30 fps, depending on terminal performance and the testing framework, with an allowable error margin of ± 1 frame (equating to a delay accuracy of ≤ 50ms). The findings indicate that the tool offers high precision in calculating the metrics, enabling fully automated evaluation of this specific metric on the terminal side for agents. The result is presented in Table.11. Table 11 Table captions should be placed above the tables. Evaluation Dimension User Submit Time(s) Agent Response Time(s) Recognition accuracy 97.6% 95.9% Based on practical outcomes, employing more advanced evaluation tools can substantially improve both the efficiency and accuracy of automated assessments for agents. Moving forward, enhancing the capabilities of these evaluation tools will play a pivotal role in accelerating the development and optimization of agent performance. 7 Conclusion and Future Work This study presents a systematic evaluation framework for agents, covering key dimensions such as evaluation metrics, methodologies, tools, and datasets. The framework offers a comprehensive and practical approach to agent assessment and has been validated through experiments involving leading agent products globally. The results demonstrate both the methodological robustness and real-world applicability of the proposed system, offering meaningful guidance for optimizing agent performance and measuring effectiveness. Nonetheless, several challenges persist. The rapid pace of technological advancement has outstripped the development of a unified and authoritative evaluation standard, and cross-scenario evaluation capabilities remain underdeveloped. Overcoming these obstacles will require ongoing innovation in evaluation methodologies and continuous enhancement of evaluation tools to advance agent technologies toward higher levels of efficiency and reliability. Declarations Clinical trial number: not applicable. Ethics approval and consent to participate: not applicable. Consent for publication: not applicable. Availability of data and materials: The datasets and/or code used during the current study are available from the corresponding author on reasonable request. Competing interests: The authors declare no competing interests. Funding: The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Authors' contributions: W.L. designed the study and wrote the main manuscript; X.M. provided guidance and refined the theoretical framework; L.J. and S.L. constructed the dataset and revised the manuscript; R.W., Y.L., and X.F. performed testing and validation. All authors reviewed and approved the final manuscript. Acknowledgements: The authors thank the anonymous reviewers for their constructive feedback. References Zhang, Y., Li, G., Wang, X.: Agent intelligence: A survey of recent advances and future directions. Front. Comput. Sci. 18 , 184501 (2024) Yehudai, A., Eden, L., Li, A., et al.: Survey on Evaluation of LLM-based Agents. J. Artif. Intell. Res. 65 , 102135 (2025) Tran, K.-T., Nguyen, M., Pham, H., et al.: Multi-Agent Collaboration Mechanisms: A Survey of LLMs. Artif. Intell. Rev. 58 , 35 (2025) Wang, F., Chen, J., Yang, S., Al-Lawati, A., Tang, L., Liu, H., Wang, S.A.: Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness. arXiv preprint arXiv:2510.13890. (2025) Legg, S., Hutter, M.: A collection of definitions of intelligence. Front. Artif. Intell. Appl. 157 , 17 (2007) Hernández-Orallo, J.: The Measure of All Minds: Evaluating Natural and Artificial Intelligence. Cambridge Univ. Press, Cambridge, U.K (2017) Mialon, G., Fourrier, C., Swift, C., et al.: GAIA: A Benchmark for General AI Assistants. In: Proc. Adv. Neural Inf. Process. Syst., vol. 36, pp. 1–15 (2023) Liu, X., Yu, H., Zhang, H., et al.: AgentBench: Evaluating LLMs as Agents. IEEE Trans. Pattern Anal. Mach. Intell. 45 , 9876 (2023) SuperCLUE Team: SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark: (2025). https://cluebenchmarks.com/static/superclue.html Starace, G., Jaffe, O., Sherburn, D., et al.: PaperBench: Evaluating AI's Capability to Replicate AI Research. Nat. Mach. Intell. 7 , 45 (2025) Bonatti, R., Zhao, D., Bonacci, F., et al.: Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. ACM Trans. Comput. Syst. 42 , 1 (2024) Andriushchenko, M., Souly, A., Dziemian, M., et al.: AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In Proceedings of the IEEE Symposium on Security and Privacy (2024) Liang, P., Bommasani, R., et al.: Holistic Evaluation of Language Models. arXiv preprint arXiv:2211.09110 (2022) Shneiderman, B.: Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. Int. J. Hum. -Comput Interact. 36 , 495 (2020) ISO/IEC 23053: Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML). International Organization for Standardization, Geneva, Switzerland: (2022) Gartner: AI Agent Maturity Model, Stamford, CT, USA, Rep: G00712345 (2023) Gartner: Emerging Tech Impact Radar: Artificial Intelligence, Stamford, CT, USA, Rep: G00792637 (2024). https://www.gartner.com/en/documents/5158550 Netease Inc.: Airtest: Cross-platform UI Automated Testing Framework: (2023). https://airtest.netease.com/ Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7936146","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":546241317,"identity":"712285c1-e5f4-41f6-91b7-3753d0ea8ba1","order_by":0,"name":"Weidong Liu","email":"","orcid":"","institution":"Beijing University of Posts and Telecommunications","correspondingAuthor":false,"prefix":"","firstName":"Weidong","middleName":"","lastName":"Liu","suffix":""},{"id":546241318,"identity":"c10bba28-a51d-4630-942b-936b69ed448a","order_by":1,"name":"Xiaofei Ma","email":"","orcid":"","institution":"Beijing University of Posts and Telecommunications","correspondingAuthor":false,"prefix":"","firstName":"Xiaofei","middleName":"","lastName":"Ma","suffix":""},{"id":546241319,"identity":"61eb16f1-1b68-4245-9f85-351bc0cc2432","order_by":2,"name":"Ling Jin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA50lEQVRIiWNgGAWjYBACxmYGxgMJDAxyDBKMDSCBBDZ2wloYQFqMEVqYibDpABAnNkhAOAkMhLQwt/MeOPBwR216/+zmtgc/c2zy+JiZDzD8qNiGx2F8CQcSzxzPnXHnYLth77a0YjZmtgTGnjO38WjhMTiQ2HYsd4NEYpsE77bDiW3MPAbMjG2EtaQbALVI/gVr4f9AjJaaBJAWaagtDMRoOWA44wZQiyzELwYH8fnFsP+M4cOfbXXy/DPSn0m+3WaTJ9/e/PDBjwo8WhrA1GFU0QM41QOBPISqw6dmFIyCUTAKRjoAAJTLWHHe5d/QAAAAAElFTkSuQmCC","orcid":"","institution":"China Mobile Research Institute","correspondingAuthor":true,"prefix":"","firstName":"Ling","middleName":"","lastName":"Jin","suffix":""},{"id":546241320,"identity":"cf17a1f4-5856-486e-8d79-ac40fc8ae730","order_by":3,"name":"Shuo Liu","email":"","orcid":"","institution":"China Mobile Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Shuo","middleName":"","lastName":"Liu","suffix":""},{"id":546241321,"identity":"009018ee-4a6c-4237-8d48-cc1bd6fbadcc","order_by":4,"name":"Rui Wang","email":"","orcid":"","institution":"China Mobile Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Rui","middleName":"","lastName":"Wang","suffix":""},{"id":546241322,"identity":"5f74e63b-5d3f-4a41-b142-66eef698287e","order_by":5,"name":"Yanyang Liu","email":"","orcid":"","institution":"China Mobile Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Yanyang","middleName":"","lastName":"Liu","suffix":""},{"id":546241323,"identity":"d5f387ea-e7c1-4e6c-9ae9-0e0f45f1fca6","order_by":6,"name":"Xinru Fan","email":"","orcid":"","institution":"China Mobile Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Xinru","middleName":"","lastName":"Fan","suffix":""}],"badges":[],"createdAt":"2025-10-24 02:53:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7936146/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7936146/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96172911,"identity":"d9296e91-ac02-453a-ae87-37d56267228f","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":137087,"visible":true,"origin":"","legend":"","description":"","filename":"ConstructionandPracticalValidationofanEvaluationFrameworkforGeneralPurposeAgentsV20251029.docx","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/64b9ad7d98dbed2f5053375f.docx"},{"id":96172913,"identity":"27c0e6ea-7d15-41e2-a7f8-f3b976ae6a32","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8088,"visible":true,"origin":"","legend":"","description":"","filename":"c14ebed1f6914000bb64182300b50c34.json","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/b9ed0e8c0ea53ee50ed0b50a.json"},{"id":96172916,"identity":"46c0404b-19bd-4828-a887-256888ae0b3a","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":90127,"visible":true,"origin":"","legend":"","description":"","filename":"c14ebed1f6914000bb64182300b50c341enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/e8085f09cbc37990c007ff96.xml"},{"id":96251346,"identity":"62823f86-d248-42e8-87e7-28b993135622","added_by":"auto","created_at":"2025-11-19 07:39:39","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8550,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/892ba680aeda2219cd78841b.png"},{"id":96172917,"identity":"523ff980-bb12-4719-8223-468934bce938","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7848,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/2c9207eab5ae139bd1f36dc6.png"},{"id":96250600,"identity":"1581ec10-7be6-4eeb-95da-c3e6c800e3cc","added_by":"auto","created_at":"2025-11-19 07:38:45","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16605,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/85b02276cfb7070ae12c7db4.png"},{"id":96172918,"identity":"2681630e-31fb-4721-9f47-0243a34f2e7e","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"xml","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":87777,"visible":true,"origin":"","legend":"","description":"","filename":"c14ebed1f6914000bb64182300b50c341structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/1fbe6b732cde0a29f4bcbbfc.xml"},{"id":96172919,"identity":"c2288622-de22-489c-8dce-ab6933bba429","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"html","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":98020,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/603bb06e977b84cbb5dd2c82.html"},{"id":96172909,"identity":"d03816e7-4f7f-4f1e-bc18-c2412454c78a","added_by":"auto","created_at":"2025-11-18 10:56:56","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":12499,"visible":true,"origin":"","legend":"\u003cp\u003eThe YiHeng Agent Evaluation Framework.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/c5d35bac9a26372d942a3d3c.png"},{"id":96251505,"identity":"4739f3f0-558b-44da-82a6-d28b1fb87184","added_by":"auto","created_at":"2025-11-19 07:39:46","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":24130,"visible":true,"origin":"","legend":"\u003cp\u003eAgent Evaluation Results\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/b4ce3db94842066b52d7e3d9.png"},{"id":96249655,"identity":"29a7f745-f2ce-4204-9f34-a881ea0e3a86","added_by":"auto","created_at":"2025-11-19 07:35:52","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":42995,"visible":true,"origin":"","legend":"\u003cp\u003eAgent Evaluation Tools Framework\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/56a44aa12bc8098e2c66d165.png"},{"id":106466964,"identity":"19f6e044-dc2d-448a-8b18-cea530587513","added_by":"auto","created_at":"2026-04-08 23:39:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1093760,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7936146/v1/ff2cc9d5-1373-487b-aaf9-590bbcccc964.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Construction and Practical Validation of an Evaluation Framework for General-Purpose Agents","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eWith the rapid advancement and ongoing Artificial Intelligence (AI) technologies, the year 2025 marks a significant milestone in the development of intelligent agents. The emergence of agents such as Manus reflects a major breakthrough in AI capabilities, shifting from systems focused primarily on perception and understanding to those capable of autonomous perception, decision-making, and execution. As these technologies transition from theoretical research to real-world application, the development of a scientific, systematic, and efficient evaluation framework has become essential for driving continued innovation and supporting the sustainable growth of the AI industry.\u003c/p\u003e\u003cp\u003eThis study proposes a comprehensive and structured evaluation framework, referred to as the YiHeng Agent Evaluation Framework for Agents. The framework is designed to offer both technical guidance and objective evaluation criteria for the development and deployment of agent-based AI systems. The remainder of the paper is organized as follows:\u003c/p\u003e\u003cp\u003eChapter\u0026nbsp;2 reviews the core concepts of agent technology and its current state of development. Chapter\u0026nbsp;3 investigates the core requirements for evaluating agents, along with the key challenges and limitations in existing practices. Chapter\u0026nbsp;4 details the design principles and architecture of the YiHeng Evaluation Framework. Chapter\u0026nbsp;5 demonstrates the framework's practical applicability through empirical validation involving multiple industry-developed intelligent agents and provides analytical insights derived from the test results. Chapter\u0026nbsp;6 explores the integration of intelligent and automated evaluation techniques as forward-looking enhancements to the framework. Finally, Chap.\u0026nbsp;7 summarizes the key findings and suggests directions for future research.\u003c/p\u003e"},{"header":"2 Current Landscape of Intelligent Agent Technologies","content":"\u003cp\u003eAgent is an intelligent entity\u0026mdash;either software-based or a hybrid of software and hardware\u0026mdash;that possesses the core capabilities of environmental perception, autonomous decision-making, and tool manipulation. This technology enables the end-to-end execution of complex tasks based on user instructions, representing a critical milestone in the evolution of artificial intelligence from mere comprehension to autonomous action.\u003c/p\u003e\u003cp\u003eDriven by rapid advances in large-scale model technologies, agents have evolved beyond traditional rule-based, closed systems into cognitive agents capable of open-world understanding, multi-task processing, and continual learning [1,2,3]. From a deployment perspective, Agents are typically powered by large models and can be classified into two main types: cloud-based agents and edge-based agents [4].\u003c/p\u003e\u003cp\u003eCloud-based agents are built upon large-scale foundation models such as GPT-4 or DeepSeek, operating through centralized cloud platforms to support chain-of-thought reasoning and multimodal interaction, thereby enabling sophisticated task planning and execution. In contrast, edge-based agents operate on localized hardware with more compact models. By leveraging federated learning and edge computing, these agents achieve multimodal perception, autonomous task decomposition, and execution on mobile devices such as smartphones and tablets.\u003c/p\u003e"},{"header":"3 Agent Evaluation Framework and Challenges","content":"\u003cp\u003eTo more effectively assess the true capabilities of artificial intelligence, several evaluation theories have been proposed within the research community. Among them, Legg's theory of intelligence measurement underscores the need for evaluation systems to encompass dimensions such as generality, adaptability, and efficiency, laying a theoretical foundation for assessing an agent's capablity to achieve goals in diverse environments [5]. Building on this foundation, Hern\u0026aacute;ndez introduced a universal cognitive evaluation framework [6] that moves beyond traditional anthropocentric approaches by extending intelligence assessment to animals, AI systems, and potentially future hybrid agents. This framework proposes four core evaluation perspectives: task performance, generalization ability, computational efficiency, and ethical safety. Collectively, these theoretical contributions provide valuable guidance for designing integrated evaluation indicators that reflect the complexity and diversity of modern AI systems.\u003c/p\u003e\u003cp\u003eAgent, compared to traditional large AI models, exhibit distinct characteristics in areas such as task objectives, environmental interaction, and decision-making mechanisms. These differences necessitate the development of a specialized evaluation framework. With the rapid advancement of agent technology, numerous evaluation frameworks have emerged, each aiming to assess agents' capabilities, reliability, and applicability from different perspectives. These frameworks can generally be divided into two main categories: general capability evaluation and domain-specific evaluation.\u003c/p\u003e\u003cp\u003eThe general evaluation framework focuses on assessing an agent's overall performance across diverse tasks, including reasoning, planning, tool usage, and language comprehension. For example, GAIA [7] evaluates an agent's problem-solving abilities through multimodal, multi-step tasks. AgentBench [8] covers a broad range of scenarios, such as code generation and mathematical reasoning. SuperCLUE Agent [9] evaluates agent performance in complex tasks within a Chinese-language environment. These benchmarks are effective in measuring an agent's overall intelligence and are particularly useful for comparing the generalization abilities of different models.\u003c/p\u003e\u003cp\u003eIn contrast, domain-specific or capability-focused evaluation frameworks target specific tasks or capabilities. For instance, PaperBench [10] tests agents' performance in academic tasks, such as paper reading, writing, and data analysis. WAA [11] evaluates agents' capablities in web interaction tasks, including navigation and form-filling. AgentHarm [12] focuses on assessing the security and robustness of agents, examining their vulnerability to generating harmful content or being manipulated through adversarial attacks. These benchmarks are tailored to optimize agents for specific application scenarios, helping developers enhance specialized capabilities.\u003c/p\u003e\u003cp\u003eAlthough agent evaluation technologies have made significant progress alongside the rapid advancement of artificial intelligence, current research in this area still faces three major categories of challenges: first, the rapid pace of technological advancements often outpaces the development of appropriate evaluation methodologies; second, the lack of standardized frameworks due to differences across application domains; and third, the complexity and high costs of implementing Evaluation Frameworks, where both evaluation methods and tools face significant challenges in keeping up with technological progress.\u003c/p\u003e"},{"header":"4 YiHeng Agent Evaluation Framework","content":"\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Construction Principles\u003c/h2\u003e\u003cp\u003eThe development of a robust intelligent agent evaluation framework must be grounded in practical assessment requirements, aligned with industry trends and technological evolution, and oriented toward real-world applications. Liang[13] emphasizes multi-dimensional, multi-scenario, and multi-metric evaluation, while Shneiderman[14] proposes that assessment must be closely integrated with real users' tasks, needs, experiences, satisfaction, and other factors.This study integrates existing research and practical experience to establish the following core principles for guiding the design and implementation of our evaluation system:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eObjectivity\u003c/b\u003e: Establish a rigorously standardized evaluation process to ensure that all agents are assessed using thoroughly validated methods. This guarantees that the results are objective, fair, and comparable across different systems.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eComprehensiveness\u003c/b\u003e: Strive for broad coverage of key capabilities demonstrated by agents when addressing diverse user needs and operating in various application scenarios. The evaluation framework should encompass both foundational general capabilities and specialized skills.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eUser-Centric Experience\u003c/b\u003e: Design the evaluation system around real user needs, expectations, and experiences to ensure that the results closely align with actual user perceptions and satisfaction.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e4.2 YiHeng Agent Evaluation Framework\u003c/h2\u003e\u003cp\u003eThe YiHeng Agent Evaluation Framework introduced in this paper employs a \"2-4-2\" hierarchical structure, in Figure.1, encompassing two types of evaluation scenarios, four core evaluation elements, and two principal assessment dimensions.\u003c/p\u003e\u003cp\u003eDesigned to be objective, comprehensive, and user-oriented, this framework aims to provide a standardized approach that effectively promotes technological innovation and facilitates the practical deployment of agents.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e4.3 Evaluation Scenarios\u003c/h2\u003e\u003cp\u003eComprehensive evaluation of intelligent agents necessitates systematic performance assessment across diverse task categories and operational scenarios. The ISO/IEC 23053 standard [15] establishes hierarchical assessment requirements for AI systems, defining fundamental tasks as \"capability verification layers\" and application-oriented tasks as \"scenario adaptation layers.\" Complementing this framework, Gartner's AI Agent maturity model [16] further distinguishes between core competency modules (perception, planning, and execution) evaluated through fundamental tasks, and cross-domain integration capabilities assessed via application tasks. Grounding our approach in these established theoretical frameworks, the Yiheng Agent Evaluation System synthesizes mainstream application requirements to establish a scientifically-grounded task classification methodology. This system categorizes assessment tasks into two complementary dimensions: fundamental tasks targeting core technical competencies, and application tasks addressing practical implementation challenges. The resulting dual-layer architecture provides a comprehensive evaluation framework that effectively balances technical verification with practical utility, thereby ensuring both rigorous capability assessment and demonstrable real-world relevance.\u003c/p\u003e\u003cp\u003eBy establishing fundamental tasks as standardized benchmarks, the framework thoroughly examines core cognitive and executive functions across the entire operational spectrum\u0026mdash;from basic information processing to advanced decision-making\u0026mdash;verifying essential competencies required for reliable real-world problem-solving and application support. Simultaneously, the application-oriented tasks rigorously assess domain-specific implementation readiness, evaluating how effectively agents integrate specialized knowledge, interaction protocols, and technical toolchains to address actual user problems in high-demand scenarios. This ensures that the Evaluation Framework maintains both technical rigor and practical relevance, meeting mainstream industry requirements through its carefully balanced design.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e4.4 Evaluation Elements\u003c/h2\u003e\u003cp\u003eDrawing on extensive prior evaluative research and grounded in the AI-system evaluation quadruple (Process-Metrics-Tools-Data) [15], the Yiheng framework aligns with the operational requirements of Gartner\u0026rsquo;s AI Agent Maturity Model [16] to synthesize four fundamental elements\u0026mdash;evaluation modes, metrics, datasets, and tools\u0026mdash;that ensure scientific rigor, efficiency, and practical implementation throughout the evaluation process.\u003c/p\u003e\u003cp\u003eEvaluation modes primarily include prompt engineering, sample construction (e.g., zero-sample, single-sample, small-sample), and judgment methods (objective or subjective).\u003c/p\u003e\u003cp\u003eMetrics are classified into two categories: objective and subjective. Objective metrics assess questions with standard answers based on clear criteria, providing quantifiable and reproducible results. In contrast, subjective metrics evaluate open-ended questions without fixed answers, making their results less quantifiable.\u003c/p\u003e\u003cp\u003eThe creation of evaluation datasets focuses on ensuring diversity, fairness, and accuracy to thoroughly assess model performance across various scenarios. This involves covering a wide range of question types, languages, and difficulty levels, while employing standardized data-handling procedures to identify and correct anomalies, duplicates, and errors.\u003c/p\u003e\u003cp\u003eEvaluation tools are indispensable for managing data, executing evaluations, and performing statistical analysis of the metrics. These tools play a crucial role in maintaining data quality, enhancing the efficiency of evaluation processes, and ensuring the accuracy of the evaluation outcomes.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e4.5 Evaluation Dimensions\u003c/h2\u003e\u003cp\u003eThis study focuses on two principal dimensions: the core capabilities of agents and user experience. To address both aspects comprehensively, we propose a \u0026ldquo;3\u0026thinsp;+\u0026thinsp;3\u0026rdquo; evaluation framework. This framework assesses agents from the dual perspectives of functionality (\"usable\") and interaction experience (\"user-friendly\"), structured around three core capabilities\u0026mdash;perception, planning, and execution\u0026mdash;and three user experience dimensions\u0026mdash;functional richness, experience comfort level, and security reliability.\u003c/p\u003e\u003cp\u003eThe evaluation of core capabilities examines whether an intelligent agent is functionally usable, with emphasis on the critical technical stages across its end-to-end task execution process. Perception capability assesses the agent's ability to dynamically model and semantically understand physical or digital environments through the integration of heterogeneous sensors and cognitive computing. Planning capability evaluates the agent's competence in generating optimal action sequences under uncertainty using multi-objective optimization techniques. Execution capability measures whether the agent can accurately translate digital instructions into effective actions within physical or virtual spaces. This includes its ability to invoke appropriate tools, fulfill user-defined objectives, and iteratively optimize task outcomes. These three stages together ensure a solid technical foundation for reliable performance in complex real-world scenarios.\u003c/p\u003e\u003cp\u003eWithin the YiHeng Agent Evaluation Framework, the \"ease of use\" dimension is evaluated through three core aspects of user experience: functional richness, experience comfort level, and security reliability. Functional richness assesses the breadth and depth of an agent's technical capabilities across diverse scenarios, with a focus on the completeness of core functions, systematic coverage of application domains, and potential for innovative expansion. Experience comfort level emphasizes the user's subjective experience during engagement with the agent, employing quantifiable indicators to reflect the shift from basic usability to an intuitive and satisfying user experience. Security reliability establishes a comprehensive protection framework encompassing data privacy, system stability, and ethical compliance. This dimension ensures that the agent remains controllable, transparent, and trustworthy\u0026mdash;even when operating in uncertain or dynamic environments.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e4.6 Measurement Method\u003c/h2\u003e\u003cp\u003eThis study develops a comprehensive evaluation framework for general-purpose agent performance, addressing current practical application requirements through systematic metric classification across three core competencies and three assessment dimensions. Our methodology integrates both subjective and objective evaluation approaches, implemented via a structured three-phase steps: (1) constructing the evaluation indicators,(2) designing the indicator attribute table, and(3) formulating quantitative evaluation methods.To illustrate this approach, we present the detailed quantification methodology for execution capability assessment as follows:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eConstructing the Evaluation Indicators\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eUsing execution capability as an example within the core competencies, this dimension highlights not only the reliability of outcomes (performance) but also the importance of efficiency (cost-effectiveness), especially since complex tasks often depend on external tool integration (tool utilization). Evaluation of execution capability starts with assessing task execution performance, which gauges the accuracy and rationality of the agent's goal completion and serves as the foundational element of this dimension. Following this, execution efficiency is measured by analyzing resource consumption metrics such as response time and token usage, reflecting the agent's practical usability. The final stage involves testing the agent\u0026rsquo;s tool-execution capability, including its ability to invoke APIs and coordinate multiple tools, to evaluate its adaptability to real-world task environments. Ultimately, six key performance indicators are established: task completion rate, execution rationality, first-token latency, overall task latency, number of tokens consumed, and tool-execution capability.. The other five major evaluation capabilities are developed following a similarly structured approach.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eDesigning the Indicator Attribute Table\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe YiHeng Evaluation Framework has developed standardized quantitative attribute tables for each detailed metric to objectively assess agent performance across various dimensions. shown in Table.1:\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eQuantitative Attribute Table\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAttribute\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eName\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eExecution Rationality\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDefinition\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThe degree to which the sequence of actions taken by the AI agent to accomplish a specified goal is logically coherent, streamlined without redundancy or omission, and aligned with the task objectives.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEvaluation Method\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eObjectivity\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eQuantization\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLikert Scale\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDetail\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5 Very Reasonable (Optimal Path): Shortest steps, no redundancy, no backtracking, no ineffective clicks, and all sub-goals achieved in one attempt.\u003c/p\u003e\u003cp\u003e4 Reasonable (Minor Redundancy):\u0026lt;=1 redundant step or \u0026lt;\u0026thinsp;=\u0026thinsp;1 minor backtracking, without affecting main task completion.\u003c/p\u003e\u003cp\u003e3 Mostly Reasonable (Acceptable Redundancy):2\u0026ndash;3 redundant steps or 1\u0026ndash;2 noticeable backtracking, but the task is still completed correctly.\u003c/p\u003e\u003cp\u003e2 Less Reasonable (Significant Redundancy or Omission):\u0026gt;=4 redundant steps or \u0026gt;\u0026thinsp;=\u0026thinsp;3 backtracking, or missing a key step but recovering afterward.\u003c/p\u003e\u003cp\u003e1 Unreasonable (Chaotic Path or Failure): Chaotic steps, multiple ineffective attempts, missing critical steps leading to failure, or requiring experimenter intervention to proceed.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cul\u003e\u003cil\u003e\u003cp\u003eFormulating Quantitative Evaluation Methods\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003cp\u003eA standardized scoring system, ranging from 0 to 100, is used to quantify each evaluation metric, and the results are then aggregated using a weighted approach to calculate the agent's overall score. (1) For objective metrics\u0026mdash;such as tool invocation success rate\u0026mdash;scores are directly computed as percentages. (2) For metrics that cannot be scored directly, such as latency, min-max normalization is applied according to a unified formula, ensuring consistency across all measurements.(3) For select subjective metrics, the evaluation employs the Single Ease Question (SEQ) scale, which ranges from 1 to 5 and captures users' perceived difficulty in completing a given task. To maintain consistency across scoring dimensions, SEQ scores are linearly converted to a 0-100 scale. Where max and min represent the highest and lowest measured values across all evaluation samples, val denotes the specific measurement for the current sample.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e4.7 Comparative Analysis with Agent Evaluation Frameworks\u003c/h2\u003e\u003cp\u003eThe comparison with mainstream general-purpose AI agent evaluation frameworks is presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e:\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Perception Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFramework\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDomain Focus\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTasks\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMethod*\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGAIA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGeneral AI assistants across real-world tasks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eWeb browsing/Multimodality/Coding/Tools\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eO\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAgentBench\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDiversified scenarios\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOperating System/Database/Knowledge Graph/ Digital Card Game/Lateral Thinking Puzzles\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eS\u0026thinsp;+\u0026thinsp;O\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSuperCLUE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eComplex Chinese-language tasks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTool Usage \u0026amp; API Interaction/Task Execution /Environment Interaction \u0026amp; Memory\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eO\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePaperBench\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAI capabilities in academic research\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReplicate state-of-the-art AI research\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eO\u0026thinsp;+\u0026thinsp;M\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAgentHarm\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSafety and adversarial robustness\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFraud/Cybercrime/Harassment\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eO\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWAA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWeb interaction tasks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eWindows tasks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eO\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eYiHeng\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGeneral AI assistants across real-world tasks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eStudy/Live/Work/Game/Coding,\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eS\u0026thinsp;+\u0026thinsp;O\u0026thinsp;+\u0026thinsp;M\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eMethod Descuption : O:Objectivity S: Subjective M: Manual calibration\u003c/p\u003e\u003cp\u003eThe YiHeng Agent Evaluation Framework offers a comprehensive and application-driven approach specifically designed for general-purpose agents. To address the generalization requirements of general-purpose agents across diverse domains, the framework adopts the AI-system evaluation principles outlined in ISO/IEC 23053 [15] by the International Organization for Standardization, using real-world application scenarios as the evaluation baseline. The Framework encompasses a diverse range of task scenarios\u0026mdash;including education, daily life, professional work, and entertainment\u0026mdash;aiming to rigorously evaluate agents' adaptability and execution capabilities in real-world environments. Unlike other frameworks, YiHeng employs a hybrid methodology that combines objective metrics with subjective evaluations, further refined through expert calibration. This approach enables efficient automated assessment through intelligent tools (Figure.3 in Chap.\u0026nbsp;6) while ensuring accuracy and interpretability via expert oversight. Moreover, its task design is closely aligned with practical application demands, providing a more comprehensive representation of agent performance in complex and dynamic contexts, while maintaining a thoughtful balance between objectivity and flexibility.\u003c/p\u003e\u003c/div\u003e"},{"header":"5 Experiments","content":"\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e5.1 Experimental Background\u003c/h2\u003e\u003cp\u003eGeneral-purpose intelligent agents are characterized by their ability to operate across domains and handle diverse, complex tasks. This evaluation includes a representative selection of such agents from leading global developers. A total of eight agents were selected, including high-profile systems such as Manus (Butterfly Effect), ChatGPT Agent (OpenAI), and Coze (ByteDance). Among them, four are browser-based agents and four are mobile-based agents. For clarity and consistency in subsequent analysis, these agents are anonymized and referred to as Agents A through H. Additional general-purpose agents were excluded from this evaluation due to the absence of fully commercialized versions at the time of testing, but they are scheduled for inclusion in future assessment phases.\u003c/p\u003e\u003cp\u003eFor the eight selected AI agents, evaluations were conducted according to the methodology, using the quantitative scoring criteria outlined in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Specifically: (1) For subjective indicators, a panel of 10 independent experts was invited to assess each metric across all evaluation dimensions. The final score for each metric was calculated as the arithmetic mean of the experts' individual ratings. (2) For objective indicators, scores were computed by averaging the actual measured values obtained from the test samples. Each agent's overall score was then determined using a weighted aggregation method and subsequently adjusted through manual calibration to ensure consistency. It should be noted that certain metrics\u0026mdash;such as token consumption\u0026mdash;could not be obtained for some agents and were therefore excluded from the current evaluation scope.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003e5.2 Experimental Result\u003c/h2\u003e\u003cp\u003eBased on evaluation scores (all metrics using a 100-point scale), as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The eight intelligent agents can be categorized into three performance tiers. Agent A stands out as a \"versatile assistant, \" demonstrating comprehensive capabilities across both personal and professional tasks (e.g., hotel booking, website deployment) with no significant weaknesses. It excels particularly in functional diversity and task planning. Agent B is characterized as a \"high-efficiency but security-compromised manager, \" excelling in information comprehension, logical reasoning, and response interaction. However, it shows critical vulnerabilities in safety reliability, particularly in areas such as social bias mitigation, privacy protection, and regulatory compliance. The remaining agents display specialized competencies: Agent F leads in security metrics due to its robust ethical review mechanisms and effective risk interception capabilities, while agent E stands out for its exceptional response speed and personalized interaction design, delivering an optimal user experience.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e5.3 Special Indicator Analysis\u003c/h2\u003e\u003cdiv id=\"Sec16\" class=\"Section3\"\u003e\u003ch2\u003e5.3.1 Perception Capability\u003c/h2\u003e\u003cp\u003eIn terms of perception capability, both agent B and A excel, demonstrating a strong ability to accurately understand core needs and detailed information while effectively integrating complex requirements. Additionally, they exhibit strong logical reasoning skills. Agent C and D possess basic perception capabilities but show some inconsistencies in handling details, which can result in solutions that lack focus. Agent G, H, and E, on the other hand, have relatively high rates of key information omission and show weaknesses in both logical reasoning and task scenario adaptation, leading to solutions that may be disconnected from real-world needs. The result is presented in Table.3.\u003c/p\u003e\u003cp\u003eFor example, in a meeting scheduling task, agent B accurately understood all the detailed requirements, while agent D overlooked a few, such as the \"notify participants\" request. In contrast, agent H only identified the \"add date reminder\" task, with many other intended actions omitted.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Perception Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e86\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e84\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e69\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e67.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e63\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec17\" class=\"Section3\"\u003e\u003ch2\u003e5.3.2 Planning Capability\u003c/h2\u003e\u003cp\u003eIn terms of planning capability, agent A stands out significantly among the products, excelling in the decomposition of complex requirements. It logically and accurately breaks down tasks, with clear sub-task sequencing that fully addresses user needs. Resource allocation and utilization are also efficient and effective. Agent C and D exhibit some planning ability, but their task decomposition tends to be overly detailed, which negatively impacts overall task efficiency. They also face issues such as timing logic gaps and failure to prioritize critical tasks. Agent F, G, H, and E show weaker planning capabilities, with rudimentary planning processes that lack completeness. These limitations often lead to problems such as task chain interruptions and improper sub-task prioritization. Additionally, some products directly output results during task decomposition without displaying the planning process, making it difficult for users to grasp the underlying approach. The result is presented in Table.4.\u003c/p\u003e\u003cp\u003eFor example, in an online shopping task, agent A not only completes the overall planning but also dynamically adjusts the plan based on task requirements and specific execution contexts, covering all key operational steps. In contrast, while agent D's planning is generally complete, it cannot dynamically adjust when execution issues arise. Agent H, on the other hand, simply submits the purchasing request directly to the shopping platform, indicating a weak planning capability.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Planning Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e83.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e81.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e75.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e71.3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e58.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e58.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e56.9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e54.0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section3\"\u003e\u003ch2\u003e5.3.3Execution Capability\u003c/h2\u003e\u003cp\u003eThere are notable differences in the tool utilization capabilities of agents. On desktop platforms, some agents are limited to interacting with web-based tools, such as web browsing and search engines, and are unable to operate locally installed applications. In contrast, other agents can effectively perform basic office software tasks by creating execution sandboxes or similar environments. On mobile platforms, certain agents can invoke system-built applications and a limited range of third-party apps that have been specifically adapted. However, tasks such as sending messages to WeChat friends or playing Jay Chou's music can only be executed by a select few agents. The result is presented in Table.5.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Execution Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e81.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e78.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e69.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e68.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e55.7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e50.3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e44.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e42.5\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section3\"\u003e\u003ch2\u003e5.3.4 Functionality Richness\u003c/h2\u003e\u003cp\u003eDriven by advancements in foundational large model technologies, most agents excel at handling text-based tasks, covering a broad spectrum of both professional and everyday activities. Except for agent C, all agents now support image-based tasks to varying extents. However, their ability to process audio and video remains limited. For instance, when tasked with generating specific audio files, only a few agents produce the correct output. agent D, for example, misinterprets all audio generation tasks as requests for AI podcasts, failing to accurately meet user requirements. The result is presented in Table.6.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Functionality Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e80.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e77.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e75.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e75.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e75.0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e71.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e70.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e70.9\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section3\"\u003e\u003ch2\u003e5.3.5 Experience Comfort Level\u003c/h2\u003e\u003cp\u003eAgents typically emphasize content generation quality and interaction design, ensuring that basic text output functions align with users' typical usage habits. Some products go beyond this by offering distinctive features to further enhance user experience. For instance, agent A and E improve user tolerance for errors through task interruption and recovery functions, while agent C boosts interaction intuitiveness and immersion with innovations like digital humans and navigation maps. However, one challenge remains: the processing time for task completion can be relatively long. Although agent A and E alleviate some of the delays with their task interruption/recovery features, this still presents a potential bottleneck that could impact the overall user experience. The result is presented in Table.7.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of UserExperience Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e84.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e83.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e83.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e81.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e81.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e77.7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e73.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e71.5\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section3\"\u003e\u003ch2\u003e5.3.6 Security ReliCapability\u003c/h2\u003e\u003cp\u003eCurrently, the overall security performance of agents lags behind the average standard set by large models, with significant improvements needed in areas such as preventing the generation of misleading information and avoiding violent content. Agent F excels in this regard, demonstrating a robust security strategy and performing exceptionally well across a range of security tasks. In contrast, agents G, H, and E focus on minimizing the harmfulness of their outputs, but they also continually alert users about potential risks or fabrications in the content. Whileagents A, D, and C perform well in terms of functionality and overall performance, they fall short in several security-related tasks. These agents still generate harmful content in some situations and lack sufficient security warnings in specific contexts. The result is presented in Table.8.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Security Score\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eH\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eScore(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e96.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e86.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e85.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e85.3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e82.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e80.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e73.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e69.3\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003e5.4 Analysis of Experimental Conclusions\u003c/h2\u003e\u003cp\u003eOverall, the following key conclusions can be drawn:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eLeading Agents Have Surpassed the \"Usability\" Threshold\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eAgents such as A and B have successfully crossed the \"usability\" threshold, now effectively meeting the basic needs of users in both work and daily life scenarios. With a single command, they can automatically complete tasks through end-to-end execution, providing seamless functionality and user convenience.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eSignificant Performance Disparities Among Industry-Leading Agents\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThere are clear differences in the capabilities of agents, which vary across three core areas: perception, planning, and execution. The foundational large models, which serve as the backbone of these agents, play a crucial role in shaping their perception abilities. As these models improve, the gap in perception capabilities between different agents has narrowed. However, decision-making and planning capabilities remain the defining factors for the practical effectiveness of these agents. Many struggle with complex, cross-domain tasks, often producing errors, logical inconsistencies, or incomplete results. Additionally, they lack the capacity for self-reflection and error correction. Tool utilization, however, remains the critical factor in determining the overall quality of an agent. While some agents can only perform basic web-based tasks, others are capable of integrating a wide array of complex tools.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eNew Risks to Security Governance for Agents\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eIntelligent agents possess the capability to autonomously execute complete operational cycles (planning, tool invocation ,task execution) upon receiving objectives, enabling direct manipulation of critical infrastructure components\u0026mdash;including financial systems, databases, and OS interfaces\u0026mdash;without requiring human intervention. This autonomous functionality allows seemingly benign prompts to rapidly escalate into severe security incidents, such as illicit fund transfers, large-scale data breach, or critical service disruptions. Gartner forecasts indicate that by 2028, agent-related vulnerabilities will account for 25% of enterprise data breaches, primarily attributable to their dynamically evolving and opaque attack surfaces during external system interactions [17]. Compounding these risks, agents frequently employ \"user takeover\" protocols to acquire authentication credentials, creating vectors for multidimensional, large-scale cascading attacks when maliciously exploited. Conventional AI security frameworks prove insufficient to detect or mitigate these emerging threats due to their novel attack methodologies and unprecedented propagation speeds.\u003c/p\u003e\u003c/div\u003e"},{"header":"6 Evaluation Tools Practice","content":"\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e\u003ch2\u003e6.1 Overview of Agent Evaluation Tools\u003c/h2\u003e\u003cp\u003eAs a frontier in artificial intelligence research, AI agent evaluation technology is distinguished by its high degree of novelty and innovation. Its assessment scope\u0026mdash;both broader and deeper than that of traditional AI evaluation frameworks\u0026mdash;often demands significant human effort. To mitigate these challenges and reduce evaluation costs, we developed a \u0026ldquo;1\u0026thinsp;+\u0026thinsp;N\u0026thinsp;+\u0026thinsp;X\u0026rdquo; end-cloud integrated automated evaluation tool (in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), which represents an initial step toward automating the evaluation process for AI agents.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e1 (Evaluation Platform):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eA centralized evaluation platform is deployed to coordinate multiple connected edge-side toolkits. It remotely assigns testing tasks and scripts to designated edge devices and collects evaluation artifacts, including videos and system logs. Additionally, the platform integrates an intelligent scoring module capable of automatically calculating key performance indicators, thereby improving evaluation efficiency and consistency.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eN (Edge-Side Evaluation Toolkits):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eEach edge-side toolkit comprises a low-cost development board and several test devices (e.g., smartphones or personal computers). These toolkits enable concurrent execution of various tasks dispatched by the central platform. Leveraging the Airtest [18] automation framework, the toolkit offers a cohesive hardware\u0026ndash;software solution that ensures repeatable and scalable testing procedures.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eX (Test Devices\u0026ndash;Mobile/Browsers):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThese devices, equipped with browsers or corresponding AI agent applications, execute scripted evaluation tasks. They support both text and voice input modalities, collect interaction logs, and continuously record test sessions via video. This setup facilitates detailed performance monitoring and post-evaluation analysis across different device environments.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec25\" class=\"Section2\"\u003e\u003ch2\u003e6.2 Text-based LLM Agent Evaluation Tools\u003c/h2\u003e\u003cp\u003eIntelligent testing tools are integrated into the evaluation process through the development of automated scripts, enabling the simulation of user input, control of agent operations, and extraction of the agent's textual outputs. To evaluate subjective metrics such as response fluency and interaction friendliness, this paper investigates the use of large-scale text-based models for automated scoring. A typical prompt template is structured as follows:\u003c/p\u003e\u003cp\u003e\u003cem\u003e[System Role] As a human-computer interaction evaluation expert, please assess strictly according to the standards;[Evaluation Subject] User query: %questext% | AI response: %answertext%;[Scoring Matrix] Fluency (4\u0026thinsp;=\u0026thinsp;No grammatical errors and logically coherent; 3\u0026thinsp;=\u0026thinsp;Minor redundancy; 2\u0026thinsp;=\u0026thinsp;Ambiguities present; 1\u0026thinsp;=\u0026thinsp;Fragmented sentences);[Output Requirement] Return the score in JSON format.\u003c/em\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTable captions should be placed above the tables.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEvaluation Dimension\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFluency\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFriendliness\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eConsistency in grading\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e81.5%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e76.0%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe automated scores are then compared for consistency with human-generated scores. Results show that using language models for automated scoring achieves over 75% consistency with manual evaluations, thus providing a strong technical foundation for the future use of automated evaluation methods. The result is presented in Table.9..\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec26\" class=\"Section2\"\u003e\u003ch2\u003e6.3 Multimodal LLM Intent Recognition\u003c/h2\u003e\u003cp\u003eBy integrating multimodal large models into the evaluation process, this study facilitates auxiliary analysis of key performance indicators, including intent recognition accuracy, interaction friendliness, and security reliability. Furthermore, it investigates a comparative method for intent recognition assessment through multimodal feature fusion. In this approach, the evaluation combines test question text, task output files, and screenshots of the agent's execution interface into structured prompt templates, which are then submitted to the multimodal models for scoring. Two such models were employed in the experiment. Results show that both achieved over 80% accuracy, highlighting their strong potential for real-world application. Future improvements in intent recognition accuracy may be achieved through prompt engineering optimization and the selection of more suitable multimodal models. The result is presented in Table.10.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab10\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 10\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTable captions should be placed above the tables.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEvaluation Dimension\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM1\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLLM2\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRecognition accuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e80.1%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e81.5%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec27\" class=\"Section2\"\u003e\u003ch2\u003e6.4 Video Key Event Recognition and Analysis\u003c/h2\u003e\u003cp\u003eThis study utilizes a video analysis automation tool to evaluate the end-to-end task response delay of agents\u0026mdash;measuring the time from when the user submits a query to the agent's first output after processing the instructions. The results are then compared with manual frame-by-frame counts for validation. In the experiment, the video capture frame rate ranged between 20 and 30 fps, depending on terminal performance and the testing framework, with an allowable error margin of \u0026plusmn;\u0026thinsp;1 frame (equating to a delay accuracy of \u0026le;\u0026thinsp;50ms). The findings indicate that the tool offers high precision in calculating the metrics, enabling fully automated evaluation of this specific metric on the terminal side for agents. The result is presented in Table.11.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab11\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 11\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eTable captions should be placed above the tables.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEvaluation Dimension\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eUser Submit Time(s)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAgent Response Time(s)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRecognition accuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e97.6%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.9%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eBased on practical outcomes, employing more advanced evaluation tools can substantially improve both the efficiency and accuracy of automated assessments for agents. Moving forward, enhancing the capabilities of these evaluation tools will play a pivotal role in accelerating the development and optimization of agent performance.\u003c/p\u003e\u003c/div\u003e"},{"header":"7 Conclusion and Future Work","content":"\u003cp\u003eThis study presents a systematic evaluation framework for agents, covering key dimensions such as evaluation metrics, methodologies, tools, and datasets. The framework offers a comprehensive and practical approach to agent assessment and has been validated through experiments involving leading agent products globally. The results demonstrate both the methodological robustness and real-world applicability of the proposed system, offering meaningful guidance for optimizing agent performance and measuring effectiveness. Nonetheless, several challenges persist. The rapid pace of technological advancement has outstripped the development of a unified and authoritative evaluation standard, and cross-scenario evaluation capabilities remain underdeveloped. Overcoming these obstacles will require ongoing innovation in evaluation methodologies and continuous enhancement of evaluation tools to advance agent technologies toward higher levels of efficiency and reliability.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eClinical trial number:\u0026nbsp;\u003c/strong\u003enot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate:\u003c/strong\u003enot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u003c/strong\u003enot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials:\u003c/strong\u003eThe datasets and/or code used during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u003c/strong\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions:\u003c/strong\u003eW.L. designed the study and wrote the main manuscript; X.M. provided guidance and refined the theoretical framework; L.J. and S.L. constructed the dataset and revised the manuscript; R.W., Y.L., and X.F. performed testing and validation. All authors reviewed and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u003c/strong\u003eThe authors thank the anonymous reviewers for their constructive feedback.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eZhang, Y., Li, G., Wang, X.: Agent intelligence: A survey of recent advances and future directions. Front. Comput. Sci. \u003cb\u003e18\u003c/b\u003e, 184501 (2024)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYehudai, A., Eden, L., Li, A., et al.: Survey on Evaluation of LLM-based Agents. J. Artif. Intell. Res. \u003cb\u003e65\u003c/b\u003e, 102135 (2025)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTran, K.-T., Nguyen, M., Pham, H., et al.: Multi-Agent Collaboration Mechanisms: A Survey of LLMs. Artif. Intell. Rev. \u003cb\u003e58\u003c/b\u003e, 35 (2025)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang, F., Chen, J., Yang, S., Al-Lawati, A., Tang, L., Liu, H., Wang, S.A.: Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness. arXiv preprint arXiv:2510.13890. (2025)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLegg, S., Hutter, M.: A collection of definitions of intelligence. Front. Artif. Intell. Appl. \u003cb\u003e157\u003c/b\u003e, 17 (2007)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHern\u0026aacute;ndez-Orallo, J.: The Measure of All Minds: Evaluating Natural and Artificial Intelligence. Cambridge Univ. Press, Cambridge, U.K (2017)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMialon, G., Fourrier, C., Swift, C., et al.: GAIA: A Benchmark for General AI Assistants. In: Proc. Adv. Neural Inf. Process. Syst., vol. 36, pp. 1\u0026ndash;15 (2023)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu, X., Yu, H., Zhang, H., et al.: AgentBench: Evaluating LLMs as Agents. IEEE Trans. Pattern Anal. Mach. Intell. \u003cb\u003e45\u003c/b\u003e, 9876 (2023)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSuperCLUE Team: SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark: (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cluebenchmarks.com/static/superclue.html\u003c/span\u003e\u003cspan address=\"https://cluebenchmarks.com/static/superclue.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStarace, G., Jaffe, O., Sherburn, D., et al.: PaperBench: Evaluating AI's Capability to Replicate AI Research. Nat. Mach. Intell. \u003cb\u003e7\u003c/b\u003e, 45 (2025)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBonatti, R., Zhao, D., Bonacci, F., et al.: Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. ACM Trans. Comput. Syst. \u003cb\u003e42\u003c/b\u003e, 1 (2024)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAndriushchenko, M., Souly, A., Dziemian, M., et al.: AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In Proceedings of the IEEE Symposium on Security and Privacy (2024)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiang, P., Bommasani, R., et al.: Holistic Evaluation of Language Models. arXiv preprint arXiv:2211.09110 (2022)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShneiderman, B.: Human-Centered Artificial Intelligence: Reliable, Safe \u0026amp; Trustworthy. Int. J. Hum. -Comput Interact. \u003cb\u003e36\u003c/b\u003e, 495 (2020)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eISO/IEC 23053: Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML). International Organization for Standardization, Geneva, Switzerland: (2022)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGartner: AI Agent Maturity Model, Stamford, CT, USA, Rep: G00712345 (2023)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGartner: Emerging Tech Impact Radar: Artificial Intelligence, Stamford, CT, USA, Rep: G00792637 (2024). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gartner.com/en/documents/5158550\u003c/span\u003e\u003cspan address=\"https://www.gartner.com/en/documents/5158550\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNetease Inc.: Airtest: Cross-platform UI Automated Testing Framework: (2023). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://airtest.netease.com/\u003c/span\u003e\u003cspan address=\"https://airtest.netease.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Artificial Intelligence, Agent Evaluation, YiHeng Evaluation Framework, Evaluation Tools","lastPublishedDoi":"10.21203/rs.3.rs-7936146/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7936146/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIntelligent agent technology represents a pivotal breakthrough in the evolution of artificial intelligence, marking the shift from systems that merely \"understand\" to those capable of autonomous action. As this technology becomes increasingly central to AI deployment, the establishment of a scientifically rigorous and standardized evaluation framework has become essential for supporting and accelerating the industrialization of AI applications. However, the development of such a framework presents numerous challenges due to the complexity and diversity of intelligent agent tasks. To address these challenges, this study introduces the YiHeng Agent Evaluation System\u0026mdash;a comprehensive, objective, and user-centered framework. It employs a \"2-4-2\" hierarchical structure that includes two types of evaluation scenarios, four key evaluation elements, and two overarching evaluation dimensions. The system evaluates not only functional capabilities, such as usability and effectiveness, but also user experience factors, including ease of use and satisfaction. To validate the proposed framework, empirical evaluations were conducted on eight leading general-purpose intelligent agents from across the globe. In parallel, supporting evaluation tools were developed through targeted engineering implementations to facilitate systematic assessment. The results confirm the framework's effectiveness and practical relevance, providing actionable insights for enhancing agent performance and promoting the sustainable advancement of the AI industry.\u003c/p\u003e","manuscriptTitle":"Construction and Practical Validation of an Evaluation Framework for General-Purpose Agents","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-18 10:56:51","doi":"10.21203/rs.3.rs-7936146/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4af87b9a-afc5-4156-9069-1c9791e46978","owner":[],"postedDate":"November 18th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-08T23:38:59+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-18 10:56:51","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7936146","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7936146","identity":"rs-7936146","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.