A Comparative Study of Six Indigenous Chinese Large Language Models' Understanding Ability: An Assessment Based on 132 College Entrance Examination Objective Test Items

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract To assist Chinese language teachers in making evidence-based choices of useful and user-friendly domestic large language models in teaching and research, the study took 132 objective questions from the national college entrance examination Chinese language papers from 2021 to 2023 as the data set to assess the performance of six domestic large language models, namely Tongyi Qianwen, GLM-4, KimiChat, Baichuan, Wenxin Yiyan, and Xunfei Spark, in semantic understanding. The assessment revealed that the overall correct rates of the responses of the above six large language models to the questions were 70%, 69%, 57%, 55%, 60%, and 62% respectively. Among them, Tongyi Qianwen and Xunfei Spark performed best in language application questions, with correct rates of 74% each; GLM-4 performed best in ancient poetry reading and modern text reading questions, with correct rates reaching 92% and 77% respectively. The performance of the six large language models in classical Chinese reading questions was not ideal. For the wrongly answered test questions, the researchers corrected and analyzed the answers using the prompt strategy. Finally, the paper put forward several suggestions for promoting the assistance of large language models in Chinese language teaching and research.
Full text 111,351 characters · extracted from preprint-html · click to expand
A Comparative Study of Six Indigenous Chinese Large Language Models' Understanding Ability: An Assessment Based on 132 College Entrance Examination Objective Test Items | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Comparative Study of Six Indigenous Chinese Large Language Models' Understanding Ability: An Assessment Based on 132 College Entrance Examination Objective Test Items Huijin Le, Qiuling Zhang, Gang Xu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5990278/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract To assist Chinese language teachers in making evidence-based choices of useful and user-friendly domestic large language models in teaching and research, the study took 132 objective questions from the national college entrance examination Chinese language papers from 2021 to 2023 as the data set to assess the performance of six domestic large language models, namely Tongyi Qianwen, GLM-4, KimiChat, Baichuan, Wenxin Yiyan, and Xunfei Spark, in semantic understanding. The assessment revealed that the overall correct rates of the responses of the above six large language models to the questions were 70%, 69%, 57%, 55%, 60%, and 62% respectively. Among them, Tongyi Qianwen and Xunfei Spark performed best in language application questions, with correct rates of 74% each; GLM-4 performed best in ancient poetry reading and modern text reading questions, with correct rates reaching 92% and 77% respectively. The performance of the six large language models in classical Chinese reading questions was not ideal. For the wrongly answered test questions, the researchers corrected and analyzed the answers using the prompt strategy. Finally, the paper put forward several suggestions for promoting the assistance of large language models in Chinese language teaching and research. Large language models Chinese comprehension skills Artificial intelligence Prompt strategy Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction In the current wave of digital transformation within education, the integration of large language models (LLMs) into language instruction presents both transformative opportunities and multifaceted challenges (N. Ahmad et al., 2023; Storey et al., 2024; Kim et al., 2025). While these advanced models demonstrate potential for enhancing pedagogical approaches, their implementation requires careful consideration of several critical dimensions (Jeon et al., 2023). Recent academic discussions emphasize that generative AI applications in education must address persistent concerns regarding output reliability, with particular attention to factual accuracy and contextual appropriateness in linguistic outputs. This becomes especially crucial in language education, where nuanced cultural sensitivities and sociolinguistic variations demand systems capable of navigating diverse communicative norms and value systems (Varsik et al., 2024). Moreover, emerging evidence highlights systemic considerations surrounding equitable implementation. Disparities in technological infrastructure and varying levels of digital literacy across different educational contexts may inadvertently exacerbate existing inequalities if not proactively addressed (Lee et al., 2024; Bulathwela et al., 2024). Concurrently, questions persist about the pedagogical effectiveness of LLM integration, particularly regarding the alignment between algorithmic outputs and established language acquisition principles. Therefore, it is imperative to establish rigorous evaluation frameworks to assess how these technologies complement rather than compromise fundamental educational objectives, ensuring they support the development of critical language competencies. As an important component of generative artificial intelligence technology, LLM is a neural network model trained based on a large amount of text data, and it possesses certain capabilities such as language understanding, generation, logic, and reasoning (Kumar, 2024). Although LLM has excellent performance in complex language processing compared to traditional text generation methods such as recurrent neural networks and long short-term memory networks, due to the technical architecture of LLM's own decoder, the results it generates still have certain hallucinations and errors (Jho, 2024; Matarazzo et al., 2024). Therefore, specific task evaluations of LLMs are needed to assess the comprehensive or specialized performance of LLMs, providing a reliable reference basis for users to select the applicable LLM according to task requirements. Relevant institutions have successively released several evaluation systems for comprehensively assessing the performance of large language models. Typical ones include: MMLU (Zhao et al., 2024), BIG-Bench (Srivastava et al., 2024), HELM (Liang et al., 2024), C-Eval (Huang et al., 2024), MMCU (Zeng et al., 2024), SuperCLUE (Xu et al., 2024), and Chinese-LLM-Benchmark (Liu et al., 2024). Among them, MMLU and C-Eval mainly adopt closed-form (multiple-choice questions) decision tasks and use the correct rate of answers as the main evaluation indicator. BIG-Bench, HELM, MMCU, and Chinese-LLM-Benchmark focus on open-form (question-answering) generation tasks. Among the evaluation systems for LLMs released by various institutions, the datasets of the first three systems primarily assess LLM's performance in English environments. In contrast, the latter four systems targeting Chinese LLM's evaluation focus on multifaceted social life domains while demonstrating limited coverage of language and literature assessment tasks. Notably, SuperCLUE distinguishes itself through diversified task formats (including multiple-choice and open-ended questions) and more comprehensive evaluation metrics, endowing it with distinct advantages in assessing Chinese large language models compared to other benchmark systems. Research shows that as the scale of model parameters increases, the accuracy of LLMs in specific tasks does not improve linearly. Instead, an emergent ability emerges, and the accuracy of the models in complex reasoning tasks (such as mathematical problems and logical reasoning) increases significantly (Zhao et al., 2023; Bean et al., 2023). However, the rankings of Chinese LLMs included in the SuperCLUE evaluation list are not consistent in terms of overall performance and semantic understanding dimensions. More importantly, the evaluation data set of the semantic understanding dimension of SuperCLUE adopts general language tasks and can only measure the general reading comprehension ability of LLMs. To better study the applicability and effectiveness of artificial intelligence technology in Chinese education and teaching, it is necessary to collect specialized Chinese reading comprehension tasks to further evaluate the performance of LLMs. Based on this, this study will select SuperCLUE as the frame of reference for the evaluation results of the semantic understanding ability of Chinese LLMs. To help teachers select LLM based on evidence, this article attempts to respond to three questions: ( 1 ) Which domestic LLMs have a better overall performance in semantic understanding ability; ( 2 ) Which one performs better when answering different types of reading comprehension tasks such as language application, modern text reading, ancient poetry reading, and classical Chinese reading; ( 3 ) When the performance of answering Chinese reading comprehension tasks is not good, how to use prompt strategies to deeply explore the semantic understanding potential of LLMs. Literature review Evaluation Content The evaluation framework for large language models (LLMs) has evolved from traditional single-task performance assessment to multidimensional competency analysis. Early research primarily focused on accuracy evaluations for natural language processing tasks such as text classification and machine translation. However, as model scales expand and societal applications deepen, evaluation dimensions have gradually extended to ethical risks, multimodal capabilities, and societal impact. For instance, Laskar et al. (2024) highlighted that LLMs may generate factual errors (e.g., "hallucinations") or embed social biases during content generation, necessitating alignment optimization through human feedback reinforcement learning (RLHF). The SEED-Bench benchmark developed by Li et al. (2023) revealed significant deficiencies in current multimodal models regarding spatial relationship comprehension and cross-modal reasoning. In recent years, evaluation frameworks have further expanded to safety assessments, encompassing content compliance, privacy protection, and adversarial attack robustness. Hendy et al. (2023) proposed a four-dimensional evaluation framework (functionality, performance, alignment, and safety), emphasizing the need for systematic analysis of potential risks such as value misalignment and maliciously induced outputs. For example, He et al. (2024) utilized clinical laboratory examination questions and case studies to demonstrate that ChatGPT-4.0 and ERNIE Bot-4.0 both achieved a 60% passing rate in medical knowledge proficiency but exhibited significant errors in complex case analyses. Evaluation Methods Evaluation methodologies are broadly categorized into automated and human assessments. Automated evaluation relies on traditional metrics (e.g., BLEU, ROUGE) and LLM-based evaluators. Liu et al. (2023) proposed enhancing interpretability through chain-of-thought prompting strategies, though such methods still suffer from low consistency with human judgments. Kocmi and Federmann (2023) found that LLM evaluators may prioritize fluency over semantic accuracy, leading to scoring biases. Human evaluation involves expert or user participation to provide contextually grounded feedback. For instance, OpenAI employed human testing during GPT-4 development to validate multistep reasoning capabilities (Achiam et al., 2023). Lundgren (2024) demonstrated that GPT-4 achieved parity with human instructors in grading political science master’s theses but exhibited conservative scoring patterns and low reliability. Recent years have witnessed the rise of comprehensive evaluation frameworks. Benchmarks like GLUE and SuperGLUE, covering core tasks such as text classification and reasoning, drove model performance from BERT’s 80.5 score to near-human baselines (Gao et al., 2024). SQuAD 2.0 introduced unanswerable scenarios to test models’ uncertainty handling (Devlin et al., 2024). OMGEval supports 56-language generation tasks, prioritizing low-resource languages (e.g., Swahili) (Liu et al., 2024). EXAMS-V integrates multidisciplinary exam questions (e.g., mathematics, biology), requiring multimodal reasoning with text and charts, evaluated via accuracy and consistency metrics (Das et al., 2024). Nevertheless, challenges persist, including the black-box nature of proprietary models and the lack of standardized multimodal evaluation criteria. Liang et al. (2022) proposed the HELM framework, spanning 42 scenarios and 7 evaluation dimensions, which exposed performance disparities in low-resource language tasks. Dong et al. (2024) developed CLR-Bench to systematically assess college-level complex reasoning. Calibration, an emerging metric, was applied by Kumar et al. (2024) to quantify alignment between model confidence and true probabilities, revealing that well-calibrated models are better suited for high-stakes scenarios like medical diagnosis. Methods Evaluating large language models According to the ease of use and usefulness of the technology acceptance model proposed by Davis to select the participating LLMs (Davis, 2024). Ease of use refers to the extent to which people feel that using a certain technology is simple and effortless; usefulness refers to the possibility that people will choose a technology or tool with better performance to increase the possibility of generating actual benefits. Given that the ease of use of the existing LLMs is very good, this study only excludes foreign LLMs such as ChatGPT and Gemini from the perspective of accessibility, and evaluates the usefulness of Chinese LLMs from the perspective of Chinese reading comprehension. Referring to the latest evaluation ranking list of Chinese LLMs released by SuperCLUE on February 27, 2024, according to the comprehensive performance ranking, the top six domestic general LLMs of Wenxin Yiyan, GLM-4, Tongyi Qianwen, Baichuan, KimiChat, and Xunfei Spark were selected in sequence as the evaluation objects. Their evaluation scores in the semantic understanding dimension of SuperCLUE, from high to low, are Tongyi Qianwen, GLM-4, Wenxin Yiyan, KimiChat, Baichuan, and Xunfei Spark. All six LLMs are free versions, and the tests are conducted through the web page to simulate the real scenarios of daily use by teachers and students from the user's perspective. Evaluation data set The National College Entrance Examination, as one of the most crucial educational evaluation systems in China, has Chinese-language multiple-choice questions that exhibit extraordinary authority and representativeness. These questions comprehensively evaluate students' diverse capabilities, such as linguistic knowledge, reading comprehension, and logical reasoning. Covering a wide range of linguistic phenomena and knowledge points, they are of great significance for assessing the problem-solving abilities of large language models within complex Chinese-language contexts. Moreover, the relatively rich publicly-available data resources of Gaokao Chinese multiple-choice questions enable large-scale data collection and analysis, thus guaranteeing the wide applicability and reliability of research results. As shown in Table 1 , To ensure that the evaluation data set is not in the pre-training corpus of LLMs as much as possible, this study selected 133 multiple-choice questions from the Chinese college entrance examination Chinese language papers of the National Volume A, National Volume B, New College Entrance Examination Volume I, and New College Entrance Examination Volume II from 2021 to 2023 as special test tasks. Since some LLMs only support pure text format input, one multiple-choice question of modern text reading with picture information was deleted, resulting in 132 test questions. Among them, there are 19 questions of language application, 36 questions of classical Chinese reading, 12 questions of ancient poetry reading, and 65 questions of modern text reading. Although the number of different types of test questions is not balanced enough, due to the structural constraints of the college entrance examination questions, no additional evaluation data will be artificially added. Table 1 Basic information about the data set Question type Quantity Example of questions Language application 19 请阅读上传的文档材料, 将下列俗语填入文中括号内, 恰当的一项是(), 并说明选择此选项的理由。 A.干打雷不下雨;B.又吃鱼又嫌腥 C.前怕狼后怕虎;D.首尾不能兼顾 Classical Chinese reading 36 请阅读上传的文档材料, 关于文中江边人们谈论季札的部分, 下列说法不正确的一项是(), 并说明选择此选项的理由。 A༎那位老人欣赏季札不就王位的高洁, 也称赞他以美好的行为感动了世人。 B༎那位年轻人认为季札不顾百姓死活, 只顾独善其身, 逃避了济世的责任。 C༎季札挂剑一事进一步说明了他的品行, 也为后文的子胥赠剑做了铺垫。 D༎季札的退耕田园, 与下文渔夫的泛舟江上, 共同表达出本文的隐逸主题。 Ancient poetry reading 12 请阅读上传的文档材料, 下列选项中, 加点的词语和文中“槐蝉”所用修辞手法不同的一项是(), 并说明选择此选项的理由。 A༎主人下马客在船, 举酒欲饮无管弦。 B༎埋骨何须桑梓地, 人生无处不青山。 C༎六军不发无奈何, 宛转娥眉马前死。 D༎心非木石岂无感, 吞声瑰躅不敢言。 Modern text reading 65 请阅读上传的文档材料, 下列句子中的“谁”和“耳机一戴, 谁也不爱”中的“谁”, 意义和用法相同的一项是(), 并说明选择此选项的理由。 A.怅寥廓, 问苍茫大地, 谁主沉浮༟ B.生活中谁都需要表达和交流。 C.我本来是跟他开玩笑的, 谁知道他竟然生气了。 D.我越来越深刻地感觉到谁是我们最可爱的人! Research design The evaluation process is shown in Fig. 1 . ( 1 ) The reading materials corresponding to each question are stored separately in PDF document format. ( 2 ) Input the prompt, the framework is: "Please read the uploaded document materials" + the question stem of the multiple-choice question + "and explain the reason for choosing this option" + the options of the multiple-choice question. ( 3 ) Each question is completed in an independent conversation window, and the LLMs will return the generated options and the reasons for the selection. ( 4 ) The generation results are divided into three situations: ① For the questions answered correctly, they are directly included in the correct rate; ② For the questions answered wrongly, the explanations and answers are regenerated by improving the prompt, and then compared and analyzed with the correct answers; ③ The questions with abnormal results are retested to exclude the contingency of abnormal occurrences. It should be noted that the correct answers obtained after improving the prompt are not included in the total correct rate. Figure 1 . The evaluation process of LLMs Statistical analysis This study employed EXCEL software to compare the test results of Chinese large language models (LLMs) on Chinese language multiple-choice questions in the National College Entrance Examination with standard answers, through which valid data were filtered. Statistical analysis of the test results was performed using SPSS 26.0 software, including descriptive statistics to examine the accuracy rates of six LLMs across different question types. Finally, prompt engineering strategies were employed to modify erroneous responses through systematic error correction attempts. Results Comparison of comprehensive semantic understanding ability Comprehensive semantic comprehension capability refers to the overall performance level of large language models (LLMs) in semantic-understanding tasks across diverse text genres. This reflects the model's ability to comprehensively grasp and interpret semantics across various textual categories. As shown in Table 2 , the overall correct rate of the six LLMs is 62% and the standard deviation is 6%, indicating that the performance of the six LLMs in semantic understanding is not much different, but the difference in the special semantic understanding is slightly larger, with a large degree of dispersion, and the standard deviation ranges from 9–18%. The overall correct rates of Tongyi Qianwen and GLM-4 are 70% and 69% respectively, and their comprehensive ability in semantic understanding performs well; the overall correct rate of Xunfei Spark is 62%, ranking third, and its comprehensive ability in semantic understanding is relatively low. The overall correct rate of Baichuan is 55%, and its comprehensive ability in semantic understanding is weak, ranking sixth. Table 2 Correct and wrong rates of Chinese large language models in Chinese reading comprehension Question type Language application Classical Chinese reading Ancient poetry reading Modern text reading Total Total 19 36 12 65 132 Tongyi Qianwen Accuracy 74% 56% 83% 75% 70% Error 26% 39% 17% 23% 27% GLM-4 Accuracy 47% 58% 92% 77% 69% Error 47% 36% 8% 22% 28% Xunfei Spark Accuracy 74% 67% 42% 60% 62% Error 21% 28% 17% 25% 24% Wenxin Yiyan Accuracy 37% 39% 83% 74% 60% Error 58% 50% 17% 26% 36% KimiChat Accuracy 58% 39% 67% 60% 57% Error 42% 58% 33% 35% 42% Baichuan Accuracy 47% 47% 67% 58% 55% Error 53% 50% 33% 42% 45% M ± SD Accuracy 56%±15% 51%±11% 72%±18% 68%±9% 62%±6% Error 41 ± 15% 44%±11% 21%±10% 29%±8% 34 ± 9% Special semantic understanding ability comparison Specialized semantic comprehension capability focuses on LLMs' performance in semantic-understanding tasks targeting specific text types, demonstrating the model's domain-specific expertise and differential competencies in processing semantic patterns characteristic of particular textual genres. In the language application test questions, the overall correct rate of the six LLMs is 56%, and the standard deviation is 15%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of language application are relatively large. The correct rates from high to low are Tongyi Qianwen (74%), Xunfei Spark (74%), KimiChat (58%), GLM-4 (47%), Baichuan (47%), and Wenxin Yiyan (37%), indicating that Tongyi Qianwen and Xunfei Spark perform well in language application. In the classical Chinese reading test questions, the overall correct rate of the six LLMs is 51%, and the standard deviation is 11%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of classical Chinese reading are relatively large. The correct rates from high to low are Xunfei Spark (67%), GLM-4 (58%), Tongyi Qianwen (56%), Baichuan (47%), Wenxin Yiyan (39%), and KimiChat (39%), indicating that Xunfei Spark performs well in classical Chinese reading. In the ancient poetry reading test questions, the overall correct rate of the six LLMs is 72%, and the standard deviation is 18%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of ancient poetry reading are very large. The correct rates from high to low are GLM-4 (92%), Tongyi Qianwen (83%), Wenxin Yiyan (83%), KimiChat (67%), Baichuan (67%), and Xunfei Spark (42%), indicating that GLM-4 performs well in ancient poetry reading. In the modern text reading test questions, the overall correct rate of the six LLMs is 68%, and the standard deviation is 9%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of modern text reading are not significant. The correct rates from high to low are GLM-4 (77%), Tongyi Qianwen (75%), Wenxin Yiyan (74%), Xunfei Spark (60%), KimiChat (60%), and Baichuan (58%), indicating that GLM-4 performs well in modern text reading. As shown in Fig. 2 , The special semantic understanding ability performance of LLMs in the four types of test questions each has its advantages. Prompt strategy can effectively correcxt the wrong results of large language models The zero-shot thinking chain strategy is a method that allows the model to handle tasks with almost no need for any additional training. That is, by adding a simple prompt-"Let's think step by step", this triggers the LLMs to expand their thinking step by step according to the input prompt, thereby triggering extensive cognitive abilities, rather than being limited to the specific skills of a particular task (Kojima et al., 2023). For example, in question 4 of the 2021 National Volume A paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of "think step by step", GLM-4 gave the correct option and analysis as shown in Fig. 3 . In view of the randomness characteristic of LLMs' responses, making them generate multiple reasoning paths or answers, and then conducting internal voting to finally select the most consistent or most common answer for output. This method can help reduce the randomness of single sampling or reasoning, thereby improving the accuracy and reliability of the entire system (Wang et al., 2022). For example, in question 3 of the 2021 National Volume B paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of "Please generate ten answers internally, and conduct internal voting on the answers, and select and output the answer with the most votes", GLM-4 gave the correct option and analysis as shown in Fig. 4 . The core idea of the expert role strategy is to invoke expert knowledge in specific domains by designing and optimizing prompts, thereby improving factual accuracy, knowledge depth, and reducing biases (Wan et al., 2024). When used, by constructing a context containing an expert role, the professional knowledge in related fields in LLMs is activated, while irrelevant information is suppressed, reducing the ambiguity of the model's understanding of the problem, and then answering the knowledge content that conforms to the characteristics of that role. For example, in question 17 of the 2021 National Volume B paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of "As an expert in the field of Chinese", GLM-4 gave the correct option and analysis as shown in Fig. 5 . Discussion The findings demonstrate that among four types of Chinese reading comprehension tasks, Qwen-Turbo and GLM-4 exhibited superior comprehensive performance (with overall accuracy rates of 70% and 69% respectively), particularly showing significant advantages in specialized tasks. For instance, GLM-4 achieved an exceptional accuracy of 92% in classical poetry comprehension, substantially outperforming other models. These results corroborate the critical role of pre-training corpora and algorithmic design in model performance (Devlin et al., 2019; Gururangan et al., 2020; Knafou et al., 2020), while aligning with GLM-4's ranking in semantic understanding dimensions within the SuperCLUE evaluation framework (a comprehensive Chinese language model evaluation benchmark). However, the study revealed suboptimal performance across all models in classical Chinese (wenyanwen) reading comprehension tasks (average accuracy rate of 51%), indicating inadequate capture of classical Chinese linguistic features by general-purpose LLMs, necessitating improvement through domain-specific corpora and algorithmic optimization (Lin et al., 2022). Regarding erroneous responses from LLMs, this research demonstrates that prompt engineering strategies-including zero-shot chain-of-thought prompting, self-consistency voting, and expert role assignment-can significantly enhance response accuracy (Wang et al., 2023; Zhao et al., 2023; Qin et al., 2023; Trad et al., 2023). The study further proposes three practical recommendations: developing domain-specific LLMs for Chinese language education, constructing interdisciplinary AI agents, and promoting pedagogical paradigm transformation. For instance, the suboptimal performance in classical Chinese reading comprehension (demonstrating an average error rate of 44%) indicates that general-purpose LLMs inadequately address instructional requirements, necessitating enhancements through classical Chinese corpora enrichment and algorithmic fine-tuning to improve cultural interpretation capabilities (Lin et al., 2022). Simultaneously, the findings suggest that establishing Chinese-language specialized AI agents could synergize the advantages of multiple models, such as integrating the language application proficiency of Tongyi Qianwen with the poetry analysis strength of GLM-4 (Ding et al., 2024). The digital-intelligent transformation of pedagogical perspectives requires balanced attention to technological empowerment and humanistic considerations, cautioning against excessive reliance on model-generated content (Heersmink et al., 2024; Faisal, 2024). The innovation of this study lies in conducting the first systematic evaluation of mainstream domestic large language models (LLMs) on Chinese language discipline-specific tasks, while proposing targeted optimization strategies based on prompt engineering. Compared with general evaluation benchmarks like SuperCLUE, this research focuses on subject-oriented scenarios, providing educators with direct references for model selection and pedagogical integration. The experimental design utilizes national college entrance examination questions as the test dataset, with multi-round prompt experiments revealing LLMs' adaptability through prompt-based interventions. These findings not only enrich the disciplinary dimensions of Chinese LLMs evaluation frameworks but also establish methodological foundations for developing educational-purpose AI agents. However, the current research is limited to objective question assessments, and future work should explore LLMs' potential in subjective task scenarios, particularly essay evaluation, through open-ended generation tasks. Conclusion Tongyi Qianwen and GLM-4 have firmly ranked in the top two in semantic understanding ability and have respectively won the top two in three special task abilities. Chinese language teachers can give priority to choosing them as teaching and research tools such as text interpretation, teaching design, test question formulation, and subject research. The overall performance of Chinese LLMs in tasks such as classical Chinese reading comprehension is not ideal. Prompt strategies such as thinking chains, using internal voting evaluations, and expert roles can be used to help LLMs understand the intentions of Chinese language teachers and enable LLMs to output content that meets the expectations of Chinese language teachers. Abbreviations LLMs Large language models Declarations Ethics declarations Ethics approval and consent to participate Not applicable. Clinical trial number Not applicable. Consent for publication Not applicable. Funding This work was supported b the priority concern project of the "13th Five-Year Plan" of Beijing Education Science in 2020 [grant number CHEA2020025]. Author Contribution Le. Conceptualization, Methodology; Xu. Writing- Original draft prepa-ration, Data curation; Zhang.Writing- Reviewing and Editing. References Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., & McGrew, B. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Ahmad, N., Murugesan, S., & Kshetri, N. (2023). Generative Artificial Intelligence and the Education Sector. Computer , 56 (6), 72–76. Bean, A. M., Korgul, K., Krones, F., McCraith, R., & Mahdi, A. (2023). Do Large Language Models have Shared Weaknesses in Medical Question Answering? arXiv preprint arXiv:2310.07225 . Bulathwela, S., Pérez-Ortiz, M., Holloway, C., Cukurova, M., & Shawe-Taylor, J. (2024). Artificial Intelligence Alone Will Not Democratise Education: On Educational Inequality, Techno-Solutionism and Inclusive Tools. Sustainability , 16 (2), 781. Das, R. J., Hristov, S. E., Li, H., Dimitrov, D. I., Koychev, I., & Nakov, P. (2024). Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models. arXiv preprint arXiv :240310378. Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology (pp. 319–340). MIS quarterly. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for compu. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171–4186). Ding, H., Fan, Z., Guehring, I., Gupta, G., Ha, W., Huan, J., … Zhou, H. (2024). Reasoning and planning with large language models in code development. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 6480–6490). Dong, J., Hong, Z., Bei, Y., Huang, F., Wang, X., & Huang, X. (2024). CLR-Bench: Evaluating Large Language Models in College-level Reasoning. arXiv preprint arXiv:2410.17558. Faisal, E. (2024). Unlock the potential for Saudi Arabian higher education: a systematic review of the benefits of ChatGPT. Frontiers in Education (Vol. 9, p. 1325601). Frontiers Media SA. Gao, M., Hu, X., Ruan, J., Pu, X., & Wan, X. (2024). Llm-based nlg evaluation: Current status and challenges. arXiv preprint arXiv :240201383. Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., & Smith, N. A. (2020). Don't stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:2004.10964. He, Z., Bhasuran, B., Jin, Q., Tian, S., Hanna, K., Shavor, C., … Lu, Z. (2024). Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study. Journal of medical Internet research, 26, e56655. Heersmink, R., de Rooij, B., Clavel Vázquez, M. J., & Colombo, M. (2024). A phenomenology and epistemology of large language models: Transparency, trust, and trustworthiness. Ethics and Information Technology , 26 (3), 41. Hendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., … Awadalla,H. H. (2023). How good are gpt models at machine translation? a comprehensive evaluation.arXiv preprint arXiv:2302.09210. Huang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., … He, J. (2023). C-eval:A multi-level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing Systems, 36, 62991–63010. Jeon, J., & Lee, S. (2023). Large Language Models in Education: A Focus on the Complementary Relationship Between Human Teachers and ChatGPT. Educ Inf Technol , 28 , 15873–15892. Jho, H. (2024). Leveraging generative AI in physics education: Addressing hallucination issues in large language models. Kim, J., Yu, S., & Detrick, R. (2025). Exploring Students’ Perspectives on Generative AI-assisted Academic Writing. Educ Inf Technol , 30 , 1265–1300. Knafou, J., Naderi, N., Copara, J., Teodoro, D., & Ruch, P. (2020). BiTeM at WNUT 2020 shared task-1: named entity recognition over wet lab protocols using an ensemble of contextual language models. In Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020) (pp. 305–313). Kocmi, T., & Federmann, C. (2023). Large language models are state-of-the-art evaluators of translation quality. arXiv preprint arXiv :230214520. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large language models are zero-shot reasoners. Advances in neural information processing systems , 35 , 22199–22213. Kumar, A., Morabito, R., Umbet, S., Kabbara, J., & Emami, A. (2024). Confidence under the hood: An investigation into the confidence-probability alignment in large language models. arXiv preprint arXiv:2405.16282. Kumar, P. (2024). Large language models (LLMs): survey, technical frameworks, and future challenges. Artificial Intelligence Review , 57 (10), 260. Laskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H.,… Huang, J. (2024). A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 13785–13816). Lee, J., Hicke, Y., Yu, R., Brooks, C., & Kizilcec, R. F. (2024). The Life Cycle of Large Language Models in Education: A Framework for Understanding Sources of Bias. British Journal of Educational Technology , 55 , 1982–2002. Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., & Shan, Y. (2024). Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 13299–13308). Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., … Koreeda,Y. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110. Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., … Koreeda,Y. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110. Lin, C. T., & Ma, W. Y. (2022). HanTrans: An Empirical Study on Cross-Era Transferability of Chinese Pre-trained Language Model. In Proceedings of the 34th Conference on Computational Linguistics and Speech Processing (ROCLING 2022) (pp. 164–173). Liu, C., Yu, L., Li, J., Jin, R., Huang, Y., Shi, L., … Xiong, D. (2024). Openeval:benchmarking Chinese LLMs across capability, alignment and safety. arXiv preprint arXiv:2403.12316. Liu, P., Yuan, W., Fu, J., Jiang, Z., & Hayashi, H. (2023). & Pre-train, G. N. prompt, and predict: A systematic survey of prompting methods in natural language processing., 55. https://doi.org/10.1145/3560815 , 1–35. Liu, Y., Xu, M., Wang, S., Yang, L., Wang, H., Liu, Z., … Yang, E. (2024). OMGEval:An Open Multilingual Generative Evaluation Benchmark for Large Language Models. arXiv preprint arXiv:2402.13524. Lundgren, M. (2024). Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders. arXiv preprint arXiv:2406.16510. Matarazzo, A., & Torlone, R. (2025). A Survey on Large Language Models with some Insights on their Capabilities and Limitations. arXiv preprint arXiv:2501.04040. Qin, L., Chen, Q., Wei, F., Huang, S., & Che, W. (2023). Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. arXiv preprint arXiv :231014799. Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., … Wang,G. (2022). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615. Storey, V. A., & Wagner, A. (2024). Integrating Artificial Intelligence (AI) Into Adult Education: Opportunities, Challenges, and Future Directions. International Journal of Adult Education and Technology (IJAET) , 15 (1), 1–15. Trad, F., & Chehab, A. (2024). Prompt engineering or fine-tuning? a case study on phishing detection with large language models. Machine Learning and Knowledge Extraction , 6 (1), 367–384. Varsik, S., & Vosberg, L. (2024). The Potential Impact of Artificial Intelligence on Equity and Inclusion in Education. OECD Artificial Intelligence Papers, No. 23 . OECD Publishing. Wan, G., Wu, Y., Chen, J., & Li, S. (2024). Dynamic self-consistency: Leveraging reasoning paths for efficient llm sampling. arXiv preprint arXiv :240817017. Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K. W., & Lim, E. P. (2023). Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091. Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., … Zhou, D. (2022).Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171. Xu, L., Li, A., Zhu, L., Xue, H., Zhu, C., Zhao, K., … Lan, Z. (2023). Superclue:A comprehensive chinese large language model benchmark. arXiv preprint arXiv:2307.15020. Zeng, H. (2023). Measuring massive multitask chinese understanding. arXiv preprint arXiv:2304.12986. Zhao, Q., Huang, Y., Lv, T., Cui, L., Sun, Q., Mao, S., … Wei, F. (2024). Mmlu-cf:A contamination-free multi-task language understanding benchmark. arXiv preprint arXiv:2412.15194. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., … Wen, J. R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2). Zhao, X., Li, M., Lu, W., Weber, C., Lee, J. H., Chu, K., & Wermter, S. (2023). Enhancing zero-shot chain-of-thought reasoning in large language models through logic. arXiv preprint arXiv:2309.13339. Additional Declarations No competing interests reported. Supplementary Files Table1.CorrectandwrongratesofChineselargelanguagemodelsinChinesereadingcomprehension.xlsx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 14 Apr, 2025 Reviews received at journal 11 Apr, 2025 Reviews received at journal 09 Apr, 2025 Reviewers agreed at journal 27 Mar, 2025 Reviewers agreed at journal 27 Mar, 2025 Reviews received at journal 26 Mar, 2025 Reviewers agreed at journal 26 Mar, 2025 Reviewers invited by journal 25 Mar, 2025 Submission checks completed at journal 24 Mar, 2025 First submitted to journal 20 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5990278","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":433562243,"identity":"1bb19878-ea35-4a1b-b1df-ab31d039e392","order_by":0,"name":"Huijin Le","email":"","orcid":"","institution":"Nanchang University","correspondingAuthor":false,"prefix":"","firstName":"Huijin","middleName":"","lastName":"Le","suffix":""},{"id":433562244,"identity":"b9ef44f4-b5a7-4dc0-a6ab-5e276397d912","order_by":1,"name":"Qiuling Zhang","email":"","orcid":"","institution":"Beijing Normal University","correspondingAuthor":false,"prefix":"","firstName":"Qiuling","middleName":"","lastName":"Zhang","suffix":""},{"id":433562245,"identity":"cb6b2db9-e625-4402-b963-40f820582229","order_by":2,"name":"Gang Xu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHklEQVRIie3RMWrDMBSAYRmBvKj2qtAmZ3ggSCmkyVVsDO6SQMZChxoM6dJ0TqCH8KRZ5oGz5AAZEwqZOhg6FQqtbTBdZAKZMugfjEH6sJ5MiM12gfmU4qEEMWDUSXRJCPcIb1acpIP0XhaxFPOR9N00z1cVYacIbLcgeBmH69ciwnrzSUJ2AYAAlLCbAt5PxzeMXBUfnIz6mabHvUE4qyDYzwEHDZmpqDqY9yA5iWWm2S0YCBWB/v/KTNF6luE1JxhmmjNhIEyEieBQbajJnXpuyW8n4RxJRZrxA3QUtkR3EuEumBTQXLLOl2rDGfXi3jtEco1saCIT9L8O5U/zK9PyWz1NfHdZiM/Hcf9tkx5NxHQj9QPaF5vNZrOd0x9z52ADKICWfwAAAABJRU5ErkJggg==","orcid":"","institution":"Beijing Normal University","correspondingAuthor":true,"prefix":"","firstName":"Gang","middleName":"","lastName":"Xu","suffix":""}],"badges":[],"createdAt":"2025-02-09 03:08:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5990278/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5990278/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79255774,"identity":"5e321484-8361-4539-9bc4-5da4294bd86d","added_by":"auto","created_at":"2025-03-26 08:54:58","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":136706,"visible":true,"origin":"","legend":"\u003cp\u003eThe evaluation process of LLMs\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/59230d3cc42d3776b168197e.png"},{"id":79255778,"identity":"1e74acca-a208-44a3-b970-1bdd7d07ba3c","added_by":"auto","created_at":"2025-03-26 08:54:58","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":164619,"visible":true,"origin":"","legend":"\u003cp\u003eComparative Evaluation of Capability Levels Across Six LLMs on Diverse Question Types\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/03e0bd5b7b833a411dd9082d.png"},{"id":79258400,"identity":"91b9871d-99e3-46de-87f4-d8daacd02b6e","added_by":"auto","created_at":"2025-03-26 09:10:58","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":203357,"visible":true,"origin":"","legend":"\u003cp\u003eExample of zero-shot thinking chain\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/0bcc5efd46448ba79609b980.png"},{"id":79255793,"identity":"5eae97d2-fbc3-41c5-b442-cffe21f944e5","added_by":"auto","created_at":"2025-03-26 08:54:58","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":178515,"visible":true,"origin":"","legend":"\u003cp\u003eSelf-consistent improvement consistency example\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/c9771eb059564b2c5f3df648.png"},{"id":79257023,"identity":"d4bc7bb0-3c45-465c-b647-0c7ef636d06d","added_by":"auto","created_at":"2025-03-26 09:02:58","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":185934,"visible":true,"origin":"","legend":"\u003cp\u003eExpert role example\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/932ae1cb2298c765c3405077.png"},{"id":79259299,"identity":"85589bcc-2c96-4fce-90f7-484324c76ed4","added_by":"auto","created_at":"2025-03-26 09:18:58","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1599022,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/ed29a170-a7af-4fa4-b62a-9733ffb74b77.pdf"},{"id":79255775,"identity":"f6d61294-99d8-4dc0-aaa6-ad6b33e754b4","added_by":"auto","created_at":"2025-03-26 08:54:58","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":10241,"visible":true,"origin":"","legend":"","description":"","filename":"Table1.CorrectandwrongratesofChineselargelanguagemodelsinChinesereadingcomprehension.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5990278/v1/534d2f0aa577dd8d27928277.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Comparative Study of Six Indigenous Chinese Large Language Models' Understanding Ability: An Assessment Based on 132 College Entrance Examination Objective Test Items","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn the current wave of digital transformation within education, the integration of large language models (LLMs) into language instruction presents both transformative opportunities and multifaceted challenges (N. Ahmad et al., 2023; Storey et al., 2024; Kim et al., 2025). While these advanced models demonstrate potential for enhancing pedagogical approaches, their implementation requires careful consideration of several critical dimensions (Jeon et al., 2023). Recent academic discussions emphasize that generative AI applications in education must address persistent concerns regarding output reliability, with particular attention to factual accuracy and contextual appropriateness in linguistic outputs. This becomes especially crucial in language education, where nuanced cultural sensitivities and sociolinguistic variations demand systems capable of navigating diverse communicative norms and value systems (Varsik et al., 2024). Moreover, emerging evidence highlights systemic considerations surrounding equitable implementation. Disparities in technological infrastructure and varying levels of digital literacy across different educational contexts may inadvertently exacerbate existing inequalities if not proactively addressed (Lee et al., 2024; Bulathwela et al., 2024). Concurrently, questions persist about the pedagogical effectiveness of LLM integration, particularly regarding the alignment between algorithmic outputs and established language acquisition principles. Therefore, it is imperative to establish rigorous evaluation frameworks to assess how these technologies complement rather than compromise fundamental educational objectives, ensuring they support the development of critical language competencies.\u003c/p\u003e \u003cp\u003eAs an important component of generative artificial intelligence technology, LLM is a neural network model trained based on a large amount of text data, and it possesses certain capabilities such as language understanding, generation, logic, and reasoning (Kumar, 2024). Although LLM has excellent performance in complex language processing compared to traditional text generation methods such as recurrent neural networks and long short-term memory networks, due to the technical architecture of LLM's own decoder, the results it generates still have certain hallucinations and errors (Jho, 2024; Matarazzo et al., 2024). Therefore, specific task evaluations of LLMs are needed to assess the comprehensive or specialized performance of LLMs, providing a reliable reference basis for users to select the applicable LLM according to task requirements.\u003c/p\u003e \u003cp\u003eRelevant institutions have successively released several evaluation systems for comprehensively assessing the performance of large language models. Typical ones include: MMLU (Zhao et al., 2024), BIG-Bench (Srivastava et al., 2024), HELM (Liang et al., 2024), C-Eval (Huang et al., 2024), MMCU (Zeng et al., 2024), SuperCLUE (Xu et al., 2024), and Chinese-LLM-Benchmark (Liu et al., 2024). Among them, MMLU and C-Eval mainly adopt closed-form (multiple-choice questions) decision tasks and use the correct rate of answers as the main evaluation indicator. BIG-Bench, HELM, MMCU, and Chinese-LLM-Benchmark focus on open-form (question-answering) generation tasks. Among the evaluation systems for LLMs released by various institutions, the datasets of the first three systems primarily assess LLM's performance in English environments. In contrast, the latter four systems targeting Chinese LLM's evaluation focus on multifaceted social life domains while demonstrating limited coverage of language and literature assessment tasks. Notably, SuperCLUE distinguishes itself through diversified task formats (including multiple-choice and open-ended questions) and more comprehensive evaluation metrics, endowing it with distinct advantages in assessing Chinese large language models compared to other benchmark systems.\u003c/p\u003e \u003cp\u003eResearch shows that as the scale of model parameters increases, the accuracy of LLMs in specific tasks does not improve linearly. Instead, an emergent ability emerges, and the accuracy of the models in complex reasoning tasks (such as mathematical problems and logical reasoning) increases significantly (Zhao et al., 2023; Bean et al., 2023). However, the rankings of Chinese LLMs included in the SuperCLUE evaluation list are not consistent in terms of overall performance and semantic understanding dimensions. More importantly, the evaluation data set of the semantic understanding dimension of SuperCLUE adopts general language tasks and can only measure the general reading comprehension ability of LLMs. To better study the applicability and effectiveness of artificial intelligence technology in Chinese education and teaching, it is necessary to collect specialized Chinese reading comprehension tasks to further evaluate the performance of LLMs. Based on this, this study will select SuperCLUE as the frame of reference for the evaluation results of the semantic understanding ability of Chinese LLMs. To help teachers select LLM based on evidence, this article attempts to respond to three questions: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) Which domestic LLMs have a better overall performance in semantic understanding ability; (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Which one performs better when answering different types of reading comprehension tasks such as language application, modern text reading, ancient poetry reading, and classical Chinese reading; (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) When the performance of answering Chinese reading comprehension tasks is not good, how to use prompt strategies to deeply explore the semantic understanding potential of LLMs.\u003c/p\u003e"},{"header":"Literature review","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eEvaluation Content\u003c/h2\u003e \u003cp\u003eThe evaluation framework for large language models (LLMs) has evolved from traditional single-task performance assessment to multidimensional competency analysis. Early research primarily focused on accuracy evaluations for natural language processing tasks such as text classification and machine translation. However, as model scales expand and societal applications deepen, evaluation dimensions have gradually extended to ethical risks, multimodal capabilities, and societal impact. For instance, Laskar et al. (2024) highlighted that LLMs may generate factual errors (e.g., \"hallucinations\") or embed social biases during content generation, necessitating alignment optimization through human feedback reinforcement learning (RLHF). The SEED-Bench benchmark developed by Li et al. (2023) revealed significant deficiencies in current multimodal models regarding spatial relationship comprehension and cross-modal reasoning. In recent years, evaluation frameworks have further expanded to safety assessments, encompassing content compliance, privacy protection, and adversarial attack robustness. Hendy et al. (2023) proposed a four-dimensional evaluation framework (functionality, performance, alignment, and safety), emphasizing the need for systematic analysis of potential risks such as value misalignment and maliciously induced outputs. For example, He et al. (2024) utilized clinical laboratory examination questions and case studies to demonstrate that ChatGPT-4.0 and ERNIE Bot-4.0 both achieved a 60% passing rate in medical knowledge proficiency but exhibited significant errors in complex case analyses.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEvaluation Methods\u003c/h3\u003e\n\u003cp\u003eEvaluation methodologies are broadly categorized into automated and human assessments. Automated evaluation relies on traditional metrics (e.g., BLEU, ROUGE) and LLM-based evaluators. Liu et al. (2023) proposed enhancing interpretability through chain-of-thought prompting strategies, though such methods still suffer from low consistency with human judgments. Kocmi and Federmann (2023) found that LLM evaluators may prioritize fluency over semantic accuracy, leading to scoring biases. Human evaluation involves expert or user participation to provide contextually grounded feedback. For instance, OpenAI employed human testing during GPT-4 development to validate multistep reasoning capabilities (Achiam et al., 2023). Lundgren (2024) demonstrated that GPT-4 achieved parity with human instructors in grading political science master\u0026rsquo;s theses but exhibited conservative scoring patterns and low reliability.\u003c/p\u003e \u003cp\u003eRecent years have witnessed the rise of comprehensive evaluation frameworks. Benchmarks like GLUE and SuperGLUE, covering core tasks such as text classification and reasoning, drove model performance from BERT\u0026rsquo;s 80.5 score to near-human baselines (Gao et al., 2024). SQuAD 2.0 introduced unanswerable scenarios to test models\u0026rsquo; uncertainty handling (Devlin et al., 2024). OMGEval supports 56-language generation tasks, prioritizing low-resource languages (e.g., Swahili) (Liu et al., 2024). EXAMS-V integrates multidisciplinary exam questions (e.g., mathematics, biology), requiring multimodal reasoning with text and charts, evaluated via accuracy and consistency metrics (Das et al., 2024). Nevertheless, challenges persist, including the black-box nature of proprietary models and the lack of standardized multimodal evaluation criteria. Liang et al. (2022) proposed the HELM framework, spanning 42 scenarios and 7 evaluation dimensions, which exposed performance disparities in low-resource language tasks. Dong et al. (2024) developed CLR-Bench to systematically assess college-level complex reasoning. Calibration, an emerging metric, was applied by Kumar et al. (2024) to quantify alignment between model confidence and true probabilities, revealing that well-calibrated models are better suited for high-stakes scenarios like medical diagnosis.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eEvaluating large language models\u003c/h2\u003e \u003cp\u003eAccording to the ease of use and usefulness of the technology acceptance model proposed by Davis to select the participating LLMs (Davis, 2024). Ease of use refers to the extent to which people feel that using a certain technology is simple and effortless; usefulness refers to the possibility that people will choose a technology or tool with better performance to increase the possibility of generating actual benefits. Given that the ease of use of the existing LLMs is very good, this study only excludes foreign LLMs such as ChatGPT and Gemini from the perspective of accessibility, and evaluates the usefulness of Chinese LLMs from the perspective of Chinese reading comprehension. Referring to the latest evaluation ranking list of Chinese LLMs released by SuperCLUE on February 27, 2024, according to the comprehensive performance ranking, the top six domestic general LLMs of Wenxin Yiyan, GLM-4, Tongyi Qianwen, Baichuan, KimiChat, and Xunfei Spark were selected in sequence as the evaluation objects. Their evaluation scores in the semantic understanding dimension of SuperCLUE, from high to low, are Tongyi Qianwen, GLM-4, Wenxin Yiyan, KimiChat, Baichuan, and Xunfei Spark. All six LLMs are free versions, and the tests are conducted through the web page to simulate the real scenarios of daily use by teachers and students from the user's perspective.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEvaluation data set\u003c/h3\u003e\n\u003cp\u003eThe National College Entrance Examination, as one of the most crucial educational evaluation systems in China, has Chinese-language multiple-choice questions that exhibit extraordinary authority and representativeness. These questions comprehensively evaluate students' diverse capabilities, such as linguistic knowledge, reading comprehension, and logical reasoning. Covering a wide range of linguistic phenomena and knowledge points, they are of great significance for assessing the problem-solving abilities of large language models within complex Chinese-language contexts. Moreover, the relatively rich publicly-available data resources of Gaokao Chinese multiple-choice questions enable large-scale data collection and analysis, thus guaranteeing the wide applicability and reliability of research results.\u003c/p\u003e \u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, To ensure that the evaluation data set is not in the pre-training corpus of LLMs as much as possible, this study selected 133 multiple-choice questions from the Chinese college entrance examination Chinese language papers of the National Volume A, National Volume B, New College Entrance Examination Volume I, and New College Entrance Examination Volume II from 2021 to 2023 as special test tasks. Since some LLMs only support pure text format input, one multiple-choice question of modern text reading with picture information was deleted, resulting in 132 test questions. Among them, there are 19 questions of language application, 36 questions of classical Chinese reading, 12 questions of ancient poetry reading, and 65 questions of modern text reading. Although the number of different types of test questions is not balanced enough, due to the structural constraints of the college entrance examination questions, no additional evaluation data will be artificially added.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBasic information about the data set\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQuestion type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eQuantity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eExample of questions\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLanguage\u003c/p\u003e \u003cp\u003eapplication\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e请阅读上传的文档材料, 将下列俗语填入文中括号内, 恰当的一项是(), 并说明选择此选项的理由。\u003c/p\u003e \u003cp\u003eA.干打雷不下雨;B.又吃鱼又嫌腥\u003c/p\u003e \u003cp\u003eC.前怕狼后怕虎;D.首尾不能兼顾\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassical\u003c/p\u003e \u003cp\u003eChinese reading\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e请阅读上传的文档材料, 关于文中江边人们谈论季札的部分, 下列说法不正确的一项是(), 并说明选择此选项的理由。\u003c/p\u003e \u003cp\u003eA༎那位老人欣赏季札不就王位的高洁, 也称赞他以美好的行为感动了世人。\u003c/p\u003e \u003cp\u003eB༎那位年轻人认为季札不顾百姓死活, 只顾独善其身, 逃避了济世的责任。\u003c/p\u003e \u003cp\u003eC༎季札挂剑一事进一步说明了他的品行, 也为后文的子胥赠剑做了铺垫。\u003c/p\u003e \u003cp\u003eD༎季札的退耕田园, 与下文渔夫的泛舟江上, 共同表达出本文的隐逸主题。\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAncient\u003c/p\u003e \u003cp\u003epoetry reading\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e请阅读上传的文档材料, 下列选项中, 加点的词语和文中\u0026ldquo;槐蝉\u0026rdquo;所用修辞手法不同的一项是(), 并说明选择此选项的理由。\u003c/p\u003e \u003cp\u003eA༎主人下马客在船, 举酒欲饮无管弦。\u003c/p\u003e \u003cp\u003eB༎埋骨何须桑梓地, 人生无处不青山。\u003c/p\u003e \u003cp\u003eC༎六军不发无奈何, 宛转娥眉马前死。\u003c/p\u003e \u003cp\u003eD༎心非木石岂无感, 吞声瑰躅不敢言。\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModern\u003c/p\u003e \u003cp\u003etext reading\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e请阅读上传的文档材料, 下列句子中的\u0026ldquo;谁\u0026rdquo;和\u0026ldquo;耳机一戴, 谁也不爱\u0026rdquo;中的\u0026ldquo;谁\u0026rdquo;, 意义和用法相同的一项是(), 并说明选择此选项的理由。\u003c/p\u003e \u003cp\u003eA.怅寥廓, 问苍茫大地, 谁主沉浮༟\u003c/p\u003e \u003cp\u003eB.生活中谁都需要表达和交流。\u003c/p\u003e \u003cp\u003eC.我本来是跟他开玩笑的, 谁知道他竟然生气了。\u003c/p\u003e \u003cp\u003eD.我越来越深刻地感觉到谁是我们最可爱的人!\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eResearch design\u003c/h2\u003e \u003cp\u003eThe evaluation process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) The reading materials corresponding to each question are stored separately in PDF document format. (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Input the prompt, the framework is: \"Please read the uploaded document materials\" + the question stem of the multiple-choice question + \"and explain the reason for choosing this option\" + the options of the multiple-choice question. (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) Each question is completed in an independent conversation window, and the LLMs will return the generated options and the reasons for the selection. (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e) The generation results are divided into three situations: ① For the questions answered correctly, they are directly included in the correct rate; ② For the questions answered wrongly, the explanations and answers are regenerated by improving the prompt, and then compared and analyzed with the correct answers; ③ The questions with abnormal results are retested to exclude the contingency of abnormal occurrences. It should be noted that the correct answers obtained after improving the prompt are not included in the total correct rate. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The evaluation process of LLMs\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eThis study employed EXCEL software to compare the test results of Chinese large language models (LLMs) on Chinese language multiple-choice questions in the National College Entrance Examination with standard answers, through which valid data were filtered. Statistical analysis of the test results was performed using SPSS 26.0 software, including descriptive statistics to examine the accuracy rates of six LLMs across different question types. Finally, prompt engineering strategies were employed to modify erroneous responses through systematic error correction attempts.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eComparison of comprehensive semantic understanding ability\u003c/h2\u003e \u003cp\u003eComprehensive semantic comprehension capability refers to the overall performance level of large language models (LLMs) in semantic-understanding tasks across diverse text genres. This reflects the model's ability to comprehensively grasp and interpret semantics across various textual categories. As shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, the overall correct rate of the six LLMs is 62% and the standard deviation is 6%, indicating that the performance of the six LLMs in semantic understanding is not much different, but the difference in the special semantic understanding is slightly larger, with a large degree of dispersion, and the standard deviation ranges from 9\u0026ndash;18%. The overall correct rates of Tongyi Qianwen and GLM-4 are 70% and 69% respectively, and their comprehensive ability in semantic understanding performs well; the overall correct rate of Xunfei Spark is 62%, ranking third, and its comprehensive ability in semantic understanding is relatively low. The overall correct rate of Baichuan is 55%, and its comprehensive ability in semantic understanding is weak, ranking sixth.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCorrect and wrong rates of Chinese large language models in Chinese reading comprehension\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eQuestion type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLanguage application\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eClassical Chinese reading\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAncient poetry reading\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eModern text reading\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e132\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTongyi Qianwen\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e56%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e83%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e75%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e70%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e39%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e23%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e27%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eGLM-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e92%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e77%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e69%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e22%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e28%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eXunfei Spark\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e74%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e67%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e42%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e60%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e62%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e21%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e28%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e25%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e24%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eWenxin Yiyan\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e37%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e39%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e83%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e74%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e60%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e26%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e36%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eKimiChat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e39%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e67%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e60%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e57%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e42%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e33%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e35%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e42%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eBaichuan\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e67%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e55%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e53%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e33%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e42%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e45%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eM\u0026thinsp;\u0026plusmn;\u0026thinsp;SD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e56%\u0026plusmn;15%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e51%\u0026plusmn;11%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e72%\u0026plusmn;18%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e68%\u0026plusmn;9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e62%\u0026plusmn;6%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eError\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e41\u0026thinsp;\u0026plusmn;\u0026thinsp;15%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e44%\u0026plusmn;11%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21%\u0026plusmn;10%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e29%\u0026plusmn;8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e34\u0026thinsp;\u0026plusmn;\u0026thinsp;9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eSpecial semantic understanding ability comparison\u003c/h2\u003e \u003cp\u003eSpecialized semantic comprehension capability focuses on LLMs' performance in semantic-understanding tasks targeting specific text types, demonstrating the model's domain-specific expertise and differential competencies in processing semantic patterns characteristic of particular textual genres. In the language application test questions, the overall correct rate of the six LLMs is 56%, and the standard deviation is 15%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of language application are relatively large. The correct rates from high to low are Tongyi Qianwen (74%), Xunfei Spark (74%), KimiChat (58%), GLM-4 (47%), Baichuan (47%), and Wenxin Yiyan (37%), indicating that Tongyi Qianwen and Xunfei Spark perform well in language application.\u003c/p\u003e \u003cp\u003eIn the classical Chinese reading test questions, the overall correct rate of the six LLMs is 51%, and the standard deviation is 11%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of classical Chinese reading are relatively large. The correct rates from high to low are Xunfei Spark (67%), GLM-4 (58%), Tongyi Qianwen (56%), Baichuan (47%), Wenxin Yiyan (39%), and KimiChat (39%), indicating that Xunfei Spark performs well in classical Chinese reading.\u003c/p\u003e \u003cp\u003eIn the ancient poetry reading test questions, the overall correct rate of the six LLMs is 72%, and the standard deviation is 18%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of ancient poetry reading are very large. The correct rates from high to low are GLM-4 (92%), Tongyi Qianwen (83%), Wenxin Yiyan (83%), KimiChat (67%), Baichuan (67%), and Xunfei Spark (42%), indicating that GLM-4 performs well in ancient poetry reading.\u003c/p\u003e \u003cp\u003eIn the modern text reading test questions, the overall correct rate of the six LLMs is 68%, and the standard deviation is 9%, indicating that the performance differences of the six LLMs in the special semantic understanding ability of modern text reading are not significant. The correct rates from high to low are GLM-4 (77%), Tongyi Qianwen (75%), Wenxin Yiyan (74%), Xunfei Spark (60%), KimiChat (60%), and Baichuan (58%), indicating that GLM-4 performs well in modern text reading. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, The special semantic understanding ability performance of LLMs in the four types of test questions each has its advantages.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003ePrompt strategy can effectively correcxt the wrong results of large language models\u003c/h2\u003e \u003cp\u003eThe zero-shot thinking chain strategy is a method that allows the model to handle tasks with almost no need for any additional training. That is, by adding a simple prompt-\"Let's think step by step\", this triggers the LLMs to expand their thinking step by step according to the input prompt, thereby triggering extensive cognitive abilities, rather than being limited to the specific skills of a particular task (Kojima et al., 2023). For example, in question 4 of the 2021 National Volume A paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of \"think step by step\", GLM-4 gave the correct option and analysis as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn view of the randomness characteristic of LLMs' responses, making them generate multiple reasoning paths or answers, and then conducting internal voting to finally select the most consistent or most common answer for output. This method can help reduce the randomness of single sampling or reasoning, thereby improving the accuracy and reliability of the entire system (Wang et al., 2022). For example, in question 3 of the 2021 National Volume B paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of \"Please generate ten answers internally, and conduct internal voting on the answers, and select and output the answer with the most votes\", GLM-4 gave the correct option and analysis as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe core idea of the expert role strategy is to invoke expert knowledge in specific domains by designing and optimizing prompts, thereby improving factual accuracy, knowledge depth, and reducing biases (Wan et al., 2024). When used, by constructing a context containing an expert role, the professional knowledge in related fields in LLMs is activated, while irrelevant information is suppressed, reducing the ambiguity of the model's understanding of the problem, and then answering the knowledge content that conforms to the characteristics of that role. For example, in question 17 of the 2021 National Volume B paper, after giving the wrong option by inputting the initial prompt, when inputting the prompt again and adding the instruction of \"As an expert in the field of Chinese\", GLM-4 gave the correct option and analysis as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe findings demonstrate that among four types of Chinese reading comprehension tasks, Qwen-Turbo and GLM-4 exhibited superior comprehensive performance (with overall accuracy rates of 70% and 69% respectively), particularly showing significant advantages in specialized tasks. For instance, GLM-4 achieved an exceptional accuracy of 92% in classical poetry comprehension, substantially outperforming other models. These results corroborate the critical role of pre-training corpora and algorithmic design in model performance (Devlin et al., 2019; Gururangan et al., 2020; Knafou et al., 2020), while aligning with GLM-4's ranking in semantic understanding dimensions within the SuperCLUE evaluation framework (a comprehensive Chinese language model evaluation benchmark). However, the study revealed suboptimal performance across all models in classical Chinese (wenyanwen) reading comprehension tasks (average accuracy rate of 51%), indicating inadequate capture of classical Chinese linguistic features by general-purpose LLMs, necessitating improvement through domain-specific corpora and algorithmic optimization (Lin et al., 2022). Regarding erroneous responses from LLMs, this research demonstrates that prompt engineering strategies-including zero-shot chain-of-thought prompting, self-consistency voting, and expert role assignment-can significantly enhance response accuracy (Wang et al., 2023; Zhao et al., 2023; Qin et al., 2023; Trad et al., 2023).\u003c/p\u003e \u003cp\u003eThe study further proposes three practical recommendations: developing domain-specific LLMs for Chinese language education, constructing interdisciplinary AI agents, and promoting pedagogical paradigm transformation. For instance, the suboptimal performance in classical Chinese reading comprehension (demonstrating an average error rate of 44%) indicates that general-purpose LLMs inadequately address instructional requirements, necessitating enhancements through classical Chinese corpora enrichment and algorithmic fine-tuning to improve cultural interpretation capabilities (Lin et al., 2022). Simultaneously, the findings suggest that establishing Chinese-language specialized AI agents could synergize the advantages of multiple models, such as integrating the language application proficiency of Tongyi Qianwen with the poetry analysis strength of GLM-4 (Ding et al., 2024). The digital-intelligent transformation of pedagogical perspectives requires balanced attention to technological empowerment and humanistic considerations, cautioning against excessive reliance on model-generated content (Heersmink et al., 2024; Faisal, 2024).\u003c/p\u003e \u003cp\u003eThe innovation of this study lies in conducting the first systematic evaluation of mainstream domestic large language models (LLMs) on Chinese language discipline-specific tasks, while proposing targeted optimization strategies based on prompt engineering. Compared with general evaluation benchmarks like SuperCLUE, this research focuses on subject-oriented scenarios, providing educators with direct references for model selection and pedagogical integration. The experimental design utilizes national college entrance examination questions as the test dataset, with multi-round prompt experiments revealing LLMs' adaptability through prompt-based interventions. These findings not only enrich the disciplinary dimensions of Chinese LLMs evaluation frameworks but also establish methodological foundations for developing educational-purpose AI agents. However, the current research is limited to objective question assessments, and future work should explore LLMs' potential in subjective task scenarios, particularly essay evaluation, through open-ended generation tasks.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eTongyi Qianwen and GLM-4 have firmly ranked in the top two in semantic understanding ability and have respectively won the top two in three special task abilities. Chinese language teachers can give priority to choosing them as teaching and research tools such as text interpretation, teaching design, test question formulation, and subject research. The overall performance of Chinese LLMs in tasks such as classical Chinese reading comprehension is not ideal. Prompt strategies such as thinking chains, using internal voting evaluations, and expert roles can be used to help LLMs understand the intentions of Chinese language teachers and enable LLMs to output content that meets the expectations of Chinese language teachers.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLLMs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLarge language models\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003e \u003cb\u003eEthics declarations\u003c/b\u003e \u003c/h2\u003e \u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eClinical trial number\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eConsent for publication\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis work was supported b the priority concern project of the \"13th Five-Year Plan\" of Beijing Education Science in 2020 [grant number CHEA2020025].\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eLe. Conceptualization, Methodology; Xu. Writing- Original draft prepa-ration, Data curation; Zhang.Writing- Reviewing and Editing.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAchiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., \u0026amp; McGrew, B. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmad, N., Murugesan, S., \u0026amp; Kshetri, N. (2023). Generative Artificial Intelligence and the Education Sector. \u003cem\u003eComputer\u003c/em\u003e, \u003cem\u003e56\u003c/em\u003e(6), 72\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBean, A. M., Korgul, K., Krones, F., McCraith, R., \u0026amp; Mahdi, A. (2023). Do Large Language Models have Shared Weaknesses in Medical Question Answering? \u003cem\u003earXiv preprint arXiv:2310.07225\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBulathwela, S., P\u0026eacute;rez-Ortiz, M., Holloway, C., Cukurova, M., \u0026amp; Shawe-Taylor, J. (2024). Artificial Intelligence Alone Will Not Democratise Education: On Educational Inequality, Techno-Solutionism and Inclusive Tools. \u003cem\u003eSustainability\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(2), 781.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDas, R. J., Hristov, S. E., Li, H., Dimitrov, D. I., Koychev, I., \u0026amp; Nakov, P. (2024). Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models. \u003cem\u003earXiv preprint arXiv\u003c/em\u003e:240310378.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDavis, F. D. (1989). \u003cem\u003ePerceived usefulness, perceived ease of use, and user acceptance of information technology\u003c/em\u003e (pp. 319\u0026ndash;340). MIS quarterly.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDevlin, J., Chang, M. W., Lee, K., \u0026amp; Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for compu.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDevlin, J., Chang, M. W., Lee, K., \u0026amp; Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171\u0026ndash;4186).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing, H., Fan, Z., Guehring, I., Gupta, G., Ha, W., Huan, J., \u0026hellip; Zhou, H. (2024). Reasoning and planning with large language models in code development. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 6480\u0026ndash;6490).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDong, J., Hong, Z., Bei, Y., Huang, F., Wang, X., \u0026amp; Huang, X. (2024). CLR-Bench: Evaluating Large Language Models in College-level Reasoning. arXiv preprint arXiv:2410.17558.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFaisal, E. (2024). Unlock the potential for Saudi Arabian higher education: a systematic review of the benefits of ChatGPT. \u003cem\u003eFrontiers in Education\u003c/em\u003e (Vol. 9, p. 1325601). Frontiers Media SA.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao, M., Hu, X., Ruan, J., Pu, X., \u0026amp; Wan, X. (2024). Llm-based nlg evaluation: Current status and challenges. \u003cem\u003earXiv preprint arXiv\u003c/em\u003e:240201383.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., \u0026amp; Smith, N. A. (2020). Don't stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:2004.10964.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe, Z., Bhasuran, B., Jin, Q., Tian, S., Hanna, K., Shavor, C., \u0026hellip; Lu, Z. (2024). Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study. Journal of medical Internet research, 26, e56655.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeersmink, R., de Rooij, B., Clavel V\u0026aacute;zquez, M. J., \u0026amp; Colombo, M. (2024). A phenomenology and epistemology of large language models: Transparency, trust, and trustworthiness. \u003cem\u003eEthics and Information Technology\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(3), 41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., \u0026hellip; Awadalla,H. H. (2023). How good are gpt models at machine translation? a comprehensive evaluation.arXiv preprint arXiv:2302.09210.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., \u0026hellip; He, J. (2023). C-eval:A multi-level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing Systems, 36, 62991\u0026ndash;63010.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJeon, J., \u0026amp; Lee, S. (2023). Large Language Models in Education: A Focus on the Complementary Relationship Between Human Teachers and ChatGPT. \u003cem\u003eEduc Inf Technol\u003c/em\u003e, \u003cem\u003e28\u003c/em\u003e, 15873\u0026ndash;15892.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJho, H. (2024). Leveraging generative AI in physics education: Addressing hallucination issues in large language models.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, J., Yu, S., \u0026amp; Detrick, R. (2025). Exploring Students\u0026rsquo; Perspectives on Generative AI-assisted Academic Writing. \u003cem\u003eEduc Inf Technol\u003c/em\u003e, \u003cem\u003e30\u003c/em\u003e, 1265\u0026ndash;1300.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKnafou, J., Naderi, N., Copara, J., Teodoro, D., \u0026amp; Ruch, P. (2020). BiTeM at WNUT 2020 shared task-1: named entity recognition over wet lab protocols using an ensemble of contextual language models. In Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020) (pp. 305\u0026ndash;313).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKocmi, T., \u0026amp; Federmann, C. (2023). Large language models are state-of-the-art evaluators of translation quality. \u003cem\u003earXiv preprint arXiv\u003c/em\u003e:230214520.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKojima, T., Gu, S. S., Reid, M., Matsuo, Y., \u0026amp; Iwasawa, Y. (2022). Large language models are zero-shot reasoners. \u003cem\u003eAdvances in neural information processing systems\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e, 22199\u0026ndash;22213.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar, A., Morabito, R., Umbet, S., Kabbara, J., \u0026amp; Emami, A. (2024). Confidence under the hood: An investigation into the confidence-probability alignment in large language models. arXiv preprint arXiv:2405.16282.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar, P. (2024). Large language models (LLMs): survey, technical frameworks, and future challenges. \u003cem\u003eArtificial Intelligence Review\u003c/em\u003e, \u003cem\u003e57\u003c/em\u003e(10), 260.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H.,\u0026hellip; Huang, J. (2024). A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 13785\u0026ndash;13816).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, J., Hicke, Y., Yu, R., Brooks, C., \u0026amp; Kizilcec, R. F. (2024). The Life Cycle of Large Language Models in Education: A Framework for Understanding Sources of Bias. \u003cem\u003eBritish Journal of Educational Technology\u003c/em\u003e, \u003cem\u003e55\u003c/em\u003e, 1982\u0026ndash;2002.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., \u0026amp; Shan, Y. (2024). Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 13299\u0026ndash;13308).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., \u0026hellip; Koreeda,Y. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., \u0026hellip; Koreeda,Y. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, C. T., \u0026amp; Ma, W. Y. (2022). HanTrans: An Empirical Study on Cross-Era Transferability of Chinese Pre-trained Language Model. In Proceedings of the 34th Conference on Computational Linguistics and Speech Processing (ROCLING 2022) (pp. 164\u0026ndash;173).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, C., Yu, L., Li, J., Jin, R., Huang, Y., Shi, L., \u0026hellip; Xiong, D. (2024). Openeval:benchmarking Chinese LLMs across capability, alignment and safety. arXiv preprint arXiv:2403.12316.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, P., Yuan, W., Fu, J., Jiang, Z., \u0026amp; Hayashi, H. (2023). \u0026amp; Pre-train, G. N. prompt, and predict: A systematic survey of prompting methods in natural language processing., 55. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3560815\u003c/span\u003e\u003cspan address=\"10.1145/3560815\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, 1\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Y., Xu, M., Wang, S., Yang, L., Wang, H., Liu, Z., \u0026hellip; Yang, E. (2024). OMGEval:An Open Multilingual Generative Evaluation Benchmark for Large Language Models. arXiv preprint arXiv:2402.13524.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundgren, M. (2024). Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders. arXiv preprint arXiv:2406.16510.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMatarazzo, A., \u0026amp; Torlone, R. (2025). A Survey on Large Language Models with some Insights on their Capabilities and Limitations. arXiv preprint arXiv:2501.04040.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQin, L., Chen, Q., Wei, F., Huang, S., \u0026amp; Che, W. (2023). Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. \u003cem\u003earXiv preprint arXiv\u003c/em\u003e:231014799.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSrivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., \u0026hellip; Wang,G. (2022). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStorey, V. A., \u0026amp; Wagner, A. (2024). Integrating Artificial Intelligence (AI) Into Adult Education: Opportunities, Challenges, and Future Directions. \u003cem\u003eInternational Journal of Adult Education and Technology (IJAET)\u003c/em\u003e, \u003cem\u003e15\u003c/em\u003e(1), 1\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrad, F., \u0026amp; Chehab, A. (2024). Prompt engineering or fine-tuning? a case study on phishing detection with large language models. \u003cem\u003eMachine Learning and Knowledge Extraction\u003c/em\u003e, \u003cem\u003e6\u003c/em\u003e(1), 367\u0026ndash;384.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVarsik, S., \u0026amp; Vosberg, L. (2024). \u003cem\u003eThe Potential Impact of Artificial Intelligence on Equity and Inclusion in Education. OECD Artificial Intelligence Papers, No. 23\u003c/em\u003e. OECD Publishing.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWan, G., Wu, Y., Chen, J., \u0026amp; Li, S. (2024). Dynamic self-consistency: Leveraging reasoning paths for efficient llm sampling. \u003cem\u003earXiv preprint arXiv\u003c/em\u003e:240817017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K. W., \u0026amp; Lim, E. P. (2023). Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., \u0026hellip; Zhou, D. (2022).Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, L., Li, A., Zhu, L., Xue, H., Zhu, C., Zhao, K., \u0026hellip; Lan, Z. (2023). Superclue:A comprehensive chinese large language model benchmark. arXiv preprint arXiv:2307.15020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeng, H. (2023). Measuring massive multitask chinese understanding. \u003cem\u003earXiv preprint\u003c/em\u003e arXiv:2304.12986.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, Q., Huang, Y., Lv, T., Cui, L., Sun, Q., Mao, S., \u0026hellip; Wei, F. (2024). Mmlu-cf:A contamination-free multi-task language understanding benchmark. arXiv preprint arXiv:2412.15194.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., \u0026hellip; Wen, J. R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, X., Li, M., Lu, W., Weber, C., Lee, J. H., Chu, K., \u0026amp; Wermter, S. (2023). Enhancing zero-shot chain-of-thought reasoning in large language models through logic. \u003cem\u003earXiv preprint\u003c/em\u003e arXiv:2309.13339.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"language-testing-in-asia","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ltia","sideBox":"Learn more about [Language Testing in Asia](http://languagetestingasia.springeropen.com)","snPcode":"40468","submissionUrl":"https://submission.springernature.com/new-submission/40468/3","title":"Language Testing in Asia","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Open","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Large language models, Chinese comprehension skills, Artificial intelligence, Prompt strategy","lastPublishedDoi":"10.21203/rs.3.rs-5990278/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5990278/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTo assist Chinese language teachers in making evidence-based choices of useful and user-friendly domestic large language models in teaching and research, the study took 132 objective questions from the national college entrance examination Chinese language papers from 2021 to 2023 as the data set to assess the performance of six domestic large language models, namely Tongyi Qianwen, GLM-4, KimiChat, Baichuan, Wenxin Yiyan, and Xunfei Spark, in semantic understanding. The assessment revealed that the overall correct rates of the responses of the above six large language models to the questions were 70%, 69%, 57%, 55%, 60%, and 62% respectively. Among them, Tongyi Qianwen and Xunfei Spark performed best in language application questions, with correct rates of 74% each; GLM-4 performed best in ancient poetry reading and modern text reading questions, with correct rates reaching 92% and 77% respectively. The performance of the six large language models in classical Chinese reading questions was not ideal. For the wrongly answered test questions, the researchers corrected and analyzed the answers using the prompt strategy. Finally, the paper put forward several suggestions for promoting the assistance of large language models in Chinese language teaching and research.\u003c/p\u003e","manuscriptTitle":"A Comparative Study of Six Indigenous Chinese Large Language Models' Understanding Ability: An Assessment Based on 132 College Entrance Examination Objective Test Items","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-26 08:54:53","doi":"10.21203/rs.3.rs-5990278/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-14T14:53:32+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-11T14:57:58+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-10T01:50:21+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"152086351623269223352307242608184132801","date":"2025-03-27T11:43:40+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"58553048506493089984974527310245365131","date":"2025-03-27T08:01:46+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-26T08:03:42+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"18387991626926382253266122090004678671","date":"2025-03-26T07:09:27+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-25T06:36:18+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-24T06:25:33+00:00","index":"","fulltext":""},{"type":"submitted","content":"Language Testing in Asia","date":"2025-03-20T15:03:34+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"language-testing-in-asia","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ltia","sideBox":"Learn more about [Language Testing in Asia](http://languagetestingasia.springeropen.com)","snPcode":"40468","submissionUrl":"https://submission.springernature.com/new-submission/40468/3","title":"Language Testing in Asia","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Open","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3e95e989-7c43-420d-ab64-c15c53c78b52","owner":[],"postedDate":"March 26th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-04-30T06:08:32+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-26 08:54:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5990278","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5990278","identity":"rs-5990278","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00