In the rapidly evolving field of artificial intelligence (AI), large language model s (LLMs) have demonstrated impressive capabilities in generating natural language. However, their proficiency in specialized domains, particularly in the field of systems engineering (SE), remains less explored and unquantified. This paper introduces SysEngBench, a novel benchmark specifically designed to evaluate LLMs in the context of SE concepts and applications. SysEngBench encompasses a comprehensive set of tasks derived from core SE processes, including requirements analysis, system architecture design, risk management, and stakeholder communication, to provide an assessment of language model abilities.Our evaluation of leading LLMs using SysEngBench reveals a strong correlation between model scale and performance, with state-of-the-art models like GPT-4o and Claude 3.5 Sonnet achieving near-human-level accuracy (above 95%) across multiple categories. Smaller models like Llama-3.2 1B exhibit significantly higher defect densities and lower accuracy scores, highlighting their limitations in handling SE tasks. However, Pareto analysis demonstrates that while larger models generally outperform smaller ones, some mid-sized models like Phi-3.5-mini-instruct achieve competitive accuracy with significantly lower computational requirements. These findings suggest pathways for practitioners and future research, particularly in optimizing smaller models for domain-specific reasoning through fine-tuning or knowledge integration. SysEngBench provides a systematic approach to assessing AI's impact on SE, offering insights into the trade-offs between model efficiency and accuracy. By establishing a benchmark for LLM evaluation in this domain, we provide a cohesive, extensible, and effective method to refine AI's role in the SE discipline.
更多
查看译文
关键词
benchmarking,large language models,performance evaluation,SysEngBench,systems engineering