arXiv ScienceSearch

arXiv subjects

Jannatul Shefa

Publications and source records attributed to Jannatul Shefa.

3 recordsLinked to original sources

Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses

Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns. Because requirements are often written in natural language, recent advances in generative AI have raised expectations that large language models (LLMs) can absorb requirement quality assessment, a task otherwise slow and human expertise-intensive. Yet empirical evidence on whether LLMs can be trusted to do so remains scarce. This study presents the first benchmarking analysis of off-the-shelf LLM performance for requirement quality evaluation. Against an expert-derived ground truth built on INCOSE quality criteria, we evaluate ten models spanning two families (OpenAI and Anthropic) and five generations each, across one hundred independent runs, two requirement sets, and five sampling temperatures. Four contributions follow. First, we quantify a strongly asymmetric error profile: across all models and runs, the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%. Second, performance degrades significantly where SE judgment is required, as necessity and correctness issues are almost always missed. Third, generational progress is non-monotonic, so newer models cannot be assumed better. Fourth, this error behavior shifts only modestly and non-monotonically across sampling temperatures, indicating characteristic model deficiencies rather than inherent stochasticity. Off-the-shelf LLMs are therefore not yet trustworthy autonomous evaluators. Findings also warrant caution for Agentic AI developers: orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them. Their defensible near-term role is human-in-the-loop decision support.

cs.SE

What is the Return on Investment of Digital Engineering for Complex Systems Development? Findings from a Mixed-Methods Study on the Post-production Design Change Process of Navy Assets

Complex engineered systems routinely face schedule and cost overruns, along with poor post-deployment performance. Championed by both INCOSE and the U.S. Department of Defense (DoD), the systems engineering (SE) community has increasingly looked to Digital Engineering (DE) as a potential remedy. Despite this growing advocacy, most of DE's purported benefits remain anecdotal, and its return on investment (ROI) remains poorly understood. This research presents findings from a case study on a Navy SE team responsible for the preliminary design phase of post-production design change projects for Navy assets. Using a mixed-methods approach, we document why complex system sustainment projects are routinely late, where and to what extent schedule slips arise, and how a DE transformation could improve schedule adherence. This study makes three contributions. First, it identifies four archetypical inefficiency modes that drive schedule overruns and explains how these mechanisms unfold in their organizational context. Second, it quantifies the magnitude and variation of schedule slips. Third, it creates a hypothetical digitally transformed version of the current process, aligned with DoD DE policy, and compares it to the current state to estimate potential schedule gains. Our findings suggest that a DE transformation could reduce the median project duration by 50.1% and reduce the standard deviation by 41.5%, leading to faster and more predictable timelines. However, the observed gains are not uniform across task categories. Overall, this study provides initial quantitative evidence of DE's potential ROI and its value in improving the efficiency and predictability of complex system sustainment projects.

cs.CY

An Analysis of Early-Stage Functional Safety Analysis Methods and Their Integration into Model-Based Systems Engineering

As systems become increasingly complex, conducting effective safety analysis in the earlier phases of a system's lifecycle is essential to identify and mitigate risks before they escalate. To that end, this paper investigates the capabilities of key safety analysis techniques, namely: Failure Mode and Effects Analysis (FMEA), Functional Hazard Analysis (FHA), and Functional Failure Identification and Propagation (FFIP), along with the current state of the literature in terms of their integration into Model-Based Systems Engineering (MBSE). A two-phase approach is adopted. The first phase is focused on contrasting FMEA, FHA, and FFIP techniques, examining their procedures, along with a documentation of their relative strengths and limitations. Our analysis highlights FFIP's capability in identifying emergent system behaviors, second-order effects, and fault propagation; thus, suggesting it is better suited for the safety needs of modern interconnected systems. Second, we review the existing research on the efforts to integrate each of these methods into MBSE. We find that MBSE integration efforts primarily focus on FMEA, and integration of FHA and FFIP is nascent. Additionally, FMEA-MBSE integration efforts could be organized into four categories: model-to-model transformation, use of external customized algorithms, built-in MBSE packages, and manual use of standard MBSE diagrams. While our findings indicate a variety of MBSE integration approaches, there is no universally established framework or standard. This leaves room for an integration approach that could support the ongoing Digital Engineering transformation efforts by enabling a more synergistic lifecycle safety management methods and tools.

cs.SE