arXiv Science⌕ Search

arXiv subjects

Wooyoung Jung

Publications and source records attributed to Wooyoung Jung.

7 recordsLinked to original sources

BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research

Occupants are a primary source of uncertainty in building energy consumption and management, yet existing occupant behavior models cannot capture adaptive and reasoning responses considering the occupant's personal history, current context, and the type of energy signal being delivered. This study presents BuildOcc, an open-source Python platform that grounds large language model agents in the American Time Use Survey (ATUS), a nationally representative diary dataset covering 16,684 respondents. Through BuildOcc, each simulated occupant agent can be instantiated with a demographic persona drawn from ATUS population statistics, a memory stream that accumulates and reflects on timestep-level observations, and an activity scheduler that samples empirically from ATUS time-at-activity distributions. The platform exposes a three-layer interface - Python library, REST API, and Model Context Protocol server - so that any building energy tool (EnergyPlus, Home Assistant) can integrate behavioral intelligence without bespoke coupling code. A plugin registry lets the community add new occupant strata, custom schedulers, and alternative memory backends as separate installable packages. Two validation tiers show that ATUS-grounded sampling reproduces empirically calibrated activity distributions and that demographic priors propagate into persona-consistent agent reasoning across timesteps, establishing internal consistency across strata. BuildOcc provides the building energy community with a reusable, openly available implementation of the occupant behavioral layer. BuildOcc is openly released at https://doi.org/10.5281/zenodo.21192895 under the Apache License 2.0 and installable via pip install buildocc.

cs.HC↗

Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying

Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), which would let building operators query these graphs through language agents, but progress is limited by the scarcity of large natural-language/SPARQL benchmarks. This paper presents Build2SPARQL, a large-scale benchmark for building KGs generated by a KG-grounded pipeline: SPARQL queries are produced and validated entirely by graph-traversal code, while large language models generate only the natural-language questions, keeping query correctness independent of model behavior. The pipeline mines six query-pattern families -- linear chains, branching, UNION, aggregation, OPTIONAL, and attribute-filtered -- and phrases each query across five vocabulary registers. Applied to 201 building KGs (180 Brick, 21 ASHRAE 223P), it yields 6,136 executable SPARQL queries and 30,680 questions. A two-rater human validation of 300 questions found 98.8% semantic fidelity, 98.8% naturalness, and 84.0% operational plausibility. A retrieval-augmented evaluation across three open-weight language models raised exact-match accuracy from 0.2-20% (zero-shot) to 56-65% (three-shot retrieved).

cs.AI↗

A Comprehensive Evaluation Framework for Conversational Home Energy Management Systems

The growing complexity in home energy management (HEM) demands advanced systems that guide occupants toward informed energy decisions reflecting their background, preferences, and context. Large language model (LLM)-integrated HEM systems (HEMS) have demonstrated promise, but previous studies relied on single-turn or single-task evaluations with response accuracy as the primary metric. Whether such systems deliver effective interactions across the extended multi-turn dialogues typical of real-world use remains an open question. This study introduces a comprehensive evaluation framework of LLM-integrated HEMS derived from the Goal-Question-Metric methodology, organized across five categories: task performance, factual accuracy, interaction quality, control capability, and system efficiency. A total of 23 metrics across multi-turn conversations are proposed and an LLM-as-judge pipeline is employed to enable scalable automated scoring. Its reliability is validated against three trained human coders: after iterative rubric calibration, twelve of the fifteen LLM-scored metrics reached strong agreement (ICC >= 0.73), three of them perfect, while the remaining three exhibited near-zero variance in human scores and are instead reported via mean absolute error (0.04-0.28). To demonstrate the framework's effectiveness, 970 dialogues -- 16 scenarios and five personas -- were generated and evaluated across four conversational HEMS configurations spanning a sophistication gradient, from a vanilla LLM with raw energy data to a multi-agent HEMS. The framework distinguished the four configurations across multiple evaluation dimensions, revealing their respective strengths and weaknesses. This study contributes to conversational HEMS by providing a reproducible, multi-dimensional evaluation methodology that comprehensively assesses sustained, context-aware system performance.

cs.HC↗

Large Language Model-Driven Context-Aware Eco-Feedback Generation and Evaluation

The objective of this study is to demonstrate the potential of generating context-aware eco-feedback - eco-feedback that reflects a household's contextual characteristics alongside its energy use patterns - through a large language model-integrated framework. Previous studies have introduced personalized eco-feedback, mostly relying on household energy use patterns; however, they frequently did not reflect distinct household characteristics, including their persona or non-negotiable routines, leaving eco-feedback ineffective and sometimes superficial. To address these limitations, we introduce a contextual engineering framework that generates eco-feedback using a self-consistency with chain-of-thought prompt, leveraging household energy analysis data, utility rate structures, and household characteristic information. We conducted a rigorous empirical validation and a combinatorial evaluation analysis to assess this framework systematically. The former tested the framework's ability to generate accurate and contextually grounded eco-feedback for three households by comparing its output against reference interventions independently derived from the same household data. The latter examined the framework's adaptability across 400 scenarios spanning 50 households, two utility rate structures, and four behavioral personas. Our framework generated eco-feedback that aligned with reference interventions at a mean rate of 92.0% and grounded its recommendations in the provided household data with 95.7% citation accuracy. It also proved highly adaptive, shifting both the appliances targeted and the energy-saving strategies recommended in response to rate structure and household context. Ultimately, this study contributes to realizing the next level of context-aware interactions between occupants and buildings which paves the way for higher occupant living quality and sustainability.

cs.HC↗

BuildGraph: A Synthetic Multi-Archetype Building Knowledge Graph Dataset

Semantic querying of building knowledge graphs (KGs) underpins the integration of artificial intelligence into building operations, from natural-language access to cross-building analytics, but such KGs are rarely public owing to proprietary, security, and cost barriers. BuildGraph is a synthetic building KG dataset of 120 buildings in the Brick Schema ontology, grounded in U.S. Department of Energy prototype models and sensor-placement patterns from real buildings. It spans eight commercial building types across three ASHRAE energy-code vintages, with five realizations per archetype varying URI naming, sensor-attachment predicates, and topology. A 75-query SPARQL benchmark confirms structural completeness: BuildGraph reaches 91.5% Query Answerability Rate versus 35.8% for 59 real-world Brick files. Independently, it reproduces real buildings' sensor-type proportions on their shared vocabulary (cosine 0.86-0.94), evidence of realistic instrumentation where measurable. A downstream text-to-SPARQL experiment with Gemma 4 (26B) reaches 27.3% Row-Matching F1 (+17.5 pp over zero-shot) on 12 held-out buildings. BuildGraph gives facility managers and digital-twin developers a testbed for portable analytics and natural-language interfaces, and its dataset, generator, and benchmark are openly available at https://github.com/humanbuildingsynergy/BuildGraph.

cs.DB↗

Multi-Agent Home Energy Management Assistant

Existing home energy management systems conceptualize occupants as passive recipients of energy information and control, which limits their ability to effectively support informed decision-making and sustained engagement. This paper presents Home Energy Management Assistant (HEMA), the first open-source, multi-agent system enabling sustained human-AI collaboration - multi-turn conversational interactions with preserved context - across diverse home energy management (HEM) tasks - from energy analysis and educational support to smart device control. HEMA combines large language model (LLM) reasoning capabilities with 36 purpose-built domain-specific tools through a three-layer architecture: a web-based conversational interface, a backend API server, and a multi-agent system. The system features three specialized agents - Analysis (energy consumption patterns and cost optimization), Knowledge (educational queries and rebate information), and Control (smart device management and scheduling) - coordinated through a self-consistency classifier that routes user queries using chain-of-thought reasoning. This architecture enables various energy analyses, adaptive explanations, and streamlined device control. HEMA also includes a comprehensive evaluation framework using an LLM-as-simulated-user methodology with 23 objective metrics across task performance, factual accuracy, interaction quality, and system efficiency, allowing systematic testing across diverse scenarios and user personas without requiring extensive human subject testing. Through demonstrations using real-world household energy consumption data, we show how HEMA supports informed decision-making and active engagement in HEM, highlighting its potential as a user-friendly, adaptable tool for residential deployment and as a research platform for HEM innovation.

cs.HC↗

Human-AI Collaboration in Large Language Model-Integrated Building Energy Management Systems: The Role of User Domain Knowledge and AI Literacy

This study aimed to comprehend how user domain knowledge and artificial intelligence (AI) literacy impact the effective use of human-AI interactive building energy management system (BEMS). While prior studies have investigated the potential of integrating large language models (LLMs) into BEMS or building energy modeling, very few studies have examined how user interact with such systems. We conducted a systematic role-playing experiment, where 85 human subjects interacted with an advanced generative pre-trained transformer (OpenAI GPT-4o). Participants were tasked with identifying the top five behavioral changes that could reduce home energy use with the GPT model that functioned as an LLM-integrated BEMS. Then, the collected prompt-response data and participant conclusions were analyzed using an analytical framework that hierarchically assessed and scored human-AI interactions and their home energy analysis approaches. Also, participants were classified into four groups based on their self-evaluated domain knowledge of building energy use and AI literacy, and Kruskal-Wallis H tests with post-hoc pairwise comparisons were conducted across 20 quantifiable metrics. Key takeaways include: most participants employed concise prompts (median: 16.2 words) and relied heavily on GPT's analytical capabilities; and notably, only 1 of 20 metrics, appliance identification rate, showed statistically significant group differences (p=0.037), driven by AI literacy rather than domain knowledge, suggesting an equalizing effect of LLMs across expertise levels. This study provides foundational insights into human-AI collaboration dynamics and promising development directions in the context of LLM-integrated BEMS and contributes to realizing human-centric LLM-integrated energy systems.

cs.HC↗