arXiv ScienceSearch

arXiv subjects

Michael Martin

Publications and source records attributed to Michael Martin.

15 recordsLinked to original sources

Exact quantum circuits for lattice Boltzmann realization of the Dirac equation

The quantum lattice Boltzmann (QLB) scheme of Succi and Dellar advances a four-component Dirac spinor on a lattice by a fixed sequence of local, exactly norm-preserving operations: a basis rotation, a collision, a streaming shift, and the inverse rotation. This unitarity is a structural property of the scheme, not an approximation, which suggests that a QLB time step should map onto a sequence of quantum gates. Here we make that mapping explicit. We give a gate-level construction of every operation of the three-dimensional Dirac QLB scheme: the fixed rotation gates, the collision gate, the streaming shift as a controlled increment on a position register, the position-dependent potential as a phase oracle, and periodic and reflecting (bounce-back) boundary conditions as unitary circuits. We then compose them into single-axis, two- and three-dimensional time steps. On a state-vector emulator the resulting circuits reproduce the classical QLB solver to machine precision (maximum density deviation between $3.7\times10^{-12}$ and $1.0\times10^{-17}$ across the one-, two-, and three-dimensional tests), so the circuits are the scheme rather than an approximation of it. The scope is narrow: we establish that the Succi-Dellar theory can be implemented on a (gate-model) quantum computer, and report the associated gate counts. We make no claim of computational advantage; state preparation, measurement, and asymptotic cost are discussed as open questions. All operators, circuits, tests, and figures are reproducible from the open-source quantumKineticMethods library.

quant-ph

Quantum NLP models on Natural Language Inference

Quantum natural language processing (QNLP) offers a novel approach to semantic modeling by embedding compositional structure directly into quantum circuits. This paper investigates the application of QNLP models to the task of Natural Language Inference (NLI), comparing quantum, hybrid, and classical transformer-based models under a constrained few-shot setting. Using the lambeq library and the DisCoCat framework, we construct parameterized quantum circuits for sentence pairs and train them for both semantic relatedness and inference classification. To assess efficiency, we introduce a novel information-theoretic metric, Information Gain per Parameter (IGPP), which quantifies learning dynamics independent of model size. Our results demonstrate that quantum models achieve performance comparable to classical baselines while operating with dramatically fewer parameters. The Quantum-based models outperform randomly initialized transformers in inference and achieve lower test error on relatedness tasks. Moreover, quantum models exhibit significantly higher per-parameter learning efficiency (up to five orders of magnitude more than classical counterparts), highlighting the promise of QNLP in low-resource, structure-sensitive settings. To address circuit-level isolation and promote parameter sharing, we also propose a novel cluster-based architecture that improves generalization by tying gate parameters to learned word clusters rather than individual tokens.

cs.CL

Characterizing Knowledge Graph Tasks in LLM Benchmarks Using Cognitive Complexity Frameworks

Large Language Models (LLMs) are increasingly used for tasks involving Knowledge Graphs (KGs), whose evaluation typically focuses on accuracy and output correctness. We propose a complementary task characterization approach using three complexity frameworks from cognitive psychology. Applying this to the LLM-KG-Bench framework, we highlight value distributions, identify underrepresented demands and motivate richer interpretation and diversity for benchmark evaluation tasks.

cs.CL

LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs

Current Large Language Models (LLMs) can assist developing program code beside many other things, but can they support working with Knowledge Graphs (KGs) as well? Which LLM is offering the best capabilities in the field of Semantic Web and Knowledge Graph Engineering (KGE)? Is this possible to determine without checking many answers manually? The LLM-KG-Bench framework in Version 3.0 is designed to answer these questions. It consists of an extensible set of tasks for automated evaluation of LLM answers and covers different aspects of working with semantic technologies. In this paper the LLM-KG-Bench framework is presented in Version 3 along with a dataset of prompts, answers and evaluations generated with it and several state-of-the-art LLMs. Significant enhancements have been made to the framework since its initial release, including an updated task API that offers greater flexibility in handling evaluation tasks, revised tasks, and extended support for various open models through the vllm library, among other improvements. A comprehensive dataset has been generated using more than 30 contemporary open and proprietary LLMs, enabling the creation of exemplary model cards that demonstrate the models' capabilities in working with RDF and SPARQL, as well as comparing their performance on Turtle and JSON-LD RDF serialization tasks.

cs.AI

Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering

As the field of Large Language Models (LLMs) evolves at an accelerated pace, the critical need to assess and monitor their performance emerges. We introduce a benchmarking framework focused on knowledge graph engineering (KGE) accompanied by three challenges addressing syntax and error correction, facts extraction and dataset generation. We show that while being a useful tool, LLMs are yet unfit to assist in knowledge graph generation with zero-shot prompting. Consequently, our LLM-KG-Bench framework provides automatic evaluation and storage of LLM responses as well as statistical data and visualization tools to support tracking of prompt engineering and model performance.

cs.AI

VERITAS and HAWC observations of unidentified source LHAASO J2108+5157

Understanding the complete nature of Galactic sources that accelerate cosmic rays up to $10^{15}$ eV energy (Galactic PeVatrons) is still an unsolved problem in high-energy astrophysics. Although supernova remnants have long been considered as the best candidates for Galactic PeVatrons, a clear association of SNRs with PeVatrons needs further exploration. Recently, the LHAASO collaboration published its first catalog of 90 very high energy (VHE) gamma-ray sources, and a few of them have no obvious counterparts at other wavelengths. Here, we will present morphology and spectral analysis of one such unassociated source LHAASO J2108+5157 using VERITAS and HAWC data.

astro-ph.HE

LLM-assisted Knowledge Graph Engineering: Experiments with ChatGPT

Knowledge Graphs (KG) provide us with a structured, flexible, transparent, cross-system, and collaborative way of organizing our knowledge and data across various domains in society and industrial as well as scientific disciplines. KGs surpass any other form of representation in terms of effectiveness. However, Knowledge Graph Engineering (KGE) requires in-depth experiences of graph structures, web technologies, existing models and vocabularies, rule sets, logic, as well as best practices. It also demands a significant amount of work. Considering the advancements in large language models (LLMs) and their interfaces and applications in recent years, we have conducted comprehensive experiments with ChatGPT to explore its potential in supporting KGE. In this paper, we present a selection of these experiments and their results to demonstrate how ChatGPT can assist us in the development and management of KGs.

cs.AI

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination of the two. Many Christian organizations are dedicated to the task of translating the Holy Bible into languages that lack a modern translation. Bible translation (BT) work is currently underway for over 3000 extremely low resource languages. We introduce the eBible corpus: a dataset containing 1009 translations of portions of the Bible with data in 833 different languages across 75 language families. In addition to a BT benchmarking dataset, we introduce model performance benchmarks built on the No Language Left Behind (NLLB) neural machine translation (NMT) models. Finally, we describe several problems specific to the domain of BT and consider how the established data and model benchmarks might be used for future translation efforts. For a BT task trained with NLLB, Austronesian and Trans-New Guinea language families achieve 35.1 and 31.6 BLEU scores respectively, which spurs future innovations for NMT for low-resource languages in Papua New Guinea.

cs.CL

Retention Is All You Need

Skilled employees are the most important pillars of an organization. Despite this, most organizations face high attrition and turnover rates. While several machine learning models have been developed to analyze attrition and its causal factors, the interpretations of those models remain opaque. In this paper, we propose the HR-DSS approach, which stands for Human Resource (HR) Decision Support System, and uses explainable AI for employee attrition problems. The system is designed to assist HR departments in interpreting the predictions provided by machine learning models. In our experiments, we employ eight machine learning models to provide predictions. We further process the results achieved by the best-performing model by the SHAP explainability process and use the SHAP values to generate natural language explanations which can be valuable for HR. Furthermore, using "What-if-analysis", we aim to observe plausible causes for attrition of an individual employee. The results show that by adjusting the specific dominant features of each individual, employee attrition can turn into employee retention through informative business decisions.

cs.AI

Open Data and the Status Quo -- A Fine-Grained Evaluation Framework for Open Data Quality and an Analysis of Open Data portals in Germany

This paper presents a framework for assessing data and metadata quality within Open Data portals. Although a few benchmark frameworks already exist for this purpose, they are not yet detailed enough in both breadth and depth to make valid statements about the actual discoverability and accessibility of publicly available data collections. To address this research gap, we have designed a quality framework that is able to evaluate data quality in Open Data portals on dedicated and fine-grained dimensions, such as interoperability, findability, uniqueness or completeness. Additionally, we propose quality measures that allow for valid assessments regarding cross-portal findability and uniqueness of dataset descriptions. We have validated our novel quality framework for the German Open Data landscape and found out that metadata often still lacks meaningful descriptions and is not yet extensively connected to the Semantic Web.

cs.IR

Decentralized Evolution and Consolidation of RDF Graphs

The World Wide Web and the Semantic Web are designed as a network of distributed services and datasets. In this network and its genesis, collaboration played and still plays a crucial role. But currently we only have central collaboration solutions for RDF data, such as SPARQL endpoints and wiki systems, while decentralized solutions can enable applications for many more use-cases. Inspired by a successful distributed source code management methodology in software engineering a framework to support distributed evolution is proposed. The system is based on Git and provides distributed collaboration on RDF graphs. This paper covers the formal expression of the evolution and consolidation of distributed datasets, the synchronization, as well as other supporting operations.

cs.DB

Decentralized Collaborative Knowledge Management using Git

The World Wide Web and the Semantic Web are designed as a network of distributed services and datasets. The distributed character of the Web brings manifold collaborative possibilities to interchange data. The commonly adopted collaborative solutions for RDF data are centralized (e.g. SPARQL endpoints and wiki systems). But to support distributed collaboration, a system is needed, that supports divergence of datasets, brings the possibility to conflate diverged states, and allows distributed datasets to be synchronized. In this paper, we present Quit Store, it was inspired by and it builds upon the successful Git system. The approach is based on a formal expression of evolution and consolidation of distributed datasets. During the collaborative curation process, the system automatically versions the RDF dataset and tracks provenance information. It also provides support to branch, merge, and synchronize distributed RDF datasets. The merging process is guarded by specific merge strategies for RDF data. Finally, we use our reference implementation to show overall good performance and demonstrate the practical usability of the system.

cs.DB

Inducing and Mitigating a Self-Reinforcing Degradation in Decision-making Teams

The models in this paper demonstrate how self-reinforcing error due to positive feedback can lead to overload and saturation of decision-making elements, and ultimately the cascading collapse of an organization due to the propagation of overload and erroneous decisions throughout the organization. We begin the paper with an analysis of the stability of the decision-making aspects of command organizations from a system-theoretic perspective. A simple dynamic model shows how an organization can enter into a self-reinforcing cycle of increasing decision workload until the demand for decisions exceeds the decision-making capacity of the organization. We then extend the model to more complex networked organizations and show that they also experience a form of self-reinforcing degradation. In particular, we find that the degradation in decision quality has a tendency to propagate through the hierarchical structure, i.e. overload at one location affects other locations by overloading the higher-level components which then in turn overload their subordinates. Our computational experiments suggest several strategies for mitigating this type of malfunction: dumping excessive load, empowering lower echelons, minimizing the need for coordination, using command-by-negation, insulating weak performers, and applying on-line diagnostics. We describe a method to allocate decision responsibility and arrange information flow dynamically within a team of decision-makers for command and control.

eess.SY

Social Network Intelligence Analysis to Combat Street Gang Violence

In this paper we introduce the Organization, Relationship, and Contact Analyzer (ORCA) that is designed to aide intelligence analysis for law enforcement operations against violent street gangs. ORCA is designed to address several police analytical needs concerning street gangs using new techniques in social network analysis. Specifically, it can determine "degree of membership" for individuals who do not admit to membership in a street gang, quickly identify sets of influential individuals (under the tipping model), and identify criminal ecosystems by decomposing gangs into sub-groups. We describe this software and the design decisions considered in building an intelligence analysis tool created specifically for countering violent street gangs as well as provide results based on conducting analysis on real-world police data provided by a major American metropolitan police department who is partnering with us and currently deploying this system for real-world use.

cs.SI

Drude Conductivity of Dirac Fermions in Graphene

Electrons moving in graphene behave as massless Dirac fermions, and they exhibit fascinating low-frequency electrical transport phenomena. Their dynamic response, however, is little known at frequencies above one terahertz (THz). Such knowledge is important not only for a deeper understanding of the Dirac electron quantum transport, but also for graphene applications in ultrahigh speed THz electronics and IR optoelectronics. In this paper, we report the first measurement of high-frequency conductivity of graphene from THz to mid-IR at different carrier concentrations. The conductivity exhibits Drude-like frequency dependence and increases dramatically at THz frequencies, but its absolute strength is substantially lower than theoretical predictions. This anomalous reduction of free electron oscillator strength is corroborated by corresponding changes in graphene interband transitions, as required by the sum rule. Our surprising observation indicates that many-body effects and Dirac fermion-impurity interactions beyond current transport theories are important for Dirac fermion electrical response in graphene.

cond-mat.mes-hall