arXiv ScienceSearch

arXiv subjects

Ruize Ma

Publications and source records attributed to Ruize Ma.

7 recordsLinked to original sources

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.

cs.CL

OpenGame: Open Agentic Coding for Games

Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across many files. While Large Language Models (LLMs) and code agents now solve isolated programming tasks with ease, they consistently stumble when asked to produce a fully playable game from a high-level design, collapsing under cross-file inconsistencies, broken scene wiring, and logical incoherence. We bridge this gap with OpenGame, the first open-source agentic framework explicitly designed for end-to-end web game creation. At its core lies Game Skill, a reusable, evolving capability composed of a Template Skill that grows a library of project skeletons from experience and a Debug Skill that maintains a living protocol of verified fixes - together enabling the agent to scaffold stable architectures and systematically repair integration errors rather than patch isolated syntax bugs. Powering this framework is GameCoder-27B, a code LLM specialized for game engine mastery through a three-stage pipeline of continual pre-training, supervised fine-tuning, and execution-grounded reinforcement learning. Since verifying interactive playability is fundamentally harder than checking static code, we further introduce OpenGame-Bench, an evaluation pipeline that scores agentic game generation along Build Health, Visual Usability, and Intent Alignment via headless browser execution and VLM judging. Across 150 diverse game prompts, OpenGame establishes a new state-of-the-art. We hope OpenGame pushes code agents beyond discrete software engineering problems and toward building complex, interactive real-world applications. Our framework will be fully open-sourced.

cs.SE

FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair

Multimodal Automated Program Repair (MAPR) extends traditional program repair by requiring models to jointly reason over source code, textual issue descriptions, and visual artifacts such as GUI screenshots. While recent LLM-based repair systems have shown promising results, existing approaches face several limitations: rigid workflow pipelines restrict exploration during debugging, visual reasoning is often performed over full-page screenshots without localized grounding, and failed repair attempts are rarely transformed into reusable knowledge. To address these challenges, we propose FailureMem, a multimodal repair framework that integrates three key mechanisms: a hybrid workflow-agent architecture that balances structured localization with flexible reasoning, active perception tools that enable region-level visual grounding, and a Failure Memory Bank that converts past repair attempts into reusable guidance. Experiments on SWE-bench Multimodal demonstrate FailureMem improves the resolved rate over GUIRepair by 3.7%.

cs.SE

ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection

Recent progress in video generative models has enabled the creation of high-quality videos from multimodal prompts that combine text and images. While these systems offer enhanced controllability, they also introduce new safety risks, as harmful content can emerge from individual modalities or their interaction. Existing safety methods are often text-only, require prior knowledge of the risk category, or operate as post-generation auditors, struggling to proactively mitigate such compositional, multimodal risks. To address this challenge, we present ConceptGuard, a unified safeguard framework for proactively detecting and mitigating unsafe semantics in multimodal video generation. ConceptGuard operates in two stages: First, a contrastive detection module identifies latent safety risks by projecting fused image-text inputs into a structured concept space; Second, a semantic suppression mechanism steers the generative process away from unsafe concepts by intervening in the prompt's multimodal conditioning. To support the development and rigorous evaluation of this framework, we introduce two novel benchmarks: ConceptRisk, a large-scale dataset for training on multimodal risks, and T2VSafetyBench-TI2V, the first benchmark adapted from T2VSafetyBench for the Text-and-Image-to-Video (TI2V) safety setting. Comprehensive experiments on both benchmarks show that ConceptGuard consistently outperforms existing baselines, achieving state-of-the-art results in both risk detection and safe video generation. Our code is available at https://github.com/Ruize-Ma/ConceptGuard.

cs.CV

Graphene Nanoribbons as a Majorana Platform

Graphene nanoribbons support a range of electronic phases that can be controlled via external stimuli. Zigzag-edged graphene nanoribbons (ZGNRs), in particular, exhibit an antiferromagnetic insulating ground state that transitions to a half-metallic phase under a transverse electric field or when embedded inside hexagonal Boron Nitride. Here, we consider a simple model of a heterostructure of a ZGNR with an Ising superconductor and show that, the Ising superconductor with a parent s-wave spin-singlet pairing can induce spin-triplet odd-parity pairing in the half-metallic phase of the ZGNR. The resulting superconducting phase is topologically nontrivial, with gate-tunable transitions that enable the emergence of Majorana zero modes.

cond-mat.mes-hall

Electrically-tunable ultra-flat bands and $\pi$-electron magnetism in graphene nanoribbons

Atomically thin crystals hosting flat electronic bands have been recently identified as a rich playground for exploring and engineering strongly correlated phases. Yet, their variety remains limited, primarily to two-dimensional moir\'e superlattices. Here, we predict the formation of reversible, electrically-induced ultra-flat bands and $\pi$-electron magnetism in one-dimensional chevron graphene nanoribbons. Our $ab$ $initio$ calculations show that the application of a transverse electric field to these nanoribbons generates a pair of isolated, nearly perfectly flat bands with widths of approximately 1 meV around the Fermi level. Upon charge doping, these flat bands undergo a Stoner-like electronic instability, resulting in the spontaneous emergence of local magnetic moments at the edges of the otherwise non-magnetic nanoribbon, akin to a one-dimensional spin-$\frac{1}{2}$ chain. Our findings expand the class of carbon-based nanostructures exhibiting flat bands and establish a novel route for inducing correlated electronic phases in chevron graphene nanoribbons.

cond-mat.mes-hall

Dirac half-semimetallicity and antiferromagnetism in graphene nanoribbon/hexagonal boron nitride heterojunctions

Half-metals have been envisioned as active components in spintronic devices by virtue of their completely spin-polarized electrical currents. Actual materials hosting half-metallic phases, however, remain scarce. Here, we predict that recently fabricated heterojunctions of zigzag nanoribbons embedded in two-dimensional hexagonal boron nitride are half-semimetallic, featuring fully spin-polarized Dirac points at the Fermi level. The half-semimetallicity originates from the transfer of charges from hexagonal boron nitride to the embedded graphene nanoribbon. These charges give rise to opposite energy shifts of the states residing at the two edges while preserving their intrinsic antiferromagnetic exchange coupling. Upon doping, an antiferromagnetic-to-ferrimagnetic phase transition occurs in these heterojunctions, with the sign of the excess charge controlling the spatial localization of the net magnetic moments. Our findings demonstrate that such heterojunctions realize tunable one-dimensional conducting channels of spin-polarized Dirac fermions that are seamlessly integrated into a two-dimensional insulator, thus holding promise for the development of carbon-based spintronics.

cond-mat.mtrl-sci