arXiv ScienceSearch

arXiv subjects

Steve Han

Publications and source records attributed to Steve Han.

4 recordsLinked to original sources

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We assess how well 54 LLMs can replicate human judgment when scoring responses from RAG (Retrieval-Augmented Generation) or Agentic pipelines against ground truth answers. Our methodology progresses from traditional correlation analysis to comprehensive Cohen's Kappa analysis that measures actual agreement patterns. The two-step approach includes: (1) a correlation test that filters judges with strong alignment, followed by (2) a human-likeness test using z-scores to identify two distinct judgment patterns: human-like judgment (|z| < 1) that mimics natural human variation, and super-consistent judgment (z > 1) that exceeds typical human-to-human agreement levels. This methodology reveals that 27 out of 54 tested LLMs achieve Tier 1 performance: 23 models exhibit human-like patterns that preserve the nuances of human judgment, while 4 models demonstrate super-consistent behavior, a pattern that could indicate either enhanced reliability or oversimplification of complex judgments. Testing 43 open-source models (1B-405B parameters) and 11 closed models (GPT, Gemini, Claude variants), we demonstrate that judge excellence is not solely dependent on model size but on specific training strategies. Our key contributions include: (1) establishing that correlation alone is insufficient for judge evaluation, (2) introducing a "Turing Test for judges" based on agreement patterns, and (3) providing a standardized benchmark for classifying LLM judges into distinct performance tiers for different evaluation needs.

cs.CL

Cube: A Roblox View of 3D Intelligence

Foundation models trained on vast amounts of data have demonstrated remarkable reasoning and generation capabilities in the domains of text, images, audio and video. Our goal at Roblox is to build such a foundation model for 3D intelligence, a model that can support developers in producing all aspects of a Roblox experience, from generating 3D objects and scenes to rigging characters for animation to producing programmatic scripts describing object behaviors. We discuss three key design requirements for such a 3D foundation model and then present our first step towards building such a model. We expect that 3D geometric shapes will be a core data type and describe our solution for 3D shape tokenizer. We show how our tokenization scheme can be used in applications for text-to-shape generation, shape-to-text generation and text-to-scene generation. We demonstrate how these applications can collaborate with existing large language models (LLMs) to perform scene analysis and reasoning. We conclude with a discussion outlining our path to building a fully unified foundation model for 3D intelligence.

cs.CV

Deep Imitation Learning for Humanoid Loco-manipulation through Human Teleoperation

We tackle the problem of developing humanoid loco-manipulation skills with deep imitation learning. The difficulty of collecting task demonstrations and training policies for humanoids with a high degree of freedom presents substantial challenges. We introduce TRILL, a data-efficient framework for training humanoid loco-manipulation policies from human demonstrations. In this framework, we collect human demonstration data through an intuitive Virtual Reality (VR) interface. We employ the whole-body control formulation to transform task-space commands by human operators into the robot's joint-torque actuation while stabilizing its dynamics. By employing high-level action abstractions tailored for humanoid loco-manipulation, our method can efficiently learn complex sensorimotor skills. We demonstrate the effectiveness of TRILL in simulation and on a real-world robot for performing various loco-manipulation tasks. Videos and additional materials can be found on the project page: https://ut-austin-rpl.github.io/TRILL.

cs.RO

Topological entanglement entropy in Gutzwiller projected spin liquids

The topological entanglement entropy(TEE) of Gutzwiller projected RVB state is studied with Monte Carlo simulation. New tricks are proposed to improve the convergence of TEE, which enable us to show that the spin liquid state studied in Ref.\cite{Vishwanath} actually does not support $Z_{2}$ topological order, a conclusion that is consistent with the information drawn from the inspection of the topological degeneracy on the same state. We find both a long ranged RVB amplitude and an approximate Marshall sign structure are at the origin of the suppression of vison gap in this spin liquid state. On the other hand, robust signature of $Z_{2}$ topological order, i.e., a TEE of $\ln2$, is clearly demonstrated for a Gutzwiiler projected RVB state on triangular lattice which is evolved from the RVB state proposed originally by Anderson\cite{Sorella}. We also find that it is the sign, rather than the amplitude of the RVB wave function, that dominates the TEE and is responsible for a positive value of TEE, which implies that the nonlocal entanglement in the RVB state is mainly encoded in the sign of the RVB wave function. Our results indicate that some information that is important for the topological property of a RVB state is missed in the effective field theory description and that a $Z_{2}$ gauge structure in the saddle point action is not enough for the RVB state to exhibit topological order, even if its spin correlation is extremely short ranged.

cond-mat.str-el