arXiv ScienceSearch

arXiv · 2508.12498

Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality

Abstract

As augmented reality (AR) applications increasingly require 3D content, generative pipelines driven by natural input such as speech offer an alternative to manual asset creation. In this work, we design a modular, edge-assisted architecture that supports both direct text-to-3D and text-image-to-3D pathways, enabling interchangeable integration of state-of-the-art components and systematic comparison of their performance in AR settings. Using this architecture, we implement and evaluate four representative pipelines through an IRB-approved user study with 11 participants, assessing six perceptual and usability metrics across three object prompts. Overall, text-image-to-3D pipelines deliver higher generation quality: the best-performing pipeline, which used FLUX for image generation and Trellis for 3D generation, achieved an average satisfaction score of 4.55 out of 5 and an intent alignment score of 4.82 out of 5. In contrast, direct text-to-3D pipelines excel in speed, with the fastest, Shap-E, completing generation in about 20 seconds. Our results suggest that perceptual quality has a greater impact on user satisfaction than latency, with users tolerating longer generation times when output quality aligns with expectations. We complement subjective ratings with system-level metrics and visual analysis, providing practical insights into the trade-offs of current 3D generation methods for real-world AR deployment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yanming Xiu, Joshua Chilukuri, Shunav Sen, Maria Gorlatova. 2025-08-17. Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality. https://arxiv.org/abs/2508.12498

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Demonstrably Informed Consent in Privacy Policy Flows: Evidence from a Randomized Experiment

Privacy policies govern how personal data is collected, used, and shared. Yet, in most privacy-policy consent flows, agreement is operationalized as a single click at the end of a long, opaque policy document. Recent privacy-law scholarship has argued for a standard of demonstrably informed consent. That is, the party drafting and designing privacy-policy consent mechanisms must generate reliable evidence that a person demonstrates comprehension of the consequential terms to which they agree. To this end, we study pedagogical friction as a design framing: minimal interventions embedded within a privacy-policy consent flow that aim to support demonstrated comprehension while keeping burden on the user low. In a randomized experiment, we tested pedagogical friction for demonstrably informed consent in the context of a privacy policy for an edtech app for young children. We recruited 293 parents of kids ages 3-8 to review the app's privacy policy under one of six conditions that varied presentation format and pacing, then complete a six-question comprehension quiz. Three conditions offered a second policy review and quiz retake for participants who did not pass this quiz on their first attempt. We find that the slide-based condition (G3) achieved the highest first-attempt threshold attainment (>=80%) (41.7%), followed by the paced, sectioned condition (G4) (30.6%). In the retake conditions, 64.9% of participants who completed a second attempt improved their score. Notably, in conditions that did not gate consent on demonstrated comprehension, 97.3% of participants who scored below the threshold still chose to consent, suggesting that ungated consent flows can record agreement without demonstrated comprehension. Our results suggest that pedagogical friction can strengthen the evidentiary basis of consent and clarify what it costs in time and burden.

cs.HC

Cyber Exodus: Burnout Symptoms, Exit Intention, and Peer Response in Online Cybersecurity Communities

Security practitioners burn out at high rates, and the resulting attrition is itself a security problem. This workforce is hard to study: security operations centers are closed to outside researchers, studies that reach practitioners recruit through employers, and those who have disengaged most may have the least reason to answer an employer's survey. The same practitioners discuss their working conditions openly in online communities. We adapt the Burnout Assessment Tool, a validated clinical instrument, into a text annotation scheme and apply it to 354,861 posts and 296,442 replies from five online communities of cybersecurity practitioners. Checked against two trained coders on 100 posts, the annotation reaches a macro F1 of 0.75 across the four symptoms and 0.98 for detecting any burnout signal. We find that the four symptoms point to different problems at work, not to the same problem at different levels of severity. Exhaustion appears in almost any complaint about staffing or workload. Mental distance, a loss of belief that the work is worthwhile, is the only symptom unrelated to operational problems, and among posts with a single symptom it is accompanied by a stated intention to leave roughly twice as often as any other. Peer responses show the opposite pattern. When a poster says they are considering leaving, the mix of replies shifts toward career advice, but this shift is smallest for mental distance. The symptom most strongly associated with leaving is thus the one peers adjust to least, and a single burnout score obscures both patterns.

cs.HC

Value Faces: Surfacing How Self-Presentation Shifts Across Relationships

People present different aspects of themselves across relationships. Computational work has captured such variation in communication style. But this variation also extends to which principles people foreground or background in a particular relationship--i.e., in the values they express and how they balance them. We conceptualize these relationship-specific expressions of values as demonstrated values. To make demonstrated values visible, we introduce Value Faces, a system that analyzes a person's existing chat histories from their everyday messaging platforms using Schwartz's ten basic human values and produces separate value profiles for their different relationships. In a mixed-methods study(N=18), we find that the resulting value profiles distinguished participants' relational contexts with twice the odds of guessing, while system-inferred differences across relationships aligned with participants' perceptions of those differences. Participants used these profiles to articulate previously implicit differences in how they presented themselves, connect them to roles and changes over time, and reconsider their self-assessments.

cs.HC