arXiv ScienceSearch

arXiv subjects

Ying Long

Publications and source records attributed to Ying Long.

At least 19 recordsLinked to original sources

Culturally uneven urban perception in large language models

Large language models (LLMs) are increasingly used to describe and evaluate cities, yet the cultural structure of their urban judgments remains understudied. Here we introduce a measurement framework for testing whether LLM-based urban perception is culturally neutral, using a globally stratified street-view image dataset. Open-ended descriptions and structured scores generated by three frontier multimodal models all show that the neutral baseline lies closer to regional framings associated with Europe and North America than to other cultural framings. Comparisons between AI and human urban perception further show that prompting can move AI responses closer to specific regional human descriptions, but fails to recover the variety and diversity of human responses, flattening observed demographic patterns and introducing sentiment-based self-favouring bias. These results indicate a systematic risk in treating AI as a neutral tool for urban tasks, especially when model outputs are used to compare, evaluate or represent cities across cultural contexts.

cs.CL

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement from Clinical Decision-making Quality

Large language models (LLMs) are increasingly adopted in clinical decision support, yet aligning them with the multifaceted reasoning pathways of real-world medicine remains a major challenge. Using more than 8,000 infertility treatment records, we systematically evaluate four alignment strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and In-Context Learning (ICL) through a dual-layer framework combining automatic benchmarks with blinded doctor-in-the-loop assessments. GRPO achieves the highest algorithmic accuracy across multiple decision layers, confirming the value of reinforcement-based optimization for structured prediction tasks. However, clinicians consistently prefer the SFT model, citing clearer reasoning processes (p = 0.035) and higher therapeutic feasibility (p = 0.019). In blinded pairwise comparisons, SFT attains the highest winning rate (51.2%), outperforming both GRPO (26.2%) and even physicians' original decisions (22.7%). These results reveal an alignment paradox: algorithmic improvements do not necessarily translate into higher clinical trust, and may diverge from human-centered preferences. Our findings highlight the need for alignment strategies that prioritize clinically interpretable and practically feasible reasoning, rather than solely optimizing decision-level accuracy.

cs.LG

Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study

Creating high-quality clinical Chains-of-Thought (CoTs) is crucial for explainable medical Artificial Intelligence (AI) while constrained by data scarcity. Although Large Language Models (LLMs) can synthesize medical data, their clinical reliability remains unverified. This study evaluates the reliability of LLM-generated CoTs and investigates prompting strategies to enhance their quality. In a blinded comparative study, senior clinicians in Assisted Reproductive Technology (ART) evaluated CoTs generated via three distinct strategies: Zero-shot, Random Few-shot (using shallow examples), and Selective Few-shot (using diverse, high-quality examples). These expert ratings were compared against evaluations from a state-of-the-art AI model (GPT-4o). The Selective Few-shot strategy significantly outperformed other strategies across all human evaluation metrics (p < .001). Critically, the Random Few-shot strategy offered no significant improvement over the Zero-shot baseline, demonstrating that low-quality examples are as ineffective as no examples. The success of the Selective strategy is attributed to two principles: "Gold-Standard Depth" (reasoning quality) and "Representative Diversity" (generalization). Notably, the AI evaluator failed to discern these critical performance differences. The clinical reliability of synthetic CoTs is dictated by strategic prompt curation, not the mere presence of examples. We propose a "Dual Principles" framework as a foundational methodology to generate trustworthy data at scale. This work offers a validated solution to the data bottleneck and confirms the indispensable role of human expertise in evaluating high-stakes clinical AI.

cs.AI

Leveraging Multi-Source Textural UGC for Neighbourhood Housing Quality Assessment: A GPT-Enhanced Framework

This study leverages GPT-4o to assess neighbourhood housing quality using multi-source textural user-generated content (UGC) from Dianping, Weibo, and the Government Message Board. The analysis involves filtering relevant texts, extracting structured evaluation units, and conducting sentiment scoring. A refined housing quality assessment system with 46 indicators across 11 categories was developed, highlighting an objective-subjective method gap and platform-specific differences in focus. GPT-4o outperformed rule-based and BERT models, achieving 92.5% accuracy in fine-tuned settings. The findings underscore the value of integrating UGC and GPT-driven analysis for scalable, resident-centric urban assessments, offering practical insights for policymakers and urban planners.

cs.CY

GenAI Models Capture Urban Science but Oversimplify Complexity

Generative artificial intelligence (GenAI) models are increasingly used for scientific data generation, yet their alignment with empirical knowledge in urban science remains unclear. We therefore ask whether generated urban data reproduce empirical regularities and support repeatable experiments. We introduce AI4US, a framework for evaluating data synthesis and conditional intervention across text and image modalities. Four cases spanning urban systems, within-city structure, neighbourhood vitality and streetscape perception are evaluated against published parameters, observed urban data and human judgements. Generated outputs recovered recognizable scaling and distance-decay patterns and yielded measurable associations between neighbourhood morphology indicators and pedestrian activity, while model judgements showed positive agreement with sampled human choices. Controlled changes to urban conditions produced repeatable output responses across all four cases. However, generated data often compressed empirical numerical coverage, local variation and visual diversity. In an inspectable GenAI model, intermediate activations linearly distinguished some urban relationships, while activation edits changed only selected output scores. AI4US provides an empirically grounded approach for assessing GenAI as the virtual urban laboratory in urban data synthesis and controlled model experiments, while revealing its tendency to simplify empirical complexity.

physics.soc-ph

Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology

Effective physician-patient communications in pre-diagnostic environments, and most specifically in complex and sensitive medical areas such as infertility, are critical but consume a lot of time and, therefore, cause clinic workflows to become inefficient. Recent advancements in Large Language Models (LLMs) offer a potential solution for automating conversational medical history-taking and improving diagnostic accuracy. This study evaluates the feasibility and performance of LLMs in those tasks for infertility cases. An AI-driven conversational system was developed to simulate physician-patient interactions with ChatGPT-4o and ChatGPT-4o-mini. A total of 70 real-world infertility cases were processed, generating 420 diagnostic histories. Model performance was assessed using F1 score, Differential Diagnosis (DDs) Accuracy, and Accuracy of Infertility Type Judgment (ITJ). ChatGPT-4o-mini outperformed ChatGPT-4o in information extraction accuracy (F1 score: 0.9258 vs. 0.9029, p = 0.045, d = 0.244) and demonstrated higher completeness in medical history-taking (97.58% vs. 77.11%), suggesting that ChatGPT-4o-mini is more effective in extracting detailed patient information, which is critical for improving diagnostic accuracy. In contrast, ChatGPT-4o performed slightly better in differential diagnosis accuracy (2.0524 vs. 2.0048, p > 0.05). ITJ accuracy was higher in ChatGPT-4o-mini (0.6476 vs. 0.5905) but with lower consistency (Cronbach's $\alpha$ = 0.562), suggesting variability in classification reliability. Both models demonstrated strong feasibility in automating infertility history-taking, with ChatGPT-4o-mini excelling in completeness and extraction accuracy. In future studies, expert validation for accuracy and dependability in a clinical setting, AI model fine-tuning, and larger datasets with a mix of cases of infertility have to be prioritized.

cs.CL

Enhancing the sensing power of bike-sharing system for urban environment

The development of smart cities requires innovative sensing solutions for efficient and low-cost urban environment monitoring. Bike-sharing systems, with their wide coverage, flexible mobility, and dense urban distribution, present a promising platform for pervasive sensing. At a relative early stage, research on bike-based sensing focuses on the application of data collected via passive sensing, without consideration of the optimization of data collection through sensor deployment or vehicle scheduling. To address this gap, this study integrates a binomial probability model with a mixed-integer linear programming model to optimize sensor allocation across bike stands. Additionally, an active scheduling strategy guides user bike selection to enhance the efficacy of data collection. A case study in Manhattan validates the proposed strategy, showing that equipping sensors on just 1\% of the bikes covers approximately 70\% of road segments in a day, highlighting the significant potential of bike-sharing systems for urban sensing.

math.OC

GaRField++: Reinforced Gaussian Radiance Fields for Large-Scale 3D Scene Reconstruction

This paper proposes a novel framework for large-scale scene reconstruction based on 3D Gaussian splatting (3DGS) and aims to address the scalability and accuracy challenges faced by existing methods. For tackling the scalability issue, we split the large scene into multiple cells, and the candidate point-cloud and camera views of each cell are correlated through a visibility-based camera selection and a progressive point-cloud extension. To reinforce the rendering quality, three highlighted improvements are made in comparison with vanilla 3DGS, which are a strategy of the ray-Gaussian intersection and the novel Gaussians density control for learning efficiency, an appearance decoupling module based on ConvKAN network to solve uneven lighting conditions in large-scale scenes, and a refined final loss with the color loss, the depth distortion loss, and the normal consistency loss. Finally, the seamless stitching procedure is executed to merge the individual Gaussian radiance field for novel view synthesis across different cells. Evaluation of Mill19, Urban3D, and MatrixCity datasets shows that our method consistently generates more high-fidelity rendering results than state-of-the-art methods of large-scale scene reconstruction. We further validate the generalizability of the proposed approach by rendering on self-collected video clips recorded by a commercial drone.

cs.CV

CMAB: A First National-Scale Multi-Attribute Building Dataset in China Derived from Open Source Data and GeoAI

Rapidly acquiring three-dimensional (3D) building data, including geometric attributes like rooftop, height and orientations, as well as indicative attributes like function, quality, and age, is essential for accurate urban analysis, simulations, and policy updates. Current building datasets suffer from incomplete coverage of building multi-attributes. This paper introduces a geospatial artificial intelligence (GeoAI) framework for large-scale building modeling, presenting the first national-scale Multi-Attribute Building dataset (CMAB), covering 3,667 spatial cities, 29 million buildings, and 21.3 billion square meters of rooftops with an F1-Score of 89.93% in OCRNet-based extraction, totaling 337.7 billion cubic meters of building stock. We trained bootstrap aggregated XGBoost models with city administrative classifications, incorporating features such as morphology, location, and function. Using multi-source data, including billions of high-resolution Google Earth images and 60 million street view images (SVIs), we generated rooftop, height, function, age, and quality attributes for each building. Accuracy was validated through model benchmarks, existing similar products, and manual SVI validation, mostly above 80%. Our dataset and results are crucial for global SDGs and urban planning.

cs.CV

Performance Evaluation of Lightweight Open-source Large Language Models in Pediatric Consultations: A Comparative Analysis

Large language models (LLMs) have demonstrated potential applications in medicine, yet data privacy and computational burden limit their deployment in healthcare institutions. Open-source and lightweight versions of LLMs emerge as potential solutions, but their performance, particularly in pediatric settings remains underexplored. In this cross-sectional study, 250 patient consultation questions were randomly selected from a public online medical forum, with 10 questions from each of 25 pediatric departments, spanning from December 1, 2022, to October 30, 2023. Two lightweight open-source LLMs, ChatGLM3-6B and Vicuna-7B, along with a larger-scale model, Vicuna-13B, and the widely-used proprietary ChatGPT-3.5, independently answered these questions in Chinese between November 1, 2023, and November 7, 2023. To assess reproducibility, each inquiry was replicated once. We found that ChatGLM3-6B demonstrated higher accuracy and completeness than Vicuna-13B and Vicuna-7B (P < .001), but all were outperformed by ChatGPT-3.5. ChatGPT-3.5 received the highest ratings in accuracy (65.2%) compared to ChatGLM3-6B (41.2%), Vicuna-13B (11.2%), and Vicuna-7B (4.4%). Similarly, in completeness, ChatGPT-3.5 led (78.4%), followed by ChatGLM3-6B (76.0%), Vicuna-13B (34.8%), and Vicuna-7B (22.0%) in highest ratings. ChatGLM3-6B matched ChatGPT-3.5 in readability, both outperforming Vicuna models (P < .001). In terms of empathy, ChatGPT-3.5 outperformed the lightweight LLMs (P < .001). In safety, all models performed comparably well (P > .05), with over 98.4% of responses being rated as safe. Repetition of inquiries confirmed these findings. In conclusion, Lightweight LLMs demonstrate promising application in pediatric healthcare. However, the observed gap between lightweight and large-scale proprietary LLMs underscores the need for continued development efforts.

cs.LG

Perspectives on stability and mobility of transit passenger's travel behaviour through smart card data

Existing studies have extensively used spatiotemporal data to discover the mobility patterns of various types of travellers. Smart card data (SCD) collected by the automated fare collection systems can reflect a general view of the mobility pattern of public transit riders. Mobility patterns of transit riders are temporally and spatially dynamic, and therefore difficult to measure. However, few existing studies measure both the mobility and stability of transit riders' travel patterns over a long period of time. To analyse the long-term changes of transit riders' travel behaviour, the authors define a metric for measuring the similarity between SCD, in this study. Also an improved density-based clustering algorithm, simplified smoothed ordering points to identify the clustering structure (SS-OPTICS), to identify transit rider clusters is proposed. Compared to the original OPTICS, SS-OPTICS needs fewer parameters and has better generalisation ability. Further, the generated clusters are categorized according to their features of regularity and occasionality. Based on the generated clusters and categories, fine- and coarse-grained travel pattern transitions of transit riders over four years from 2010 to 2014 are measured. By combining socioeconomic data of Beijing in the year of 2010 and 2014, the interdependence between stability and mobility of transit riders' travel behaviour is also discussed.

cs.CY

Automated identification and characterization of parcels (AICP) with OpenStreetMap and Points of Interest

Against the paucity of urban parcels in China, this paper proposes a method to automatically identify and characterize parcels (AICP) with OpenStreetMap (OSM) and Points of Interest (POI) data. Parcels are the basic spatial units for fine-scale urban modeling, urban studies, as well as spatial planning. Conventional ways of identification and characterization of parcels rely on remote sensing and field surveys, which are labor intensive and resource-consuming. Poorly developed digital infrastructure, limited resources, and institutional barriers have all hampered the gathering and application of parcel data in developing countries. Against this backdrop, we employ OSM road networks to identify parcel geometries and POI data to infer parcel characteristics. A vector-based CA model is adopted to select urban parcels. The method is applied to the entire state of China and identifies 82,645 urban parcels in 297 cities. Notwithstanding all the caveats of open and/or crowd-sourced data, our approach could produce reasonably good approximation of parcels identified from conventional methods, thus having the potential to become a useful supplement.

cs.CY

Discovering functional zones using bus smart card data and points of interest in Beijing

Cities comprise various functional zones, including residential, educational, commercial zones, etc. It is important for urban planners to identify different functional zones and understand their spatial structure within the city in order to make better urban plans. In this research, we used 77976010 bus smart card records of Beijing City in one week in April 2008 and converted them into two-dimensional time series data of each bus platform, Then, through data mining in the big database system and previous studies on citizens' trip behavior, we established the DZoF (discovering zones of different functions) model based on SCD (smart card Data) and POIs (points of interest), and pooled the results at the TAZ (traffic analysis zone) level. The results suggested that DzoF model and cluster analysis based on dimensionality reduction and EM (expectation-maximization) algorithm can identify functional zones that well match the actual land uses in Beijing. The methodology in the present research can help urban planners and the public understand the complex urban spatial structure and contribute to the academia of urban geography and urban planning.

cs.CY

Early Birds, Night Owls,and Tireless/Recurring Itinerants: An Exploratory Analysis of Extreme Transit Behaviors in Beijing, China

This paper seeks to understand extreme public transit riders in Beijing using both traditional household survey and emerging new data sources such as Smart Card Data (SCD). We focus on four types of extreme transit behaviors: public transit riders who (1) travel significantly earlier than average riders (the 'early birds'); (2) ride in unusual late hours (the 'night owls'); and (3) commute in excessively long distance (the 'tireless itinerants'); (4) travel over frequently in a day (the 'recurring itinerants). SCD are used to identify the spatiotemporal patterns of these three extreme transit behaviors. In addition, household survey data are employed to supplement the socioeconomic background and provide a tentative profiling of extreme travelers. While the research findings are useful to guide urban governance and planning in Beijing, the methods developed in this paper can be applied to understand travel patterns elsewhere.

physics.soc-ph

Profiling underprivileged residents with mid-term public transit smartcard data of Beijing

Mobility of economically underprivileged residents in China has seldom been well profiled due to privacy issue and the characteristics of Chinese over poverty. In this paper, we identify and characterize underprivileged residents in Beijing using ubiquitous public transport smartcard transactions in 2008 and 2010, respectively. We regard these frequent bus/metro riders (FRs) in China, especially in Beijing, as economically underprivileged residents. Our argument is tested against (1) the household travel survey in 2010, (2) a small-scale survey in 2012, as well as (3) our interviews with local residents in Beijing. Cardholders' job and residence locations are identified using Smart Card Data (SCD) in 2008 and 2010. Our analysis is restricted to cardholders that use the same cards in both years. We then classify all identified FRs into 20 groups by residence changes (change, no change), workplace changes (change, no change, finding a job, losing a job, and all-time employed) during 2008-2010 and housing place in 2010 (within the fourth ring road or not). The underprivileged degree of each FR is then evaluated using the 2014 SCD. To the best of our knowledge, this is one of the first studies for understanding long- or mid-term urban dynamics using immediate "big data", and also for profiling underprivileged residents in Beijing in a fine-scale.

cs.OH

Population spatialization and synthesis with open data

Individuals together with their locations & attributes are essential to feed micro-level applied urban models (for example, spatial micro-simulation and agent-based modeling) for policy evaluation. Existed studies on population spatialization and population synthesis are generally separated. In developing countries like China, population distribution in a fine scale, as the input for population synthesis, is not universally available. With the open-government initiatives in China and the emerging Web 2.0 techniques, more and more open data are becoming achievable. In this paper, we propose an automatic process using open data for population spatialization and synthesis. Specifically, the road network in OpenStreetMap is used to identify and delineate parcel geometries, while crowd-sourced POIs are gathered to infer urban parcels with a vector cellular automata model. Housing-related online Check-in records are then applied to distinguish residential parcels from all of the identified urban parcels. Finally the published census data, in which the sub-district level of attributes distribution and relationships are available, is used for synthesizing population attributes with a previously developed tool Agenter (Long and Shen, 2013). The results are validated with ground truth manually-prepared dataset by planners from Beijing Institute of City Planning.

cs.OH

Big Models: From Beijing to the whole China

This paper propose the concept of big model as a novel research paradigm for regional and urban studies. Big models are fine-scale regional/urban simulation models for a large geographical area, and they overcome the trade-off between simulated scale and spatial unit by tackling both of them at the same time enabled by emerging big/open data, increasing computation power and matured regional/urban modeling methods. The concept, characteristics, and potential applications of big models have been elaborated. We addresse several case studies to illustrate the progress of research and utilization on big models, including mapping urban areas for all Chinese cities, performing parcel-level urban simulation, and several ongoing research projects. Most of these applications can be adopted across the country, and all of them are focusing on a fine-scale level, such as a parcel, a block, or a township (sub-district), which is not the same with the existing studies using conventional models that are only suitable for a certain single or two cities or regions, or for a larger area but have to significantly sacrifice the data resolution. It is expected that big models will mark a promising new era for the urban and regional study in the age of big data.

cs.OH

Mapping parcel-level urban areas for a large geographical area

As a vital indicator for measuring urban development, urban areas are expected to be identified explicitly and conveniently with widely available dataset thereby benefiting the planning decisions and relevant urban studies. Existing approaches to identify urban areas normally based on mid-resolution sensing dataset, socioeconomic information (e.g. population density) generally associate with low-resolution in space, e.g. cells with several square kilometers or even larger towns/wards. Yet, few of them pay attention to defining urban areas with micro data in a fine-scaled manner with large extend scale by incorporating the morphological and functional characteristics. This paper investigates an automated framework to delineate urban areas in the parcel level, using increasingly available ordnance surveys for generating all parcels (or geo-units) and ubiquitous points of interest (POIs) for inferring density of each parcel. A vector cellular automata model was adopted for identifying urban parcels from all generated parcels, taking into account density, neighborhood condition, and other spatial variables of each parcel. We applied this approach for mapping urban areas of all 654 Chinese cities and compared them with those interpreted from mid-resolution remote sensing images and inferred by population density and road intersections. Our proposed framework is proved to be more straight-forward, time-saving and fine-scaled, compared with other existing ones, and reclaim the need for consistency, efficiency and availability in defining urban areas with well-consideration of omnipresent spatial and functional factors across cities.

cs.OH