arXiv · 2509.19329
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
Abstract
We examined how model size, temperature, and prompt style affect Large Language Models' (LLMs) alignment within itself, between models, and with human in assessing clinical reasoning skills. Model size emerged as a key factor in LLM-human score alignment. Study highlights the importance of checking alignments across multiple levels.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Julie Jung, Max Lu, Sina Chole Benker, Dogus Darici. 2025-09-14. How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment. https://arxiv.org/abs/2509.19329
Cite the original work for its findings. Save a collection to share your selection of sources.