arXiv ScienceSearch

arXiv · 2609.05160

Students' Perception of Big Data Engineering in Higher Education Curricula: Expectations, Interest and Ethical Implications

Abstract

The study investigates students' interest and expectations in a Big Data Engineering course integrated with a Master curricula, as well as ethical implications of using Big Data. An anonymous online survey was conducted with 42 of the 67 students enrolled in the Big Data course offered to Computer Science and Bioinformatics Master's programs. The responses were analyzed and interpreted using thematic analysis, highlighting interesting aspects related to students' expectations, interest, and their perspective of the ethical implications of working with Big Data. The study concludes that, even though there is significant difference in students' background, the majority are interested in learning Big Data, for practical and personal reasons related to the potential for career growth and their passion for the field. The main expectation expressed is related to enhancing their knowledge related to Big Data via practical activities. All students demonstrate awareness of potential ethical threats related to security and privacy, while Computer Science students are aware of the possibility of introducing bias in data during acquisition and analysis and of potential abusive data usage.

Explore related subjects

Keep this discovery

BibTeXRIS

Ioana-Georgiana Ciuciu, Petrescu Manuela-Andreea. 2026-09-04. Students' Perception of Big Data Engineering in Higher Education Curricula: Expectations, Interest and Ethical Implications. https://doi.org/10.5220/0013475000003928

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

cs.CL

Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority

Artificial intelligence is being rapidly incorporated into traditional search systems, yet scant work audits the information disparities across platforms, geography, and languages. We address this gap by comparing Google and Baidu's AI Overview systems for health queries, and measure informational anchors that emerge. Auditing 1,920 health queries across 12 countries and 4 languages, we find that Google and Baidu exhibit vertical integration, routing users toward their own company platforms in AI Overviews rather than a diverse set of primary sources. Smaller, lower-localization countries receive fewer domestically sourced references for health queries. Issuing the same query in a country's official language rather than English raises the share of locally sourced citations approximately 3.5- to 13.5-fold. Comparing queries across health topics of varying severity and controversy, including Traditional Chinese Medicine as an example, we also show that health disclaimers are multidimensional and vary across language and culture. We discuss how generative search influences access to health information, and the urgent need for culturally-aware oversight of these systems that influence critical health decisions.

cs.IR

Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos

Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a semester-long deployment of VideoPoints platform with a retrieval-augmented chatbot that answers from course lecture materials and returns timestamped citations. The chatbot retrieves only from the active course, uses chapter summaries to guide transcript ranking, and returns clickable timestamped citations. Students used it for quick lookups and exam review. Across 833 messages, 70.5% included citations, none crossed a course boundary, and when no lecture evidence matched, the chatbot usually declined rather than answering. Among the users, citations were the most consistently useful feature, while practice-question generation was the strongest unmet request. We also evaluated the design on the real-world test split of EduVidQA, a public multimodal benchmark for lecture-video question answering. Our design improved correct-lecture retrieval by 6.3 percentage points over dense-only retrieval. Together, the results show that effective deployment depends on course isolation, supported citations, and alignment with students' study practices.

cs.CL