12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Investigating Novice Researchers' Perceptions of Research Privacy Within LLM-Assisted Workflows

Shuning Zhang, Changxi Wen, Eve He, Ying Ma, Robert Xiao, Xin Yi et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Semi-structured interviews with thematic analysis.

Sample

N = 44, 10 groups

Primary method

Thematic analysis. Four authors independently reviewed initial subset of four randomly selected transcripts to generate initial codes, met to compare codes and establish unified codebook. Four authors applied codebook to remaining transcripts iteratively, with each author coding 10 transcripts. Calculated code frequency post-coding. No inter-rater reliability calculated or reported (as suggested by prior guidelines for inductive/exploratory analysis). Original Chinese quotes manually translated by one primary author and checked by three other primary authors fluent in both English and Chinese.

Main result

For RQ1, we identified that researchers valued the privacy of unpublished ideas and data. Specifically, they noted that ideas are more vague and can leak indirectly. Moreover, they hold misunderstandings around research privacy risks. They worried about the theft of unpublished ideas by service providers and other peers, causing professional consequences that hindered progress. Simultaneously, they underestimated the risks of model memorization and training, generally assuming that their data would be diluted within LLM training sets, or assuming that model refusals to output their sensitive data would protect their privacy. For RQ2, we found that novice researchers, operating without institutional sandboxes, employed ad-hoc mitigation strategies. They fragmented proprietary ideas or frameworks into different sessions or models to prevent data leakage, and used adversarial probing techniques such as asking whether LLMs memorize their data to empirically test the boundaries of profiling. However, they perceived these practices as mostly ineffective, and had to sacrifice effective assistance for privacy.

Reports effect sizes.

Research paradigm

Qualitative empirical research (interpretive/constructivist)

Author conclusions

This paper examines novice researchers' privacy perceptions regarding LLM-assisted research workflows through semi-structured interviews (N=44). Findings reveal a prioritization of productivity over ideas' protection. Users rely on ad-hoc mitigation strategies, such as data fragmentation, but feel that they serve merely as psychological placebos against backend retention. A critical discrepancy exists between users' mental models, which falsely assume input dilution prevents memorization, and actual technical risks. We therefore recommend actions such as institutional automated screening and separated tools, transparent risk visualization, and interactive privacy training.

Risk of bias

Selection bias from convenience sampling via online platforms (LinkedIn, RedBook, WeChat); Self-reporting bias including recall bias; Social desirability bias in interview responses; Geographic bias toward Mainland China-based researchers (majority of sample); Disciplinary imbalance: 28/44 from social sciences/humanities vs. 11/44 from technical fields; Self-reporting bias including recall bias and social desirability bias; Potential lack of generalizability to global research population; Geographic concentration bias (primarily Mainland China participants: 32/44); Disciplinary representation bias (28/44 from social sciences/humanities); Participant recruitment predominantly from Chinese platforms may not represent global research populations; Lack of inter-rater reliability calculation, though authors cite inductive exploratory nature as justification

Limitations

  • "First, the reliance on convenience sampling via online platforms introduces selection bias
  • While we recruited a diverse pool spanning multiple countries and backgrounds, the findings, especially code frequencies, may not fully generalize to the global research population
  • Future research should use large-scale longitudinal studies to observe the evolution of user mental models
  • Second, as with most qualitative studies, our data is subject to self-reporting biases, such as recall bias and social desirability bias." Additionally, "our paper focuses on novice researchers, whose workflows and risk perceptions may differ from established scholars."

Open questions raised

  • Large-scale longitudinal studies needed to observe evolution of user mental models
  • Empirical technical audits to objectively measure discrepancy between user-perceived mitigations and actual usage
  • Comparative studies involving senior researchers to reveal how career stages and institutional responsibilities influence research risk assessments
  • Research-specific privacy contexts lacking in existing literature
  • Large-scale longitudinal studies to observe the evolution of user mental models
  • Limited research on research-specific privacy contexts (as opposed to general personal privacy or educational privacy)
Extracted from: pdfAgreement 67%

Explore related topics

Related papers