THURSDAY, SEPTEMBER 3, 2026|No. 13662
Technology · AI Impact

LLMs Linked to Reduced Linguistic Diversity, Study Finds

A new study suggests that the widespread use of large language models as writing assistants is contributing to a homogenization of writing styles, potentially impacting societal and psychological insights derived from language.

1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

Abstract

Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21–50% across datasets and models ( P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.

Data availability

Of the seven datasets analysed in this Article, six are publicly available and one is access restricted. Publicly available datasets: Reddit r/WritingPrompts posts were obtained from Project Arctic Shift ( https://github.com/ArthurHeitmann/arctic_shift), which mirrors publicly accessible Reddit content (excluding private subreddits and content for which the original poster has filed a removal request) and is itself a successor archive to the Pushshift Reddit dataset144. arXiv abstracts were obtained from the publicly released metadata collection92, distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Patch News articles were accessed via the public Patch.com archive. United States Congressional Records floor speeches were obtained from the publicly released parsed corpus119, with the underlying speeches in the public domain as works of the US federal government. YourMorals Facebook posts and Moral Foundations Questionnaire responses are publicly released by the original investigators20; the data were collected from participants who voluntarily completed self-report measures on yourmorals.org and consented at registration to having their Facebook posts accessed for research purposes, under University of Southern California Institutional Review Board protocol UP-07-00393-AM019 (M. Dehghani, PI). Empathic Conversations are publicly released by the original investigators128; the dataset was collected under University of Pennsylvania IRB protocol #826448 from Amazon Mechanical Turk workers who were informed that their conversations, demographic information, personality surveys and essays would be used for academic research and distributed externally for research purposes, and were compensated per HIT. Access-restricted dataset: The Essays corpus124 is not publicly redistributable and is governed by access conditions imposed by the dataset owner. The corpus can be requested directly from Prof. James W. Pennebaker (University of Texas at Austin; pennebaker@utexas.edu or jwpennebaker@gmail.com); the original data were collected at UT-Austin from undergraduate psychology students who participated for course credit, with informed consent and institutional research approval reported in the original paper. Access is granted at the discretion of the dataset owner and does not require additional ethics approval from the requester beyond standard institutional review. We obtained the corpus via direct request and were not permitted to redistribute it. Use of Reddit and Patch News content: In line with the secondary-analysis posture taken by recent Nature Human Behaviour work on public text corpora56, all use of Reddit and Patch.com content in this Article is restricted to aggregate quantitative analysis of linguistic patterns over time. No individual Reddit post or Patch.com article is reproduced, republished or redistributed in this paper or its accompanying data release. Usernames, post identifiers and article identifiers are not retained in the analysis pipeline. Use is academic and non-commercial. Future deletion requests submitted through Arctic Shift’s public removal mechanism will be honoured in any continuing analyses. Source data are provided with this paper.

Code availability

All custom code used in this Article (implemented in Python 3.11.7 and R 4.4.2) is publicly available under the MIT license at https://doi.org/10.5281/zenodo.21633457 (ref. 145). The repository includes scripts for data preprocessing, LLM rewriting, classifier training and evaluation, statistical analyses, and figure generation. No proprietary software is required to reproduce the analyses.

References

  1. Orwell, G. Nineteen Eighty-Four 69–70 (Penguin Books, 1949).
  2. Park, G. et al. Automatic personality assessment through social media language. J. Pers. Soc. Psychol. 108, 934–952 (2015).

Article PubMed Google Scholar 003. Oberlander, J. & Gill, A. J. Language with character: a stratified corpus comparison of individual differences in e-mail communication. Discourse Process. 42, 239–270 (2006).

Article Google Scholar 004. Moreno, J. D., Martinez-Huertas, J. A., Olmos, R., Jorge-Botana, G. & Botella, J. Can personality traits be measured analyzing written language? A meta-analytic study on computational methods. Pers. Individ. Dif. 177, 110818 (2021).

Article Google Scholar 005. Mairesse, F., Walker, M. A., Mehl, M. R. & Moore, R. K. Using linguistic cues for the automatic recognition of personality in conversation and text. J. Artif. Intell. Res. 30, 457–500 (2007).

Article Google Scholar 006. Schwartz, H. A. et al. Personality, gender, and age in the language of social media: the open-vocabulary approach. PLoS ONE 8, 73791 (2013).

Article Google Scholar 007. Kramsch, C. Language and culture. AILA Rev. 27, 30–55 (2014).

Article Google Scholar 008. Gumperz, J. The speech community. Int. Encycl. Soc. Sci. 9, 381–386 (1968).

Google Scholar 009. Nguyen, D. & Rosé, C. P. Language use as a reflection of socialization in online communities. In Proc. Workshop on Language in Social Media (eds Nagarajan, M. & Gamon, M.) 76–85 (Association for Computational Linguistics, 2011). 010. Bamman, D., Eisenstein, J. & Schnoebelen, T. Gender identity and lexical variation in social media. J. Socioling. 18, 135–160 (2014).

Article Google Scholar 011. Pennebaker, J. W. The secret life of pronouns. New Sci. 211, 42–45 (2011).

Article Google Scholar 012. Robinson, M. D., Boyd, R. L., Fetterman, A. K. & Persich, M. R. The mind versus the body in political (and nonpolitical) discourse: linguistic evidence for an ideological signature in US politics. J. Lang. Soc. Psychol. 36, 438–461 (2017).

Article Google Scholar 013. Huang, Y., Guo, D., Kasakoff, A. & Grieve, J. Understanding US regional linguistic variation with Twitter data analysis. Comput. Environ. Urban Syst. 59, 244–255 (2016).

Article Google Scholar 014. Eisenstein, J., O’Connor, B., Smith, N. A. & Xing, E. A latent variable model for geographic lexical variation. In Proc. 2010 Conference on Empirical Methods in Natural Language Processing (eds Li, H. & Màrquez, L.) 1277–1287 (Association for Computational Linguistics, 2010). 015. Peterson, K., Hohensee, M. & Xia, F. Email formality in the workplace: a case study on the enron corpus. In Proc. Workshop on Language in Social Media (LSM 2011) (eds Nagarajan, M. & Gamon, M.) 86–95 (Association for Computational Linguistics, 2011). 016. Stamatatos, E. A survey of modern authorship attribution methods. J. Assoc. Inf. Sci. Technol. 60, 538–556 (2009).

Article Google Scholar 017. Grieve, J. Quantitative authorship attribution: an evaluation of techniques. Lit. Ling. Comput. 22, 251–270 (2007).

Article Google Scholar 018. Cassell, J. & Tversky, D. The language of online intercultural community formation. J. Comput. Mediat. Commun. 10, 1027 (2005).

Google Scholar 019. Danet, B. & Herring, S. C. Introduction: the multilingual internet. J. Comput. Mediat. Commun. 9, 9110 (2003).

Google Scholar 020. Kennedy, B. et al. Moral concerns are differentially observable in language. Cognition 212, 104696 (2021).

Article PubMed Google Scholar 021. Jackson, J. C., Gelfand, M., De, S. & Fox, A. The loosening of American culture over 200 years is associated with a creativity–order trade-off. Nat. Hum. Behav. 3, 244–250 (2019).

Article PubMed Google Scholar 022. Hofmann, V., Kalluri, P. R., Jurafsky, D. & King, S. AI generates covertly racist decisions about people based on their dialect. Nature 633, 147–154 (2024).

Article CAS PubMed PubMed Central Google Scholar 023. Richard, A. B., Lelandais, M., Reilly, K. T. & Jacquin-Courtois, S. Linguistic markers of subtle cognitive impairment in connected speech: a systematic review. J. Speech Lang. Hear. Res. 67, 4714–4733 (2024).

Article PubMed Google Scholar 024. Eyigoz, E., Mathur, S., Santamaria, M., Cecchi, G. & Naylor, M. Linguistic markers predict onset of Alzheimer’s disease. EClinicalMedicine 28, 100583 (2020). 025. Roark, B., Mitchell, M., Hollingshead, K. & Kaye, J. Spoken language derived measures for detecting mild cognitive impairment. IEEE Trans. Audio Speech Lang. Process. 19, 2081–2090 (2011).

Article PubMed PubMed Central Google Scholar 026. Trifu, R. N. et al. Linguistic markers for major depressive disorder: a cross-sectional study using an automated procedure. Front. Psychol. 15, 1355734 (2024).

Article PubMed PubMed Central Google Scholar 027. Weerasinghe, J., Morales, K. & Greenstadt, R. "Because… i was told… so much”: linguistic indicators of mental health status on Twitter. Proc. Priv. Enhanc. Technol. 4, 152–171 (2019). 028. Whorf, B. L. Language, Thought, and Reality: Selected Writings of Benjamin Lee Whorf (MIT Press, 2012). 029. Eckert, P. Three waves of variation study: the emergence of meaning in the study of sociolinguistic variation. Annu. Rev. Anthropol. 41, 87–100 (2012).

Article Google Scholar 030. Hofstede, G. Culture’s Consequences: Comparing Values, Behaviors, Institutions and Organizations Across Nations 2nd edn (Sage, 2001). 031. OpenAI. Introducing ChatGPT https://openai.com/blog/chatgpt (2022). 032. Gemini Team et al. Gemini: a family of highly capable multimodal models. Preprint at https://doi.org/10.48550/arXiv.2312.11805 (2023). 033. Nearly 1 in 3 College Students Have Used ChatGPT on Written Assignments https://www.intelligent.com/nearly-1-in-3-college-students-have-used-chatgpt-on-written-assignments/ (Intelligent, 2024). 034. McClain, C. Americans’ Use of ChatGPT is Ticking Up, But Few Trust Its Election Information https://www.pewresearch.org/short-reads/2024/03/26/americans-use-of-chatgpt-is-ticking-up-but-few-trust-its-election-information/ (Pew Research Center, 2024). 035. Handa, K. et al. Which economic tasks are performed with AI? Evidence from millions of Claude conversations. Preprint at https://doi.org/10.48550/arXiv.2503.04761 (2025). 036. Mizrahi, M. et al. State of what art? A call for multi-prompt LLM evaluation. Trans. Assoc. Comput. Linguist. 12, 933–949 (2024).

Article Google Scholar 037. Serapio-García, G. et al. A psychometric framework for evaluating and shaping personality traits in large language models. Nat. Mach. Intell. 7, 1954–1968 (2025).

Article PubMed PubMed Central Google Scholar 038. Ghosh, S. et al. A closer look at the limitations of instruction tuning. In Proc. 41st International Conference on Machine Learning 624 (JMLR, 2024). 039. Santurkar, S. et al. Whose opinions do language models reflect? In International Conference on Machine Learning 29971–30004 (JMLR, 2023). 040. Ireland, M.E. & Mehl, M. R. in The Oxford Handbook of Language and Social Psychology (ed. Holtgraves, T. M.) 201–218 https://doi.org/10.1093/oxfordhb/9780199838639.013.034 (Oxford Univ. Press, 2014). 041. Corona Hernández, H. et al. Natural language processing markers for psychosis and other psychiatric disorders: emerging themes and research agenda from a cross-linguistic workshop. Schizophr. Bull. 49, 86–92 (2023).

Article Google Scholar 042. Rude, S., Gortner, E.-M. & Pennebaker, J. Language use of depressed and depression-vulnerable college students. Cogn. Emot. 18, 1121–1133 (2004).

Article Google Scholar 043. Coppersmith, G., Leary, R., Crutchley, P. & Fine, A. Natural language processing of social media as screening for suicide risk. Biomed. Inform. Insights 10, 1178222618792860 (2018).

Article PubMed PubMed Central Google Scholar 044. Matz, S. C. & Netzer, O. Using big data as a window into consumers’ psychology. Curr. Opin. Behav. Sci. 18, 7–12 (2017).

Article Google Scholar 045. Winter, S., Maslowska, E. & Vos, A. L. The effects of trait-based personalization in social media advertising. Comput. Hum. Behav. 114, 106525 (2021).

Article Google Scholar 046. Sundar, S. S. & Marathe, S. S. Personalization versus customization: the importance of agency, privacy, and power usage. Hum. Commun. Res. 36, 298–322 (2010).

Article Google Scholar 047. Ryan, M. J., Held, W. & Yang, D. Unintended impacts of LLM alignment on global representation. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics 1 , 16121–16140 (2024). 048. Navigli, R., Conia, S. & Ross, B. Biases in large language models: origins, inventory, and discussion. ACMJ. Data Inf. Qual. 15, 1–21 (2023).

Article Google Scholar 049. Atari, M., Xue, M. J., Park, P. S., Blasi, D. & Henrich, J. Which humans? Preprint at PsyArXiv https://doi.org/10.31234/osf.io/5b26t (2023). 050. Wang, A., Morgenstern, J. & Dickerson, J. P. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nat. Mach. Intell. 7, 400–411 (2025).

Article Google Scholar 051. Rozado, D. The political preferences of LLMs. PLoS ONE 19, 0306621 (2024).

Article Google Scholar 052. Pan, K. & Zeng, Y. Do LLMs possess a personality? Making the MBTI test an amazing evaluation for large language models. Preprint at https://doi.org/10.48550/arXiv.2307.16180 (2023). 053. Abdurahman, S. et al. Perils and opportunities in using large language models in psychological research. PNAS Nexus 3, 245 (2024).

Article Google Scholar 054. Kobak, D., González-Márquez, R., Horvát, E. -Á & Lause, J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Sci. Adv. 11, 3813 (2025).

Article Google Scholar 055. Liang, W. et al. Quantifying large language model usage in scientific papers. Nat. Hum. Behav. 9, 2599–2609 (2025).

Article PubMed Google Scholar 056. Liang, W. et al. The widespread adoption of large language model-assisted writing across society. Patterns 6, 101366 (2025). 057. Bao, T., Zhao, Y., Mao, J. & Zhang, C. Examining linguistic shifts in academic writing before and after the launch of chatGPT: a study on preprint papers. Scientometrics 130, 3597–3627 (2025).

Article Google Scholar 058. Guo, Y., Shang, G., Vazirgiannis, M. & Clavel, C. The curious decline of linguistic diversity: training language models on synthetic text. In Findings of the Association for Computational Linguistics: NAACL 2024 3589–3604 (ACL, 2024). 059. Muñoz-Ortiz, A., Gómez-Rodríguez, C. & Vilares, D. Contrasting linguistic patterns in human and LLM-generated news text. Artif. Intell. Rev. 57, 265 (2024).

Article PubMed PubMed Central Google Scholar 060. Xu, W., Jojic, N., Rao, S., Brockett, C. & Dolan, B. Echoes in AI: quantifying lack of plot diversity in LLM outputs. Proc. Natl Acad. Sci. USA 122, 2504966122 (2025).

Article Google Scholar 061. Padmakumar, V. & He, H. Does writing with language models reduce content diversity? In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=Feiz5HtCD0 062. Doshi, A. R. & Hauser, O. P. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 10, 5290 (2024).

Article Google Scholar 063. Anderson, B. R., Shah, J. H. & Kreminski, M. Homogenization effects of large language models on human creative ideation. In Proc. 16th Conference on Creativity & Cognition 413–425 (Association for Computing Machinery, 2024). 064. Agarwal, D., Naaman, M. & Vashistha, A. AI suggestions homogenize writing toward western styles and diminish cultural nuances. In Proc. 2025 CHI Conference on Human Factors in Computing Systems 1–21 (Association for Computing Machinery, 2025). 065. Moon, K., Green, A. E. & Kushlev, K. Homogenizing effect of large language models (LLMs) on creative diversity: an empirical comparison of human and chatGPT writing. Comput. Hum. Behav. Artif. Hum. 6, 100207 (2025).

Article Google Scholar 066. Alvero, A. et al. Large language models, social demography, and hegemony: comparing authorship in human and synthetic text. J. Big Data 11, 138 (2024).

Article Google Scholar 067. Zhang, S., Xu, J. & Alvero, A. Generative AI meets open-ended survey responses: research participant use of AI and homogenization. Sociol. Methods Res. https://doi.org/10.1177/00491241251327130 (2025). 068. Lee, J., Alvero, A., Joachims, T. & Kizilcec, R. Poor alignment and steerability of large language models: evidence from college admission essays. In Proc. Workshop on Social Simulation with LLMs at COLM (2025). 069. Hans, A. et al. Spotting LLMs with binoculars: zero-shot detection of machine-generated text. In Proc. 41st International Conference on Machine Learning 698 (JMLR, 2024). 070. Tufts, B., Zhao, X. & Li, L. A practical examination of AI-generated text detectors for large language models. In Findings of the Association for Computational Linguistics: NAACL 2025 (eds Chiruzzo, L. et al.) 4839–4856 https://doi.org/10.18653/v1/2025.findings-naacl.271 (Association for Computational Linguistics, 2025). 071. Dugan, L. et al. Raid: a shared benchmark for robust evaluation of machine-generated text detectors. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics 1, 12463–12492 (ACL, 2024). 072. Buchert, J.-M. The 6 Best AI Detectors Based on Objective Studies and Usage https://intellectualead.com/best-ai-detectors-guide/ (Intellectual Lead, 2025). 073. Singer, J. D., Willett, J. B. Applied Longitudinal Data Analysis: Modeling Change and Event Occurrence (Oxford Univ, 2003). 074. Linden, A. Conducting interrupted time-series analysis for single- and multiple-group comparisons. Stata J. 15, 480–500 (2015).

Article Google Scholar 075. Bell, D., Kay, J. & Malley, J. A non-parametric approach to non-linear causality testing. Econ. Lett. 51, 7–18 (1996).

Article Google Scholar 076. Dubey, A. et al. The llama 3 herd of models. Preprint at https://doi.org/10.48550/arXiv.2407.21783 (2024). 077. Kosinski, M., Stillwell, D. & Graepel, T. Private traits and attributes are predictable from digital records of human behavior. Proc. Natl Acad. Sci. USA 110, 5802–5805 (2013).

Article CAS PubMed PubMed Central Google Scholar 078. Tufekci, Z. Engineering the public: big data, surveillance and computational politics. First Monday https://doi.org/10.5210/fm.v19i7.4901 (2014). 079. Garg, N., Schiebinger, L., Jurafsky, D. & Zou, J. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proc. Natl Acad. Sci. USA 115, 3635–3644 (2018).

Article Google Scholar 080. Kennedy, B., Ashokkumar, A., Boyd, R. L. & Dehghani, M. in Handbook of Language Analysis in Psychology (eds Dehghani, M. & Boyd, R. L.) 3–62 (Guilford Publications, 2022). 081. Hirsh, J. B. & Peterson, J. B. Personality and language use in self-narratives. J. Res. Pers. 43, 524–527 (2009).

Article Google Scholar 082. Mehl, M. R., Gosling, S. D. & Pennebaker, J. W. Personality in its natural habitat: manifestations and implicit folk theories of personality in daily life. J. Pers. Soc. Psychol. 90, 862–877 (2006).

Article PubMed Google Scholar 083. Boyd, R. L., Ashokkumar, A., Seraj, S. & Pennebaker, J. W. The Development and Psychometric Properties of LIWC-22 Technical report https://www.liwc.app (Univ. Texas at Austin, 2022). 084. Frimer, J. A., Boghrati, R., Haidt, J., Graham, J. & Dehgani, M. Moral Foundations Dictionary for Linguistic Analyses 2.0 (Provalis Research, 2019). 085. Ishikawa, Y. Gender differences in vocabulary use in essay writing by university students. ProcediaSoc. Behav. Sci. 192, 593–600 (2015).

Article Google Scholar 086. Chen, J., Qiu, L. & Ho, M.-H. R. A meta-analysis of linguistic markers of extraversion: positive emotion and social process words. J. Res. Pers. 89, 104035 (2020).

Article Google Scholar 087. Li, L. & Tomasello, M. On the moral functions of language. Soc. Cogn. https://doi.org/10.1521/soco.2021.39.1.99 (2021). 088. Nguyen, D., Doğruöz, A. S., Rosé, C. P. & De Jong, F. Computational sociolinguistics: a survey. Comput. Linguist. 42, 537–593 (2016).

Article Google Scholar 089. Tausczik, Y. R. & Pennebaker, J. W. The psychological meaning of words: LIWC and computerized text analysis methods. J. Lang. Soc. Psychol. 29, 24–54 (2010).

Article Google Scholar 090. Evans, N. & Levinson, S. C. The myth of language universals: language diversity and its importance for cognitive science. Behav. Brain Sci. 32, 429–448 (2009).

Article PubMed Google Scholar 091. arXiv.org submitters. arXiv dataset. kaggle https://doi.org/10.34740/KAGGLE/DSV/7548853 (2024). 092. Verma, V., Fleisig, E., Tomlin, N. & Klein, D. Ghostbuster: detecting text ghostwritten by large language models. In Proc. 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (eds Duh, K. et al.) 1702–1717 https://doi.org/10.18653/v1/2024.naacl-long.95 (Association for Computational Linguistics, 2024). 093. Adam, G. A. et al. GPTZero: robust detection of LLM-generated texts. Preprint at https://doi.org/10.48550/arXiv.2602.13042 (2026). 094. Zhan, H., He, X., Xu, Q., Wu, Y. & Stenetorp, P. G3detector: general GPT-generated text detector. Preprint at https://doi.org/10.48550/arXiv.2305.12680 (2023). 095. Zellers, R. et al. Defending against neural fake news. Adv. Neural Inf. Process. Syst. 32, 9051–9062 (2019).

Google Scholar 096. SIMPSON, E. H. Measurement of diversity. Nature 163, 688 (1949).

Article Google Scholar 097. Shannon, C. E. A mathematical theory of communication. Bell Syst. Tech. J. 27, 379–423 (1948).

Article Google Scholar 098. Gibson, E. Linguistic complexity: locality of syntactic dependencies. Cognition 68, 1–76 (1998).

Article CAS PubMed Google Scholar 099. Johnson, W. Studies in language behavior: a program of research. Psychol. Monogr. 56, 1–15 (1944).

Article Google Scholar 100. Mardaga, H. Hapax legomena: a neglected field in biblical studies. Curr. Biblic. Res. 10, 264–274 (2012).

Article Google Scholar 101. Cronbach, L. J. Coefficient alpha and the internal structure of tests. Psychometrika 16, 297–334 (1951).

Article Google Scholar 102. West, S. G., Finch, J. F. & Curran, P. J. in Structural Equation Modeling: Concepts, Issues, and Applications (ed. Hoyle, R. H.) 56–75 (Sage, 1995). 103. Curran, P. J., West, S. G. & Finch, J. F. The robustness of test statistics to nonnormality and specification error in confirmatory factor analysis. Psychol. Methods 1, 16–29 (1996).

Article Google Scholar 104. Kilian, L. New introduction to multiple time series analysis, by Helmut Lütkepohl, Springer, 2005. Econ. Theory 22, 961–967 (2006).

Article Google Scholar 105. Ivanov, V. & Kilian, L. A practitioner’s guide to lag order selection for VAR impulse response analysis. Stud. Nonlinear Dyn. Econom. 9, 1 (2005). 106. Bruns, S. B. & Stern, D. I. Lag length selection and p-hacking in Granger causality testing: prevalence and performance of meta-regression models. Empir. Econ. 56, 797–830 (2019).

Article Google Scholar 107. Dickey, D. A. & Fuller, W. A. Distribution of the estimators for autoregressive time series with a unit root. J. Am. Stat. Assoc. 74, 427–431 (1979).

Google Scholar 108. Box, G. E., Jenkins, G. M., Reinsel, G. C. & Ljung, G. M. Time Series Analysis: Forecasting and Control (John Wiley & Sons, 2015). 109. Wood, S. N. Generalized Additive Models: An Introduction with R 2nd edn https://doi.org/10.1201/9781315370279 (Chapman and Hall/CRC, 2017). 110. Hastie, T. & Tibshirani, R. Generalized additive models. Stat. Sci. 1, 297–310 (1986).

Google Scholar 111. Wieling, M. Analyzing dynamic phonetic data using generalized additive mixed modeling: a tutorial focusing on articulatory differences between l1 and l2 speakers of English. J. Phon. 70, 86–116 (2018).

Article Google Scholar 112. Greene, R. et al. New and Improved Embedding Model https://openai.com/index/new-and-improved-embedding-model (OpenAI, 2024). 113. Conneau, A. & Kiela, D. SentEval: an evaluation toolkit for universal sentence representations. In Proc. Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (eds Calzolari, N. et al.) https://aclanthology.org/L18-1269/ (European Language Resources Association (ELRA), 2018). 114. Gilardi, F., Alizadeh, M. & Kubli, M. ChatGPT outperforms crowd workers for text-annotation tasks. Proc. Natl Acad. Sci. USA 120, 2305016120 (2023).

Article Google Scholar 115. Alizadeh, M. et al. Open-source LLMs for text annotation: a practical guide for model setting and fine-tuning. J. Comput. Soc. Sci. 8, 17 (2025).

Article PubMed Google Scholar 116. Gwet, K. L. Computing inter-rater reliability and its variance in the presence of high agreement. Br. J. Math. Stat. Psychol. 61, 29–48 (2008).

Article PubMed Google Scholar 117. Levene, H. in Contributions to Probability and Statistics (ed. Olkin, I.) 278–292 (Stanford Univ. Press, 1960). 118. Gentzkow, M., Shapiro, J. M. & Taddy, M. Congressional Record for the 43rd–114th Congresses: Parsed Speeches and Phrase Counts https://data.stanford.edu/congress_text (Stanford Libraries, 2018). 119. Silver, N. & Mehta, D. Both Republicans and Democrats Have an Age Problem https://web.archive.org/web/20221216043420/; [https://fivethirtyeight.com/features/both-republicans-and-democrats-have-an-ag

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →