LLMs时代:语言多样性的日渐式微
摘要
本研究论文探讨了大型语言模型如何导致语言多样性下降,并强调了潜在的文化和社会风险。
暂无内容
查看缓存全文
缓存时间: 2026/09/03 02:53
# 大语言模型时代下语言多样性的萎缩图景
来源:https://www.nature.com/articles/s41562-026-02550-0?error=cookies_not_supported&code=6d4e7e1c-1fe0-42ca-a421-83ae07346d91
## 参考文献
1. 奥威尔,G. 《一九八四》69–70页(企鹅出版社,1949年)。
2. Park, G. 等。基于社交媒体语言的自动化人格评估。《人格与社会心理学杂志》**108**,934–952(2015年)。文章(https://doi.org/10.1037%2Fpspp0000020)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=25365036)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Automatic%20personality%20assessment%20through%20social%20media%20language&journal=J.%20Pers.%20Soc.%20Psychol.&doi=10.1037%2Fpspp0000020&volume=108&pages=934-952&publication_year=2015&author=Park%2CG)
3. Oberlander, J. & Gill, A. J. 富有个性的语言:电子邮件沟通中个体差异的分层语料库比较。《话语过程》**42**,239–270(2006年)。文章(https://doi.org/10.1207%2Fs15326950dp4203_1)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Language%20with%20character%3A%20a%20stratified%20corpus%20comparison%20of%20individual%20differences%20in%20e-mail%20communication&journal=Discourse%20Process.&doi=10.1207%2Fs15326950dp4203_1&volume=42&pages=239-270&publication_year=2006&author=Oberlander%2CJ&author=Gill%2CAJ)
4. Moreno, J. D., Martinez-Huertas, J. A., Olmos, R., Jorge-Botana, G. & Botella, J. 人格特质能否通过分析书面语言来衡量?一项关于计算方法的元分析研究。《人格与个体差异》**177**,110818(2021年)。文章(https://doi.org/10.1016%2Fj.paid.2021.110818)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Can%20personality%20traits%20be%20measured%20analyzing%20written%20language%3F%20A%20meta-analytic%20study%20on%20computational%20methods&journal=Pers.%20Individ.%20Dif.&doi=10.1016%2Fj.paid.2021.110818&volume=177&publication_year=2021&author=Moreno%2CJD&author=Martinez-Huertas%2CJA&author=Olmos%2CR&author=Jorge-Botana%2CG&author=Botella%2CJ)
5. Mairesse, F., Walker, M. A., Mehl, M. R. & Moore, R. K. 使用语言线索自动识别对话和文本中的人格。《人工智能研究杂志》**30**,457–500(2007年)。文章(https://doi.org/10.1613%2Fjair.2349)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Using%20linguistic%20cues%20for%20the%20automatic%20recognition%20of%20personality%20in%20conversation%20and%20text&journal=J.%20Artif.%20Intell.%20Res.&doi=10.1613%2Fjair.2349&volume=30&pages=457-500&publication_year=2007&author=Mairesse%2CF&author=Walker%2CMA&author=Mehl%2CMR&author=Moore%2CRK)
6. Schwartz, H. A. 等。社交媒体语言中的人格、性别与年龄:开放词汇方法。《PLoS ONE》**8**,73791(2013年)。文章(https://doi.org/10.1371%2Fjournal.pone.0073791)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Personality%2C%20gender%2C%20and%20age%20in%20the%20language%20of%20social%20media%3A%20the%20open-vocabulary%20approach&journal=PLoS%20ONE&doi=10.1371%2Fjournal.pone.0073791&volume=8&publication_year=2013&author=Schwartz%2CHA)
7. Kramsch, C. 语言与文化。《AILA评论》**27**,30–55(2014年)。文章(https://doi.org/10.1075%2Faila.27.02kra)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Language%20and%20culture&journal=AILA%20Rev.&doi=10.1075%2Faila.27.02kra&volume=27&pages=30-55&publication_year=2014&author=Kramsch%2CC)
8. Gumperz, J. 言语社区。《国际社会科学百科全书》**9**,381–386(1968年)。Google Scholar (http://scholar.google.com/scholar_lookup?&title=The%20speech%20community&journal=Int.%20Encycl.%20Soc.%20Sci.&volume=9&pages=381-386&publication_year=1968&author=Gumperz%2CJ)
9. Nguyen, D. & Rosé, C. P. 语言使用作为网络社区社会化的反映。载于《社交媒体语言研讨会论文集》(Nagarajan, M. & Gamon, M. 编)76–85页(计算语言学协会,2011年)。
10. Bamman, D., Eisenstein, J. & Schnoebelen, T. 性别认同与社交媒体中的词汇变异。《社会语言学杂志》**18**,135–160(2014年)。文章(https://doi.org/10.1111%2Fjosl.12080)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Gender%20identity%20and%20lexical%20variation%20in%20social%20media&journal=J.%20Socioling.&doi=10.1111%2Fjosl.12080&volume=18&pages=135-160&publication_year=2014&author=Bamman%2CD&author=Eisenstein%2CJ&author=Schnoebelen%2CT)
11. Pennebaker, J. W. 代词的秘密生活。《新科学家》**211**,42–45(2011年)。文章(https://doi.org/10.1016%2FS0262-4079%2811%2962167-2)Google Scholar (http://scholar.google.com/scholar_lookup?&title=The%20secret%20life%20of%20pronouns&journal=New%20Sci.&doi=10.1016%2FS0262-4079%2811%2962167-2&volume=211&pages=42-45&publication_year=2011&author=Pennebaker%2CJW)
12. Robinson, M. D., Boyd, R. L., Fetterman, A. K. & Persich, M. R. 政治(与非政治)话语中的心智与身体:美国政治中意识形态印记的语言证据。《语言与社会心理杂志》**36**,438–461(2017年)。文章(https://doi.org/10.1177%2F0261927X16668376)Google Scholar (http://scholar.google.com/scholar_lookup?&title=The%20mind%20versus%20the%20body%20in%20political%20%28and%20nonpolitical%29%20discourse%3A%20linguistic%20evidence%20for%20an%20ideological%20signature%20in%20US%20politics&journal=J.%20Lang.%20Soc.%20Psychol.&doi=10.1177%2F0261927X16668376&volume=36&pages=438-461&publication_year=2017&author=Robinson%2CMD&author=Boyd%2CRL&author=Fetterman%2CAK&author=Persich%2CMR)
13. Huang, Y., Guo, D., Kasakoff, A. & Grieve, J. 利用Twitter数据分析理解美国区域语言变异。《计算、环境与城市系统》**59**,244–255(2016年)。文章(https://doi.org/10.1016%2Fj.compenvurbsys.2015.12.003)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Understanding%20US%20regional%20linguistic%20variation%20with%20Twitter%20data%20analysis&journal=Comput.%20Environ.%20Urban%20Syst.&doi=10.1016%2Fj.compenvurbsys.2015.12.003&volume=59&pages=244-255&publication_year=2016&author=Huang%2CY&author=Guo%2CD&author=Kasakoff%2CA&author=Grieve%2CJ)
14. Eisenstein, J., O’Connor, B., Smith, N. A. & Xing, E. 用于地理词汇变异的潜变量模型。载于《2010年自然语言处理实证方法会议论文集》(Li, H. & Màrquez, L. 编)1277–1287页(计算语言学协会,2010年)。
15. Peterson, K., Hohensee, M. & Xia, F. 职场中的电子邮件正式性:基于安然语料库的案例研究。载于《社交媒体语言研讨会(LSM 2011)论文集》(Nagarajan, M. & Gamon, M. 编)86–95页(计算语言学协会,2011年)。
16. Stamatatos, E. 现代作者归属方法综述。《信息科学与技术学会杂志》**60**,538–556(2009年)。文章(https://doi.org/10.1002%2Fasi.21001)Google Scholar (http://scholar.google.com/scholar_lookup?&title=A%20survey%20of%20modern%20authorship%20attribution%20methods&journal=J.%20Assoc.%20Inf.%20Sci.%20Technol.&doi=10.1002%2Fasi.21001&volume=60&pages=538-556&publication_year=2009&author=Stamatatos%2CE)
17. Grieve, J. 定量作者归属:技术评估。《文学与语言计算》**22**,251–270(2007年)。文章(https://doi.org/10.1093%2Fllc%2Ffqm020)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Quantitative%20authorship%20attribution%3A%20an%20evaluation%20of%20techniques&journal=Lit.%20Ling.%20Comput.&doi=10.1093%2Fllc%2Ffqm020&volume=22&pages=251-270&publication_year=2007&author=Grieve%2CJ)
18. Cassell, J. & Tversky, D. 在线跨文化社区形成的语言。《计算机媒介传播杂志》**10**,1027(2005年)。Google Scholar (http://scholar.google.com/scholar_lookup?&title=The%20language%20of%20online%20intercultural%20community%20formation&journal=J.%20Comput.%20Mediat.%20Commun.&volume=10&publication_year=2005&author=Cassell%2CJ&author=Tversky%2CD)
19. Danet, B. & Herring, S. C. 引言:多语言互联网。《计算机媒介传播杂志》**9**,9110(2003年)。Google Scholar (http://scholar.google.com/scholar_lookup?&title=Introduction%3A%20the%20multilingual%20internet&journal=J.%20Comput.%20Mediat.%20Commun.&volume=9&publication_year=2003&author=Danet%2CB&author=Herring%2CSC)
20. Kennedy, B. 等。道德关切在语言中存在差异化显现。《认知》**212**,104696(2021年)。文章(https://doi.org/10.1016%2Fj.cognition.2021.104696)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=33812153)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Moral%20concerns%20are%20differentially%20observable%20in%20language&journal=Cognition&doi=10.1016%2Fj.cognition.2021.104696&volume=212&publication_year=2021&author=Kennedy%2CB)
21. Jackson, J. C., Gelfand, M., De, S. & Fox, A. 美国文化在200年间的松动与创造力-秩序权衡有关。《自然-人类行为》**3**,244–250(2019年)。文章(https://doi.org/10.1038%2Fs41562-018-0516-z)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=30953010)Google Scholar (http://scholar.google.com/scholar_lookup?&title=The%20loosening%20of%20American%20culture%20over%20200%20years%20is%20associated%20with%20a%20creativity%E2%80%93order%20trade-off&journal=Nat.%20Hum.%20Behav.&doi=10.1038%2Fs41562-018-0516-z&volume=3&pages=244-250&publication_year=2019&author=Jackson%2CJC&author=Gelfand%2CM&author=De%2CS&author=Fox%2CA)
22. Hofmann, V., Kalluri, P. R., Jurafsky, D. & King, S. AI基于方言对人产生隐蔽的种族主义决策。《自然》**633**,147–154(2024年)。文章(https://doi.org/10.1038%2Fs41586-024-07856-5)CAS (https://www.nature.com/articles/cas-redirect/1:CAS:528:DC%2BB2cXhvVKit7%2FM)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=39198640)PubMed Central (http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11374696)Google Scholar (http://scholar.google.com/scholar_lookup?&title=AI%20generates%20covertly%20racist%20decisions%20about%20people%20based%20on%20their%20dialect&journal=Nature&doi=10.1038%2Fs41586-024-07856-5&volume=633&pages=147-154&publication_year=2024&author=Hofmann%2CV&author=Kalluri%2CPR&author=Jurafsky%2CD&author=King%2CS)
23. Richard, A. B., Lelandais, M., Reilly, K. T. & Jacquin-Courtois, S. 连续言语中细微认知障碍的语言标记:系统综述。《言语、语言与听力研究杂志》**67**,4714–4733(2024年)。文章(https://doi.org/10.1044%2F2024_JSLHR-24-00274)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=39546411)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Linguistic%20markers%20of%20subtle%20cognitive%20impairment%20in%20connected%20speech%3A%20a%20systematic%20review&journal=J.%20Speech%20Lang.%20Hear.%20Res.&doi=10.1044%2F2024_JSLHR-24-00274&volume=67&pages=4714-4733&publication_year=2024&author=Richard%2CAB&author=Lelandais%2CM&author=Reilly%2CKT&author=Jacquin-Courtois%2CS)
24. Eyigoz, E., Mathur, S., Santamaria, M., Cecchi, G. & Naylor, M. 语言标记预测阿尔茨海默病的发病。《EClinicalMedicine》**28**,100583(2020年)。
25. Roark, B., Mitchell, M., Hosom, J.-P., Hollingshead, K. & Kaye, J. 用于检测轻度认知障碍的口语派生测量。《IEEE音频、语音与语言处理汇刊》**19**,2081–2090(2011年)。文章(https://doi.org/10.1109%2FTASL.2011.2112351)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=22199464)PubMed Central (http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3244269)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Spoken%20language%20derived%20measures%20for%20detecting%20mild%20cognitive%20impairment&journal=IEEE%20Trans.%20Audio%20Speech%20Lang.%20Process.&doi=10.1109%2FTASL.2011.2112351&volume=19&pages=2081-2090&publication_year=2011&author=Roark%2CB&author=Mitchell%2CM&author=Hosom%2CJ-P&author=Hollingshead%2CK&author=Kaye%2CJ)
26. Trifu, R. N. 等。重度抑郁症的语言标记:使用自动化程序的横断面研究。《心理学前沿》**15**,1355734(2024年)。文章(https://doi.org/10.3389%2Ffpsyg.2024.1355734)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=38510303)PubMed Central (http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10953917)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Linguistic%20markers%20for%20major%20depressive%20disorder%3A%20a%20cross-sectional%20study%20using%20an%20automated%20procedure&journal=Front.%20Psychol.&doi=10.3389%2Ffpsyg.2024.1355734&volume=15&publication_year=2024&author=Trifu%2CRN)
27. Weerasinghe, J., Morales, K. & Greenstadt, R. "因为...我被告知...如此之多”:Twitter上心理健康状态的语言指标。《隐私增强技术会议录》**4**,152–171(2019年)。
28. 沃尔夫,B. L. 《语言、思想与现实:本杰明·李·沃尔夫文集选》(MIT出版社,2012年)。
29. Eckert, P. 变异研究的三波浪潮:社会语言学变异研究中意义的涌现。《人类学年鉴》**41**,87–100(2012年)。文章(https://doi.org/10.1146%2Fannurev-anthro-092611-145828)Google Scholar (http://scholar.google.com/scholar_lookup?&title=Three%20waves%20of%20variation%20study%3A%20the%20emergence%20of%20meaning%20in%20the%20study%20of%20sociolinguistic%20variation&journal=Annu.%20Rev.%20Anthropol.&doi=10.1146%2Fannurev-anthro-092611-145828&volume=41&pages=87-100&publication_year=2012&author=Eckert%2CP)
30. Hofstede, G. 《文化的后果:跨国价值观、行为、制度与组织的比较》第二版(Sage出版社,2001年)。
31. OpenAI. 《介绍ChatGPT》 https://openai.com/blog/chatgpt(2022年)。
32. Gemini Team 等。Gemini:一个能力强大的多模态模型家族。预印本见 https://doi.org/10.48550/arXiv.2312.11805(2023年)。
33. Bailyn, E. 《ChatGPT使用统计:2026年3月》 https://firstpagesage.com/seo-blog/chatgpt-usage-statistics/(FirstPageSage, 2026年)。
34. 《近三分之一大学生曾在书面作业中使用ChatGPT》 https://www.intelligent.com/nearly-1-in-3-college-students-have-used-chatgpt-on-written-assignments/(Intelligent, 2024年)。
35. McClain, C. 《美国人对ChatGPT的使用正在增加,但很少有人信任其选举信息》 https://www.pewresearch.org/short-reads/2024/03/26/americans-use-of-chatgpt-is-ticking-up-but-few-trust-its-election-information/(皮尤研究中心, 2024年)。
36. Handa, K. 等。哪些经济任务由AI执行?来自数百万次Claude对话的证据。预印本见 https://doi.org/10.48550/arXiv.2503.04761(2025年)。
37. Mizrahi, M. 等。什么样的“艺术现状”?呼吁对大语言模型进行多提示评估。《计算语言学协会汇刊》**12**,933–949(2024年)。文章(https://doi.org/10.1162%2Ftacl_a_00681)Google Scholar (http://scholar.google.com/scholar_lookup?&title=State%20of%20what%20art%3F%20A%20call%20for%20multi-prompt%20LLM%20evaluation&journal=Trans.%20Assoc.%20Comput.%20Linguist.&doi=10.1162%2Ftacl_a_00681&volume=12&pages=933-949&publication_year=2024&author=Mizrahi%2CM)
38. Serapio-García, G. 等。大语言模型中人格特质评估与塑造的心理测量框架。《自然-机器智能》**7**,1954–1968(2025年)。文章(https://doi.org/10.1038%2Fs42256-025-01115-6)PubMed (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=39546411)Google Scholar (http://scholar.google.com/scholar_lookup?&title=A%20psychometric%20framework%20for%20evaluating%20and%20shaping%20personality%20traits%20in%20large%20language%20models&journal=Nat.%20Mach.%20Intell.&doi=10.1038%2Fs42256-025-01115-6&volume=7&pages=1954-1968&publication_year=2025&author=Serapio-Garcia%2CG&author=... )
相似文章
Linguistic Monoculture in LLM-Assisted Language Use
This paper introduces a mathematical framework to study how reliance on shared LLMs for writing may reduce population-level linguistic diversity, analyzing fixed, recursive, and personalized interaction mechanisms and characterizing equilibria and convergence rates.
对齐更优,多样性下降?分析两代大语言模型的语法与词汇特征
这篇学术论文分析了两代大语言模型与人类撰写新闻文本相比的句法和词汇多样性,发现较新的对齐模型表现出多样性降低的现象。
超越“AI语言”:论LLM输出的个人言语特征本质
本文认为,大语言模型的输出展现出类似人类个人言语的模型特定语言特征,通过分析2024年和2026年的语料库,揭示了代际变化以及稳定的个体模型画像。
迈向超越英语中心化开发的大语言模型
本文证明了大语言模型严重偏向英语,并表明持续预训练在将模型适配到其他语言(尤其是文化理解方面)时,并不比从头训练更具成本优势。
LLMs是否变得同样具有创造力?来自三年模型的证据
本文分析了三年来LLM的输出,使用开放式任务和创造力评估,发现多样性在统计上显著下降,这表明创造性内容在不同模型间趋于一致,这可能会削弱人类在协作创意工作中的自主性。