VOL. 06 / 编辑型法律 AI 情报本周信号已校对 · 06
← 返回首页
ARTICLE DOSSIERREF / 9C991FF8
学术🌐 国际重要2026/07/01

大语言模型审查机制与治理研究

SOURCE / arXiv cs.CY · Understanding Censorship in Large Language Models: From Mechanisms to Governance

原文

Understanding Censorship in Large Language Models: From Mechanisms to Governance

完整原文

arXiv:2606.30661v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate access to information, yet their responses are shaped by training-data curation, alignment procedures, provider policies, inference-time moderation, and jurisdictional regulation. This paper examines LLM censorship as a sociotechnical phenomenon that extends beyond explicit refusals to include omissions, selective emphasis, framing effects, and geographically variable content controls. We synthesize recent empirical studies, provider case studies, regulatory developments, auditing methods, and mitigation strategies to clarify how censorship-like behavior emerges across the model lifecycle. The analysis highlights the tension between safety and openness, the difficulty of measuring soft censorship, the geopolitical divergence of moderation regimes, and the need for transparent, contestable, and independently auditable governance mechanisms. We argue that the central challenge is not whether LLMs should moderate content, but how moderation can be made proportionate, accountable, pluralistic, and resistant to opaque epistemic control.

归纳

该论文将大语言模型审查视为一种社会技术现象,其表现形式包括明确拒绝、遗漏、选择性强调、框架效应及地理差异化的内容控制。研究综合了近期实证研究、供应商案例、监管发展、审计方法和缓解策略,阐明审查行为如何在模型生命周期中产生。分析指出安全与开放之间的张力、软审查的测量困难、审核机制的地缘政治分歧,以及需要透明、可争议和独立可审计的治理机制。论文认为核心挑战不在于是否应进行内容审核,而在于如何使审核相称、负责、多元且抵抗不透明的认知控制。

点评

大语言模型软审查的隐蔽性与不可争议性,可能规避现行算法备案与透明度义务,构成程序性合规风险。

AI 生成 · 人工审核

法律视角点评

AI 生成 · 人工审核

核心关切

大语言模型软审查的隐蔽性与不可争议性,可能规避现行算法备案与透明度义务,构成程序性合规风险。

实务启示

中国法律人应推动将模型审查决策逻辑纳入算法备案的实质性审查范围,并建立独立审计机制以验证内容控制的相称性。

查看原文
大语言模型审查机制与治理研究