英文标题:Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models
作者:Luka Debevc, Nishan Chatterjee, Antoine Doucet, Senja Pollak, Matej Martinc
arXiv ID:2609.08637 | 分类:cs.CL | 发表:2026-09-08
许可:CC-BY
摘要 大型语言模型正日益被部署为新闻、政策与内容审核的信息中介,然而衡量其政治行为仍然十分脆弱。单次问卷施测会将模型真实的倾向与测量伪影及回答诱发偏差混为一谈。我们引入一个稳健的评估框架,利用政治罗盘测试(PCT)来测量大型语言模型的政治偏见,通过在八维扰动空间中系统采样300种实验配置,变化语言、框架、指令、答案格式、选项顺序和人格措辞。我们对八个Gemma 3和Qwen 3模型以14种语言原生评估,并覆盖三个量化级别,提取出带有量化不确定性的设计平均政治坐标。我们表明,尽管大多数模型平均而言偏向自由意志主义-左翼,但关键评估因素(指令措辞、语言和答案格式)对恢复出的坐标产生显著影响,使得所
Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。