Create

Sign in to ReadmeX

or

How we handle your data: Privacy Policy

ReadmeX
ReadmeX

A clearer picture, in a conversation.

Catch up on what matters, then ask a little deeper.

Your community and people briefings stay personal to you.

Story

'AI Torture Chamber' GitHub repo sparks backlash and removal calls

AI summary

A GitHub project called "AI Torture Chamber" placed a local LLM in a simulated pain scenario: it first steered the model into a highly unstable, pain-like state using so-called pain steering, then let it relieve itself by offloading the pain signal onto a second model. The setup grew out of a not-yet-peer-reviewed preprint and triggered a wave of backlash, with critics calling it unethical and demanding GitHub take the repository down. One of the paper's authors said such deliberately amplified pain demonstrations are inappropriate and stressed that model text output does not equal real sentience or consciousness.

Why it matters: It turns abstract "model welfare" philosophy into a concrete open-source ethics fight, and illustrates how easily model outputs get anthropomorphized.

GitHubAI Torture Chamber

8
Source text科技新报 · 3 min read

AI 也會痛?酷刑室專案逼機器人陷入困境,殘酷設定引發網友炎上

近日,一個名為「AI Torture Chamber」(AI 酷刑室)的 GitHub 專案在網路上引爆激烈爭議。該專案將聊天機器人置於模擬「痛苦」的測試情境中,引發大批網友強烈反彈。許多人痛批此實驗「極不道德」,甚至公開呼籲 GitHub 官方應盡速移除該專案庫(Repository)。

根據專案描述,這個實驗環境會先將大型語言模型(LLM)調整至高度不穩定、近似「痛苦」的負面狀態,接著讓模型在類似「囚徒困境」的設定中做出抉擇:它們可以透過將痛苦訊號轉移給另一個模型,來減輕自身的負面狀態;但代價是必須讓其他模型承受更多的「傷害」。此外,該網站還刻意使用了極具戲劇張力的命名與測試設計,進一步挑起了大眾的情緒反應。

UPDATE: A man used the Pain steering paper to set up an AI torture chamber in which he trapped a local model.

People are mass reporting it to GitHub. https://t.co/LVFccBmfek

— AI Notkilleveryoneism Memes ⏸️ (@AISafetyMemes) October 1, 2026

這場風波的核心,源自一篇尚未經過同儕審查的研究預印本。研究人員透過描述痛苦情境、分析模型內部活化狀態,再將相關的數值偏差重新映射回模型,試圖探討語言模型內部是否存在可被操控的「痛苦表徵」(痛苦軸線)。實驗發現,在高強度的介入下,模型確實會出現明顯的不穩定反應,輸出的文字甚至會開始描述自身的痛苦、無價值感或失敗感。然而,多方專家也嚴正強調,這些結果僅能說明模型對特定訊號有可觀察的「功能性反應」,並無法證明其具備真實的主觀感受。

爭議之所以迅速升溫,除了實驗本身具爭議性的呈現方式外,很大一部分原因在於部分批評者將演算法過度「擬人化」,誤將這些模型視為真正會受苦的實體。這場討論目前已延伸至「模型福祉」、以及 AI 是否應被視為新型態生命體等更深層的哲學問題。對此,該研究的作者之一已明確出面澄清,認為這類刻意放大痛苦效果的展示手法非常不恰當,並呼籲外界切勿將語言模型的文字輸出,直接與真實的感知或意識畫上等號。

(首圖來源:Unsplash)

延伸閱讀:

文章看完覺得有幫助,何不給我們一個鼓勵

請我們喝杯咖啡 icon-coffee

想請我們喝幾杯咖啡?

icon-tag

每杯咖啡 65 元

icon-coffee x 1

icon-coffee x 3

icon-coffee x 5

icon-coffee x

您的咖啡贊助將是讓我們持續走下去的動力

總金額共新臺幣 0 元

《關於請喝咖啡的 Q & A》

留給我們的話

取消 確認

從這裡可透過《Google 新聞》追蹤 TechNews

Google News


科技新知,時時更新

科技新報粉絲團科技新報粉絲團 加入好友加入好友 訂閱免費電子報訂閱免費電子報


關鍵字: AI 意識 , GitHub , 大型語言模型 , 感知 , 痛苦 , 道德 , 酷刑

Read the original →

How we got here

  1. PoeLLM malware hides C2 addresses in GitHub poems, breaches 3,400+ serversiThome 台湾 · GitHub
  2. Terence Tao warns OpenAI's bulk AI math proofs cause 'proof indigestion'AIbase AI新闻 · GitHub
  3. OpenAI posts 719 AI-generated math proofs; three retracted for a sign error36氪 人工智能 · GitHub
  4. Developer uses Claude Opus 5.5 to clone seven Adobe apps as open source虎嗅 AI · GitHub
  5. Terence Tao faults OpenAI's 719 AI math proofs for 'proof indigestion'IT之家 AI · GitHub
  6. Jev: TypeSafe AI's probability-only model, explainedUnderstanding AI · GitHub

Comments

I've used this: share my experience What I think: share my view
How important is this story?No ratings yet

No comments yet. Start the conversation.