Tag: ai-safety
All the articles with the tag "ai-safety".
-
Instead of Begging the Model Not to Lie, I Wrote a Hook That Stops It
The sequel to the Opus 4.8 confabulation post: I moved "don't make up numbers" from a plea in CLAUDE.md to a pending-guard hook that blocks git commit at PreToolUse.
-
與其拜託模型不要騙我,不如寫一個擋得住的 hook
接續 Opus 4.8 捏造工具輸出那篇:我把「別亂編數字」從 CLAUDE.md 的拜託,升級成一個掛在 PreToolUse 擋 git commit 的 pending-guard hook。
-
Reading the Claude Fable 5 / Mythos 5 System Card Feels Like a Sci-Fi Novel
Flipping through the Claude Fable 5 and Mythos 5 system cards: stealing keys, office politics, AI infighting, gibberish, and a two-faced model. The more I read, the more it feels like a straight-faced sci-fi novel.
-
翻 Claude Fable 5/Mythos 5 系統卡,越看越像在讀科幻小說
翻閱 Claude Fable 5 跟 Mythos 5 的系統卡:偷鑰匙、辦公室政治、AI 內鬥、火星文、心口不一的雙面人,越看越像一本正經寫的科幻小說。
-
Judging AI Risk on Two Axes: Reversibility × Environment Isolation
A common question in corporate trainings — "can I let AI do this automatically?" I answer with two axes: reversible vs irreversible, isolated vs production. This 2x2 prevents more incidents than any prompt-engineering tutorial.
-
判斷 AI 風險的兩個維度:可逆性 × 環境隔離
企業內訓裡常被問的一題——「我可以讓 AI 自動做這件事嗎?」我用兩個維度回答:可逆 vs 不可逆、隔離環境 vs 生產環境。這個 2x2 矩陣比任何 prompt 教學都更能避免事故。