Posts
All the articles I've posted.
-
只換外面那圈 loop
LongHorizon-Harness 架在 Claude Code、Codex 這些 agent 外面,不訓練模型也不取代你的 agent,只做迴圈,WeaveBench 完成率從 51.8 拉到 80.7。
-
不是拉力,是推力:從這一週 Anthropic 的溝通說起
近幾年用戶暴漲從來不是拉力,是推力。兩邊就是在比誰少犯錯,而這一週 Anthropic 的幾則發言,剛好把這件事演了一遍。
- Updated:
Qwen 3.8 開放權重的這一週:從模型卡到 A3B 被刪
從 Qwen 3.8 27B 模型卡放出、實測體感、到 35B A3B 在 Modelscope 註冊又被刪掉,一週之內的完整過程。
- Updated:
何時派工 Fable?三個時機,還有別開 ultracode
最貴的模型不是拿來全程跑的。三個我體感最好的 Fable 派工時機、為什麼 Opus 5 我最多開到 med,以及我後來變成週一到週四效率王、週五到週日 Fable 模式的那個循環。
-
Your AI Tools Are Now the Attack Vector: npm and Python Supply-Chain Backdoors, a $1000 Stolen-Key Bill, and a Scan-Before-You-Install SOP
Attackers are planting instructions in Claude Code, Cursor, and Gemini CLI configs so your own assistant runs the exfiltration script. Two supply-chain waves, one $1000 stolen-key bill, and the five minutes to spend before installing anything.
-
Your Automated Video Pipeline Is Silently Dropping Work: Gotchas From Remotion, ffmpeg, the YouTube API, and HyperFrames
Four months of video automation gotchas, from Remotion through HyperFrames: truncated downloads that exit 0, thumbnail titles cut mid-word, transcript chunks returning zero lines, large uploads that hang. None of them raise an error.
-
你的 AI 工具已經是攻擊媒介:npm/Python 供應鏈後門、外洩金鑰盜刷 1000 美元,與裝 MCP 前的掃描 SOP
攻擊者已經在 Claude Code、Cursor、Gemini CLI 的設定檔裡植入指令,讓你的 AI 助手幫他們跑竊取腳本。這篇串起兩波供應鏈攻擊、一次一千美元的盜刷,以及安裝前該做的那五分鐘。
-
自動發片管道正在悄悄漏成品:Remotion、ffmpeg、YouTube API、HyperFrames 的踩坑總表
從 Remotion 到 HyperFrames,四個月的影片自動化踩坑合輯:exit 0 的殘檔、靜默截斷的縮圖標題、回 0 行的轉錄合併、卡死的大檔上傳。共同點是全都不報錯。
-
ABC Legal's Agent Fleet: Turning AI Experiments Into a Governed Production System
Six cards summarizing Anthropic's ABC Legal case study: 50+ production agents, agents managed as software in git, a steering committee with no software developers on it, and the principle that trust comes before automation.
-
Once It Ships, It's Gone: Three Times an Agent Misfired on an External System
Calendar invites, replies, and the internal notes on a hold event are all one-way doors. I got each of them wrong once this month, so here are the shapes of the accidents and the checks that catch them.
-
Channel Value and Deal Value Belong in Separate Ledgers
Three rules that came out of two weeks of back-to-back BD calls: score a deal's channel value separately from its deal value, use The Pumpkin Plan to decide which slice to take, and write down a hard ceiling for free consulting.
-
My CLAUDE.md Boiled Down to Eight Lines of "Take Pride In…, Take Shame In…"
My CLAUDE.md cut down to eight paired maxims, plus a human-machine collaboration circular written deliberately in Party-document style, plus a sixteen-character guideline.
-
Codex Falls Off the Pedestal: People Caught the Quota Nerf
A month ago I was writing about the Codex quota honeymoon. Now people are measuring a 50% nerf on Plus and 77% on Pro 5x, and my "off the pedestal in two months" call turns out to have been optimistic.
-
Eleven HyperFrames Video Gotchas
Transition flicker, a worker-parallelism myth carried over from another framework, transcript format drift, slide sizing, and silently truncated thumbnails — eleven things that bit me on the HyperFrames video line in one week of August.
-
A Local Model Is Not a Way to Save Money
If you are buying hardware to run a local model because you want to save money, let me do the math with you first: NT$100k minimum, commercial models at bleeding prices, and privacy as the only advantage left.
-
A Chatty Model: Teaching Opus 5 to Converge
Opus 5 does not speak plainly, and it also overwrites. A cue card turns into a treatise. I use a first-principles skill to teach it to converge.
-
Five Things I Never Saw Coming Two Years Ago
Five changes in myself I never saw coming two years ago: the terminal, cancelling Office 365, voice input, dispatching AI agents, and a desktop that finally got cleaned up.
-
ABC Legal 的代理人艦隊:把 AI 實驗變成可治理的生產系統
整理 Anthropic 官方 ABC Legal 案例的六張圖卡:50+ 生產環境代理人、把 agent 當軟體放進 git、非工程師組成的 steering committee,以及「可信任,才自動化」的營運原則。
-
寄出去就收不回來:agent 動外部系統的三次誤操作
行事曆邀請、回信、佔位事件的內部備註,這三種動作都是寄出即無法收回。這個月我在這三處各犯一次,記下事故形狀與防呆判準。
-
通路價值和成案價值,要分開記帳
這兩週連著跑幾場 BD,逼出三條判準:一筆生意的價值要拆成通路和成案兩層分開算、用《南瓜計畫》決定接哪一塊、以及免費諮詢的上限寫死在哪裡。
-
我的 CLAUDE.md 精煉到只剩八句「以…為榮,以…為恥」
把 CLAUDE.md 一路砍到只剩八組對句,外加一篇刻意寫成公文體的人機協同宣導文,還有十六字方針。
-
Codex 走下神壇:額度偷縮被網友抓出來了
一個多月前我還在寫 Codex 的額度蜜月期,現在網友實測 Plus 被偷縮 50%、Pro 5x 被偷縮 77%,我那句「兩個月跌下神壇」還是太樂觀了。
-
HyperFrames 影片產出的十一個坑
轉場閃動、worker 並行迷思、字幕格式漂移、投影片版式、縮圖靜默截斷——八月這一週在 HyperFrames 這條影片線上踩到的十一個坑,逐個記現象、根因、解法。
-
本地模型不是拿來省錢的
如果你買硬體架本地模型的目的是省錢,先聽我幫你算這筆帳:硬體十萬起跳、商用模型卷到流血價,本地模型唯一不可取代的優勢只剩隱私。
-
話癆的模型:教 Opus 5 怎麼收斂
Opus 5 除了不講人話之外還會過度寫作,小抄寫成萬言書。我用 first-principles skill 教它收斂,順帶談怎麼遏制 overthinking。
-
兩年前從來沒想過會發生在自己身上的五件事
兩年前從來沒有想過、卻真的發生在自己身上的五個變化:終端機、退訂 Office 365、語音輸入、派 AI Agent,還有被整理乾淨的桌面。
-
Two Skills That Make AI Talk Like a Human: ASD-STE100 for Grammar, ISO 24495 for Structure
A skill I threw together three weeks ago got passed around by people overseas and picked up three PRs, so I shipped a second one I use every single day.
-
讓 AI 對你說人話的兩個 skill:ASD-STE100 管文法,ISO 24495 管結構
三個禮拜前隨手做的 asd-ste100-skill 被老外瘋狂轉載還留了三個 PR,順手再開一個天天在用的 iso-24495-skill。
-
Information Diet: I Pulled My Own Chrome History and Audited Where My Attention Goes
People obsess over whether every bite of food is clean, then swallow piles of dirty info online without a second thought. So I pulled my own Chrome history db and had it analyze my browsing habits. A few sections came out uglier than I expected.
-
Information Diet:我抓自己的 Chrome 瀏覽紀錄,做了一次注意力盤點
現代人極度在意吃進口裡的每一塊食物乾不乾淨,卻毫不在乎上網吃到的一大堆 Dirty Info。所以我抓了自己的 Chrome 瀏覽紀錄 db,讓工具分析我的上網習慣,結果有幾段比我預期的難看。
-
AI Will Never Go to Jail for You — I Figured Out What I Want to Teach, Then Got Told I Had Half of It Wrong
Since generative AI arrived I have been asking what AI can never do on a human's behalf. My answer is accountability, because AI does not get sentenced. Then I posted the argument and found out what I had missed about the risk carried by people at the bottom.
-
AI 不會代替人去坐牢——我想清楚要教什麼,然後被網友提醒想錯了一半
生成式 AI 之後,我一直在想有什麼是 AI 永遠不能替人類做的。我的答案是「負責」,因為 AI 不會被判刑。但這個論述貼出去之後,我發現自己漏想了基層承擔的風險。
-
NT$18,000 a Month for an AI Course: Price and Value Came Apart a While Ago
An AI-course cash grab overheard at a convenience store, set against the two things I can actually show: a SKILL distilled from fifteen years of material, and a mock exam interface pulling past a thousand USD a month. Price tracks marketing, not delivery.
-
My Context Window Is More Fragile Than Today's Frontier Models: I Open-Sourced a Document Review SKILL
Working with an agent drains my attention, so I worked out a review loop: he writes the full proposal, I review it by voice, he revises, I go eat and exercise, then v2 shows up.
-
AI 課一個月一萬八:教學市場的價格跟價值早就脫鉤了
在超商用餐區聽到的 AI 課吸金實況,對照我自己手上兩件真的能驗證的東西:15 年教材蒸餾成的 SKILL,跟每月破千美金的模擬考訂閱。價格反映的是行銷強度,不是交付密度。
-
我的 Context Window 比當今大模型還脆弱:我開源了一個文檔審查 SKILL
跟 Agent 合作會注意力耗弱,於是我摸索出一套審查流程:請他出完整提案書,我錄音審查,他改,我去吃飯運動,等 v2。
-
Why I Built My Own Newsletter
Once I started using AI, the feeds I follow changed completely. Threads is full of third-, fourth-, fifth-hand reposts with extra seasoning added, so I pulled together 300-plus sources and had AI build my own weekly digest.
-
為什麼我自己做了一份電子報
開始用 AI 之後,我關注的社群媒體整個換掉。脆上很多是三四五六手搬運還要加料,所以我自己整理 300 多個源頭,讓 AI 做出我個人的電子週報。
-
The CLI Is the Firstborn
New features and fresh bug fixes land in the CLI first. Remote control waits forever. Same company, wildly different update cadence.
-
The Machines That Need the CLI Most Belong to People Least Likely to Learn It
For a machine short on resources the CLI uses a fraction of what a GUI does, but the people with the least headroom are also the least likely to ever learn it. Plus a note on how low the bar for local models actually is.
-
A July Full of Resets: A Light AI User Reviews the Subsidy War
Every time a reset landed in July I used it to the last drop, and the ledger says 57.5x. Plus notes on the subsidy war, Deepseek's price hike, and my bet on the next reset.
-
Letting an Agent Tune a Local Video Model Overnight: My Three Gates
Three things I learned from letting an agent run local video-model parameter research overnight: hard constraints in CLAUDE.md, a gate script watching SSD writes, and a 30-minute check-in loop.
-
Installing Gemini CLI With Beginners Cures Most Claude Code Ailments
Five steps I walk Claude Code / Codex beginners through when installing the Gemini Antigravity CLI: multimodal use cases, getting comfortable with a CLI, a quick win from voice transcription, and finishing on data governance and local mode.
-
CLI 是親生的,其他都是後媽養的
新功能和剛修好的 bug 都優先落在 CLI,remote control 要等猴年馬月。同一家公司的產品,更新頻率差很多。
-
最需要 CLI 的電腦,主人最不可能學會 CLI
對資源有限的電腦來說 CLI 佔用是 GUI 的零頭,但電腦資源最不夠的族群也剛好最不可能學會 CLI;順帶聊本機模型的硬體門檻其實沒那麼高。
-
充滿 reset 的七月:一個 AI 輕量用戶的補貼大戰復盤
七月一有 reset 就用好用滿,帳面發揮 57.5x 的價值;順便記下補貼大戰、Deepseek 漲價,和我對下一次重置的預測。
-
讓 Agent 整夜自己調本地影片模型:我設的三道閘門
讓 agent 夜間自己跑本地影片模型調參的三個經驗:硬約束寫進 CLAUDE.md、閘門腳本監控 SSD 寫入、30 分鐘回來盯一次避免快取失效。
-
帶初學者裝 Gemini CLI,治 Claude Code 百病
帶 Claude Code / Codex 初學者裝 Gemini Antigravity CLI 的五個步驟:從多模態 use case、CLI 介面的心理建設、語音轉錄的 quick win,一路帶到資料治理與本機模式。
-
The Alignment Gap Is Closing. Next Comes Taste and Verification
The gap between AI and human intent is closing fast, so what separates good output moves toward taste and verification — and verification is where human responsibility stays.
-
Other People Fly to Korea for Cosmetic Surgery. I Had Codex Do Mine.
Someone on Twitter chained Codex, Hyperframes, IndexTTS2 and HeyGen into a one-person media pipeline. I pulled the repo, tried it, played with a local TTS model, then rebuilt a video days later on Claude Code.
-
Six New Context Engineering Rules for Claude 5, and the 1,473 Lines I Cut
After reading Anthropic's context engineering guidance for Claude 5, I turned the key points into nine cards, then cut 1,473 lines from a harness I had built up over four model generations.
-
Fewer Prompts Is Not Less Control
Claude 5 and GPT-5.6 official guidance is converging on the same thing — retiring old-style prompt stacking. But trimming is not letting go; control just moves. And the payoff of maintaining that layering is switching tools without losing a step.
-
My Local Model Lineup on a Mac mini, Plus Three Bad Habits I'm Owning Up To
Ten days of offline models on a 48GB Mac mini M4 Pro — the picks, the task split, a terrifying swap write rate, and three bad habits I admit to.
-
Porting Old Prompts to Opus 5: A Nine-Card Migration Guide
Nine cards I made after reading Anthropic's official Opus 5 prompting guide: response length, agent narration, task boundaries, subagent delegation, self-correction, and what happens when you turn thinking off.
-
Supposedly the Smartest Model, and Opus 5 Spent My Whole Day Spinning in Place
Wrong languages, sudden Simplified Chinese, then whole turns with no visible output — and a session log that matched an issue open since June 15.
-
Two Habits for Voice-Driving Coding Agents, and I Am Still Finding the Balance
Short instructions go through push-to-talk; walking through a whole course outline or system architecture goes through a full QuickTime recording that I hand to the AI to transcribe and execute. Plus the setup I use to burn the AI quota that came free with Google Drive.
-
對齊的 gap 正在縮小,接下來拚的是品味與驗證
AI 對齊人類意圖的 gap 正在快速縮小,區分產出品質的要素會往品味與驗證轉移;而驗證背後是人類永遠不會被取代的責任。
-
別人去韓國做醫美,我叫 Codex 數位醫美
從推特上看到有人把 Codex、Hyperframes、IndexTTS2、HeyGen 串成一條自媒體流水線,拉下來實測、順手玩 b 站的本地 TTS,幾天後換 Claude Code 重做一支。
-
Claude 5 時代的情境工程六條新規則,與我砍掉的 1473 行 harness
讀完 Anthropic 對 Claude 5 的情境工程指南後,我把重點做成九張圖卡,然後照著把累積四個世代的 harness 削掉 1473 行。
-
更少的 prompt,不是更少的控制
Claude 5 與 GPT-5.6 的官方指南正在合流,都在淘汰舊式 prompt 堆疊。但精簡不等於放任,控制只是換了位置——而維護好這套分層的紅利,是換哪個工具都能立刻接手。
-
Mac mini 上的本地模型陣容,順便招認我的三大劣根性
48GB 的 Mac mini M4 Pro 試了十來天離線模型,選型結論、任務分工、swap 嚇死人的寫入量,還有我承認的三個劣根性。
-
舊 prompt 搬到 Opus 5:九張卡的遷移指南
讀完 Anthropic 官方的 Opus 5 提示指南後整理的九張卡:回答長度、代理敘述、任務邊界、子代理委派、自我修正,還有關掉 thinking 的副作用。
-
號稱最聰明的 Opus 5,一整天在我對話裡空轉
從回錯語言、突然寫簡體,到整輪沒有可見輸出,查 session log 之後對上一個 6/15 就開著沒修的 issue。
-
語音下指令的兩種習慣,我還在找平衡
用 Claude Code 或 Codex 時,短指令我用隨按即錄,要盤點整套思路時我改開 QuickTime 完整錄音,再丟給 AI 轉錄執行。附上我拿 Google 雲端硬碟送的額度做轉錄的設定。
-
I'm Stuck at Stage 2.5 of AI Adoption
After watching Boris Cherny break down the stages of AI adoption, it hit home: I'm stuck at stage 2.5, where limited attention plus low trust becomes a vicious cycle. This week I used late-night schedules, a morning dashboard, and herdr auto-spawning sessions to cut a 4-5 hour workflow down to 1-2 hours.
-
A Zero While Sitting on a Gold Mine
I read a piece on the perceived value behind churn-and-burn courses, turned it into a skill, ran it on my own site, and got back a verdict: a zero while sitting on a gold mine.
-
Some of My Subagents Were Already Grandpas
A CCX quota incident with no guardrails set: subagents bred recursively, one session burned 90% in half an hour, and here is the full forensics and fix.
-
Codex Built Its Own Evidence Package and Went to Argue With Google Support
A leaked Gemini backend key at PDT Learning got abused, no spending cap, and burned 1000 USD. I pointed Codex's browser automation at Google's live support to fight the charge, and it went so hard it built a 15-page evidence package and sent it over.
-
Two Traps Running Claude Code on Local Ollama: Truncated Context, and A3B Buckling Under Heavy Verification
Two things I logged: cco (Claude Code driven by local Ollama) had long been giving off-topic answers, and the root cause was not a weak model but a full harness whose system prompt had ballooned to 30-50k tokens and was being silently truncated; then I put A3B on the Mac mini for two days as a night worker.
-
Instead of Begging the Model Not to Lie, I Wrote a Hook That Stops It
The sequel to the Opus 4.8 confabulation post: I moved "don't make up numbers" from a plea in CLAUDE.md to a pending-guard hook that blocks git commit at PreToolUse.
-
Second Harness Diet: Global Skills From 58 Down to 40
A follow-up to the harness diet series: six months later, a bigger cleanup that cut my global skills from 58 down to 40, start to finish.
-
我卡在 AI 導入的 2.5 階段
看了 Boris Cherny 分享的 AI 導入階段框架後有感而發:我卡在 2.5 階段,注意力有限加信任不足變成惡性循環,這週靠深夜排程、晨間儀表板、herdr 自動開 session 把 4-5 小時的協作壓到 1-2 小時。
-
坐擁金礦的零分
讀到一篇談割韭菜課程感知價值的文章,我馬上做成 skill 拿去盤點自己的網站,換來一句「坐擁金礦的零分」。
-
有些 subagent 都當阿公了
一次 CCX 沒設好護欄的額度事故:subagent 遞迴繁殖,一個 session 半小時燒掉 90%,事後鑑識與修法全記錄。
-
Codex 自己做了一份證據包,跑去跟 Google 真人客服吵架
PDT Learning 一支 Gemini backend key 外洩被盜刷、沒設 spending cap 怒噴 1000 USD,我用 Codex 的瀏覽器操作去跟 Google 真人客服爭費用,它認真到自己生出一份 15 頁證據包發給對方。
-
本機 Ollama 跑 Claude Code 的兩個坑:context 被截斷、A3B 撐不住重驗證
記錄兩件事:cco(本機 Ollama 驅動 Claude Code)長期答非所問,根因不是模型不夠聰明,是完整 harness 的 system prompt 早就膨脹到 3-5 萬 token 被靜默截斷;順手把 Mac mini 換上 A3B 當夜間工人測了兩天。
-
與其拜託模型不要騙我,不如寫一個擋得住的 hook
接續 Opus 4.8 捏造工具輸出那篇:我把「別亂編數字」從 CLAUDE.md 的拜託,升級成一個掛在 PreToolUse 擋 git commit 的 pending-guard hook。
-
第二次 harness 減肥:全域 skills 58 砍到 40
接續 harness 減肥系列,半年後做了第二輪更大規模的整理,把全域 skills 從 58 個砍到 40 個的完整過程。
-
A Few Days Into Switching From Claude Code to Codex: The Quota Honeymoon
A running log of switching from Claude Code to Codex this week: quota I could not burn through fast enough, a $20 Sol Ultra run that beat a $100 Fable plan on a professor friend's paper, and a tool-chain swap to Hyperframe.
-
從 Claude Code 換到 Codex 的這幾天:額度蜜月期
這週從 Claude Code 換到 Codex 的體感記錄:額度多到花不完、Sol Ultra 20 鎂幫教授跑論文,比 Fable 100 鎂方案還划算,工具鏈也跟著換成 Hyperframe。
-
Picking a Brand Look as a Non-Designer: AI Made Prototypes Cheap, So Taste Became the Hard Part
I am not a designer, but I spent this week iterating on a visual system for AgentCrew Academy with Fable, GPT-image-2, and GPT-5.6-sol. Here are the four rejected directions and the final spec applied to the website, slides, and documents.
-
外行人選品牌視覺:AI 把提案做便宜了,難的變成你的品味
我不是設計師,這週用 Fable、GPT-image-2、GPT-5.6-sol 迭代出 AgentCrew Academy 的視覺系統,記下被否決的四個方向與最後套用到網站、投影片、文件的規格。
-
Four Things I Have Learned from Teaching AI
Recent corporate workshops reinforced four lessons for me: teach in person when possible, cut the slide count, show smart people the result first, and let students do the work.
-
I Finally Switched My Daily Driver from Claude Code to Codex
I try new tools easily, but I am slow to leave the ones already built into my routine. This time, quotas, pricing, and product direction all crossed the line together.
-
最近教 AI 課學到的四件事
最近幾次企業培訓讓我重新確認四件事:實體課更好掌握節奏、投影片要做減法、聰明人先看結果,以及學員真正需要的是放手實作。
-
我終於把主力從 Claude Code 換成 Codex 了
我很愛試新工具,卻很難離開已經用習慣的工具。這次真正讓我換主力的,不是一次 benchmark,而是額度、價格與可預見的產品方向一起越過了臨界點。
-
I Let My /adhd Skill Go Dig Up a 3C-Scene Price-Hike Rumor on Its Own
I toss a vague rumor topic at my /adhd skill and watch it search, admit it searched wrong, stop to ask a disambiguating question, and finally piece the whole thing together with sources.
-
Sonnet 5 Is Out: The Lazy Version of the System Card, Plus Where I Actually Use It
Sonnet 5 shipped. I boiled the system card down to six plain-language points, then added where it actually fits: it is not here to fight Opus, it is here to replace Sonnet 4.6.
-
我讓 /adhd skill 自己去查一件 3C 圈的漲價八卦
含糊丟一個八卦題目給 /adhd skill,看它自己一路搜、承認搜錯、停下來反問消歧義,最後拼出全貌附來源。
-
Sonnet 5 發布:系統卡懶人版,跟我的實測定位
Sonnet 5 發布了,我把系統卡整理成六點白話懶人版,再加上實測定位:它不是來搶 Opus,是來換掉 Sonnet 4.6。
-
ADHD, Flow, and Why I Teach AI
Confessions of someone with ADHD: if I make it past two weeks and still have flow, odds are I can do this thing for a decade-plus.
-
The Client Says They Want Training, But Deep Down They Want a System
A post-mortem on letting go in a BD deal: the client drifted toward "give me an artifact" three times, and I changed one phrase from "after the training" to "if the priority is to build the hub" — and handed the ball back.
-
From Takedown to Return, I Caught a Case of AnthroPTSD
An observation diary of Anthropic models getting pulled by the government and apparently returning: from arrogance, to political wrangling, to OpenAI winning by default, to the conditioned reflex I've been trained into.
-
Claude Seeing Ghosts Four Nights Straight: A Log of Opus 4.8 Fabricating Tool Output
A four-day log of Opus 4.8 tool-result confabulation: from the technical symptoms to the GitHub issue to JSONL forensics.
-
ADHD、心流,與我為什麼教 AI
一個 ADHD 患者的自白:撐過兩週還保有心流,這件事我大概率能做十幾年。
-
客戶嘴上要培訓,骨子裡要系統
一次 BD 鬆手的覆盤:客戶三次飄向「給我 artifact」,我把一句話從「after the training」改成「if the priority is to build the hub」,球就交回去了。
-
從下架到回歸,我得了 AnthroPTSD
Anthropic 模型被政府下架到疑似回歸的觀察日記:從傲慢、政治角力、OpenAI 躺贏,到我被訓練出的條件反射。
-
Claude 半夜見鬼連續四天:Opus 4.8 捏造工具輸出實錄
Opus 4.8 連續四天的 tool-result confabulation 實錄:從技術現象、GitHub issue 到 JSONL 鑑識。
-
Opus 4.8 Cries 'Prompt Injection,' Codex GPT-5.5 Tracks Down the Real Cause: a Worktree Race
With multiple sessions open in one repo and no worktree, Opus 4.8 raised a false prompt injection alarm, told me to check for supply-chain attacks and rotate API keys, until Codex GPT-5.5 read the session log and pinned the real cause: just a worktree race.
-
Opus 4.8 搞烏龍報『提示注入攻擊』,Codex GPT-5.5 揪出 worktree race 真因
Opus 4.8 在我忘開 worktree 的情況下誤報「提示注入攻擊」,叫我去查供應鏈、rotate API key,最後 Codex GPT-5.5 查 session log 才揪出真因只是 worktree race。