Tag: model-comparison
All the articles with the tag "model-comparison".
-
A Few Days Into Switching From Claude Code to Codex: The Quota Honeymoon
A running log of switching from Claude Code to Codex this week: quota I could not burn through fast enough, a $20 Sol Ultra run that beat a $100 Fable plan on a professor friend's paper, and a tool-chain swap to Hyperframe.
-
從 Claude Code 換到 Codex 的這幾天:額度蜜月期
這週從 Claude Code 換到 Codex 的體感記錄:額度多到花不完、Sol Ultra 20 鎂幫教授跑論文,比 Fable 100 鎂方案還划算,工具鏈也跟著換成 Hyperframe。
-
Model Picks Beyond the Main One: Notes I Left at Different Times
Pulling together the model-picking notes I left on Threads at different times: Codex for complex bugs, Sonnet for documents, Haiku for full-stack, DeepSeek as the value alternative, Qwen on the sidelines, small models for local deployment, Gemini Skills wait-and-see.
-
Letting AI Write CAD From an Engineering Drawing: A Single Vision Model Can't Be Trusted to Read Topology
A hands-on lesson: let AI look at an engineering drawing and write CadQuery directly, and a single vision model will confidently get the topology wrong. Two more independent models and a 2:1 veto are what caught it.
-
主力之外,各模型的定位:我在不同時間留下的選型碎念
把我在 Threads 上不同時間留下的選型碎念整理成一篇:Codex 解複雜 bug、Sonnet 做文書、Haiku 做全端、DeepSeek 當平替、Qwen 觀望、小模型本地部署、Gemini Skill 觀望。
-
讓 AI 看工程圖寫 CAD:單一視覺模型裸寫拓樸不可信
一個實機教訓:讓 AI 看著工程圖直接寫 CadQuery,單一視覺模型會自信地把拓樸讀錯。加兩個獨立模型交叉、2:1 否決才抓出來。
-
Racing for the Fastest Summary — Opus 4.8's 244-Page System Card, and How I Read It With 20 Agents in Half an Hour
The night Opus 4.8 launched, I split the 244-page system card into 20 chunks, handed them to 20 gemini agents to summarize in parallel, and pieced together the fastest rundown online. Plus a digest of Reddit's hands-on reactions within two hours of launch.
-
拼全網最速——Opus 4.8 系統卡 244 頁重點,外加我怎麼用 20 個 agent 半小時讀完
Opus 4.8 發布當晚,我把 244 頁系統卡切成 20 份、丟給 20 個 gemini agent 並行摘要,拼出全網最速重點。附上發布兩小時內的 Reddit 實測整理。
-
Every AI Leader Starts Cutting Corners — The GPT, Claude, Gemini Cycle
GPT image generation degrading, Claude quietly shrinking rate limits, Gemini Flash hiking prices. Every model provider starts cutting corners once they reach the top. Annual subscriptions are the worst bet.
-
坐上第一就開始拿翹——GPT、Claude、Gemini 的降智循環
GPT 生圖降智、Claude 額度縮水、Gemini Flash 漲價。三家輪流坐莊,坐上去就開始偷料。按年訂閱是最傻的事,沒消息才是最好的消息。
-
Stop Dismissing Gemini — Four Use Cases Where Nothing Else Comes Close
Everyone seems to be dismissing Gemini now that Codex and Claude dominate the agent space. But Gemini has four use cases other models cannot match: Flash Lite cost efficiency, audio multimodal, video understanding, and book scanning OCR.
-
別一味貶低 Gemini——四個其他家打不過的 Use Case
最近 Codex 跟 Claude 搶盡風頭,Gemini 好像被嫌棄了。但 Gemini 有四個其他模型打不過的場景:Flash Lite 性價比、音訊多模態、影片理解、書籍掃描 OCR。
-
Gemini 3.5 Flash Reddit Reviews — 3x Price, Vision Regression, Tool Calling Disaster
Reddit user reviews after Gemini 3.5 Flash launch: 3x price increase over 3 Flash, vision regression, tool calling running 32 calls before forced stop. Speed is genuinely fast, but overall reception skews negative.
-
Gemini 3.5 Flash Reddit 實測彙整——貴三倍、Vision 退步、Tool Calling 災難
Gemini 3.5 Flash 上線後 Reddit 用戶實測回報彙整:價格比 3 Flash 貴三倍、Vision 退步、Tool Calling 跑到 32 次被中斷。正面是速度快且程式碼風格好,但整體評價偏負面。
-
2026 Model Personality Watch: Gemini, Claude, Codex Compared
A year in, the three flagships have developed very visible "personalities" — Gemini 3 is the dramatic PhD, Claude 4.7 is the slick veteran, and GPT-5.5 turns out to be the most pragmatic colleague of the wave. Plus a fun trick for guessing the version from "sass density."
-
2026 模型脾氣觀察:Gemini、Claude、Codex 的個性對比
用了一整年下來,三家的旗艦模型各自有很明顯的「脾氣」——Gemini 3 像戲精博士、Claude 4.7 像油條前輩、GPT-5.5 反而是這波最務實的同事。把累積的觀察整理成一篇對比,順便講一個從「貧嘴密度」反推版本號的玩法。
-
Opus 4.7 After One Week: From the System Card to the Roasting, Where Did It Go Wrong?
Opus 4.7 launched on 4/16 and the Chinese and English communities ended up with opposite takes. The system card reads like a full win, but real usage burns quota at roughly 2x what Anthropic claimed, and Reddit spent the week roasting it. Here's what I saw, and why I rolled back to 4.6.
-
Opus 4.7 上線一週:從系統卡到火烤文,到底哪邊出了問題
Opus 4.7 上線一週,系統卡數據看起來全面超越 4.6,但實測下來燒 quota 是官方宣稱的 2 倍,Reddit 一片火烤文。整理這週的災情、逆向分析、以及我自己退回 4.6 的路徑。