AI SystemsRead in English

Benchmark 飽和,其實是個驗證問題

GSM8k 99%、MMLU 90 出頭、HLE 在 2026 年中已進入 40 分檔。每出一份『更難的 benchmark』看起來都在解決問題,但結構性的事沒變:我們從來沒在驗證模型學會了什麼,只是在量它有沒有看過。

2026-06-01 · 6 min read · 2690 words · KbWen · ZH
AI Systems看中文版

LLM Benchmark Saturation Is a Verification Problem

GSM8k at 99%, MMLU at the 88-94% noise band, HLE already in the mid-40s by mid-2026. Each round of harder benchmarks looks like progress, but the field never solved the underlying problem: we measure correlation with a test distribution and call it capability.

2026-06-01 · 3 min read · 1486 words · KbWen · EN
Python List Comprehensions: Read Them as For-Loops
Python看中文版

Python List Comprehensions: Read Them as For-Loops

A relaxed take on Python list comprehensions: translate them back into the equivalent for-loop, and check what's actually true about variable leaking and speed on Python 3.14.

2026-05-31 · 9 min read · 1867 words · KbWen · EN
Python 列表推導式:一行取代 for 迴圈
Python輕鬆讀Read in English

Python 列表推導式:一行取代 for 迴圈

用比較白話的方式聊 Python 列表推導式:把它翻回普通的 for 迴圈來看,順便用 Python 3.14 實測一下變數外洩跟效能到底是怎樣。

2026-05-31 · 7 min read · 3348 words · KbWen · ZH
A three-stage evolution diagram: a small four-line atomic skill on the left, a cluster of overlapping skills in the middle (pattern emerging), and a taller seventeen-line production skill on the right, connected by dashed timeline arrows
AI Systems

The Skill Your Annoyed Prompt Becomes

Your first Claude Code skill won't look like the polished examples in tutorials. It'll look like a prompt you've typed three times in a row, saved into a four-line markdown file. This post walks that minimum shape, shows the three things that break, and compares it to a real seventeen-line production-grade skill from the framework I use daily.

2026-05-28 · 7 min read · 1478 words · KbWen · EN
三層演化圖:左邊一個 4 行 atomic skill 草稿,中間幾個 atomic skill 群聚,右邊一個成熟的 17 行 production skill,細線串成時間軸
AI Systems

怎麼寫你的第一個 skill — 從一個煩躁的 prompt 開始

你的第一個 skill 不會長得像書裡那些 production-grade 的成熟形態,它會長得像「你重複打三次的同一個 prompt」。從那裡開始,比從一個成熟框架的 skill 倒著學容易很多。

2026-05-28 · 7 min read · 3018 words · KbWen · ZH
A cutaway view of two files side by side: a small 13-line dispatcher on the left, a heavier protocol file on the right, connected by a thin line
AI Systems

What a 13-Line Skill Leaves Out

I asked Claude to draft me a skill that calls OpenAI's Codex CLI. It came back as thirteen lines of markdown. The thirteen lines aren't the skill — they point to where the skill actually lives. That split between dispatcher and contract is what separates a skill from a prompt.

2026-05-27 · 3 min read · 1112 words · KbWen · EN
兩份檔案的剖視對照:左邊是 13 行的 /codex-cli skill,右邊是它指向的、比較厚的 workflow
AI Systems

13 行的 skill:AI 起稿,我事後才看懂

我請 AI 幫我寫一個能從 Claude Code 呼叫 Codex CLI 的 skill,它給我 13 行 markdown。13 行很小,但 skill 跟 prompt 真正的差別不在這 13 行裡——在它指過去的那一份東西裡。

2026-05-27 · 4 min read · 1852 words · KbWen · ZH
MCP 資安危機:問題出在治理
AI Systems偏技術,給工程的Read in English

MCP 資安危機:問題出在治理

MCP(Model Context Protocol)一年內成為 AI 業界標準,2026 年卻接連爆出 RCE、tool poisoning、rug pull 等資安漏洞。本文整理多方專家觀點,並提出我的看法:真正要補的是治理這一層。

2026-05-25 · 7 min read · 3084 words · KbWen · ZH
MCP Security Is a Governance Problem
AI Systems看中文版

MCP Security Is a Governance Problem

MCP became the industry's default agent-to-tool interface in barely a year, then 2026 brought a wave of RCE, tool poisoning, and rug-pull disclosures. Weighing the expert debate, my take: the real exposure is a governance gap that better protocol design alone won't close.

2026-05-25 · 3 min read · 1447 words · KbWen · EN