subagent 分派:第二個 agent 該拿到什麼
開一個 agent 去看另一個 agent 做完的東西很容易,難的是後面那個看得見前面漏掉了什麼。從 Agentic OS 的 review 規定看下去:要讓第二個 agent 看得出東西,靠的是刻意不給它 session、對話記錄和實作理由。
開一個 agent 去看另一個 agent 做完的東西很容易,難的是後面那個看得見前面漏掉了什麼。從 Agentic OS 的 review 規定看下去:要讓第二個 agent 看得出東西,靠的是刻意不給它 session、對話記錄和實作理由。
A transformer redoes the same attention projections for every past token at each decoding step. The KV cache stores those keys and values so they get reused instead of recomputed, and the one cost it adds is memory that grows with the sequence.
把 softmax 加交叉熵對 logit 的導數一路算出來,結果剛好是預測機率減去標籤。這篇從一個三類別的小例子走進這個梯度,看它為什麼乾淨、又為什麼信心錯得越離譜就修得越用力。
Moonshot's model card scores Kimi K3 against Claude Fable 5 and GPT-5.6 Sol across 45 benchmarks. Some rows go to K3 and some to the others, with margins running from fifteen points down to a tenth — plus what $3/$15 per million tokens buys.
Moonshot 七月發表的開源模型 Kimi K3,README 附了一張跟 Fable 5、GPT-5.6 Sol 的對照表。這篇挑六項分數來看,也整理了價格與輸出速度,談哪些任務可以交給它。
A decorator replaces your function with a wrapper, so its name, docstring, and signature change. Here is exactly what functools.wraps copies back and how it records __wrapped__.
Hugging Face's forensics were refused by the commercial APIs it tried first. The block is an access setting, so test what your account does with attack data.
同一串討論底下,有人量到 tokenization 不到總推論時間的 0.1%,也有人量到九成以上的 CPU 時間都花在這裡。這篇看這個落差怎麼來的:算的窗口不同、模型大小不同,還有一些工作根本沒有模型在裡面。
A worked walk through the three main LLM sampling knobs: temperature reshapes the whole next-token distribution, while top-k and top-p truncate which tokens you may sample from.
用 list 或 dict 當函式的預設參數值,資料會跨呼叫累積,因為預設值在 def 執行時就算好一次並掛在函式物件上。本文示範現象、用 __defaults__ 驗證,並給出 None 哨兵修法。