Machine Learning

Fitting preprocessing before the split inflates your accuracy

Fitting a preprocessing or feature-selection step on the whole dataset before the train/test split leaks the labels and inflates a model's estimated accuracy. A pure-noise scikit-learn run shows the gap, and the pipeline fix closes it.

2026-07-20 · 5 min read · 918 words · KbWen · EN
k-NN 是什麼:手刻最近鄰,以及向量檢索為什麼改用近似搜尋
Machine Learning

k-NN 是什麼:手刻最近鄰,以及向量檢索為什麼改用近似搜尋

從 2017 年那支用 Scipy 手刻的 k-NN 出發:iris 對半切、for 迴圈掃過全部訓練資料,再看 KD tree 怎麼把 75 萬筆的精確查詢壓到毫秒等級,以及維度上到 768 之後樹狀索引為什麼失效、向量檢索為什麼改用 HNSW 這類近似索引。

2017-06-30 · 6 min read · 2557 words · KbWen · ZH