Google NLP API parsing
Machine Learning

Google NLP API parsing

使用 Google Cloud Natural Language API 進行中英文語意分析:從 GCP API Key 設定、JSON 請求格式,到解讀 partOfSpeech、lemma 和 dependencyEdge 結果。

2020-09-04 · 2 min read · 712 words · KbWen · ZH
Before Data processing: ELT
Python

Before Data processing: ELT

ETL vs ELT: what each of the three steps does, why modern warehouses push the transform step to the target for performance, and where PAAS/SAAS fits in.

2020-09-02 · 2 min read · 253 words · KbWen · EN
Python Comments
Python

Python Comments

三引號字串只有放在模組、函式或類別的第一個陳述式才會變成 docstring 進到 __doc__,放在其他位置就是一段求值完丟掉的運算式,而 # 註解在語法階段就被忽略。用 __doc__ 和 dis 實際跑一次看兩者的差別。

2020-05-04 · 5 min read · 2055 words · KbWen · ZH
Python context manager
Python

Python context manager

學習 Python with 語句與 context manager:如何用 __enter__ 和 __exit__ 管理資源,確保文件、連線等資源被正確釋放,避免資源洩漏。

2020-04-14 · 1 min read · 277 words · KbWen · ZH
Python Iterable
Python

Python Iterable

只實作 __getitem__ 的物件,for 迴圈跑得動、iter() 也拿得到 iterator,但 isinstance(obj, Iterable) 會回 False。本文用跑得出來的程式碼示範這個落差,並說明 iterator 為什麼只能走一遍。

2020-04-11 · 5 min read · 2331 words · KbWen · ZH
python pdb
Python

python pdb

介紹 Python 內建除錯工具 pdb 的三種使用方式:命令列直接執行、設定斷點以及 Python shell 中使用 pdb.pm(),讓你告別 print 除錯法。

2020-04-10 · 1 min read · 396 words · KbWen · ZH
OpenCV 人臉偵測
Machine Learning

OpenCV 人臉偵測

OpenCV Haar Cascade 人臉偵測:AdaBoost 弱分類器級聯架構,scaleFactor 與 minNeighbors 到底在數什麼、為什麼要一起調,用 detectMultiScale2 讀出鄰居數量,以及這些 cascade 檔案在現在的 OpenCV 裡的狀態。

2017-07-12 · 5 min read · 2038 words · KbWen · ZH
k-NN 是什麼:手刻最近鄰,以及向量檢索為什麼改用近似搜尋
Machine Learning

k-NN 是什麼:手刻最近鄰,以及向量檢索為什麼改用近似搜尋

從 2017 年那支用 Scipy 手刻的 k-NN 出發:iris 對半切、for 迴圈掃過全部訓練資料,再看 KD tree 怎麼把 75 萬筆的精確查詢壓到毫秒等級,以及維度上到 768 之後樹狀索引為什麼失效、向量檢索為什麼改用 HNSW 這類近似索引。

2017-06-30 · 6 min read · 2557 words · KbWen · ZH
LSTM 是什麼:三個閘門怎麼運作
Machine Learning

LSTM 是什麼:三個閘門怎麼運作

LSTM 用遺忘門、輸入門、輸出門三道閘控制 cell state,讓資訊靠加法而不是連乘往前傳,避開一般 RNN 的梯度消失。這篇把當年只列了名字的閘門部分補完,並交代它在 Transformer 之後的位置。

2017-06-29 · 4 min read · 1561 words · KbWen · ZH
Kaggle Titanic 入門:資料裡的訊號在哪裡
Machine Learning

Kaggle Titanic 入門:資料裡的訊號在哪裡

Kaggle Titanic 解題紀錄:2017 年用 Logistic Regression 拿到 0.76555,加上回頭讀那支程式的筆記。preprocessing 先把 Age 正規化才跑 Age 小於 16 的規則,Sex 整欄變成同一個值;0.76555 換算回去是 418 題裡答對 320 題。

2017-06-09 · 7 min read · 3036 words · KbWen · ZH