停用詞
What?
停用詞(Stop Words)是在自然語言處理(NLP)中被認為信息量較少、可以忽略的常見詞彙。這些詞包括冠詞(a, the)、代詞(he, she)、介詞(in, on)和連接詞(and, but)等。在中文中,停用詞包括「的」、「是」、「在」、「被」等虛詞。
停用詞的移除是文本預處理中的重要步驟。儘管這些詞在自然語言中頻繁出現,但它們對於信息檢索和文本分類任務的區分能力有限。例如,在搜尋「最好的咖啡機」時,「最」、「的」等虛詞對搜尋意圖的貢獻遠小於「咖啡」和「機」。
停用詞列表不是通用的,不同領域和應用可能會有不同的定義。某些任務(如情感分析)可能需要保留某些停用詞,因為它們可能改變句子的含義(如 not)。
Who?
- 自然語言處理研究人員和工程師
- 搜尋引擎和信息檢索系統開發者
- 文本分類和Machine Learning從業者
- 語言學家和計算語言學家
When?
- 準備文本用於Machine Learning模型訓練時進行預處理
- 構建倒排索引(Inverted Index)以提高搜尋效率
- 進行文本相似度計算時去除無意義詞
- 提取關鍵詞和主題建模時提高結果質量
Where?
- 文本清理和規範化流程的早期階段
- Tokenization和詞頻(TF-IDF)計算前
- 信息檢索系統的索引構建階段
- Embedding和Word2Vec模型的訓練前預處理
Why?
- 降低維度:移除停用詞減少特徵空間,加快處理速度
- 提高信號質量:移除低信息量詞彙,使Machine Learning模型專注於有意義的特徵
- 改善效率:減少存儲和計算開銷,尤其在大規模文本處理時
How?
🛠️ 建立階段
# 使用 NLTK 的英文停用詞
from nltk.corpus import stopwords
import nltk
# 下載停用詞列表(首次需要)
nltk.download('stopwords')
# 獲取英文停用詞列表
stop_words_en = set(stopwords.words('english'))
print(f"英文停用詞數量: {len(stop_words_en)}")
print(f"示例: {list(stop_words_en)[:10]}")
# 中文停用詞(自定義或使用開源列表)
stop_words_zh = {
'的', '是', '在', '了', '和', '不', '有', '被', '我', '他',
'她', '它', '我們', '他們', '一個', '這', '那', '但是', '因為', '所以'
}
# 自定義停用詞(特定領域)
custom_stop_words = stop_words_en.copy()
custom_stop_words.add('http') # 移除 URL
custom_stop_words.add('click') # 移除特定詞彙
🔍 查詢階段
from nltk.tokenize import word_tokenize
import string
def remove_stopwords(text, stop_words):
# 分詞
tokens = word_tokenize(text.lower())
# 移除停用詞和標點
filtered_tokens = [
token for token in tokens
if token not in stop_words and token not in string.punctuation
]
return filtered_tokens
# 使用示例
text = "Machine learning is a subset of artificial intelligence"
filtered = remove_stopwords(text, stop_words_en)
print(f"原文詞數: {len(word_tokenize(text.lower()))}")
print(f"過濾後詞數: {len(filtered)}")
print(f"關鍵詞: {filtered}")
# 中文文本處理
import jieba
chinese_text = "機器學習是人工智能的一個重要分支"
stop_words_zh = {'是', '的', '一個', '在'}
# 使用 jieba 分詞
tokens = jieba.cut(chinese_text)
filtered_cn = [t for t in tokens if t not in stop_words_zh]
print(f"中文關鍵詞: {filtered_cn}")
補充說明
📌 範例比較
| 方法 | 優點 | 缺點 | 適用場景 |
|---|---|---|---|
| 預定義列表 | 快速、簡單 | 不夠靈活、可能遺漏 | 通用任務 |
| TF-IDF 篩選 | 自適應、基於數據 | 需要訓練集 | Machine Learning |
| 語言模型 | 上下文感知 | 計算開銷大 | 高精度需求 |
| 領域自定義 | 針對性強 | 需要領域知識 | 專業應用 |
🧠 延伸/常見誤解
- 誤解:應該總是移除停用詞。實際上,在某些任務如Sentiment Analysis(情感分析)中,「not」等詞至關重要,因為它們改變句子極性。
- 延伸:現代深度學習模型如 BERT 和 GPT 在預訓練時已經學習了停用詞的含義,因此在使用這些模型時通常不需要手動移除停用詞。