跳至主要内容

停用詞

What?​

停用詞(Stop Words)是在自然語言處理(NLP)中被認為信息量較少、可以忽略的常見詞彙。這些詞包括冠詞(a, the)、代詞(he, she)、介詞(in, on)和連接詞(and, but)等。在中文中,停用詞包括「的」、「是」、「在」、「被」等虛詞。

停用詞的移除是文本預處理中的重要步驟。儘管這些詞在自然語言中頻繁出現,但它們對於信息檢索和文本分類任務的區分能力有限。例如,在搜尋「最好的咖啡機」時,「最」、「的」等虛詞對搜尋意圖的貢獻遠小於「咖啡」和「機」。

停用詞列表不是通用的,不同領域和應用可能會有不同的定義。某些任務(如情感分析)可能需要保留某些停用詞,因為它們可能改變句子的含義(如 not)。


Who?​

  • 自然語言處理研究人員和工程師
  • 搜尋引擎和信息檢索系統開發者
  • 文本分類和Machine Learning從業者
  • 語言學家和計算語言學家

When?​

  1. 準備文本用於Machine Learning模型訓練時進行預處理
  2. 構建倒排索引(Inverted Index)以提高搜尋效率
  3. 進行文本相似度計算時去除無意義詞
  4. 提取關鍵詞和主題建模時提高結果質量

Where?​

  1. 文本清理和規範化流程的早期階段
  2. Tokenization和詞頻(TF-IDF)計算前
  3. 信息檢索系統的索引構建階段
  4. Embedding和Word2Vec模型的訓練前預處理

Why?​

  • 降低維度:移除停用詞減少特徵空間,加快處理速度
  • 提高信號質量:移除低信息量詞彙,使Machine Learning模型專注於有意義的特徵
  • 改善效率:減少存儲和計算開銷,尤其在大規模文本處理時

How?​

🛠️ 建立階段​

# 使用 NLTK 的英文停用詞
from nltk.corpus import stopwords
import nltk

# 下載停用詞列表(首次需要)
nltk.download('stopwords')

# 獲取英文停用詞列表
stop_words_en = set(stopwords.words('english'))
print(f"英文停用詞數量: {len(stop_words_en)}")
print(f"示例: {list(stop_words_en)[:10]}")

# 中文停用詞(自定義或使用開源列表)
stop_words_zh = {
'的', '是', '在', '了', '和', '不', '有', '被', '我', '他',
'她', '它', '我們', '他們', '一個', '這', '那', '但是', '因為', '所以'
}

# 自定義停用詞(特定領域)
custom_stop_words = stop_words_en.copy()
custom_stop_words.add('http') # 移除 URL
custom_stop_words.add('click') # 移除特定詞彙

🔍 查詢階段​

from nltk.tokenize import word_tokenize
import string

def remove_stopwords(text, stop_words):
# 分詞
tokens = word_tokenize(text.lower())

# 移除停用詞和標點
filtered_tokens = [
token for token in tokens
if token not in stop_words and token not in string.punctuation
]

return filtered_tokens

# 使用示例
text = "Machine learning is a subset of artificial intelligence"
filtered = remove_stopwords(text, stop_words_en)
print(f"原文詞數: {len(word_tokenize(text.lower()))}")
print(f"過濾後詞數: {len(filtered)}")
print(f"關鍵詞: {filtered}")

# 中文文本處理
import jieba

chinese_text = "機器學習是人工智能的一個重要分支"
stop_words_zh = {'是', '的', '一個', '在'}

# 使用 jieba 分詞
tokens = jieba.cut(chinese_text)
filtered_cn = [t for t in tokens if t not in stop_words_zh]
print(f"中文關鍵詞: {filtered_cn}")

補充說明​

📌 範例比較​

方法優點缺點適用場景
預定義列表快速、簡單不夠靈活、可能遺漏通用任務
TF-IDF 篩選自適應、基於數據需要訓練集Machine Learning
語言模型上下文感知計算開銷大高精度需求
領域自定義針對性強需要領域知識專業應用

🧠 延伸/常見誤解​

  • 誤解:應該總是移除停用詞。實際上,在某些任務如Sentiment Analysis(情感分析)中,「not」等詞至關重要,因為它們改變句子極性。
  • 延伸:現代深度學習模型如 BERT 和 GPT 在預訓練時已經學習了停用詞的含義,因此在使用這些模型時通常不需要手動移除停用詞。