Running environment: Local — Windows 11 (10.0.26200), CPU: Intel Core i7-11800H (8 cores / 16 threads), RAM 16 GB Python version: 3.12.3 > 說明:以下所有數字都來自本筆記本的執行輸出,括號內標出是哪一格的輸出,用 Jupyter 畫面左邊的執行序號表示(例如〔In [23]〕)。比較實驗一律先對照第 5 題 (a) 量到的雜訊範圍;標「(推論)」的是根據原理的解釋,沒有另外用實驗驗證。 ##### 1. Which embedding model do you use? What are the pre-processing steps? What are the hyperparameter settings? (5%) Answer: **模型** - Part II:`glove-wiki-gigaword-100`(Stanford GloVe,Wikipedia+Gigaword 約 60 億字訓練,40 萬字、100 維、全部小寫;https://nlp.stanford.edu/projects/glove/ )。 - Part III:Gensim `Word2Vec`(Skip-gram, `sg=1`),用 20% Wikipedia 自己訓練。 **抽樣**〔In [20]〕:以「篇」為單位,`random.seed(42)`,每篇擲一個亂數 `r`,`r < 0.2` 就留下,共 1,124,733 / 5,623,655 篇(20.00%)。同一個 `r` 同時產生 5%、10% 樣本,所以 5% ⊂ 10% ⊂ 20%,三份只差在資料量。直接串流讀取 11 個 `.gz` 檔,不解壓、不合併,記憶體裡只放一篇文章。 **前處理**〔In [21]〕 - 只保留 `[a-z]+` 的字。助教的資料已經是小寫、去標點、斷好字,所以這一步只刪掉 0.72% 的字(4,828,696 / 670,881,822)。 - 字典最多保留 30 萬個最常見的字(`max_final_vocab=300000`)。 - **刻意不移除停用詞**:考卷有 206 題含 NLTK 停用詞(family 86 題,如 he/she/his/her;currency 58 題,如 Korea won 的 won)〔In [35]〕,移除後這些題目必錯;老師上課講搜尋引擎建索引時也提到,移除停用詞「以前很重要,但現在一點都不重要也不會去做」,而且可能刪掉重要的東西(W2 01:39:50;投影片 W1 p47 "Avoid to broke the semantics");本作業的實驗結果也支持不刪。第 5 題 (b) 有實驗。 - **刻意不做詞形還原**:語法類考的就是字形變化(bigger、biggest、cars、danced),還原成原形會讓答案消失。 **超參數**:`vector_size=100`(與 GloVe-100 相同,方便比較)、`window=5`、`sg=1`、`negative=5`、`sample=1e-3`、`min_count=5`、`max_final_vocab=300000`、`epochs=5`、`workers=15`、`seed=42`。 - 注意:因為設了 `max_final_vocab`,gensim 會自動把門檻往上調。20% 維基的實際門檻是出現 **30** 次以上才收進字典(`effective min_count 30`);5%、10% 分別是 9、16〔In [21], In [31]〕。 - `seed` 固定初始值,但多執行緒訓練的計算順序不固定,結果仍有小幅變動(第 5 題 (a))。 - 20% 維基處理後 666,053,126 字,訓練 37.9 分鐘,字典 299,405 字〔In [21]〕。 **答題方式**:3CosAdd,`most_similar(positive=[b, c], negative=[a])`(gensim 會排除題目的 a、b、c 三個字);題目先轉小寫;題目的字不在字典裡就算答錯(本次 0 題)。 ##### 2. What will the performance be like if you sample 5%, 10% and 20% of wiki text in TODO4? (10%, 3% for each) Answer: 〔In [31]〕 | | 5% | 10% | 20% | 5%→20% | 雜訊範圍* | |---|---|---|---|---|---| | 處理後字數 | 166,194,401 | 333,052,711 | 666,053,126 | | | | 實際字典門檻 | 9 | 16 | 30 | | | | 訓練時間 | 9.1 min | 19.7 min | 37.9 min | | | | **Overall** | 46.12 | 47.15 | **48.39** | +2.27 | 1.05 | | Semantic | 55.35 | 55.60 | 56.70 | +1.35 | 1.80 | | Syntactic | 38.45 | 40.13 | 41.48 | +3.03 | 1.26 | | family | 68.77 | 73.32 | 74.31 | +5.54 | 4.74 | | gram4-superlative | 17.02 | 22.37 | 24.33 | +7.31 | 2.76 | | gram2-opposite | 11.82 | 14.29 | 16.01 | +4.19 | 2.22 | | capital-common-countries | 81.82 | 76.48 | 77.87 | −3.95 | 2.77 | | OOV 題數 | 107 | 0 | 0 | | | \* 雜訊範圍=同樣用 5% 維基、只換亂數種子訓練 4 次,最高與最低的差(第 5 題 (a)〔In [34]〕)。差距的絕對值大於它,才當作真的有差。 觀察: - **整體隨資料量上升,但幅度小**:資料每多一倍(5→10→20%),整體約多 1 分(+1.03、+1.24),訓練時間卻跟著加倍。整體 +2.27 大於雜訊 1.05,是可信的上升。 - **語法類的上升明確,語意類加總不明確**:Syntactic +3.03(雜訊 1.26),最高級 +7.31、反義字首 +4.19 都超過雜訊;Semantic 加總只 +1.35,**小於雜訊 1.80,不能下結論**。但語意類內部有升有降:capital-world +2.98(雜訊 2.39)、currency +2.66(雜訊 1.50)超過雜訊,capital-common-countries 卻 −3.95,互相抵消。(推論)biggest、worst、unaware 這類字形比較少見,需要更多出現次數才學得穩。 - **5% 有 107 題 OOV,10% 以後沒有**:5% 的資料裡有些考題的字出現次數低於實際門檻(9 次),沒有收進字典。 - capital-common-countries 下降 3.95,只略大於雜訊 2.77,而且這一類題數較少,不當作「資料多反而變差」的證據。 - 限制:雜訊是在 5% 量的,20% 的雜訊可能不同。 ##### 3. What is the performance for different categories or sub-categories when trained on different corpora? (15%) 3.1. Present your results. (5%) Answer: 語料:SemEval 2016/2017 Task 3 的論壇問答(Qatar Living 生活論壇)。為了把「資料量」與「內容種類」分開,另外從 5% 維基依序切出**同樣字數**(72,920,923 字,論壇是 72,919,300 字)的對照組;兩者用相同前處理與相同參數訓練〔In [27], In [28]〕。 | | Wiki control | Forum | Forum − Wiki | |---|---|---|---| | **Overall** | **41.46** | 25.78 | −15.68 | | Semantic | 45.77 | 12.84 | −32.93 | | Syntactic | 37.87 | 36.53 | −1.34 | | capital-common-countries | 76.48 | 33.60 | −42.88 | | capital-world | 54.00 | 12.51 | −41.49 | | currency | 6.93 | 1.27 | −5.66 | | city-in-state | 33.81 | 2.92 | −30.89 | | family | 66.21 | 63.24 | −2.97 | | gram1-adjective-to-adverb | 10.28 | 13.10 | +2.82 | | gram2-opposite | 8.87 | 14.04 | +5.17 | | gram3-comparative | 53.08 | 64.94 | +11.86 | | gram4-superlative | 17.20 | 40.55 | +23.35 | | gram5-present-participle | 31.82 | 51.33 | +19.51 | | gram6-nationality-adjective | 76.55 | 21.70 | −54.85 | | gram7-past-tense | 40.96 | 33.78 | −7.18 | | gram8-plural | 36.26 | 41.97 | +5.71 | | gram9-plural-verbs | 32.99 | 41.49 | +8.50 | | OOV 題數 | 169 | 3,510 | | 答案覆蓋率(四個字都在字典裡的比例):capital-world 98.3% vs **48.6%**、currency 86.8% vs **39.0%**、city-in-state 100% vs 68.3%、family 100% vs 91.3%。 為了排除「論壇字典比較小、很多字查不到」的影響,另外只算**四個字在兩本字典都查得到**的 15,684 題〔In [28]〕:Overall 42.80 vs 32.12、Semantic 52.80 vs 22.10、Syntactic 37.91 vs 37.01;論壇仍在 superlative(42.90 vs 18.28)、present-participle(51.33 vs 31.82)勝出,在 nationality(22.81 vs 77.12)、past-tense(33.78 vs 40.96)落後。 判斷規則:論壇和對照組沒有重跑,所以借用第 5 題 (a) 在 5% 維基量到的各小類雜訊範圍,差距超過該小類的雜訊才判勝負(例如 family −2.97 小於雜訊 4.74,不判勝負;Syntactic −1.34 只略大於雜訊 1.26,視為差不多)。 3.2. Introduce the corpus you selected and explain what are the differences between the Wikipedia corpus and your corpus. (including data size, topic difference, structural difference … ) (5%) Answer: - **來源**:gensim-data 的 `semeval-2016-2017-task3-subtaskA-unannotated`(https://github.com/RaRe-Technologies/gensim-data ;SemEval-2016 Task 3, https://alt.qcri.org/semeval2016/task3/ )。內容是卡達 Qatar Living 論壇的發問與留言,共 189,941 個討論串〔In [27]〕。只取標題、內文、每則留言三種文字,刪除 HTML 標籤與網址後,用與維基相同的規則斷詞(小寫、2~15 個字母),並刪掉少於 3 個字的段落,最後留下 2,118,254 段文字。 - **資料量**:72,919,300 字,約是 20% 維基(666,053,126 字)的 1/9,所以另做同字數的維基對照組。字典大小:論壇 100,310 字、對照組 224,050 字〔In [28]〕。 - **主題**:維基涵蓋歷史、地理、人物、科學;論壇集中在簽證、工作、薪水、租屋、購物等在卡達生活的問題。 - **文體與結構**:維基是編輯校稿過的長篇說明文,20% 樣本平均每篇約 592 字(666,053,126 / 1,124,733);論壇是短貼文,平均每段約 34 字(72,919,300 / 2,118,254),口語、問句多、錯字與縮寫多。例如論壇裡 salary 最像的字是 slary、salry、salery、sallary、salaray;visa 最像的是 rp、viza、iqama(卡達居留證)〔In [29]〕。 3.3. Explain why the performance increases or decreases. (5%) Answer: 在資料量相同的前提下,差異主要來自內容。下表是同一批字在兩份語料中每百萬字的出現次數〔In [30]〕: | 字 | Wiki control | Forum | Forum / Wiki | |---|---|---|---| | cheapest | 0.96 | 27.50 | 28.64 | | biggest | 37.26 | 59.97 | 1.61 | | best | 553.01 | 988.65 | 1.79 | | looking | 56.31 | 680.74 | 12.09 | | cheaper | 7.78 | 80.27 | 10.32 | | going | 104.22 | 914.05 | 8.77 | | better | 116.98 | 990.40 | 8.47 | | goes | 66.47 | 235.33 | 3.54 | | capital | 160.38 | 43.23 | 0.27 | | brazilian | 42.43 | 33.46 | 0.79 | | albanian | 13.37 | 0.75 | 0.06 | | illinois | 90.02 | 1.10 | 0.01 | - **語意類大幅落後**:論壇很少談世界首都、美國州名、國籍形容詞(capital 只有維基的 0.27 倍、illinois 0.01 倍)。很多答案根本不在論壇字典裡(capital-world 覆蓋率 48.6%);而且就算只看兩邊都查得到的題目,論壇的 Semantic 仍只有 22.10 vs 52.80,代表這些字即使出現,也很少出現在「首都/國家」的前後文中,關係學不起來。 - **比較級、最高級、-ing、第三人稱動詞反而勝出**:網友常寫 cheapest(28.64 倍)、looking(12.09 倍)、better(8.47 倍)、goes(3.54 倍)。(推論)這些字形在論壇中出現得更多、前後文更一致,所以字形關係學得比較好。但頻率不能解釋全部:biggest 只多 1.61 倍、best 1.79 倍,最高級卻大贏 +23.35;brazilian 在論壇也有維基的 0.79 倍,國籍類卻大輸 −54.85。(推論)關鍵可能不只是單字出現次數,而是「國家–國籍」「原級–最高級」這種成對關係有沒有在同樣的前後文裡反覆出現。另一個可能的因素是論壇字典只有對照組的一半大,干擾的候選字比較少;排除查不到的題目後論壇仍在 superlative、present-participle 勝出,但字典大小的影響沒有被完全排除。 - **錯字變成「最像的字」**:錯字與正確字出現的前後文高度相似,而詞向量是用前後文學出來的,所以 slary 成了 salary 的近鄰。這也呼應投影片 W1 p94 "Vector search is not robust to typos"。 - 結論:**語料的主題與文體決定詞向量學到哪些關係**——想讓某類關係學好,語料裡就要有大量該類的用法(投影片 W1 p93 "What vectors should I use? It depends.")。 - 限制:第 5 題 (a) 的雜訊是在 5% 維基量的,論壇與對照組沒有重跑;差距小於該小類雜訊的(例如 family −2.97)不下結論。 ##### 4. Select a few words and use their embeddings to retrieve the five most similar words. What do you observe? (10%) Answer: 挑了 7 個字,各有目的:king(考卷中的字)、cat(老師上課的例子)、apple 與 bank(一字多義)、good(看反義詞)、taiwan(專有名詞)、computer(一般名詞)〔In [32], In [33]〕。 | 字 | GloVe | Wiki 20% | |---|---|---| | king | prince, queen, son, brother, monarch | nangklao, chlothar, queen, suryavarman, bodawpaya | | cat | dog, rabbit, cats, monkey, pet | rabbit, dog, sourpuss, mouse, pet | | apple | microsoft, ibm, intel, software, dell | blackberry, iphone, goldieblox, tvos, raspberry | | bank | banks, banking, credit, investment, financial | guaranty, savings, ameriprise, indymac, onewest | | good | better, sure, really, kind, very | sure, decent, lovely, bad, thankful | | taiwan | mainland, china, taiwanese, taipei, hong | taipei, china, guangdong, hainan, fujian | | computer | computers, software, technology, pc, hardware | computing, computers, software, mainframe, hardware | 觀察: 1. **冷門字效應**:Wiki 20% 的鄰居常是冷門字。king 的鄰居 nangklao(泰國國王)只出現 34 次、suryavarman 53 次、chlothar 76 次;bank 的 onewest 38 次、ameriprise 49 次〔In [33]〕——都只出現幾十次,在 30 萬字的字典裡屬於最冷門的一群(字典門檻是 30 次)。(推論)一個字只出現幾十次、而且幾乎都在國王或銀行的上下文裡,它的向量就被拉到 king、bank 旁邊。改成只在最常見的 5 萬字裡找,king 的鄰居變成 queen, throne, prince, reigned, ruler;bank 變成 savings, bancorp, banking, lenders, jpmorgan,合理得多。GloVe 的鄰居也比較「一般」。 2. **反義詞很近**:good 的鄰居有 bad,similarity(good, bad) = 0.70。兩者出現在相似的前後文("a ___ movie"),詞向量分不出正反。 3. **一字多義被常見意思主導**:apple 前 15 名全是科技相關(blackberry, iphone, tvos, ipad, android…;raspberry 在這裡是 Raspberry Pi),similarity(apple, banana) = 0.41、(apple, fruit) = 0.37,遠低於 (apple, microsoft) = 0.65。bank 前 5 名都是金融義,但 similarity(bank, river) = 0.49 反而高於 (bank, money) = 0.40。(推論)河岸義可能沒有消失,而是和金融義混在同一個向量裡;不過 money 是很泛用的字,這組比較只是旁證。這是「一個字只有一個向量」的靜態詞向量的限制(投影片 W1 p60 正是用 bank 舉例;W2 p39 上下文詞向量才能解決)。 4. **鄰居反映語料內容**:cat 的鄰居 sourpuss 只出現 64 次。回查訓練語料印出的前 5 處上下文〔In [33]〕,有 2 處是卡通角色名單("toon characters including … gandy goose sourpuss dinky duck …"),其餘是字源說明、電視集名、樂團名——它不是「貓」這個概念的近義詞,而是少量、特定的上下文把它拉到 cat 旁邊,呼應老師說的「貓的鄰居不一定是貓」。 5. **taiwan 與 computer**:taiwan 在兩個模型的鄰居都是中國的地名與省份(taipei、china、guangdong、fujian),反映語料中的地理與歷史脈絡;computer 的鄰居兩邊都是同類詞(computers、software、hardware),是 7 個字裡最「乖」的一組——一般名詞、用法單一時,詞向量的表現最符合直覺。 6. 兩個模型的 cosine 分數尺度不同,只比較各自的排名,不拿分數互相比大小。 ##### 5. … Anything that can strengthen your report. (5%) Answer: **(a) 雜訊有多大**〔In [34]〕:同樣用 5% 維基、只換亂數種子訓練 4 次(seed 42, 1, 2, 3),Overall 在 46.05~47.10 之間(範圍 1.05),Semantic 範圍 1.80、Syntactic 1.26,小類範圍 1.50~4.74(family、comparative、present-participle 最大,約 4.7)。之後所有比較都用這個範圍判斷差距是否可信。 **(b) 移除停用詞**〔In [35]〕:同一份 5% 維基,只差在有沒有移除 NLTK 停用詞。 | | 保留 | 移除 | 差 | 超過雜訊? | |---|---|---|---|---| | Overall | 46.12 | 45.25 | −0.87 | 否 | | Semantic | 55.35 | 55.41 | +0.06 | 否 | | Syntactic | 38.45 | 36.81 | −1.65 | 是 | | family | 68.77 | 53.95 | −14.82 | 是 | | gram9-plural-verbs | 37.82 | 26.55 | −11.26 | 是 | | gram3-comparative | 44.14 | 37.01 | −7.13 | 是 | | gram8-plural | 38.59 | 41.89 | +3.30 | 是 | | OOV 題數 | 107 | 284 | | | 整體分數的變化在雜訊範圍內,但小類有明顯的此消彼長:family 大跌(he/she/his/her 被刪掉,題目必錯),plural-verbs 與 comparative 也下跌。(推論)停用詞(is、has、more、than…)是判斷字形的線索。只看整體分數會以為「刪不刪都沒差」。另外,`sample=1e-3` 本來就會隨機略過高頻字,效果和部分移除停用詞相似,可能也減小了整體差異。這支持不移除停用詞的決定。 **(c) 超參數(5% 維基,一次只改一個)**〔In [36]〕 | 設定 | Overall | Semantic | Syntactic | Overall 差 | |---|---|---|---|---| | 基準:Skip-gram, window 5, 100 維 | 46.12 | 55.35 | 38.45 | — | | CBOW (`sg=0`) | 52.58 | 58.60 | 47.58 | +6.46 | | `window=10` | 41.75 | 51.22 | 33.87 | −4.37 | | `vector_size=300` | 57.44 | 68.54 | 48.22 | +11.32 | 三個 Overall 差距都遠大於雜訊 1.05。 - **300 維**:只用 5% 資料就超過 100 維的 20%(48.39),city-in-state +25.78、capital-common-countries +13.64。在本設定下,向量維度的影響比資料量(5%→20%:+2.27)大。但只做了一次,而且老師提醒維度不是越大越好(「設一萬……效果不會比較好」,W3 01:22:17),也不和 GloVe-300 比較,所以不推論到其他設定。 - **CBOW**:Overall、Semantic、Syntactic 都較好,語法類大幅領先(comparative +25.00、superlative +20.59);但 14 個小類裡有 3 類較差,其中 nationality-adjective −1.81 超過雜訊。老師上課以 Skip-gram 為主(W3 01:12:12),原論文(Mikolov et al., 2013, https://arxiv.org/abs/1301.3781 )也是 Skip-gram 語意類較好;這是「這份資料、這組參數」下的結果。 - **window=10**:family −18.18、capital-common-countries −9.29。(推論)一般認為較大的視窗偏向「主題相關」而非「用法相似」;但語意類整體也下降,與這個說法不完全吻合,所以只當作可能的解釋。 **(d) Top-k 命中率**〔In [37]〕:Wiki 20% 的 top-1 / 3 / 5 / 10 為 48.39 / 61.53 / 66.32 / 71.74(GloVe:63.11 / 73.68 / 77.73 / 82.01)。正確答案常排在第 2~3 名,呼應老師說的 queen「可能在前三名」。 **(e) 錯題分析**〔In [38]〕:Wiki 20% 答錯 10,087 題,其中 0 題是 OOV,4,563 題(45.2%)的正確答案在第 2~10 名。錯法可以分成幾類: - 拼法變體:Argentina → argentinian(考卷答案 argentinean)、Belarus → belarusian(belorussian)、Slovakia → slovak(slovakian)——這些其實是考卷的拼法限制。 - 方向相反:large → smaller(larger)、long → shorter(longer)——反義詞在向量空間中太近。 - 同類選錯:Baghdad → syria(iraq)、Armenia → hryvnia(dram)、Cambodia → kyat(riel)。 - 近親字:father → stepmother(mother)、groom → bridesmaid(bride)。 - 被其他詞義帶走:great → britain(greater)、child → childbearing(children)。 **(f) 換算答案的公式**〔In [39]〕:同一份 Wiki 20% 向量改用 3CosMul(Levy & Goldberg, 2014, https://aclanthology.org/W14-1618/ ),Overall 48.39 → 44.66,14 個小類全部下降。原因未驗證;(推論)可能與字典中大量冷門字有關(見第 4 題觀察 1)。 **(g) PCA 與 t-SNE**〔In [40]〕:PCA 的縱軸只有微弱的性別傾向:he、his、king 在最上,mom、grandma、aunt 在下,但 she、her、queen 也在上半,dad、grandpa、husband 在下半,分不乾淨;t-SNE 則把每一對(boy/girl, groom/bride, dad/mom, king/queen)貼得很近、分群清楚。PCA 是線性投影,較能保留整體方向;t-SNE 是非線性,主要保留局部鄰居。 **(h) 覆蓋率**〔In [41]〕:GloVe 與 Wiki 20% 在 14 個小類都是 100%,所以 Wiki 20% 的錯誤都不是「查不到字」造成的。 ##### 6. Generative AI Usage Answer: I used Claude (Anthropic) through Claude Code on my own computer for this assignment. - **Code**: All code added to the template was generated by Claude — TODO1–TODO7, the data-loading changes, the code for Question 3, all extra experiments below the report, and the section headings below the report. Every AI-generated code cell has a `# [Generative AI]` comment at the top. - **Experiment design**: The choice of the forum corpus, the same-size Wikipedia control, the noise measurement and the other extra experiments were proposed by Claude. - **Execution**: Claude ran the whole notebook on my computer (environment above). All numbers and figures in this report come from the outputs of this notebook. - **Report**: The report text, including the analysis and tables, was written by Claude based on these outputs. - **What I did**: ⟦WHAT_I_DID⟧