Running environment: Local — Windows 11 (10.0.26200), CPU: Intel Core i7-11800H (8 cores / 16 threads), RAM 16 GB Python version: 3.12.3 > 說明:以下所有數字都來自本筆記本的執行輸出,括號內標出是哪一格(例如〔cell 27〕)。比較實驗一律先對照第 5 題 (a) 量到的雜訊範圍;標「(推論)」的是根據原理的解釋,沒有另外用實驗驗證。 ##### 1. Which embedding model do you use? What are the pre-processing steps? What are the hyperparameter settings? (5%) Answer: **模型** - Part II:`glove-wiki-gigaword-100`(Stanford GloVe,Wikipedia+Gigaword 約 60 億字訓練,40 萬字、100 維、全部小寫)。 - Part III:Gensim `Word2Vec`(Skip-gram, `sg=1`),用 20% Wikipedia 自己訓練。 **抽樣**〔cell 24〕:以「篇」為單位,`random.seed(42)`,每篇擲一個亂數 `r`,`r < 0.2` 就留下,共 1,124,733 / 5,623,655 篇(20.00%)。同一個 `r` 同時產生 5%、10% 樣本,所以 5% ⊂ 10% ⊂ 20%,三份只差在資料量。直接串流讀取 11 個 `.gz` 檔,不解壓、不合併,記憶體裡只放一篇文章。 **前處理**〔cell 25〕 - 只保留 `[a-z]+` 的字。助教的資料已經是小寫、去標點、斷好字,所以這一步只刪掉 0.72% 的字(4,828,696 / 670,881,822)。 - 字典最多保留 30 萬個最常見的字(`max_final_vocab=300000`)。 - **刻意不移除停用詞**:考卷有 206 題含 NLTK 停用詞(family 86 題,如 he/she/his/her;currency 58 題,如 Korea won 的 won)〔cell 44〕,移除後這些題目必錯;老師上課也提到停用詞移除「以前很重要,但現在一點都不重要也不會去做」,而且可能刪掉重要的東西(W2 01:39:50;投影片 W1 p47 "Avoid to broke the semantics")。第 5 題 (b) 有實驗。 - **刻意不做詞形還原**:語法類考的就是字形變化(bigger、biggest、cars、danced),還原成原形會讓答案消失。 **超參數**:`vector_size=100`(與 GloVe-100 相同,方便比較)、`window=5`、`sg=1`、`negative=5`、`sample=1e-3`、`min_count=5`、`max_final_vocab=300000`、`epochs=5`、`workers=15`、`seed=42`。 - 注意:因為設了 `max_final_vocab`,gensim 會自動把門檻往上調。20% 維基的實際門檻是出現 **30** 次以上才收進字典(`effective min_count 30`);5%、10% 分別是 9、16〔cell 25, 38〕。 - `seed` 固定初始值,但多執行緒訓練的計算順序不固定,結果仍有小幅變動(第 5 題 (a))。 - 20% 維基處理後 666,053,126 字,訓練 37.9 分鐘,字典 299,405 字〔cell 25〕。 **答題方式**:3CosAdd,`most_similar(positive=[b, c], negative=[a])`(gensim 會排除題目的 a、b、c 三個字);題目先轉小寫;題目的字不在字典裡就算答錯(本次 0 題)。 ##### 2. What will the performance be like if you sample 5%, 10% and 20% of wiki text in TODO4? (10%, 3% for each) Answer: 〔cell 38〕 | | 5% | 10% | 20% | 5%→20% | 雜訊範圍* | |---|---|---|---|---|---| | 處理後字數 | 166,194,401 | 333,052,711 | 666,053,126 | | | | 實際字典門檻 | 9 | 16 | 30 | | | | 訓練時間 | 9.1 min | 19.7 min | 37.9 min | | | | **Overall** | 46.12 | 47.15 | **48.39** | +2.27 | 1.05 | | Semantic | 55.35 | 55.60 | 56.70 | +1.35 | 1.80 | | Syntactic | 38.45 | 40.13 | 41.48 | +3.03 | 1.26 | | family | 68.77 | 73.32 | 74.31 | +5.54 | 4.74 | | gram4-superlative | 17.02 | 22.37 | 24.33 | +7.31 | 2.76 | | gram2-opposite | 11.82 | 14.29 | 16.01 | +4.19 | 2.22 | | capital-common-countries | 81.82 | 76.48 | 77.87 | −3.95 | 2.77 | | OOV 題數 | 107 | 0 | 0 | | | \* 雜訊範圍=同樣用 5% 維基、只換亂數種子訓練 4 次,最高與最低的差(第 5 題 (a)〔cell 43〕)。差距的絕對值大於它,才當作真的有差。 觀察: - **整體隨資料量上升,但幅度小**:資料每多一倍(5→10→20%),整體約多 1 分(+1.03、+1.24),訓練時間卻跟著加倍。整體 +2.27 大於雜訊 1.05,是可信的上升。 - **語法類的上升明確,語意類不明確**:Syntactic +3.03(雜訊 1.26),最高級 +7.31、反義字首 +4.19 都超過雜訊;Semantic 只 +1.35,**小於雜訊 1.80,不能下結論**。(推論)biggest、worst、unaware 這類字形本來就少見,需要更多出現次數才學得穩;首都、國家這些常見字在 5% 時就已出現夠多次。 - **5% 有 107 題 OOV,10% 以後沒有**:5% 的資料裡有些考題的字出現次數低於實際門檻(9 次),沒有收進字典。 - capital-common-countries 下降 3.95,只略大於雜訊 2.77,而且這一類只有 506 題,不當作「資料多反而變差」的證據。 - 限制:雜訊是在 5% 量的,20% 的雜訊可能不同。 ##### 3. What is the performance for different categories or sub-categories when trained on different corpora? (15%) 3.1. Present your results. (5%) Answer: 語料:SemEval 2016/2017 Task 3 的論壇問答(Qatar Living 生活論壇)。為了把「資料量」與「內容種類」分開,另外從 5% 維基依序切出**同樣字數**(72,920,923 字,論壇是 72,919,300 字)的對照組;兩者用相同前處理與相同參數訓練〔cell 33, 34〕。 | | Wiki control | Forum | Forum − Wiki | |---|---|---|---| | **Overall** | **41.46** | 25.78 | −15.68 | | Semantic | 45.77 | 12.84 | −32.93 | | Syntactic | 37.87 | 36.53 | −1.34 | | capital-common-countries | 76.48 | 33.60 | −42.88 | | capital-world | 54.00 | 12.51 | −41.49 | | currency | 6.93 | 1.27 | −5.66 | | city-in-state | 33.81 | 2.92 | −30.89 | | family | 66.21 | 63.24 | −2.97 | | gram1-adjective-to-adverb | 10.28 | 13.10 | +2.82 | | gram2-opposite | 8.87 | 14.04 | +5.17 | | gram3-comparative | 53.08 | 64.94 | +11.86 | | gram4-superlative | 17.20 | 40.55 | +23.35 | | gram5-present-participle | 31.82 | 51.33 | +19.51 | | gram6-nationality-adjective | 76.55 | 21.70 | −54.85 | | gram7-past-tense | 40.96 | 33.78 | −7.18 | | gram8-plural | 36.26 | 41.97 | +5.71 | | gram9-plural-verbs | 32.99 | 41.49 | +8.50 | | OOV 題數 | 169 | 3,510 | | 答案覆蓋率(四個字都在字典裡的比例):capital-world 98.3% vs **48.6%**、currency 86.8% vs **39.0%**、city-in-state 100% vs 68.3%、family 100% vs 91.3%。 為了排除「論壇字典比較小、很多字查不到」的影響,另外只算**四個字在兩本字典都查得到**的 15,684 題〔cell 34〕:Overall 42.80 vs 32.12、Semantic 52.80 vs 22.10、Syntactic 37.91 vs 37.01;論壇仍在 superlative(42.90 vs 18.28)、present-participle(51.33 vs 31.82)勝出,在 nationality(22.81 vs 77.12)、past-tense(33.78 vs 40.96)落後。 3.2. Introduce the corpus you selected and explain what are the differences between the Wikipedia corpus and your corpus. (including data size, topic difference, structural difference … ) (5%) Answer: - **來源**:gensim-data 的 `semeval-2016-2017-task3-subtaskA-unannotated`(https://github.com/RaRe-Technologies/gensim-data ;SemEval-2016 Task 3, https://alt.qcri.org/semeval2016/task3/ )。內容是卡達 Qatar Living 論壇的發問與留言,共 189,941 個討論串、2,118,254 則貼文〔cell 33〕。只取標題、內文、留言三個欄位,刪除 HTML 標籤與網址後,用與維基相同的規則斷詞(小寫、2~15 個字母)。 - **資料量**:72,919,300 字,約是 20% 維基(666,053,126 字)的 1/9,所以另做同字數的維基對照組。字典大小:論壇 100,310 字、對照組 224,050 字〔cell 34〕。 - **主題**:維基涵蓋歷史、地理、人物、科學;論壇集中在簽證、工作、薪水、租屋、購物等在卡達生活的問題。 - **文體與結構**:維基是編輯校稿過的長篇說明文;論壇是短貼文、口語、問句多、錯字與縮寫多。例如論壇裡 salary 最像的字是 slary、salry、salery、sallary、salaray;visa 最像的是 rp、viza、iqama(卡達居留證)〔cell 35〕。 3.3. Explain why the performance increases or decreases. (5%) Answer: 在資料量相同的前提下,差異主要來自內容。下表是同一批字在兩份語料中每百萬字的出現次數〔cell 36〕: | 字 | Wiki control | Forum | Forum / Wiki | |---|---|---|---| | cheapest | 0.96 | 27.50 | 28.64 | | looking | 56.31 | 680.74 | 12.09 | | cheaper | 7.78 | 80.27 | 10.32 | | going | 104.22 | 914.05 | 8.77 | | better | 116.98 | 990.40 | 8.47 | | goes | 66.47 | 235.33 | 3.54 | | capital | 160.38 | 43.23 | 0.27 | | albanian | 13.37 | 0.75 | 0.06 | | illinois | 90.02 | 1.10 | 0.01 | - **語意類大幅落後**:論壇很少談世界首都、美國州名、國籍形容詞(capital 只有維基的 0.27 倍、illinois 0.01 倍)。很多答案根本不在論壇字典裡(capital-world 覆蓋率 48.6%);而且就算只看兩邊都查得到的題目,論壇的 Semantic 仍只有 22.10 vs 52.80,代表這些字即使出現,也很少出現在「首都/國家」的前後文中,關係學不起來。 - **比較級、最高級、-ing、第三人稱動詞反而勝出**:網友常寫 cheapest(28.64 倍)、looking(12.09 倍)、better(8.47 倍)、goes(3.54 倍)。(推論)這些字形在論壇中出現得更多、前後文更一致,所以字形關係學得比較好。另一個可能的因素是論壇字典只有對照組的一半大,干擾的候選字比較少;排除查不到的題目後論壇仍在 superlative、present-participle 勝出,但字典大小的影響沒有被完全排除。 - **錯字變成「最像的字」**:錯字與正確字的前後文高度相似,而詞向量是用前後文學出來的,所以 slary 成了 salary 的近鄰。這也呼應投影片 W1 p94 "Vector search is not robust to typos"。 - 結論:**語料的主題與文體決定詞向量學到哪些關係**——想讓某類關係學好,語料裡就要有大量該類的用法(投影片 W1 p93 "What vectors should I use? It depends.")。 - 限制:第 5 題 (a) 的雜訊是在 5% 維基量的,論壇與對照組沒有重跑,個位數的小類差距(例如 family −2.97)不下結論。 ##### 4. Select a few words and use their embeddings to retrieve the five most similar words. What do you observe? (10%) Answer: 挑了 7 個字,各有目的:king(考卷中的字)、cat(老師上課的例子)、apple 與 bank(一字多義)、good(看反義詞)、taiwan(專有名詞)、computer(一般名詞)〔cell 40, 41〕。 | 字 | GloVe | Wiki 20% | |---|---|---| | king | prince, queen, son, brother, monarch | nangklao, chlothar, queen, suryavarman, bodawpaya | | cat | dog, rabbit, cats, monkey, pet | rabbit, dog, sourpuss, mouse, pet | | apple | microsoft, ibm, intel, software, dell | blackberry, iphone, goldieblox, tvos, raspberry | | bank | banks, banking, credit, investment, financial | guaranty, savings, ameriprise, indymac, onewest | | good | better, sure, really, kind, very | sure, decent, lovely, bad, thankful | | taiwan | mainland, china, taiwanese, taipei, hong | taipei, china, guangdong, hainan, fujian | | computer | computers, software, technology, pc, hardware | computing, computers, software, mainframe, hardware | 觀察: 1. **冷門字效應**:Wiki 20% 的鄰居常是冷門字。king 的鄰居 nangklao(泰國國王)只出現 34 次、suryavarman 53 次、chlothar 76 次;bank 的 onewest 38 次、ameriprise 49 次〔cell 41〕,都接近字典門檻 30。(推論)一個字只出現幾十次、而且幾乎都在國王或銀行的上下文裡,它的向量就被拉到 king、bank 旁邊。改成只在最常見的 5 萬字裡找,king 的鄰居變成 queen, throne, prince, reigned, ruler;bank 變成 savings, bancorp, banking, lenders, jpmorgan,合理得多。GloVe 的鄰居也比較「一般」。 2. **反義詞很近**:good 的鄰居有 bad,similarity(good, bad) = 0.70。兩者出現在相同的前後文("a ___ movie"),詞向量分不出正反。 3. **一字多義被常見意思主導**:apple 前 15 名全是科技相關(blackberry, iphone, tvos, ipad, android…;raspberry 在這裡是 Raspberry Pi),similarity(apple, banana) = 0.41、(apple, fruit) = 0.37,遠低於 (apple, microsoft) = 0.65。bank 前 5 名都是金融義,但 similarity(bank, river) = 0.49 反而高於 (bank, money) = 0.40——河岸義並沒有消失,而是和金融義混在同一個向量裡。這是「一個字只有一個向量」的靜態詞向量的限制(投影片 W1 p60 正是用 bank 舉例;W2 p39 上下文詞向量才能解決)。 4. **鄰居反映語料內容**:cat 的鄰居 sourpuss,回查訓練語料發現是卡通角色("toon characters including … gandy goose sourpuss dinky duck …")〔cell 41〕——呼應老師說的「貓的鄰居不一定是貓」。 5. 兩個模型的 cosine 分數尺度不同,只比較各自的排名,不拿分數互相比大小。 ##### 5. … Anything that can strengthen your report. (5%) Answer: **(a) 雜訊有多大**〔cell 43〕:同樣用 5% 維基、只換亂數種子訓練 4 次(seed 42, 1, 2, 3),Overall 在 46.05~47.10 之間(範圍 1.05),Semantic 範圍 1.80、Syntactic 1.26,小類範圍 1.50~4.74(family 與 comparative 最大)。之後所有比較都用這個範圍判斷差距是否可信。 **(b) 移除停用詞**〔cell 44〕:同一份 5% 維基,只差在有沒有移除 NLTK 停用詞。 | | 保留 | 移除 | 差 | 超過雜訊? | |---|---|---|---|---| | Overall | 46.12 | 45.25 | −0.87 | 否 | | Semantic | 55.35 | 55.41 | +0.06 | 否 | | Syntactic | 38.45 | 36.81 | −1.65 | 是 | | family | 68.77 | 53.95 | −14.82 | 是 | | gram9-plural-verbs | 37.82 | 26.55 | −11.26 | 是 | | gram3-comparative | 44.14 | 37.01 | −7.13 | 是 | | gram8-plural | 38.59 | 41.89 | +3.30 | 是 | | OOV 題數 | 107 | 284 | | | 整體分數的變化在雜訊範圍內,但小類有明顯的此消彼長:family 大跌(he/she/his/her 被刪掉,題目必錯),plural-verbs 與 comparative 也下跌。(推論)停用詞(is、has、more、than…)是判斷字形的線索。只看整體分數會以為「刪不刪都沒差」。另外,`sample=1e-3` 本來就會隨機略過高頻字,效果和部分移除停用詞相似,可能也減小了整體差異。這支持不移除停用詞的決定。 **(c) 超參數(5% 維基,一次只改一個)**〔cell 45〕 | 設定 | Overall | Semantic | Syntactic | Overall 差 | |---|---|---|---|---| | 基準:Skip-gram, window 5, 100 維 | 46.12 | 55.35 | 38.45 | — | | CBOW (`sg=0`) | 52.58 | 58.60 | 47.58 | +6.46 | | `window=10` | 41.75 | 51.22 | 33.87 | −4.37 | | `vector_size=300` | 57.44 | 68.54 | 48.22 | +11.32 | 三個 Overall 差距都遠大於雜訊 1.05。 - **300 維**:只用 5% 資料就超過 100 維的 20%(48.39),city-in-state +25.78、capital-common-countries +13.64。在本設定下,向量維度的影響比資料量(5%→20%:+2.27)大。但只做了一次,而且老師提醒維度不是越大越好(「設一萬……效果不會比較好」,W3 01:22:17),也不和 GloVe-300 比較,所以不推論到其他設定。 - **CBOW**:語法類大幅領先(comparative +25.00、superlative +20.59)。老師上課以 Skip-gram 為主(W3 01:12:12),原論文也是 Skip-gram 語意類較好;本設定 CBOW 全面較好,是「這份資料、這組參數」下的結果。 - **window=10**:family −18.18、capital-common-countries −9.29。(推論)一般認為較大的視窗偏向「主題相關」而非「用法相似」;但語意類整體也下降,與這個說法不完全吻合,所以只當作可能的解釋。 **(d) Top-k 命中率**〔cell 46〕:Wiki 20% 的 top-1 / 3 / 5 / 10 為 48.39 / 61.53 / 66.32 / 71.74(GloVe:63.11 / 73.68 / 77.73 / 82.01)。正確答案常排在第 2~3 名,呼應老師說的 queen「可能在前三名」。 **(e) 錯題分析**〔cell 47〕:Wiki 20% 答錯 10,087 題,其中 0 題是 OOV,4,563 題(45.2%)的正確答案在第 2~10 名。錯法可以分成幾類: - 拼法變體:Argentina → argentinian(考卷答案 argentinean)、Belarus → belarusian(belorussian)、Slovakia → slovak(slovakian)——這些其實是考卷的拼法限制。 - 方向相反:large → smaller(larger)、long → shorter(longer)——反義詞在向量空間中太近。 - 同類選錯:Baghdad → syria(iraq)、Armenia → hryvnia(dram)、Cambodia → kyat(riel)。 - 近親字:father → stepmother(mother)、groom → bridesmaid(bride)。 - 被其他詞義帶走:great → britain(greater)、child → childbearing(children)。 **(f) 換算答案的公式**〔cell 48〕:同一份 Wiki 20% 向量改用 3CosMul(Levy & Goldberg, 2014, https://aclanthology.org/W14-1618/ ),Overall 48.39 → 44.66,14 個小類全部下降。原因未驗證;(推論)可能與字典中大量冷門字有關(見第 4 題觀察 1)。 **(g) PCA 與 t-SNE**〔cell 49〕:PCA 的縱軸大致區分男女(he, his, king, prince, brothers 在上;mom, grandma, aunt, mother, stepmother 在下),但並不乾淨(she, her, woman 也在上半);t-SNE 則把每一對(boy/girl, groom/bride, dad/mom, king/queen)貼得很近、分群清楚。PCA 是線性投影,較能保留整體方向;t-SNE 是非線性,主要保留局部鄰居。 **(h) 覆蓋率**〔cell 50〕:GloVe 與 Wiki 20% 在 14 個小類都是 100%,所以 Wiki 20% 的錯誤都不是「查不到字」造成的。 ##### 6. Generative AI Usage Answer: I used Claude (Anthropic) through Claude Code on my own computer for this assignment. - **Code**: All code in this notebook was generated by Claude — TODO1–TODO7, the data-loading changes, the code for Question 3, and all extra experiments below the report. Every AI-generated cell starts with a `# [Generative AI]` comment. - **Execution**: Claude ran the whole notebook on my computer (environment above). All numbers and figures in this report come from the outputs of this notebook. - **Report**: The report text, including the analysis and tables, was written by Claude based on these outputs. - **What I did**: ⟦WHAT_I_DID⟧