=== p1 === Natural Language Processing Introduction === p2 === Outline NLP Applications General Text Processing ◦Indexing ◦Lexical processing ◦Document representation Latent semantic analysis Word Embedding 2 === p3 === 3 “Knowledgeable Machine Learning for Natural Language Processing”, H. Han, Communications of the ACM, 2021. === p4 === What is Natural Language Processing (NLP, 自然語言處理)? Natural language processing (NLP) is the use of human languages by a computer and is viewed as a branch of machine learning. 4 Translation Information Retrieval Chatbot Information Verification === p5 === An example 5 Indexing / Search Classification Event detection News influence Summarization Article relationship Sentiment Unsupervised Training Corpus === p6 === 6 [IMAGE-ONLY] === p7 === 7 [IMAGE-ONLY] === p8 === How? Splitting text into sentences. Find the part-of-speech for words inside a sentence. Determine different types of subclauses. Determine the subject and the direct object of a sentence. Easy for human, hard for computer 8 https://aclanthology.org/W09-3410.pdf, ACL-IJCNLP, 2009. === p9 === Word sense ambiguity 9 Watch for kids ? [IMAGE-ONLY] === p10 === Context Ambiguity 10 [IMAGE-ONLY] === p11 === How teach computer to understand this? Q:曾有一項調查發現,很多員工生病的時候不敢請假,因為他們擔心老闆 會不高興,覺得他們沒有責任感。有人認為,員工會這麼想是公司的責任。 一個好的公司應該能照顧員工,而不是讓他們拿健康去換錢。因此,讓員工 有幸福感,應該是未來企業努力的方向。 這篇文章說了什麼內容? 1. 老闆應該給員工多一點兒假 2. 常關心別人的人更有責任感 3. 對公司有意見要勇敢說出來 4. 照顧身體比認真工作更重要 11 科技大擂台, 2017 === p12 === English is difficult I saw a man on a hill with a telescope.  There’s a man on a hill, and I’m watching him with my telescope. (山上有一個人,我正用望遠鏡觀察他。) There’s a man on a hill, who I’m seeing, and he has a telescope. (我看到山上有一個人,他有一架望遠鏡。) There’s a man, and he’s on a hill that also has a telescope on it. (有一個男人,他在一座山上,山上也有一架望遠鏡。) I’m on a hill, and I saw a man using a telescope. (我在一座山上,看到一個男人正在使用望遠鏡。) There’s a man on a hill, and I’m sawing him with a telescope. (山上有一個男人,我用望遠鏡觀察他。) 12 (ROOT (S (NP (PRP I)) (VP (VBD saw) (NP (DT a) (NN man)) (PP (IN on) (NP (DT a) (NN hill))) (PP (IN with) (NP (DT a) (NN telescope)))) (. .))) 我用望遠鏡在山上看到一個男人 === p13 === syntactic ambiguity 13 [IMAGE-ONLY] === p14 === 14 [IMAGE-ONLY] === p15 === Is reasoning difficult for LLM? 大舅去二舅家找三舅說四舅被五舅騙去六舅家偷七舅放在八舅櫃子裡九舅借十舅發給十一 舅的1000元薪水。 ◦問:1.究竟誰是小偷?2.錢本來是誰的? 15 1 2 3 === p16 === Is reasoning difficult for LLM? 16 4 5 [IMAGE-ONLY] === p17 === 情境理解 冬天:能穿多少穿多少;夏天:能穿多少穿多少。 17 [IMAGE-ONLY] === p18 === 《季姬擊雞記》 18 幾乎所有字都是聲母「j」、韻母「i」的音(拼音如ji, ji, ji...) 字義不同,但發音非常類似,靠上下文推敲邏輯 季姬寂,雞棲笈。笈既棄,姬擊雞。雞既棄,姬寂,既擊既棄。 趙元任, Yuen Ren Chao === p19 === more 19 [IMAGE-ONLY] === p20 === 20 NLP Based Applications (before LLM) UGC based predication e.g., stock, voting, … Representation Semantics Understanding Event/Trend Detection Event alarm… Breaking News Detection 重要新聞偵測 Language Translation Knowledge Base Construction IBM Watson Sentiment Mining Depression detection… Event Threading News recommendation… Summarization Search Engine Information Extraction Price comparisons… Hot keyword Extraction Deep Web Mining Auto ticketing… Siri, 小冰 Medical Suggestion Fake News Detection / Conclusion Opinion / Sentiment Mining Competitor product opinion comparison… === p21 === Text mining in Twitter 21 [IMAGE-ONLY] === p22 === Flu Prediction from Twitter 22 [IMAGE-ONLY] === p23 === More applications Japan FOODMOOD 23 2. ramen 1. curry 4.sushi === p24 === Powerful application Facbook PNAS paper ◦Private traits and attributes are predictable from digital records of human behavior ◦http://www.pnas.org/content/early/2013/03/06/1218772110 24 === p25 === A medical text mining example Research objective: ◦Follow chains of causal implication to discover a relationship between migraines and biochemical levels. Data: ◦medical research papers, medical news (unstructured text information) Key concept types: ◦symptoms, drugs, diseases, chemicals… 25 === p26 === Example Application: Medical Research stress is associated with migraines stress can lead to loss of magnesium calcium channel blockers prevent some migraines magnesium is a natural calcium channel blocker spreading cortical depression (SCD) is implicated in some migraines high levels of magnesium inhibit SCD migraine patients have high platelet aggregability magnesium can suppress platelet aggregability (source: Swanson and Smalheiser, 1994) 26 === p27 === Unstructured Data Management Natural language processing is a subset of Unstructured Data Management. UDM can be broken down into: ◦Content and Document Management ◦Search and Retrieval ◦XML database and tools ◦Categorization, Classification, and Visualization 80%-90% of Data is Unstructured 27 === p28 === Data Warehouse vs. Document Warehouse Data warehouse ◦Who, what, when, where, how much ◦Internally focused ◦Operational information ◦Rarely include external information 28 Document warehouse ◦Why ◦May not be internally focused ◦May contain a range of information ◦Often integrate external information === p29 === NLP Levels 29 • Prefix / Suffix • Lemmatization / Stemming • Spelling Checking Morphology • Part-of-Speech Tagging • Syntax trees • Dependency Trees Syntax • Named Entity Recognition / Normalization • Relation Extraction • Word Sense Disambiguation Semantics • Co-reference resolution • Topic Segmentation • Summarization Pragmatics “Advancing Multi- Criteria Chinese Word Segmentation Through Criterion Classification and Denoising”, ACL 2023 === p30 === The Building Blocks of Language Morphology ◦the study of the structure and form of words Syntax ◦the study of how words and phrases form sentences Semantics ◦relates to the meaning of words and statements Phonology ◦the study of sounds in the language Pragmatics ◦the study of idiomatic phrases that cannot be analyzed with strict semantic analysis 30 === p31 === Morphology Understanding words ◦Stems ◦Affixes ◦Prefix ◦Suffix ◦Inflectional elements 31 ▪ Reducing complexity of analysis ▪ Reduces complexity of representation ▪ Supports text mining Noun Prefix Noun Stem Suffix - able dispute in - === p32 === Morphology level representation Traditional: Words are made of morphemes ◦“unfortunately” = un + fortunate + ly (prefix + stem + suffix) In DL (Luong, 2013) ◦every morpheme is a vector ◦A neural network combines two vectors into one vector Extend to word, parsing, semantics level representations 32 === p33 === Syntax Ex: The Bank of Canada will curb inflation with higher interest rates. 33 Prepositional phrase Adjective Sentence Noun phrase Verb phrase Noun Verb Aux Noun phrase Noun Adjective Noun The Bank of Canada inflation curb will Interest rates higher with 中研院平衡語料庫詞類標記集 === p34 === Problems with NLP Limitations of Natural Language Processing ◦Correctly identifying the role of noun phrases ◦Representing abstract concepts ◦Classifying synonyms ◦Representing the number of concepts 34 === p35 === Underlying Technology is Based on Linguistics The Linguistic Approach: ◦Does not treat a document as a bag of words ◦Removes ambiguity by extracting structured concepts Concepts are the DNA of text. 35 Text is unstructured, ambiguous, and language dependent. === p36 === What is NLP? Natural language processing is a field at the intersection of ◦computer science ◦artificial intelligence ◦and linguistics. Fully understanding and representing the meaning of language (or even defining it) is a difficult goal. Perfect language understanding is AI-complete. (Christopher Manning) 36 Christopher @ACL2023 === p37 === AI-enhanced Natural Language Processing 37 Data Preprocessing ● Representation Learning ○ Word2Vec. GloVe. ELMo. BERT... ● Named Entity Recognition and Normalization ● ... Model Building ● Seq2Seq ● AutoEncoders ● Attention Mechanism ● Reinforcement Learning ● Generative Adversarial Networks (GAN) ● ... Applications ● News/Document Summarization ● Toxic Comment Classification ● Rumor Detection ● Text/Report Generation ● Sentiment Analysis ● Latent Aspect Mining ● Dialogue System === p38 === First meet with text FROM THE VIEW OF INFORMATION RETRIEVAL 38 === p39 === Information Retrieval 39 Analyzing the textual content of individual Web pages (documents) ◦given user’s query ◦determine a maximally related subset of documents Retrieval ◦index a collection of documents (access efficiency) ◦rank documents by importance (accuracy) Categorization (classification) ◦assign a document to one or more categories ◦Man-made (Yahoo & Dmoz ) vs. automation === p40 === Indexing 40 Inverted index ◦effective for very large collections of documents ◦associates lexical items to their occurrences in the collection Terms  ◦lexical items: words or expressions Vocabulary V ◦the set of terms of interest LLM vocabulary size 32k ~ 256k (LLaMA 1&2 (32K), Mistral 7B (32K), GPT- 3 (50K), GPT-4 (128K), Qwen (152K)) Bigger = Better ? === p41 === Inverted Index 41 The simplest example ◦a dictionary ◦each key is a term   V ◦associated value b() points to a bucket (posting list) ◦a bucket is a list of pointers marking all occurrences of  in the text collection === p42 === Inverted Index 42 Bucket entries: 1. document identifier (DID) ◦ the ordinal number within the collection 2. separate entry for each occurrence of the term ◦ DID ◦ offset (in characters) of term’s occurrence within this document ◦ present a user with a short context ◦ E.g., google results ◦ enables vicinity queries === p43 === Lexical Processing 43 Performed prior to indexing or converting documents to vector representations ◦Tokenization ◦extraction of terms from a document ◦Text conflation and vocabulary reduction ◦Stemming ◦reducing words to their root forms ◦Porter algorithm, http://www.tartarus.org/~martin/PorterStemmer/ ◦Removing stop words ◦common words, such as articles, prepositions, non-informative adverbs ◦20-30% index size reduction 詞根詞幹 === p44 === Tokenization 44 Extraction of terms from a document ◦stripping out ◦administrative metadata ◦structural or formatting elements Example ◦removing HTML tags ◦removing punctuation and special characters ◦folding character case (e.g. all to lower case) === p45 === Stemming 45 Want to reduce all morphological variants of a word to a single index term ◦e.g. a document containing words like fish and fisher may not be retrieved by a query containing fishing (no fishing explicitly contained in the document) ◦But what for fishing rod Stemming - reduce words to their root form ◦e.g. fish – becomes a new index term Porter stemming algorithm (1980) ◦relies on a preconstructed suffix list with associated rules ◦e.g. if suffix=IZATION and prefix contains at least one vowel followed by a consonant, replace with suffix=IZE ◦BINARIZATION => BINARIZE === p46 === Stemming vs. Lemmatization 46 Stemming Lemmatization Method Rule-based Corpus + Syntactic Output Not always real word Real word Performance fast slow Usage Search engine, fast preprocessing Semantics understanding, QA, etc. studies studi study === p47 === Issues about stop words 47 What are stop words ? ◦A static set or dynamic one ? ◦Depend on your application ? How to remove stop words ? ◦Dictionary-based matching When to remove stop words ? ◦Avoid to broke the semantics To be or not to be! === p48 === Document Representation: Vector-space Model 48 Text documents are mapped to a high-dimensional vector space Each document d ◦represented as a sequence of terms (t) ◦d = ((1), (2), (3), …, (|d|)) 1, 2 and 3 are terms in document, x and x are document vectors Vector-space representations are sparse, |V| >> |d| (the number of distinct terms in any single document) === p49 === Term frequency (TF) A term that appears many times within a document is likely to be more important than a term that appears only once nij - Number of occurrences of a term j in a document di Term frequency i ij d n TFij = 49 === p50 === Inverse document frequency (IDF) A term that occurs in a few documents is likely to be a better discriminator than a term that appears in most or all documents nj - Number of documents which contain the term j n - total number of documents in the set Inverse document frequency j j n n IDF log = 50 absolute measure of term importance A theoretical justification of IDF, Papineni 2001 === p51 === Document Similarity 51 Ranks documents by measuring the similarity between each document and the query Similarity between two documents d and d is a function s(d, d) R In a vector-space representation the cosine coefficient of two document vectors is a measure of similarity ' ' ') , cos( x x x x x x T  = O (|d| + |d’|) Cosine similarity vs. dot product === p52 === tf-idf weighting has many variants 52 Columns headed ‘n’ are acronyms for weight schemes. Many search engines allow for different weightings for queries v.s. documents ltn.lnc ? === p53 === BM25 (Best Matching 25) Okapi BM25 (1980s) ◦A ranking function used by search engines to rank matching documents according to their relevance to a given query. BM11, BM15 k1, b are free parameters generally k1=2, b=0.75 Alternative (improved) TF-IDF 𝑇𝐹𝐵𝑀25 ≈𝑓∗(𝑘1+1) 𝑓+ 𝑘1 f from 1 to 2 v.s. f from 100 to 101 === p54 === Word Representation -- Motivating Continuous, Distributed Representations 54 === p55 === Bag of Words 錢, 不是問題 不, 錢是問題 [IMAGE-ONLY] === p56 === Example: Opinion Mining 今天的牛排很讚又便宜 這家餐廳的前菜很有名 牛小排煮的有點太老 肉質很嫩 難吃又貴 價格太貴服務又差 便宜 讚 有名 嫩 老難吃 太貴 差 肉新鮮且價位低? === p57 === One-hot encoding 便宜 有名 讚 嫩 難吃 太貴 差 老 便宜 1 0 0 0 0 0 0 0 有名 0 1 0 0 0 0 0 0 讚 0 0 1 0 0 0 0 0 嫩 0 0 0 1 0 0 0 0 難吃 0 0 0 0 1 0 0 0 太貴 0 0 0 0 0 1 0 0 差 0 0 0 0 0 0 1 0 老 0 0 0 0 0 0 0 1 Vocabulary space Term vector = (1,1,1,1,0,0,0,0) 便宜 讚 有名 嫩 (0,1,0,0,0,0,0,0) (1,0,0,0,0,0,0,0) (0,0,1,0,0,0,0,0) (0,0,0,1,0,0,0,0) === p58 === Word embedding (dense space) d1 d2 d3 d4 d5 d6 d7 d8 便宜 1 3 2 3 0 -2 2 0 有名 3 1 4 2 0 2 0 1 讚 0 0 1 0 0 1 -1 0 嫩 1 3 0 1 0 0 0 0 難吃 0 0 0 0 1 0 2 0 太貴 0 0 0 0 3 1 0 1 差 -1 0 1 -1 0 2 1 0 老 0 -2 -1 0 1 3 2 1 Concept space Term vector === p59 === 59 Feature space Data instance [IMAGE-ONLY] === p60 === Synonymy and Polysemy 60 Synonymy (同義詞) ◦the same concept can be expressed using different sets of terms ◦e.g. bandit, brigand, thief ◦negatively affects recall ? Polysemy (一詞多義) ◦identical terms can be used in very different semantic contexts ◦e.g. bank ◦repository where important material is saved ◦the slope beside a body of water ◦negatively affects precision ? Hybrid cases === p61 === Concept matching v.s. term matching 61 A query may conceptually be very close to a given set of documents, but corresponding consine similarity is small, because … ◦Using different language ◦(計程車, 出租车, 的士) ◦Using different domain lexicons ◦(智慧型行動運算裝置, 手機), (AI伺服器, 運算伺服器) ◦(高效能溫度調節器, 冰箱), (互動多媒體視訊裝置, 80吋電視) ◦A small fraction of the entire dictionary for present terms ◦(COVID-19, 新冠病毒, SARS-CoV-2病毒) === p62 === How do we represent the meaning of a word? Dictionary based solutions WordNet: a resource containing list of synonyms sets and hypernyms (“is a” relationships) 62 Cited from Stanford cs224n === p63 === How do we represent the meaning of a word? Issues ◦context is not considered ◦New words missing ◦Maintenance cost 63 Synset A honorable estimable respectable good Synset B moral honourable ethical honorable Synset C dishonorable inglorious shameful disgraceful, The same string with different sense Synonyms Antonyms Hard to measure the similarity! -- one-hot vector -- every different words are orthogonal === p64 === Solution: Continuous, Distributed Representations By continuous we mean real-valued Distributed means a vector in a space, like dog cat dog cat truck 64 === p65 === Breakthrough Paper: Bengio et al. 2003: A Neural Probabilistic Language Model We can also pre-train our vectors with encoded world knowledge (e.g. similarity) http://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf 65 3 Gods @AAAI 2019 === p66 === 語言模型(Language Model) 66 Claude Shannon 1916 – 2001 Andrey Markov 1856 - 1922 [1913] The chance of a letter appearing depends on the letter before it. [1951] Prediction and Entropy of Printed English If a character was a vowel, the probability of the next being a consonant was roughly 87%; if it was a consonant, a vowel followed 66% of the time. Markov demonstrated that the next step depends only on your current state, not your entire past history. He coined Information Entropy: the more surprising or unpredictable a message is, the more information it carries. === p67 === Why Does This Work? 1. Similar words are expected to have similar vector representations (and be closer in vector space) 2. The probability function is a smooth function of feature values ○ Small changes in the feature values will induce a small change in the value of the function ○ This is not true for unorganized discrete spaces, where small changes in input can lead to large changes in the function value 3. Therefore, the presence of one sentence in the training set can effectively distribute probability density in the vector space for a combinatorial number of unseen sentences 67 === p68 === Learning Vector Representation of Words 68 === p69 === Representation Vector Usage 69 [IMAGE-ONLY] === p70 === Distributional Hypothesis In other words, similar words will appear in similar contexts The cat licked its fur The dog licked its fur No surprise “cat” and “dog” both appear near “lick” and “fur”. We should find they also often appear near “eat”, “run”, “bite” and so on… but not near words like “read”, “sophisticated”, or “wheel”. “A word is characterized by the company it keeps.” (Firth 1957) 70 A Synopsis of Linguistic Theory, 1930–1955 === p71 === Latent Semantic Analysis Using a corpus of text as input, draw a window of a defined length around each word and count co-occurrence statistics. The resulting matrix contains our word vectors. Count Co-Occurrence 71 “Indexing by latent semantic analysis”, S. Deerwester, 1990 Words with similar meanings tend to appear in similar textual contexts or co-occur across similar documents. === p72 === Simple example for word co-occurrence usage Training corpus ◦“I like deep learning.”, “I like NLP.”, “I enjoy programming." 72 cooccurence I like enjoy deep learning NLP programming I 0 2 1 0 0 0 0 like 2 0 0 1 0 1 0 enjoy 1 0 0 0 0 0 1 deep 0 1 0 0 1 0 0 learning 0 0 0 1 0 0 0 NLP 0 1 0 0 0 0 0 programming 0 0 1 0 0 0 0 === p73 === Simple example for word co-occurrence usage SVD on this co-occurrence matrix ◦Reduce dimension Use the 2 biggest singular value to represent words 73 like enjoy learning deep programming NLP I === p74 === Latent Semantic Analysis 74 Why need it? ◦serious problems for retrieval methods based on term matching ◦vector-space similarity approach works only if the terms of the query are explicitly presented in the relevant documents ◦有出現不見得是相關 ◦沒出現不見得是無關 ◦the rich expressive power of natural language ◦often queries contain terms that express concepts related to text to be retrieved TK Landauer “An Introduction to Latent Semantic Analysis” === p75 === Latent Semantic Indexing(LSI) 75 A statistical technique Uses a linear algebra technique called singular value decomposition (SVD) ◦attempts to estimate the hidden structure that generates terms given concepts ◦discovers the most important associative patterns between words and concepts Data driven ◦A large collection of sentences or documents is employed === p76 === LSI and Text Documents 76 Let X denote a term-document matrix X = [x1 . . . xn]T ◦each row is the vector-space representation of a document ◦each column contains occurrences of a term in each document in the dataset Latent semantic indexing ◦compute the SVD of X: ◦ - singular value matrix (diagonal matrix) ◦set to zero all but largest K singular values - ◦obtain the reconstruction of X by: ˆ T V U =  ˆ ˆ T V U =  === p77 === Genome news LSI Example 77 A collection of documents: d1: Indian government goes for open-source software d2: Debian 3.0 Woody released d3: Wine 2.0 released with fixes for Gentoo 1.4 and Debian 3.0 d4: gnuPOD released: iPOD on Linux… with GPLed software d5: Gentoo servers running at open-source mySQL database d6: Dolly the sheep not totally identical clone d7: DNA news: introduced low-cost human genome DNA chip d8: Malaria-parasite genome database on the Web d9: UK sets up genome bank to protect rare sheep breeds d10: Dolly’s DNA damaged Linux OS === p78 === LSI Example 78 The term-document matrix XT d1 d2 d3 d4 d5 d6 d7 d8 d9 d10 open-source 1 0 0 0 1 0 0 0 0 0 software 1 0 0 1 0 0 0 0 0 0 Linux 0 0 0 1 0 0 0 0 0 0 released 0 1 1 1 0 0 0 0 0 0 Debian 0 1 1 0 0 0 0 0 0 0 Gentoo 0 0 1 0 1 0 0 0 0 0 database 0 0 0 0 1 0 0 1 0 0 Dolly 0 0 0 0 0 1 0 0 0 1 sheep 0 0 0 0 0 1 0 0 0 0 genome 0 0 0 0 0 0 1 1 1 0 DNA 0 0 0 0 0 0 2 0 0 1 === p79 === LSI Example 79 The reconstructed term-document matrix after projecting on a subspace of dimension K=2  = diag(2.57, 2.49, 1.99, 1.9, 1.68, 1.53, 0.94, 0.66, 0.36, 0.10) = diag(2.57, 2.49, 0, 0, 0, 0, 0, 0, 0, 0) d1 d2 d3 d4 d5 d6 d7 d8 d9 d10 open-source 0.34 0.28 0.38 0.42 0.24 0.00 0.04 0.07 0.02 0.01 software 0.44 0.37 0.50 0.55 0.31 -0.01 -0.03 0.06 0.00 -0.02 Linux 0.44 0.37 0.50 0.55 0.31 -0.01 -0.03 0.06 0.00 -0.02 released 0.63 0.53 0.72 0.79 0.45 -0.01 -0.05 0.09 -0.00 -0.04 Debian 0.39 0.33 0.44 0.48 0.28 -0.01 -0.03 0.06 0.00 -0.02 Gentoo 0.36 0.30 0.41 0.45 0.26 0.00 0.03 0.07 0.02 0.01 database 0.17 0.14 0.19 0.21 0.14 0.04 0.25 0.11 0.09 0.12 Dolly -0.01 -0.01 -0.01 -0.02 0.03 0.08 0.45 0.13 0.14 0.21 sheep -0.00 -0.00 -0.00 -0.01 0.03 0.06 0.34 0.10 0.11 0.16 genome 0.02 0.01 0.02 0.01 0.10 0.19 1.11 0.34 0.36 0.53 DNA -0.03 -0.04 -0.04 -0.06 0.11 0.30 1.70 0.51 0.55 0.81 T ˆ ˆ Negative value! === p80 === Problems ●The original matrix still has the same dimension as our vocabulary size - potentially 100,000 x 100,000 ●The matrix is extremely sparse ●We can perform Singular Value Decomposition (SVD) for dimensionality reduction, however at a quadratic computational cost ●Adding a new word changes the entire matrix 80 === p81 === Probabilistic Model: Some Choices Unary language model Binary language model word2vec models (using window to get more context) Continuous Bag of Words (CBOW) Skip-Gram (which we will focus on) Ridiculous not to consider word order Better but still limited by short context distance The cat [center word] its fur [left context] licked [right context] 81 === p82 === Skip-Gram: Training Data Corpus Context window size = 1 Training Word Pairs What if context size > 1? ● More word pairs ● Only pass two words to the model at a time The cat licked its fur. The truck moved. Center Context the cat cat the cat licked licked cat licked its ... ... 82 === p83 === Model Parameters word2vec uses a single hidden layer feedforward neural network 1. Input to hidden layer matrix has our target word embeddings 2. Hidden to output layer matrix has another set of word embeddings called output embeddings http://mccormickml.com/2016/04/19/word2vec-tutorial-the-skip-gram-model/ 83 === p84 === Recap: Skip-Gram Ingredients: 1. A probabilistic model 2. Small chunks of data for training 3. Model parameters Iterative process: 1. Generate probability estimates 2. Calculate the error in the estimates 3. Update the parameters until optimized Training word pairs Neural network parameters V and U 84 === p85 === Generate Probability Estimates (1) Input vector selects input embedding from hidden layer matrix (2) Softmax over multiplication with output matrix creates probability distribution over the vocabulary This should be high for context words, low for others 85 Output weight for “ability” === p86 === Network Input We imagine the words are encoded as “one-hot” vectors (all zeros except a one at the the word index in the embedding matrix) the = cat = licked = etc... 86 === p87 === Embedding Lookup In practice we don’t do matrix multiplication, but just use the word index to pick out the vector. 87 V === p88 === Probability Over Vocab Our word vector does a dot product with every word’s output embedding, then applying softmax gives a probability distribution over the vocabulary 88 Output weight for cat === p89 === Calculate Error in Estimates 89 [IMAGE-ONLY] === p90 === Loss Function Cross entropy measures the distance between probability distributions * 90 === p91 === GloVe (Global Vectors) 2014 ●word2vec is excellent at learning similarity and linear regularities ●But it doesn’t make full use of co-occurrence statistics ●GloVe improves word2vec by considering global statistical information ●Pre-Trained vectors available here: https://nlp.stanford.edu/projects/glove/ ●The paper is worth reading 91 === p92 === FastText FastText is much newer and has also seen good results. Considers character-level features, and can thus take a guess and generate vectors for out-of-vocabulary words, taking advantage of patterns in morphology Pre-trained vectors are available in many languages: https://fasttext.cc/ 92 For Chinese:CW2Vec, Stroke-rich FastText === p93 === What Vectors Should I Use? --It depends. GloVe is very widely used with good results. FastText is new and adoption may be slow, but can also achieve good results. Training your own with word2vec may be indicated when you have many “out-of-vocabulary” words - i.e. your dictionary has lots of words for which there are no pre-trained vectors. This can happen in special areas - for example when dealing with classical Chinese text. A good idea is to try different vectors on specific tasks. Evaluation: Example dataset: WordSim353 http://www.cs.technion.ac.il/~gabr/resources/data/wordsim353/ 93 https://www.kaggle.com/datasets/julianschelb/wordsim353-crowd === p94 === 37 Things I Learned About Information Retrieval in Two Years at a Vector Database Company BM25 is a strong baseline for search MTEB (Massive Text Embedding Benchmark, https://huggingface.co/spaces/mteb/leaderboard) contextual embeddings (e.g., BERT), vs. static embeddings (e.g., Word2Vec, GloVe). 768 dimensions vs. 1536 dimensions Similar does not necessarily mean relevant. “How to fix a faucet” and “Where to buy a kitchen faucet” keyword-based search vs. vector-based search Vector search is not robust to typos Out-of-domain is not the same as out-of-vocabulary Information retrieval is so hot right now! 94 https://www.leoniemonigatti.com/blog/what_i_learned.html https://weaviate.io/ === p95 === Summary 95 SOP in traditional NLP Structural analysis, document representation Basic word understanding Word vector