=== p1 === Natural Language Processing Word Embeddings and Language Modeling === p2 === GAI Motivation (for NLP tasks) Supervised learning ◦Text classification ◦QA system ◦… Issues ◦Lack of training data ◦Limitation of domain knowledge 2 Think about your learning strategies ! === p3 === Natural Language Generation The commonest way to generate sentences is by writing the words down, one after another. e.g. Please turn your homework … The next could be ? 3 in ? over ? === p4 === 語言模型(Language Model) 4 Claude Shannon 1916 – 2001 Andrey Markov 1856 - 1922 [1913] The chance of a letter appearing depends on the letter before it. [1951] Prediction and Entropy of Printed English === p5 === Language Model 5 [IMAGE-ONLY] === p6 === 6 童年的紙飛機現在終於飛回我 手裡 天青色等煙雨而我在 等妳 怎麼這樣子 雨還沒停妳就撐 傘要走 妳說這一句 很有夏天的感覺 我用幾行字形容妳是我 的誰 為妳翹課的那一天 花落的 那一天 消失的下雨天 === p7 === 7 童年的紙飛機現在終於飛回我 手裡 天青色等煙雨 而 在等 怎麼這樣子 雨還沒停妳 就撐 傘要走 說這一句 很有夏天的 感覺 用幾行字形容 翹課的那一天 花落 消失的下 天 為 是 的誰 P (b | a) P (c | ab) . . . P (説 | 妳) = 1/4 P (説 | 沒停妳) = 0 === p8 === N-grams An n-gram is a sequence of n words: e.g. Please turn your homework … 1-gram (unigram): "please", "turn", "your", or "homework" 2-gram (bigram): "please turn", "turn your", or "your homework" 3-gram (trigram): "please turn your", or "turn your homework" ... 8 === p9 === linguistic signature JK Rowling vs. Robert Galbraith Harry Potter vs. The Cuckoo’s Calling Text Analysis Software: Patrick Juola's JGAAP Stylometric Analysis: Sequences of four-character strings (four-grams), which are strong indicators of authorship. The frequency of the most common words. The distribution of word lengths. Pairs of words that frequently appeared together. 9 https://www.scientificamerican.com/article/how-a-computer-program-helped-show-jk-rowling-write-a-cuckoos-calling/ === p10 === N-gram Language Models We can use a naïve statistic method to model the language 10 I want to eat lunch. I want to eat Chinese food. I don't want to spend time cooking ... In a bi-gram model we have to count the occurrences of each bi- gram. e.g. C(I want) = 2 C(want to) = 3 C(spend time) = 1 ... === p11 === N-gram Language Models An example of bi-gram LM. Draw a table of bigram counts for eight of the words in all of the sentences 11 Word Embeddings and Language Modeling I want to eat Chinese food launch speed I 5 827 0 9 0 0 0 2 want 2 0 608 1 6 6 5 1 to 2 0 4 686 2 0 6 211 eat 0 0 2 0 16 2 42 0 Chinese 1 0 0 0 0 82 1 0 food 15 0 15 0 1 4 0 0 launch 2 0 0 0 0 1 0 0 speed 1 0 1 0 0 0 0 0 === p12 === N-gram Language Models Apply add-k smoothing (k=1): 12 Compute Probability (relative frequency): I want to eat Chinese food launch speed I 6 828 1 10 1 1 1 3 want 3 1 609 2 7 7 6 2 to 3 1 5 687 3 1 7 212 eat 1 1 3 1 17 3 43 1 Chinese 2 1 1 1 1 83 2 1 food 16 1 16 1 2 5 1 1 launch 3 1 1 1 1 2 1 1 speed 2 1 2 1 1 1 1 1 I want to eat Chinese food launch speed I 0.0015 0.21 0.00025 0.0025 0.00025 0.00025 0.00025 0.0075 want 0.0013 0.00042 0.26 0.0084 0.0029 0.0029 0.0025 0.0015 to 0.00078 0.00026 0.0013 0.18 0.00078 0.00026 0.0018 0.055 eat 0.00046 0.00046 0.0014 0.00046 0.0078 0.0014 0.02 0.00046 Chinese 0.0012 0.00062 0.00062 0.00062 0.00062 0.052 0.0012 0.00062 food 0.0063 0.00039 0.0063 0.00039 0.00079 0.002 0.00039 0.00039 launch 0.0017 0.0056 0.00056 0.0056 0.00056 0.0011 0.0056 0.00056 speed 0.0012 0.0058 0.0012 0.0058 0.0058 0.0058 0.0058 0.0058 Word Embeddings and Language Modeling === p13 === Probability of a Sequence Let's begin with the task of computing the probability e.g., Compute the probability of the word "the" given the history "its water is so transparent that". 13 w: the word to be generated​ h: some history C: the times the pattern show up in the dataset Word Embeddings and Language Modeling === p14 === Probability of a Sequence Compute probabilities of entire sequences like in N-gram model(Chain Rule of Probabilities) 14 e.g., Bi-gram model Word Embeddings and Language Modeling === p15 === Language Models Language Models (LMs) : Models that assign probabilities to sequences of words are called LMs. Including: N-Gram Language Models: A purely statistical model of language. Neural Language Models: Use neural networks to predict the likelihood of sequences. 15 Word Embeddings and Language Modeling === p16 === LM evaluation: Perplexity (困惑度) Perplexity (PPL) is a quantitative criterion used to evaluate the capacities/effectiveness of language modeling models. • Given the sequence of words and an n-gram model. The PPL of the model was computed by: The lower the value of perplexity, the better the language modeling capability of the model. 16 1~∞, 1~|𝑉| Word Embeddings and Language Modeling === p17 === Perplexity meaning A Measure of Uncertainty: Perplexity quantifies the level of uncertainty or unpredictability that a model experiences when making predictions. An Average Branching Factor: Perplexity can be viewed as an "average branching factor" of possible choices at each step in the sequence. Quantification of Model Performance: In language modeling, perplexity reflects how well the model understands the rules and structure of the language. ◦A lower perplexity indicates better language comprehension and prediction ability, signifying the model's efficiency in capturing language patterns. Compression Efficiency Indicator: Perplexity can also be seen as a measure of how well the model compresses the test data. ◦Lower perplexity means that the model predicts with higher probabilities, reducing the uncertainty and effectively compressing the information. 17 Can we use perplexity to evaluate whether text is generated by AI or not? === p18 === N-gram Language Models Bigram model Approximates the probability of a word given all the previous words by using only the conditional probability of the preceding word . The assumption that the probability of a word depends only on the previous word is called a Markov assumption. 18 i.e., Instead of computing the probability​ Bigram model approximates it with the probability​ Word Embeddings and Language Modeling === p19 === Shortcomings of N-gram LMs Limited context ◦N-gram models are unable to capture longer-distance (>>N) language dependencies. Data sparsity (High time/space complexity) ◦As the N value increases, the number of parameters to store and compute grows exponentially. Ignoring word order / context information ◦N-gram models assume independence between words, neglecting the influence of word order on semantics. Low flexibility N-gram language models struggle with synonyms and have limited ability to adapt to varying conditions. (e.g. dialogue). 19 Rare/Unseen words? New contexts or domains? https://books.google.com/ngrams/info Word Embeddings and Language Modeling === p20 === Sparse Vectors Sparse vector embeddings represent words as high-dimensional vectors with mostly zero values. Each dimension corresponds to a unique feature (word), measured by some well-designed methods. ◦TF-IDF ◦PPMI 20 Word Embeddings and Language Modeling === p21 === TF-IDF (recap) The mathematical representation of TF-IDF: ◦Where is the i-th word in j-th text in the dataset. TF (Term Frequency) ◦Represents the "frequency" of a term appearing in a text. IDF (Inverse Document Frequency) ◦Aims for terms to have higher specificity, meaning the fewer texts in the dataset contain the term, the better. 21 where Word Embeddings and Language Modeling === p22 === TF-IDF (recap) After computing the scores of every terms, we get a sparse vector to represent the text. 22 Word Embeddings and Language Modeling === p23 === TF-IDF (recap) Preprocessing text (optional) Stemming ◦By removing the suffixes from words (e.g."cats," "catlike," "catty" all have "cat" as their base), we can revert the words back to their root forms. Feature Selection ◦Filter and select which parts of speech to retain, such as verbs or nouns. ◦Analyze the frequency of the terms using statistical methods or algorithms like TF-IDF. 23 Word Embeddings and Language Modeling === p24 === Distributional Hypothesis Words that occur in similar contexts tend to have similar meanings. The NLP approaches utilize the context around the word to define its meaning. e.g. ◦I enjoy coding and I do it everyday! ◦I like coding and I do it everyday! "enjoy" and "like" are synonym and they are in the same context. 24 Word Embeddings and Language Modeling === p25 === PPMI Mutual Information (MI) ◦It is a measure of how often two events x and y occur: PMI (Pointwise MI) ◦The mutual information between a target word w and a context word c: 25 Word Embeddings and Language Modeling === p26 === PPMI PPMI (Positive PMI) ◦PPMI replaces all negative PMI values with zero e.g. Co-occurrence counts for 4 words in 5 contexts 26 computer data result pie sugar count(w) cherry 2 8 9 442 25 486 strawberry 0 0 1 60 19 80 digital 1670 1683 85 5 4 3447 information 3325 3972 378 5 13 7003 count(context) 4997 5673 473 512 61 11716 Word Embeddings and Language Modeling === p27 === computer data result pie sugar cherry 0 0 0 4.38 3.30 strawberry 0 0 0 4.10 5.51 digital 0.18 0.01 0 0 0 information 0.02 0.09 0.28 0 0 PPMI 27 Replacing the counts in with joint probabilities: Compute the PPMI matrix Sparse vectors are obtained =log2(0.0021/(0.0415*0.0052)) cherry=(0,0,0,4.38,3.30) P(w, context) p(w) computer data result pie sugar p(w) cherry 0.0002 0.0007 0.0008 0.0377 0.0021 0.0415 strawberry 0.0000 0.0000 0.0001 0.0051 0.0016 0.0068 digital 0.1425 0.1436 0.0073 0.0004 0.0003 0.2942 information 0.2838 0.3399 0.0323 0.0004 0.0011 0.6575 p(context) 0.4265 0.4842 0.0404 0.0437 0.0052 Word Embeddings and Language Modeling === p28 === Word embedding d1 d2 d3 d4 d5 d6 d7 d8 便宜 1 3 2 3 0 -2 2 0 有名 3 1 4 2 0 2 0 1 讚 0 0 1 0 0 1 -1 0 嫩 1 3 0 1 0 0 0 0 難吃 0 0 0 0 1 0 2 0 太貴 0 0 0 0 3 1 0 1 差 -1 0 1 -1 0 2 1 0 老 0 -2 -1 0 1 3 2 1 Concept space Term vector === p29 === Dense Vectors Dense vector embeddings represent words in a continuous vector space Semantically similar words are closer together. Word2Vec Contextualized Embeddings 29 Word Embeddings and Language Modeling === p30 === Properties of Embeddings 30 Vectors for representing words are called embeddings • Analogy/Relational Similarity Washinton - U.S. = London - U.K. Washinton - U.S. + U.K. = London Word Embeddings and Language Modeling === p31 === Properties of Embeddings 31 • Historical Semantics ➢Embeddings can also be a useful tool for studying how meaning changes over time spread Broadcast (1850s) Broadcast (1900s) Broadcast (1990s) sow sows seed scatter circulated newspapers television radio bbc Network (1920s) Network (1960s) Network (1990s) Network (2020s) Telegraph Electricity Transport Wires Computer Data Information Communication Connection Internet Browser Website Email TCP/IP IoT Social Media Blockchain AI Cloud Computing Word Embeddings and Language Modeling Google Bomb Data Poisoning === p32 === Word2Vec 32 ➢Skip-gram ▪Use words to predict their contexts. ➢CBOW ▪Use the context to predict the target word. Word Embeddings and Language Modeling === p33 === Word2Vec 33 1. Treat the target word and a neighboring context word as positive examples. 2. Randomly sample other words in the lexicon to get negative samples. 3. Use logistic regression to train a classifier to distinguish those two cases. 4. Use the learned weights as the embeddings. Word Embeddings and Language Modeling === p34 === Word2Vec 34 Calculate the co-occurrence matrix from the corpus, where each element represents how often a word appears in the context of another word within a certain window size. Assume window size = 2 Word Embeddings and Language Modeling === p35 === Word2Vec 35 Skip-gram: Assume c is the word id in the context, w is the center word, u and v are vectors of c and w. Word Embeddings and Language Modeling === p36 === Word2Vec 36 [IMAGE-ONLY] === p37 === Model Parameters word2vec uses a single hidden layer feedforward neural network Input to hidden layer matrix has our target word embeddings Hidden to output layer matrix has another set of word embeddings called output embeddings http://mccormickml.com/2016/04/19/word2vec-tutorial-the-skip-gram-model/ 37 === p38 === Word2Vec 38 ➢Large updates leads to inefficient o Stochastic Gradient Descent ▪only update word vectors in context window o Assume window size is m, In each window, there is only 2m+1 words. Sparse Word Embeddings and Language Modeling === p39 === Contextualized Embeddings 40 ➢Contextualized embeddings o capture the meaning of a word in its context within a sentence or a larger body of text. ➢Contextualized embeddings allows words to have different representations depending on their usage in different contexts. o Traditional word embeddings assign a fixed vector to each word. o Contextualized embeddings generate a vectors based on the surrounding words. ➢Contextualized embeddings are widely used in various NLP tasks. (e.g., BERT, GPT) Word Embeddings and Language Modeling === p40 === Word context 41 蘋果 蘋果 公司 派 蘋果改變了 他的一生, 對牛頓來 說 蘋果改變了 他的一生, 對賈伯斯 來說 === p41 === Contextualized Word Embeddings 42 ➢ Learn word vectors using long contexts instead of a context window ➢ Learn a deep Neural LM and use all its layers in prediction The probability of a sequence becomes: Word Embeddings and Language Modeling === p42 === Contextualized Word Embeddings 43 ➢ The embedding is designed to be a part of the model. ➢ When the downstream task provides a preferred context, it can be adjusted during the training process, incorporating the relevant information from the task context. Word Embeddings and Language Modeling LM Layers Embedding → Inputs Neural Network Architecture forzen? train together? === p43 === Neural Language Models A model structure is needed to process the hidden feature in the embeddings. Neural LMs leverage neural networks to learn and represent complex language patterns. Basic Definitions ◦FFN (Feedforward network) ◦RNN (Recurrent Neural Network) 44 Word Embeddings and Language Modeling === p44 === Basic Definitions 45 ➢In a deep learning project, there are some fundamental components. 1. Model ▪ A model refers to the architecture or structure used to represent relationships between input data and output predictions. 2. Optimizer ▪ An optimizer is an algorithm used to adjust the parameters of the model during training in order to minimize the error between predicted and actual output values. 3. Loss function ▪ A loss function (objective function) measures the difference between the predicted output of a model and the true target output. Word Embeddings and Language Modeling === p45 === Training Neural Networks 46 The training process typically involves the following steps: 1. Data Preparation: Prepare training and testing datasets. 2. Model Construction: Construct the model using a deep learning framework (TensorFlow, PyTorch, ...) 3. Loss Function Definition: Select an appropriate loss function. (Cross-entropy, Logloss, ...) 4. Optimizer Selection: Choose a suitable optimization algorithm. (Adam, SGD, ...) ◦ Adam: Adaptive Moment Estimation 5. Model Training: Train the model using the training dataset. 6. Model Evaluation: Evaluate the trained model using the testing dataset. (F1, LCS, ...) Word Embeddings and Language Modeling === p46 === Activation functions 47 ➢The core idea of using activation functions is to introduce nonlinearity into neural networks. ➢Neural network models aim to avoid the final processing stage being merely a linear transformation of the inputs, so that the model is able to have good performance on complex problems. ➢Commonly used activation functions include: Sigmoid, ReLU, tanh, GeLU... Word Embeddings and Language Modeling === p47 === Activation functions 48 Range Applications softmax [0, 1] (sum up to 1) It is commonly used in multi-class classification tasks where the model needs to predict the probability distribution over classes. sigmoid [0, 1] It often used in binary classification tasks where a threshold is needed to predict probabilities of belonging to one of the two classes. tanh [-1, 1] It is often preferred over sigmoid for hidden layers as it produces zero-centered output ReLU [0, +inf] It introduce nonlinearity and avoid gradient vanishing Word Embeddings and Language Modeling === p48 === Training Neural Networks 49 Word Embeddings and Language Modeling In text viewpoint === p49 === FFN 50 ➢Feedforward network (FFN) is a multilayer feedforward network in which the units are connected with no cycles. Word Embeddings and Language Modeling x1 x2 xn ... h1 h2 h3 hn y1 y2 ... yn ... Input layer Hidden layer Output layer W U === p50 === Short comings of FFN 51 ➢Lack of Sequence Modeling: o In NLP tasks, understanding the sequence of words and their dependencies is crucial for accurate predictions. ➢Fixed Input Size: o FFNs require fixed-size inputs, which can be problematic for NLP tasks where input sequences vary in length. ➢Limited Contextual Information: o Many NLP tasks benefit from capturing long-range dependencies and understanding the broader context of a text, which is better addressed by models capable of modeling sequential data effectively. Word Embeddings and Language Modeling === p51 === RNN 52 ➢Recurrent Neural Networks (RNNs) are a type of artificial neural network designed to handle sequential data by capturing temporal dependencies. o The same set of weights and biases are used across all time steps Word Embeddings and Language Modeling Moving average 進階版 Learnable weight Non-linear transformation === p52 === RNN History 1982 1986 BPTT (Backpropagation Through Time) Hopfield Network 2000s Elman RNN 1997 LSTM (Long Short-Term Memory) Bidirectional RNN, Deep LSTM 1990 2014-2017 GRU (Gated Recurrent Unit), Seq2seq Attention (2015) Transformer (2017) === p53 === RNN 54 ➢The equation on step t is: where f and g are activation functions Word Embeddings and Language Modeling xt yt W U V + ht-1 ht === p54 === Properties of RNNs 55 ➢Sequential Processing: o RNNs handle sequences, allowing them to model temporal dependencies in data. ➢Recurrent Connections: o RNNs maintain internal memory, facilitating the capture of long-term dependencies. ➢Parameter Sharing: o RNNs share parameters across time steps, enhancing efficiency in learning sequential data. ➢Vanishing Gradient Problem: o Traditional RNNs may face vanishing gradient issues, hindering learning of long-term dependencies. Word Embeddings and Language Modeling === p55 === Example: Name Entity Recognition 56 ➢Name Entity Recognition (NER) is a fundamental task in NLP. ➢The model needs to identify named entities within the sequence, such as countries, organizations, and individuals. Word Embeddings and Language Modeling === p56 === RNNs for NER 57 ➢In token classification task, every output should be mapped to a one-hot vector. ➢A feed forward network is added to the RNN model. Word Embeddings and Language Modeling === p57 === RNNs for Sequence Classification 58 ➢RNNs classify the entire sequences rather than the tokens within them. ➢Take the hidden layer for the last token of the text. Word Embeddings and Language Modeling RNN RNN RNN RNN x1 x2 x3 xn Softmax hn FFN === p58 === Stacked RNNs 59 ➢Stacked RNNs consist of multiple networks where the output of one layer serves as the input to a subsequent layer Word Embeddings and Language Modeling Figures are from https://reurl.cc/1389mD x1 x2 x3 xn RNN1 RNN2 RNN3 y1 y2 y3 yn === p59 === Stacked RNNs 60 ➢Stacked RNNs generally outperform single-layer networks. o The network induces representations at differing levels of abstraction across layers o The initial layers of stacked networks induce representations that serve as useful abstractions for further layers ➢However, as the number of stacks is increased the training costs rise quickly. Word Embeddings and Language Modeling === p60 === Bidirectional RNNs 61 ➢In many applications, RNNs have to access the entire input sequence. ➢Bidirectional RNN was introduced. oIt combines two independent bidirectional RNNs, one where the input is processed from the start to the end, and the other from the end to the start Word Embeddings and Language Modeling x1 x2 x3 xn y1 y2 y3 yn RNN1 RNN2 === p61 === Bidirectional RNNs for Sequence Classification 62 ➢The final hidden units from the forward and backward passes are combined to represent the entire sequence. ➢This combined representation serves as input to the subsequent classifier. Word Embeddings and Language Modeling x1 x2 x3 xn RNN1 RNN2 FFN Softmax === p62 === Summary 63 ➢LM: Models that assign probabilities to sequences of words ➢N-gram LM: Traditional statistical model of language Embedding approaches: ➢Sparse vectors: Use contexts to encode the word embedding. (TF-IDF, PPMI) ➢Dense vectors: Apply self-supervised training. (Word2Vec, Contextualized Embeddings) Neural LMs: ➢FFN: Fully-connected dense neural network. ➢RNN: Recurrent, store temporal hidden-state. Word Embeddings and Language Modeling