=== p1 === Natural Language Processing Syllabus Slido: 9433372 === p2 === Hung-Yu Kao (高宏宇) 2 Data Competition Award • BioCreative, Rank 1, 2010, 2011, 2015 • TREC Blog, Rank 3, 2010 • ACL SemEval Rumor Detection Competition, Score Rank1, 2017 • ACL SemEval Argument Reasoning Comprehension Test, Rank 2, 2018 • CIKM AnalyticCup Short Text Matching, Rank 2, 2018 • WSDM Fake News Classification, Rank 3, 2019 • ACL/Google AI Gender Bias for Natural Language Processing Rank 4, 2019 • WSDM Visual Question Answering Challenge, Rank 4, 2022 NLP related publication (2019-2026) • ACLx4, EMNLPx4, AAAIx3, ICML, Coling, WSDM, TKDE, TASLP, TAI • 清華大學資訊工程系教授 • 清華大學計算與通訊中心主任 • 國科會人工智慧驅動藥物開發工程處召集人 • 曾任成功大學資訊工程系特聘教授 • 曾任成功大學電機資訊學院副院長 • 曾任中華民國人工智慧學會 理事長 • 曾任中華民國 計算語言學會 理事長 • 曾任人工智慧學校講師 • 曾多次獲成功大學教學傑出獎, 優良導師 • 曾獲中國電機工程學會112年度「傑出電機工程教授獎」 === p3 === 115-1 自然語言處理 開設學校:清華大學 開授教師:高宏宇 班級人數:1200人 開課級別:研究所 同步遠距上課時間:星期四9:00-12:00 課程目標: ◦本課程涵蓋自然語言處理(NLP)與大型語言模型(LLM)的基礎 與前瞻技術 ◦針對生成式人工智慧技術的快速發展,探討NLP在各領域的廣泛應 用 ◦提供理論基礎與實際應用,幫助學生掌握最新NLP與LLM技術 3 === p4 === 自然語言處理 課程簡介 探索自然語言處理領域的基礎知識。 了解電腦如何理解和處理人類語言。 清華大學資訊工程系高宏宇 大型語言模型時代的技術進化與挑戰 。 === p5 === Why Study Natural Language Processing? [IMAGE-ONLY] === p6 === Worries in the LLM era, Eduard Hovy, CMU, Rocling 2024. 6 Look what an LLM can do! Why can it do that? I have no idea / that’s future work / I’ve never thought about it Look what an LLM cannot do! Why not? I have no idea / that’s future work / I’ve never thought about it I don’t care about LLMs – here’s what I did Why are you doing that? Can’t an LLM do it already? I have no idea / that’s future work / I’ve never thought about it === p7 === Three major directions for New NLP, Eduard Hovy, CMU, Rocling 2024. 7 NLP engineering: make LLMs usable • Build smaller and cheaper LLMs • Systematize prompt engineering • Integrate language, images, and other media and functions NLP applications: make LLMs useful • Tune LLMs to domains and companies for enterprise processing • Add functionality and agency in the world • Tailor LLMs to people to be their personal daemons / amanuenses in everyday life NLP Research: make LLMs understandable (or at least, be solid engineering) • Fix the problems with LLMs • Ger explanations how LLMs do what they do • Formalize them well enough for autonomy, assurance, and ethics === p8 === Language is the most natural way humans communicate. Every day, the world generates massive text: news, social media, research papers, and customer records. Manually processing this data is impossible and inefficient. 8 NLP = enabling computers to “understand” and “use” language. === p9 === Why Is NLP Important? 9 Everyday life: Google Translate, Siri/ChatGPT, speech-to-text. Industry: chatbots for customer service, medical record analysis, market sentiment detection. Academia: automatic summarization, knowledge retrieval, scientific discovery. Future: human–AI interaction, intelligent assistants, cross-language communication. === p10 === 10 [IMAGE-ONLY] === p11 === The revolution of ChatGPT (2022.11) 11 [IMAGE-ONLY] === p12 === What is Generative Artificial Intelligence? Generative Artificial Intelligence (AI) describes algorithms (such as ChatGPT) that can be used to create novel content, including: 12 • Audio • Code • Images • Text (article, translation, ...) • Videos • And so on ... Method Company Image GAN, Stable Diffusion, Transformer NVIDIA (StyleGAN2), OpenAI (DALL-E, CLIP), DeepMind (BigGAN), LeepMotion (Midjourney), StabilityAI (Stable Diffusion) Text GPT (Generative Pre- trained Transformer) OpenAI (ChatGPT, GPT-4), Meta (LLaMA2), Google (T5, XLNet, PaLM), Speech GAN, VAE Google (WaveNet), OpenAI (WaveGlow) Code Seq2Seq, GPT OpenAI (Copilot) === p13 === 13 User: Design a modern, visually appealing 16: 9 banner for a Generative AI course. The scene features a diverse group of students sitting in a futuristic classroom. DALL E: Here are the images designed for your Generative AI course banner. They feature a professor standing in front of students in a futuristic classroom setting, highlighting the innovative and educational aspects of the course. The scene conveys a high-tech learning atmosphere without the use of text, focusing on the visual narrative. Text-to-Image === p14 === 14 Code Llama by Meta Try it: https://huggingface.co/spaces/codellama/codellama-13b-chat === p15 === Coding copilot 15 [IMAGE-ONLY] === p16 === ChatGPT as a fake News generator 16 [IMAGE-ONLY] === p17 === 17 Multi-modal Understanding using Large Language Model Liu, Haotian, et al. "Improved Baselines with Visual Instruction Tuning." NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following. 2023. === p18 === Robot Manipulation with Multimodal Prompts 18 Watch video here: https://vimalabs.github.io/ Jiang, Yunfan, et al. "VIMA: Robot Manipulation with Multimodal Prompts." ICML (2023). === p19 === Multimodality with Text 19 VQA Applications Vision-Language-Model Applications === p20 === 20 https://finance.yahoo.com/technology/ai/articles/sam-altman-says-ai-triggers-093000057.html https://openai.com/index/navier-stokes-solution/ === p21 === Generation by AI 21 [IMAGE-ONLY] === p22 === 生成對抗網路 (Generative Adversarial Network,GAN) 22 https://github.com/jonbruner/generative-adversarial-networks/blob/master/gan-notebook.ipynb === p23 === Power of Image GAI Training data generation / augmentation Style transform Super-resolution image 23 ●Image compression NIPS 2016 tutorial: http://auai.org/uai2017/media/tutorials/shakir.pdf === p24 === Power of Image GAI 24 NIPS 2016 tutorial: http://auai.org/uai2017/media/tutorials/shakir.pdf === p25 === How about text? 25 [IMAGE-ONLY] === p26 === NLP Levels 26 • Prefix / Suffix • Lemmatization / Stemming • Spelling Checking Morphology • Part-of-Speech Tagging • Syntax trees • Dependency Trees Syntax • Named Entity Recognition / Normalization • Relation Extraction • Word Sense Disambiguation Semantics • Co-reference resolution • Topic Segmentation • Summarization Pragmatics === p27 === How teach computer to understand this? Q:曾有一項調查發現,很多員工生病的時候不敢請假,因為他們擔心 老闆會不高興,覺得他們沒有責任感。有人認為,員工會這麼想是公司 的責任。一個好的公司應該能照顧員工,而不是讓他們拿健康去換錢。 因此,讓員工有幸福感,應該是未來企業努力的方向。 這篇文章說了什麼內容? 1.老闆應該給員工多一點兒假 2.常關心別人的人更有責任感 3.對公司有意見要勇敢說出來 4.照顧身體比認真工作更重要 27 科技大擂台, 2017 === p28 === 語言模型(Language Model) 28 Claude Shannon 1916 – 2001 Andrey Markov 1856 - 1922 [1913] The chance of a letter appearing depends on the letter before it. [1951] Prediction and Entropy of Printed English === p29 === Language Model 29 [IMAGE-ONLY] === p30 === 30 童年的紙飛機現在終於飛回我 手裡 天青色等煙雨而我在 等妳 怎麼這樣子 雨還沒停妳就撐 傘要走 妳說這一句 很有夏天的感覺 我用幾行字形容妳是我 的誰 為妳翹課的那一天 花落的 那一天 消失的下雨天 === p31 === 31 童年的紙飛機現在終於飛回我 手裡 天青色等煙雨而 在 等 怎麼這樣子 雨還沒停妳 就撐 傘要走 說這一句 很有夏天的感覺 用幾行字形容 翹課的那一天 花落 消失的下 天 為 是 的誰 P (b | a) P (c | ab) . . . P (説 | 妳) = 1/4 P (説 | 沒停妳) = 0 === p32 === NLP is not NLP before [IMAGE-ONLY] === p33 === History of Language Science (AI view) 33 === p34 === 34 [IMAGE-ONLY] === p35 === 35 “pre-train (預訓練), fine-tune(微調)” paradigm 轉移學習 Transfer Learning === p36 === 36 https://github.com/Mooler0410/LLMsPracticalGuide === p37 === 37 https://jjanes.ca/llm-tree/ LLM Family Tree === p38 === Traditional NLP applications (text) 38 文字語料 Collection, annotation, … 前處理 Segmentation, stopping word removal, stemming, parsing tree, … 前處理 - 2 NER, NEN, Relation extraction, …. Operation I One-hot vector, TD-IDF, Word2Vec Operation II Similarity measurement, LSI, semantics representation Operation III Domain knowledge === p39 === Large Language Model (LLM) enabled NLP applications 39 Pre-train BERT GPT ChatGPT GPT-4 . . . Fine-tune doc summary 文件摘要模型 doc Question/ answer 問答系統 xB ~ xxxB 參數 + docs 知識問答系統 Retrieval / QA === p40 === Large Language Model (LLM) enabled NLP applications (no fine-tune) 40 Pre-train BERT GPT ChatGPT GPT-4 . . . prompt-usage Prompt: doc, 摘要 Summary Prompt: doc, Question xB ~ xxxB 參數 docs Retrieval / QA Answer Answer Length limitation === p41 === 自然語言處理 基本概念與方法 將介紹NLP的核心概念,包括字詞表示、文字處理、字詞向量表示 、 語意 意分析和情感分析等。 1 詞彙分析 分析文字結構,將句子拆 解為詞彙並標記詞性。 2 語法分析 解析句子的文法結構,理 解詞彙之間的關係。 3 語意理解 理解文字的含義,並分析 語句之間的邏輯關係。 4 情感分析 識別文字的情緒和主觀性 ,了解作者的立場和態度 。 === p42 === 深度學習與 大型預訓練語言模型的應用 例如BERT、GPT-3等,解決不同語言處理任務。 1 Transformer 重要基礎模型原理。 BERT :雙向編碼器表示模型,善 於理解上下文資料。 2 序列類神經網路 RNN, LSTM, Seq2seq 。 3 生成式AI T5 GPT 4 大型語言模型系統架構流程與應用 預訓練 微調 PEFT 提示學習 / In-context learning 將探討大型預訓練語言模型 (LLMs) 如何革新NLP領域,並展示其在各項應用中的優勢。 === p43 === In this course Basics ◦Python usage NLP Basics NLP with DL ◦Sequential model ◦Generative Models (text) NLP issues in LLM era ◦Reasoning ◦Agentic AI Not included ◦Speech ◦Prompt usage ◦Develop new models ◦Solve GAI problems === p44 === Schedule Week Topics Note W1 課程簡介Syllabus / Introduction to NLP W2 自然語言處理簡介(1/2) Introduction to NLP (vector space, indexing, parts of speech, phrase structure) W3 自然語言處理簡介(2/2) Introduction to NLP (Language model) W4 基礎文字資料機器學習(1/2) Basic machine learning for text (Text Classification, NB, NN) W5 基礎文字資料機器學習(2/2) Basic machine learning for text (word embedding, text representation) W6 Python for text tutorial (1/2) W7 Python for text tutorial (2/2) W8 文字生成式AI簡介(1/3) Introduction to GAI (text): Word Embeddings, Language Modeling (RNN), Sequence-to-sequence Models, and Attention Mechanisms, Sub-word Tokenization; Transformers W9 文字生成式AI簡介(2/3) Introduction to GAI (text): ELMo, BERT, GPT, and T5 (BERT and its Family) W10 文字生成式AI簡介(3/3) Introduction to GAI (text): Decoding Strategies and Evaluations for Natural Language Generation W11 大語言模型簡介與訓練(1/3): Large language model concept and training (GPT-3, InstructGPT, RLHF) W12 大語言模型簡介與訓練(2/3): Parameter Efficient Fine-Tuning (PEFT) W13 大語言模型簡介與訓練(2/3): Introduction and Review technique of Retrieval Augmented Generation (RAG) I W14 Exam W15 大語言模型簡介與訓練(2/3): Introduction and Review technique of Retrieval Augmented Generation (RAG) II W16 Reasoning / Agent Lecture Tutorial Reporting Exam The course content is subject to minor adjustments. === p45 === Grading Assignments 75 % ◦4 assignments for each student ◦Coding needed ◦The loading of Single assignments will be ”fine-tuned” according to the alternative designs. Their credits will also be slightly changed (-3%~+3%) Midterm Exam 25 % (week 14, Tentative date) This course is X-class in NTHU, you can enroll other courses with the same time slots in NTHU. The exam will be held in person. === p46 === Grade in 2025 46 [IMAGE-ONLY] === p47 === Computation Resources from TAICA  提供期間:9 月至12 月(3 個月),每個帳號自啟用日起90 天內有效, 若期間未用完將自動失效。 名額與預算:總額9 萬元,您可以參考以下組合,或做任何調整: - 100 運算元帳號(每個帳號約300 元):100 名 - 500 運算元帳號(每個帳號約1500 元):40 名 Wait for announcement 47 === p48 === Homework evaluation 48 Human-TA批改 AI-TA 協助 Code / Results Evaluation Discussion Evaluation Your Insight, not GPT insight === p49 === This course 49 Is designed for ◦對自然語言處理有高度興趣的學生 ◦有一些深度學習模型訓練基礎的學生 ◦想了解why LLM so power? === p50 === The hardships of learning 50 [IMAGE-ONLY] === p51 === Instructors and TAs ▪Instructors: ◦Hung-Yu Kao 台達館 633 (hykao@cs.nthu.edu.tw) ▪TAs: ▪台達館 7F, Room 714 IKM lab. ▪Course website ▪NTU Cool (https://cool.ntu.edu.tw/courses/41436 ) ▪Github (https://github.com/IKMLab/NTHU_Natural_Language_Processing )