68 Action Model BERT BERT vs GPT (Encoder vs Decoder) ChatGPT (2022.11) Co-occurrence Code generation Conditional probability Course prerequisites Data-driven statistical models Decoder Decoding (next-token generation) Deep Learning Dependency relation Expert System GAN (Generative Adversarial Network) GPT (Generative Pre-trained Transformer) GPT-3 General Text Processing Generative AI Hallucination Hidden-prompt trap in peer review In-context learning InstructGPT and ChatGPT Instruction Fine-tuning Knowledge graph LLM + downstream model fine-tuning LLM fake news generation Language Model Large Language Model (LLM) Latent Semantic Analysis (LSA) Length limitation (context length) Lost in the middle Machine Learning Masked Language Model (MLM) Multimodality N-gram conditional probability NLP Levels (Morphology / Syntax / Semantics / Pragmatics) NLP applications (make LLMs useful) NLP engineering (make LLMs usable) NLP research (make LLMs understandable) Named Entity Recognition (NER) Neural Language Model Neural Network Neural models One-hot encoding PEFT (LoRA) Parse Tree (Syntax Tree) Positional embedding Pretrained Language Model (PLM) Prompt usage (no fine-tune) RNN / LSTM Rare-symbol watermark Sampling-based watermark Seq2seq (Sequence to sequence model) Stemming vs. Lemmatization Supervised learning TF-IDF Text watermarking Traditional NLP pipeline Training data generation / augmentation Training-data watermark Transfer Learning (pre-train, fine-tune paradigm) Transformer Vibe Coding Word embedding (dense space) Worries in the LLM era (Eduard Hovy) Zero-probability (unseen N-gram) problem word2vec