[人工智慧導論](https://app.notion.com/p/3e6fc631b0308131b213f0a3d93fe91d) › [W2(9/17)](https://app.notion.com/p/3e6fc631b0308175a433e5fa4bc500df) › 02|影片 [0:13:28–0:40:28](https://www.youtube.com/watch?v=hNZQIO0q74o&t=808s)|投影片 Ch2 p.2、p.9–26(p.2–12 是上週複習)|上一章 [01 HW1 提案寫法與分組(0:01–0:13)](https://app.notion.com/p/3e6fc631b030816cb5e8d1b405b4349e)|下一章 [03 搜尋問題的定義(0:53–1:08)](https://app.notion.com/p/3e6fc631b03081c8a7e0e152c12bbd25) ## 重點 - Agent 是任何「用 sensor(感測器)感知環境、用 actuator(致動器)對環境做動作」的東西;收到資料就算感知,送出資料也算動作。描述一個 agent,要先寫出它的任務環境 PEAS。 - 任務環境用七組性質分類(看得全不全、有沒有別的 agent、結果確不確定、這步會不會影響之後、思考時環境會不會變、離散或連續、知不知道規則)。環境不同,agent 的設計方法就不同;開計程車幾乎每組都落在難的那邊。 - AI 的工作是設計 agent program(把感知對應到動作的程式)。由簡單到複雜有五種:simple reflex(只看現在)、model-based(加記憶)、goal-based(加目標、想未來)、utility-based(算期望效益)、learning(靠回饋自己變好)。 ## Exam-ready - **Agent**: "An agent is anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators (促動器)."(Ch2 p.2)【老師強調】(0:14:14) - 中文:agent 是任何靠感測器(sensors)感知環境、靠致動器(actuators,促動器)對環境做動作的東西。白話:收到資料就算感知,送出資料就算動作,不一定要是機器人才算 agent。 - **Task environment (PEAS)**: "Task environment: we have to specify the performance measure, the environment, and the agent’s actuators and sensors."(Ch2 p.9) - 中文:要描述一個任務環境,得寫出四件事:績效怎麼算(performance measure)、環境本身(environment)、agent 能做的動作(actuators)、agent 收資訊的方式(sensors)。白話:PEAS 就是幫 agent 寫規格書要填的四個欄位。 - **Observable / Multiagent**: "Fully observable: if the sensors detect all aspects that are relevant to the choice of action" "The key distinction is whether B’s behavior is best described as maximizing a performance measure whose value depends on agent A’s behavior."(Ch2 p.11) - 中文:fully observable(完全可觀察)指感測器看得到做決定所需的全部資訊;判斷對方算不算另一個 agent,關鍵是它的行為是不是在最大化一個「會被你的行為影響」的績效分數。白話:能不能看到全局、身邊有沒有「對手」,是兩個不同的問題。 - **Deterministic**: "If the next state of the environment is completely determined by the current state and the action executed by the agent, then we say the environment is deterministic; otherwise, it is stochastic." "“Stochastic” generally implies that uncertainty about outcomes is quantified in terms of probabilities"(Ch2 p.12) - 中文:deterministic(確定性)指下一個狀態完全由現在狀態+agent 的動作決定;否則是 stochastic(隨機性),這時結果的不確定性會用機率描述。白話:同一個動作做兩次,會不會保證得到一樣的結果。 - **Episodic / Dynamic**: "Episodic: the next episode does not depend on the actions taken in previous episodes." "If the environment can change while an agent is deliberating (思考), then we say the environment is dynamic for that agent."(Ch2 p.13) - 中文:episodic(情節式)指這一輪的決定不會影響下一輪;dynamic(動態)指 agent 還在思考(deliberating)的時候,環境會自己變。白話:這步跟下一步有沒有關聯、想事情的時候世界會不會偷偷動。 - **Discrete / Known**: "The chess environment has a finite number of distinct states." "Taxi driving is a continuous-state and continuous-time problem." Known vs. unknown: "The agent’s (or designer’s) state of knowledge about the “laws of physics” of the environment."(Ch2 p.14) - 中文:discrete(離散)指狀態數量數得完(像西洋棋);continuous(連續)指狀態、時間都連續變化(像開車)。known/unknown 指 agent(或設計者)知不知道環境背後的規則,像物理定律一樣的東西。白話:格子數不數得完是一件事,規則是不是黑箱是另一件事,兩者互相獨立。 - **Agent program**: "The job of AI is to design an agent program that implements the agent function— the mapping from percepts to actions." "...computing device with physical sensors and actuators—we call this the architecture." "agent = architecture + program"(Ch2 p.16)"The table-driven approach to agent construction is doomed to failure – tables could be too huge"(Ch2 p.17) - 中文:AI 的工作是設計 agent program(把感知對應到動作的程式),它跑在有實體感測器、致動器的裝置(architecture)上,所以 agent=architecture+program。用查表法把每一種可能的感知序列都列出來,表會大到列不完,注定失敗。白話:一個像硬體、一個像軟體,硬湊一張列出所有情況的表不可行。 - **Simple reflex agent**: "Select actions on the basis of the current percept, ignoring the rest of the percept history."(Ch2 p.18)"Work only if the correct decision can be made on the basis of only the current percept—that is, only if the environment is fully observable."(Ch2 p.19) - 中文:只看現在這一刻收到的 percept(感知資料)來選動作,不管過去發生過什麼;只有在光靠目前這筆資訊就能做對決定時才有效,也就是環境要 fully observable(完全可觀察)。白話:像膝跳反射,看到什麼就反應什麼,不記憶。 - **Model-based reflex agent**: "Handle partial observability: the agent keeps track of the part of the world it can’t see now. The agent should maintain some sort of internal state that depends on the percept history."(Ch2 p.20) - 中文:靠 internal state(內部狀態)處理只能看到部分環境的情況:agent 心裡記著一份「世界現在長什麼樣」的狀態,這份狀態是從過去收到的一連串感知(percept history)算出來的,會不斷更新。白話:看不到的部分靠記憶和推理補上。 - **Goal-based agent**: "Knowing the current state of environment is not always enough to decide what to do." "The goal-based agent’s behavior can easily be changed ... simply by specifying that destination as the goal." "Consideration of the future – both “What will happen if I do such-and-such?” and “Will that make me happy?”"(Ch2 p.22) - 中文:光知道世界現在的狀態,不足以決定要做什麼,還要有目標(goal);這種 agent 會考慮「做了這件事之後會怎樣」以及「那樣我會不會滿意」,也就是會想未來。白話:多了「往哪裡去」的方向感,不只是反射動作,換個目的地就能自動換行為。 - **Utility-based agent**: "Goals alone are not enough to generate high-quality behavior in most environments." "When there are conflicting goals, only some of which can be achieved" "A rational utility-based agent chooses the action that maximizes the expected utility of the action outcomes."(Ch2 p.24)"Utility-based agent programs handle the uncertainty inherent in stochastic or partially observable environments."(Ch2 p.25) - 中文:光有目標還不夠讓 agent 表現好——目標可能互相衝突,或只能達成一部分;理性的 utility-based agent 會選能讓 expected utility(期望效益)最大的動作,這種做法也能處理隨機(stochastic)或看不全(partially observable)環境裡的不確定性。白話:目標只有「達成/沒達成」兩種結果,utility 幫你在幾個目標之間打分數、做取捨。 - **Learning agent**: "Learning element is responsible for making improvements, and the performance element is responsible for selecting external actions." "The learning element uses feedback from the critic ... and determines how the performance component should be modified to do better in the future."(Ch2 p.26) - 中文:learning element(學習元件)負責讓 agent 進步,performance element(績效元件)負責選出真正要執行的動作;learning element 靠 critic(評判者)給的回饋,決定要怎麼修改 performance element,讓它以後表現更好。白話:一塊負責「做」,一塊負責「靠回饋改進做法」。 ## [0:13:28](https://www.youtube.com/watch?v=hNZQIO0q74o&t=808s) 複習:agent、PEAS 與前三組環境性質 Agent 感知環境、做出動作,動作又改變環境,下一刻再感知,一直循環。老師提醒 sensor 不一定是物聯網感測器,收到資料就算;動作也不一定是實體移動,送出一筆資料也算【老師強調】(0:14:14)。描述 agent 要寫出 PEAS:Performance(怎麼評分,看需求定)、Environment(身處的環境)、Actuators(能做的動作)、Sensors(資訊從哪來:人類司機靠眼睛耳朵,自駕車靠攝影機、聲納、光達)。老師補充:agent 的概念幾十年前就有了,現在常讓 LLM(大型語言模型)當 agent 的「大腦」,決定收到訊息後做什麼 (0:17:37)。環境性質一變,agent 的設計方法就跟著變 (0:18:39),所以接著複習前三組性質(見摺疊)。下表是課本的計程車 PEAS(p.10 另列醫療診斷等五種)。 (第 1 週學過:agent 就是「感知→決定→動作」一直循環的東西,PEAS 是描述它的四個欄位。這段是複習,忘了可以回 [W1 週頁](https://app.notion.com/p/3e7fc631b03081309a2adad7f8ae1e27) 看。)
| Agent | Performance | Environment | Actuators | Sensors |
| Taxi driver | Safe, fast, legal, comfortable trip, maximize profits | Roads, other traffic, pedestrians, customers | Steering, accelerator, brake, signal, horn, display | Cameras, sonar, speedometer, GPS, odometer, accelerometer, engine sensors, keyboard |
| 類型 | 做決定時看什麼 | 例子 |
| **Simple reflex** | 這一刻的 percept + condition-action rule | 髒就吸,乾淨就換一格 |
| **Model-based reflex** | 記憶中的世界狀態(internal state)+規則 | 記得 A 已經吸過,兩格都乾淨就停 |
| **Goal-based** | 預測「做了會怎樣」,挑能達成目標的 | 路口轉哪邊,看乘客要去哪 |
| **Utility-based** | 每種結果多好,挑期望效益最大的 | 又要快又要安全 |
| **Learning** | 上面任一種+critic 的回饋,自己變好 | 上次三點這條路很塞,這次避開 |
| 步 | Simple reflex | 分 | Model-based | 分 |
| 1 | [A, Dirty]:Suck | 1 | 同左 | 1 |
| 2 | [A, Clean]:Right | 0 | 同左 | 0 |
| 3 | [B, Dirty]:Suck | 2 | 同左 | 2 |
| 4 | [B, Clean]:Left | 1 | [B, Clean]:NoOp | 2 |
| 5–8 | Right、Left 來回跑 | 每步 1 | 停在 B 不動 | 每步 2 |
| **合計** | **8** | **13** |