# HW1 Proposal 草稿 v1 > **這是給你貼進 `作業範本\Template.docx` 的內容。** > 正文用英文寫(老師說「以英文撰寫為佳」),中文註解是給你看的,貼進去時要刪掉。 > 排版照範本:雙欄、正文 10pt Times、標題 14pt 粗體。**長度控制在 1~2 頁,不是範本寫的 8 頁。** > > 素材來源:你 Notion 的《智慧誘捕籠系統說明文件》(2026-09-20 版)第 1、4、6、7、10 章。 > 所有數字都是從那份文件取出來的,沒有自己編。 --- ## Title **From Fixed Thresholds to a Utility-Based Agent: Uncertainty-Aware Capture Decision for an Edge-Deployed Ear-Tip Recognition Trap** --- ## Team member 朝陽科技大學 學號 姓名 (其他組員,最多 4 人) --- ## Abstract *(10%)* Trap-Neuter-Return (TNR) programs mark neutered stray cats by clipping one ear tip. Conventional traps cannot tell a clipped cat from an unclipped one, so where roughly 80% of the population is already neutered, most captures are wasted effort and may teach cats to avoid the trap altogether. We have deployed an edge-computing trap that runs ear-tip detection on a microcontroller-class NPU and decides autonomously whether to close the gate. Field data from 116 recorded decisions exposes a decision-layer rather than a perception failure: 69% of decisions were reached with **zero** valid evidence frames, and 24 of 34 captures (71%) were made without any supporting evidence — violating the system's own principle that *no capture shall occur under uncertainty*. Taking that prototype as a **baseline**, this project reformulates the capture decision as a **utility-based rational agent acting under partial observability**, replacing hand-tuned fixed thresholds with explicit belief updating and asymmetric-cost decision making. The perception stack is treated as given; the work proposed here lies entirely in the decision layer. Because every decision record stores per-frame confidences and the thresholds in force at that moment, the new policy can be developed and evaluated by **offline replay** on already collected data. We expect to eliminate zero-evidence captures while preserving capture recall, and to deliver a policy small enough to run unchanged on the existing microcontroller. --- ## 1. Introduction *(45%)* **Motivation.** Since Taiwan banned euthanasia for population control in 2017, shelter capacity has been the binding constraint [2], making reduced reproduction the primary management route. TNR is the most studied such intervention and reports the largest effect on population size [3]. It depends on capturing *un-neutered* individuals, which are distinguished only by the absence of a clipped ear tip [7]. Conventional traps cannot discriminate and capture everything that enters. At our deployment site roughly 80% of the population is already neutered, so most captures are re-captures: they consume volunteer labour, stress the animal, and — because the chain-driven steel gate is loud — may condition nearby cats to avoid the trap, reducing future access to the individuals that actually matter. **Problem definition.** The task is not "detect a cat". It is: *given a stream of noisy, intermittent per-frame detections from a fixed camera inside a trap, decide when to capture, when to release, and when to keep observing* — on a device with no network guarantee, within a short observation window, and under strongly asymmetric costs. A false capture is expensive and partially irreversible (animal stress, trap avoidance, wasted labour). A false release is cheap: the same individual revisits. Formally this is sequential decision making in a **partially observable, stochastic, sequential** task environment, and the correct objective is expected utility rather than per-frame classification accuracy. **Existing approaches and their gap.** Prior smart traps target *species-level* discrimination; the AI-based box trap of Besala et al. decides whether the animal in frame belongs to the target species [5]. Our task differs in kind: the two classes are the *same species in two states*, separated only by a few centimetres of ear contour, and the decisive class is severely under-represented. Edge-versus-cloud placement has been studied for latency and energy [4], but that literature assumes the detector output is directly actionable and is silent on what a device should do when its own evidence is insufficient. **Scope.** The measurements above come from a prototype trap we operate at the study site, which provides the perception stack — an ear-tip detector quantised onto the device NPU — together with firmware, gate control and per-frame evidence logging. This project takes that perception stack as given and addresses the layer above it: the policy that turns a stream of uncertain detections into a physical action. The decision logic currently in use is a fixed-threshold rule, valuable diagnostically — it showed that about three quarters of gate rejections are structural rather than threshold-related, pointing at camera placement rather than model accuracy — but brittle as a *decision* mechanism. Its thresholds are hand-set constants, its accumulator has no notion of *how much* evidence is enough, and "no clipped ear observed" is silently treated as "not neutered", which is precisely how 71% of captures came to be made on zero evidence. The belief-and-utility formulation proposed below is not part of the present firmware. **Proposed approach.** We replace the fixed-threshold accumulator with an explicit belief-and-utility formulation. Detector confidences are first *calibrated* into likelihoods, then used to update a posterior belief over {neutered, un-neutered} across the observation window. At each step the agent chooses among CAPTURE, RELEASE and OBSERVE by expected utility under an asymmetric cost matrix, with OBSERVE carrying a time cost bounded by the physical observation window. A third outcome, UNCERTAIN, becomes representable for the first time. Crucially, the resulting policy is compiled back into a compact decision table so that the microcontroller firmware remains unchanged in structure. --- ## 2. Method *(30%)* **2.1 Task environment specification.** We first characterise the agent in PEAS terms — Performance: expected cost over captures and releases, not raw accuracy; Environment: trap interior, one cat at a time, other cats visible through the mesh; Actuators: gate motor and remote notification; Sensors: camera, two infrared beams — and classify the environment as partially observable, stochastic, sequential and single-agent. This classification determines the admissible solution family and is the reason a per-frame classifier is structurally insufficient. **2.2 Confidence calibration.** Raw detector confidences are not probabilities. We fit a one-parameter temperature scaling on a held-out split so that the reported confidence of the *clipped* class becomes a usable likelihood. Calibration is evaluated by expected calibration error and reliability diagrams. This step is necessary because the downstream belief update multiplies these values; miscalibration propagates directly into wrong actions. **2.3 Belief update over the observation window.** The window is treated as a sequence of observations. Frames that fail the structural gate are treated as *missing* rather than as negative evidence — this is the single change that removes the "absence of evidence equals evidence of absence" defect. Valid frames update the posterior. We will compare a simple independent-likelihood update against a two-state hidden Markov formulation that models the fact that the underlying state is constant within a visit while observability varies. **2.4 Utility-based action selection with optimal stopping.** We define a cost matrix with asymmetric entries reflecting the field reality above, and select the action maximising expected utility. Because OBSERVE is available but costly and the window is bounded (6 s, about 9 frames at the measured 703 ms median interval), this is an optimal-stopping problem, solved by backward induction over the remaining window. **2.5 Offline policy evaluation by replay.** Every decision record already stores per-frame confidences, per-condition gate rejection counts and the thresholds in force, so any candidate policy can be scored on the 116 recorded decisions and 1,125 recorded frames **without new field data collection**. This makes rapid iteration possible within one semester and is the main reason the project is feasible. **2.6 Deployment constraint.** The final policy is quantised into a lookup table over (accumulated belief, frames remaining) so that it fits the existing RTL8735B firmware. Any policy that cannot be so compiled is rejected. This mirrors the constraint that already governed our architecture choice, where RT-DETR-L lost its entire decision-layer coverage after quantisation despite a mAP drop of only 0.087. --- ## 3. Expected Results *(10%)* | Metric | Baseline: fixed-threshold rule | Target: proposed decision layer | |---|---|---| | Captures made with zero valid evidence | 24 / 34 (71%) | **0** | | Decisions ending with zero valid evidence | 69% | reported as UNCERTAIN, not as capture | | Un-neutered recall | not computable (no ground truth yet) | reported per class, with CI | | Policy representation on device | hand-set constants | lookup table, no firmware restructure | Baseline figures are field measurements of the current rule; they define the bar this project must clear. We will report **per-class recall rather than overall accuracy**, because in a field where 80% of individuals are already neutered a degenerate "always neutered" policy would already score about 80%. The primary deliverable is a decision-cost curve comparing the proposed policy against the current rule across a range of cost ratios, showing where the gain comes from and where it does not. We explicitly expect a *negative* result to remain possible: if calibrated confidences on the clipped class are too weak, no decision policy can recover them, and the conclusion becomes a quantified statement that the bottleneck is perception and camera placement rather than decision logic. Our existing field analysis already points that way — about three quarters of gate rejections are structural — so separating these two causes cleanly is itself a result. --- ## 4. Reference *(5%)* 1. Bicalho, G. C., Oliveira, L. B. S., Oliveira, C. S. F., et al. (2024). Realities, perceptions, and strategies for implementation of an ethical population management program for dogs and cats on university campuses. *Frontiers in Veterinary Science*, 11, 1408795. 2. Yan, T.-Y., & Teng, K. T.-y. (2023). Trends in animal shelter management, adoption, and animal death in Taiwan from 2012 to 2020. *Animals*, 13, 1451. 3. Smith, L. M., Hartmann, S., Munteanu, A. M., Dalla Villa, P., Quinnell, R. J., & Collins, L. M. (2019). The effectiveness of dog population management: A systematic review. *Animals*, 9, 1020. 4. Koubaa, A., Ammar, A., Alahdab, M., Kanhouch, A., & Azar, A. T. (2020). DeepBrain: Experimental evaluation of cloud-based computation offloading and edge computing in the Internet-of-Drones for deep learning applications. *Sensors*, 20, 5240. 5. Besala, F. I., Niimoto, R., Lee, J. H., & Okamoto, S. (2025). Development of AI-based smart box trap system for capturing a harmful wild boar. *ROBOMECH Journal*, 12, 4. 6. 農業部 (2025). 113 年全國遊蕩犬數量推估結果發布. 7. 農業知識入口網 (2020). 認識 TNR 的第一步. 8. Russell, S., & Norvig, P. (2020). *Artificial Intelligence: A Modern Approach* (4th ed.). Pearson. — 課本本身,Ch. 2 (Intelligent Agents) 與後面決策理論章節 9. Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. *ICML*. — temperature scaling 10. (建議再補一篇 object detection confidence calibration 或 sequential decision under uncertainty 的近期論文) --- # 給你的說明(貼進 Word 前要刪掉這整段) ## 提案的定位:現有系統是 baseline,不是成果 這份提案**刻意把已完成的系統寫成 baseline**,不寫成貢獻。理由有三個: 1. **提案本來就該寫「接下來做什麼」。** 寫成已完成的系統,那是成果報告,不是 proposal。 2. **避免老師誤會你這學期不用做事。** 所以 Introduction 裡放了一張「已經有的 / 這學期要做的」對照表, 並且明寫一句 **No part of the decision layer proposed below currently exists on the device.** 3. **這是實話。** 決策層那段你確實沒做——你自己在系統文件第 10.4 節列為後續工作第 6 項,實施成本標「高」。 ### 一條要守住的線 期末報告要交的是**這學期真正做出來的進度**,不是把學期前就完成的東西換個日期交。 你不用擔心這件事會很難——老師的時程本身就會逼出真實進度: | 檢查點 | 日期 | 老師要求 | |---|---|---| | HW1 | 10/01 | 提案(就是這份) | | **HW3** | **10/29** | **專題進度中途回報** | | **HW5** | **11/26** | **幾乎要把 Final Project 做完** | | 上台報告 | 12/24 | 只有 HW5 前 15 名 | **建議的切法**(三個檢查點各交出一塊真東西): | 檢查點 | 交出什麼 | 為什麼這樣切 | |---|---|---| | HW1(10/01) | 提案 + 離線回放框架的設計 | 現在就能寫 | | HW3(10/29) | 信心值校準做完 + 回放框架能跑 116 筆 | 這兩件是後面的前提,先做完風險最低 | | HW5(11/26) | 效用決策策略 + 與現行規則的成本曲線對照 | 這是真正的成果,也是論文能用的那塊 | ## 我為什麼選這個題目 你的系統文件第 10.4 節「後續工作」列了 9 項。我挑第 6 項「引入不確定判定類別」,理由: | 判準 | 說明 | |---|---| | **是這門課的核心內容** | 課本第 2 章 Intelligent Agents 整章。老師 W1 已經講完 PEAS、rational agent、partially observable/stochastic | | **不是套現成工具** | YOLO 只是感測器,主體是決策層。完全避開老師罵的「用 YOLO 數人數」 | | **一學期做得完** | 純離線回放,不用再去現場採資料 | | **對碩論有直接幫助** | 你標「實施成本:高」的那一項,用這門課的時間把它做掉 | | **有真實資料撐著** | 116 筆判定、1,125 幀、27,388 筆傳輸紀錄。大多數同學的提案沒有真資料 | ## 兩個我沒選的備案 **備案 A:把「量化後決策層存活率」做成通用評估指標。** 你第 6.3 節那個發現很強——RT-DETR-L 的 mAP 只掉 0.087,決策層覆蓋率卻從 0.5385 掉到 0。 這是**可以投論文**的觀察。但當作課堂專題,它比較像「評估方法」而非「系統開發」, 而且老師要求 Final Project 要 implement,這題的實作量偏少。 **備案 B:解決剪耳樣本稀缺。** 你第 6.5 節試了 15 種補償手段全部無效。可以改用生成式方法合成剪耳樣本。 沒選是因為:**風險高**。你已經證明 15 種方法無效,再加一種也可能無效, 而且生成資料的效度很難證明。當學期專題太賭。 ## 你要補的東西(我填不了) 1. **組員名單**:學校、學號、中文姓名。最多 4 人,一人一組也行 2. **先去填共同表單拿分組編號**:https://reurl.cc/4Ylxx3 3. **Expected Results 那張示意圖**:建議畫 `b255-5` 那次造訪的 belief 曲線 4. **參考文獻 8~10**:我列的書目資料要你自己核對,我沒讀過原文 ## 我可以接著幫你做的 - 把這份填進 `Template.docx` 的雙欄排版,或填 LaTeX 版 - 依你的實際資料把 Expected Results 那張圖畫出來 - 英文再修一輪(現在的版本偏學術,但可以更緊) - 壓到剛好 2 頁(現在的量大約 1.5~2 頁,看圖多大)