Saturday, July 25, 2026

中國產烤鰻供給過剩並導致日本市場烤鰻價格暴跌

 由於去年日本玻璃鰻苗因豐收致養殖成鰻大量增產,加上今年玻璃鰻漁況也佳,使中國產烤鰻供給過剩並導致市場價格暴跌。目前廣受日本喜愛每盒10公斤含30尾肥大鰻魚的價格一度跌到每公斤2,000日圓,是五年來最低水準。即便如此,零售量仍缺乏成長力道,以總零售額來計算,年增率不到10%。解決當前此銷售不振瓶頸問題,關鍵在日本超商與量販店大規模積極銷售瘦小之小型鰻魚。

中國產烤鰻製品在日本之零售而言,近年來由於其鰻魚是日本產鰻魚所沒有的肥大而有優質感備受好評,加上前年日本鰻玻璃鰻苗極端不足,中國以每年持續入池美洲鰻作為烤鰻原料魚以替代日本鰻而迅速成長。去年日本鰻玻璃鰻苗大豐收,光中國就有85公噸以上漁獲,故今年成鰻大幅增產已成定局。此外,今年新鰻苗漁期也比平常年多,有40公噸以上的入池量,導致中國面臨鰻魚養殖池不足狀態。養殖池放養密度已過高,養殖環境不再適合養殖肥大的鰻魚。當地養殖業者越來越希望增加每箱10公斤、50尾裝瘦小鰻魚銷往日本的數量,以調整生產水準。然而,目前日本市場對肥大者有強勁需求,但比起美洲鰻製成之烤鰻,其口味與口感均更勝一籌,對其評價程度超過預期。特別是最近由日本鰻為原料魚製成之烤鰻,比美洲鰻烤鰻之價格還便宜。

中國希望強化其鰻魚之輸日量,但目前的情況使得過剩瘦小鰻魚銷日變得困難。因此中國產燒烤鰻魚一度在餐飲業爆炸性成長,現在已趨穩定。目前只得靠專門銷售一尾分包裝的量販店及超市以低價競爭來維持銷售量。但以目前的價格來看1尾包裝者,即使無法以單價500日圓銷售,也可以在超市以低於1,000日圓的價格售出。而且有跡象顯示持續下滑的市場行情有觸底反彈可能,有人認為「從4月26日那週開始,其物流將會活躍起來」,但除非有實際需求出現,否則供給過剩問題將無法解決,日本相關業界希望盡快恢復市場正常化,期望量販店與超市能採取更積極主動方式來擴大銷售量,普遍認為量販店與超商應大膽擴大銷售攤位,並為瘦小成鰻提供特價優惠。

鰻苗產量增加引發中國鰻魚市場行情嚴重下滑

東亞四國的日本鰻玻璃鰻苗採捕在繼2025年漁期創紀錄超過100公噸後,2026年管理期玻璃鰻苗漁汛期又是一個豐收年,漁獲量已接近70公噸,由於鰻苗產量增加引發世界最大鰻魚養殖國中國鰻魚市場行情嚴重下滑。中國有關產業團體不得不嚴肅因應之,在今年3月下旬舉行的產業團體會議上通過四項決議,其中包括禁止捕撈日本鰻稚魚一年及禁止在未來一年內進口美洲鰻鰻苗,這些決定對日本市場的影響為何?日刊新聞記者訪問日本鰻魚進口協會理事長松浦信也。

 

記  者:最近中國養鰻業界發生哪些變化。

理事長:3月27日在中國福建省召開之「中國漁業協會鰻漁業工作委員會」會議,作成4項決議,根據當地報導包括有1)禁止捕撈日本鰻稚魚一年,並禁止進口美洲鰻鰻苗一年;2)建立全國最低採購價格指導制度;3)積極推動自主加工1萬公噸蒲燒鰻,以供應國內市場;4)全部有關人員參與促銷方式以開拓國內市場等。儘管政府沒有參加此一會議,但一群傑出的民間企業相關成員聚集在一起,旨在維護因產量增加導致價格持續下跌而陷入困境的鰻魚養殖戶與相關企業。

記  者:這對於依賴進口日本玻璃鰻苗及依賴本地活成鰻與加工鰻的日本而言,也事關重大吧?

理事長:本協會已將上述相關資訊分享給協會團體會員高層,但我個人並不感到緊迫性,因為該4項決議要執行的可能性都很低,即我懷疑這些協議的執行力。倒是中國養殖業者陷入困境是錯不了。

            就作為日本鰻魚進口商協會而言,現在反而擔心去年玻璃鰻大豐收而養成之成鰻即將全面收成,勢必對日本國內成鰻消費市場帶來巨大壓力,導致更多企業湧入中國鰻魚市場,可能給日本國內市場帶來混亂。原因是如果無秩序的進口,可能會出現重大問題,甚至可能導致整個日本鰻魚生產系統癱瘓,這是一個嚴重後果問題。

記  者:是否希望相關進口企業盡可能加入協會,以秩序化因應中國產鰻魚進口問題?

理事長:這是鰻魚進口商所擁有的共同課題當然要共同來解決,本協會自去年CITES締約國大會以來,已能夠與國內6個組織建構強大合作體系,目前不只在協會內部,而且在與其他協會間資訊交流共享等充分合作,這是本協會以前所沒有的優點。因此強烈歡迎新日本鰻魚進口商考量加入本協會。

記  者:如何消費日本市場上大量供應的活成鰻與加工鰻?

理事長:就日本市場之活成鰻而言,國產與中國產的價格每公斤只差150日圓而已,因此進口之中國活成鰻銷售量有限。且活成鰻在烤鰻餐廳的實際價格基本上是一致的,使得進口鰻的消費數量更難以增加,再說日本國產雌鰻受歡迎也是原因之一。更由於過去幾年一直有提供重新加熱的冷凍蒲燒鰻連鎖餐廳,從去年下半年開始就一直減少,感覺上日本的成鰻消費量似乎已觸頂。

記  者:未來如果想增加鰻魚消費量,降低價格是唯一的選擇嗎?

理事長:至少到今年夏天,零售商店貨架上烤鰻魚價格範圍無疑將與去年夏天大不相同。我原以為烤鰻市場原本一向以美洲鰻為主要商品,可能會替換為今年增產的日本鰻,然而烤美洲鰻特有的口感比預期受到一群消費群喜愛,其肉質鮮嫩也逐漸獲得認可,價格又便宜,消費群被取代是想多了。

記  者:從今年玻璃鰻苗的採捕量來看,是否感覺日本的入池量累積太慢了?

理事長:目前入池量累積速度雖然尚未加快,但港邊玻璃鰻苗採捕量正在增加,有相當數量輸出。日本養鰻池之所以不放養是因為去年的鰻魚還在池中未出貨。而且兩年前因玻璃鰻苗採捕量不佳、價格高,許多鰻魚養殖場就開始停止放養也是原因之一。日本的入池量上限為21.7公噸,個人認為今年入池量不會滿15公噸。

記  者:也就是如果考慮其他瞄準日本鰻魚市場的3個東亞國家,是否認為其今年也有充分玻璃鰻苗可入池?

理事長:個人希望日本鰻玻璃鰻苗不只是連續2年連續豐收,往後也能恢復常態。也就是大力宣導日本鰻放流計畫,降河鰻魚的保護等努力取得成功,可有助阻止2年後CITES締約國大會再次提出將日本鰻納入該公約附錄二的提案。

「中日鰻魚貿易會議」,第37次會議在中國福建省廈門市召開

由日本與中國鰻魚貿易相關民間人士輪流主辦的「中日鰻魚貿易會議」,第37次會議最近在中國福建省廈門市召開。日本鰻魚進口協會(理事長松浦信也)理事龜山時宏等約10人,中國方面則約有60人出席此會議並進行意見交換。

受到2年連續玻璃鰻苗豐收,雙方對今年夏天成鰻之銷售戰預測進行整合與理解。有關華盛頓公約(CITES)已決於2028年在巴拿馬召開下屆締約國會議,兩國於本次會議中已決定再度攜手合作,正如上次CITES大會否決歐盟提案一樣,以民間層級合作展開鰻苗資源管理合作。會議當天,協會松浦理事長在住處受傷,未能出席,由龜山時宏理事(活鰻部會長)與諸川祐輔理事(加工部會長)代理出席會議,由龜山時宏代讀松浦信也理事長之致詞稿,中國則由食品土畜進口商會副會長于露致詞。雙方一致認為在CITES問題上合作深化了兩國間的聯繫,並表示希望此合作關係能成為克服鰻魚市場行情下跌所帶給養殖戶困境等緊急問題的突破口。

接著日本由諸川祐輔部會長就近年來業界最關心的資源保護與管理日本作法進行介紹,他也進一步表示日本正加速致力實現鰻魚完全養殖,也將玻璃鰻苗指定為水產流通適正化適用對象魚種,嚴格管理其生產履歷。據諸川祐輔表示,中方對日本此一舉措給予很高評價,並希望「兩國將保持密切資訊交換,並同意兩國分別採取必要措施」。

其次針對兩國今年玻璃鰻苗的採捕與入池情況進行報告。中方表示目前為止今年廣東省之入池量有20公噸,福建省也有20公噸,合計入池量為40公噸,平均價格為每尾4元人民幣(每公斤平均以5,000尾玻璃鰻苗計,每公斤約為47萬日圓)。另外中國今年之美洲鰻入池量有25公噸。日本方面,迄會議當天為止,國內玻璃鰻苗的採捕量約有8公噸,入池量約達14公噸,然而,肇因於去年玻璃鰻苗創紀錄豐收,因此今年日本成鰻市場行情低迷,日中兩國養殖業者對今年玻璃鰻苗之入池意願均低迷。至於加工鰻方面,今年由於日本量販店與超市將加工鰻視為擴大銷售商品,其銷售額將超過去年水準,估計2026年鰻魚管理期(2025年9月-2026年8月)最終銷售實績可望達2萬5,000公噸。

其次則是有關今年夏天成鰻市場商戰預測問題,中方表示今年中國不論是日本鰻還是美洲鰻的養殖業者均必須面對出貨產生的虧損問題,雖然努力調整出貨,但目前還是面臨艱困局面。其次中方估計今年4-8月輸往日本的活鰻約有3,000公噸,鰻魚加工製品約在1萬到1萬5,000公噸間。

日方則表示,其國內鰻魚的產量約為1萬公噸,且由於今年以大豆異黃酮飼料投餵的影響,產量有可能上調,至於活鰻的進口方面,由於目前其國產鰻魚的供應狀況充足以及日本物價飆升,導致其國內對鰻魚消費量也下滑,估計今年5-8月間活鰻進口量約為3,000公噸而已。

日本將鰻玻璃鰻苗納入鰻魚產業價值鏈可追溯性支持系統

 由日本全國永續鰻魚養殖協會(保科正樹理事長,以下簡稱養鰻協會)為迎合日本鰻玻璃鰻苗納入水產流通適正法(以下簡稱流適法)而主導之鰻魚產業價值鏈可追溯性支持系統實施委員會於去年12月開始提供此服務。此溯源系統運作第一年即將結束。該系統於設計時特別注重減輕交易量大的玻璃鰻苗採捕人,將鰻苗轉移給收集鰻苗之收購人時的行政負擔,為驗證該系統之實用性,去年的實證試驗除在日本主要鰻魚養殖縣(鹿兒島、宮崎與愛知)外,也在鰻苗採捕人數眾多的高知縣與三重縣,對近9,000鰻苗採捕人、約200名收集鰻苗的人進行利用實證試驗。

就日本2026年玻璃鰻苗漁期(2025年11月-2026年10月)之鰻苗漁況而言,雖然不及去年之豐收年,但比豐收年之前連續4年漁況欠佳之年,今年之採捕量多得多,在試用該系統的人中,保科正樹理事長表示「雖然只有在採捕多的主要縣進行試用,即有些地方還沒採用該系統,而且在易用性方面顯示該系統還是存在些不足之處,但整體而言第一年的實證實驗,算是該追溯系統之實用性已達及格分數」。

協議委員會開發的這套系統由一臺平板電腦、一個二維碼閱讀器、一個數位鍵盤及一臺收據印表機等組成,玻璃鰻苗採捕人帶著其捕獲的鰻苗與用戶卡前往收貨點,收貨員讀取用戶卡並將鰻苗的計算結果輸入平板電腦。採捕人的出貨交易紀錄、收貨人的收貨交易紀錄以及鰻苗漁獲編號等都會被保存在系統中。漁獲編號由系統自動產生,當其交付給批發商或鰻魚養殖戶時,系統會自動產生與漁獲編號關連的貨物編號,以達到順利交貨。

日本養殖鰻魚生產量第四的靜岡縣與玻璃鰻苗採捕量最多的千葉縣,第一年度並沒有導入此系統,而採捕人也相當多的德島縣,第一年度導入該系統也不多,根據2024年日本全國鰻魚總產量1萬6,674公噸,該系統在佔總產量86%的縣中使用。

水產廳2025年財政年度追加預算中列有經費將對此可追溯性系統進一步加以改良。根據秘書處收到的用戶回饋,系統將新增一項連動功能,即用於向都府縣知事報告鰻苗捕撈量,並將鰻苗的入池量報告給水產廳。其目標是從2027年管理期(2026年11月-2027年10月)開始提供改良版之追溯系統。

由於報告義務的行政負擔也可以減輕,保科正樹理事長也表達他對未來該系統的期望,表示「期望來年使用該系統的人比第一年更多」。

Wednesday, July 22, 2026

高表達不等於作用大,且高表達基因確實往往是『功能的末端執行者』;而低表達基因則多半扮演『上游的指揮開關』

 「高表達不等於作用大,且高表達基因確實往往是『功能的末端執行者』;而低表達基因則多半扮演『上游的指揮開關』。」

💡 一、高表達 = 作用大?高表達 = 作用末端?

1. 高表達 ≠ 作用大(表現量高只是因為「消耗量大」)

在細胞內,RNA 拷貝數高往往是因為該蛋白質需要大量作為結構組成大規模參與催化

  • 例如:核糖體蛋白、組蛋白、細胞骨架(Actin)、DNA 複製相關蛋白。

  • 這些基因就像工廠裡的工人與建材,雖然數量極多,但牠們只是在「執行命令」,並不是決定細胞命運的「決策者」。

2. 高表達 ≒ 作用末端 (Terminal Effectors)

你的直覺非常準確!高表達基因(如分位數 $0.75 \sim 0.95$ 的基因)通常是細胞狀態轉變後的最終表型產物(Phenotypic Endpoint)

  • 當細胞決定要進行有絲分裂時,它會大量轉錄 DNA 複製與細胞週期蛋白(這就是為什麼在上調高分位數中,CELL_CYCLE$-\log_{10}P$ 會高達 24.92)。

  • 這些基因告訴你的是:「細胞目前正在發生什麼結果(What is happening right now)」

🧬 二、如何界定與研究「較低表達基因」的作用?

如果高表達基因是「執行者」,那麼較低表達基因(如分位數 $0.01 \sim 0.50$)通常就是「開關與指揮官」(如轉錄因子、膜受體、激酶、細胞因子)。

細胞只需要極少量的受體或轉錄因子 mRNA,經由信號放大機制(Signal Amplification),就能驅動整個細胞產生巨大的表型變化。要界定並挖出這些低表達基因的作用,生物資訊學通常採用以下 4 種策略:

1. 分位數過濾(Quantile Stratification)—— 即我們剛才做的事

  • 作法: 將高表達的「背景噪音/終端執行者」(如核糖體、細胞週期蛋白)剝離,單獨對低/中分位數區間進行 GO 富集或 GSEA。

  • 作用: 讓原本被掩蓋的微環境訊號(如 GPCR 訊號轉導、離子通道、細胞間通訊)浮出水面。

2. 重視「相對變化量(Fold Change)」而非「絕對表現量」

  • 作法: 觀察基因的 $\Delta asinh$$\log_2\text{FC}$

  • 作用: 一個轉錄因子的絕對表現量可能只有 $asinh = 0.5$,但若它從 $0.1 \rightarrow 0.5$(激增 5 倍),其生物學衝擊力遠大於一個從 $8.0 \rightarrow 8.5$ 的高表達結構蛋白。

3. 蛋白質互作網路(PPI Network)的「瓶頸節點(Bottleneck/Hub)」分析

  • 作法: 將篩選出的低表達基因繪製成 PPI 網絡(如 STRING 資料庫),計算其 Betweenness Centrality(介數中心性)

  • 作用: 許多低表達蛋白質位於網絡的核心樞紐(Hub),連接不同的路徑。一旦刪除這個低表達基因,整個傳遞網路就會癱瘓。

4. 反推上游調控者分析(Upstream Regulator / Master Regulator Analysis)

  • 作法: 利用如 DoRothEA、ChEA 或 SCENIC 等工具,透過「下游大量改變的高表達基因群」,反推是哪一個上游轉錄因子(TF)在操控牠們。

  • 作用: 即使該轉錄因子本身的 mRNA 表達量極低(甚至在 RNA-seq 中抓不到),也能精準界定出它的核心控制地位。

📝 總結

  • 高表達區間 $[0.51 \sim 0.95]$ 代表表型結果 (Phenotype),解答「細胞演變成什麼狀態?」(如:大量增殖、癌化、DNA 損傷)。

  • 低/中表達區間 $[0.01 \sim 0.50]$ 代表上游機制與微環境 (Upstream & Environment),解答「是什麼訊號/受體啟動了這個變化?」(如:激素刺激、離子通道改變、免疫微環境誘導)。


「高表達不等於作用大,且高表達基因確實往往是『功能的末端執行者』;而低表達基因則多半扮演『上游的指揮開關』。」
那傳統生物統計如何解釋?(由 Gemini 回答)
這是轉錄組學與生物統計學中極為經典的「**統計顯著性(Statistical Significance)與生物重要性(Biological Relevance)的悖論**」。
傳統生物統計(如 DESeq2、edgeR、limma)的數學模型,本質上是建立在「訊號強度(Signal-to-Noise Ratio, SNR)」**與**「線型效應(Linear Effect Size)」的假設之上。面對「低表達是上游開關、高表達是末端執行者」的生物學事實,傳統生物統計從數學機制與模型設計上有以下幾層解釋與局限:
---
### 1. 抽樣噪訊與統計檢定力的天然限制(Statistical Power & Noise)
RNA-seq 的定序計數(Counts)通常採用負二項分佈(Negative Binomial Distribution)進行建模:
$$\text{Var}(Y) = \mu + \alpha \mu^2$$
* $\mu$:基因的平均表達量
* $\alpha$:離散參數(Dispersion)
**統計上的困境:**
* **低表達基因(上游開關):** 當平均表達量 $\mu \to 0$ 時,泊松抽樣噪訊(Poisson Shot Noise)佔據主導地位,相對變異度(Coefficient of Variation)極高。這會導致 Wald Test 或 Likelihood Ratio Test 中的標準差偏大,**直接壓低 $Z$ 分數或 $t$ 分數,使得 $p$-value 不顯著**。
* **傳統統計的處置:** 為了防止多重檢定(Multiple Testing)的 False Discovery Rate (FDR) 被無意義的低表達噪訊稀釋,傳統統計工具會採用**獨立過濾(Independent Filtering)**,在檢定前直接將低表達基因剔除。這意味著**許多低表達的轉錄因子(TF)或受體,在傳統統計第一關就被當作「噪訊」扔掉了**。
---
### 2. 廣義線性模型(GLM)與非線性級聯放大的錯位
傳統微分表達分析(DEG)的核心是廣義線性模型:
$$\log(\mu_{ij}) = \beta_0 + \beta_1 \cdot \text{Condition}_j$$
該模型隱含了一個假設:**對數表達量的變化量($\beta_1$,即 Log Fold Change, LFC)直接對應於生物學效應的大小**。
**生物學與統計的錯位:**
* **級聯放大(Signal Cascade):** 生物系統是非線性催化系統。上游轉錄因子或激酶只需要微量的 mRNA 波動($\Delta \text{RNA}$ 極小),透過蛋白質翻譯與酶催化放大,就能啟動下游數百個結構蛋白或代謝酶的大量轉錄($\Delta \text{RNA}$ 極大)。
* **傳統統計的盲點:** 傳統 GLM 只能量化「轉錄本濃度的改變量」,無法量化「該分子在調控網路中的**控制權重(Control Coefficient)**」。因此,統計上最顯著(LFC 大、$p$-value 小)的基因,往往都是幾何級數增加的末端執行者(如熱休克蛋白、結構蛋白、代謝酶)。
---
### 3. 傳統生物統計為了修正此問題所作的努力
為了減輕「低表達基因因估計不準而產生高偽陽性或高偽陰性」的問題,傳統統計學家引入了幾項核心技術:
| 統計修正技術 | 數學/邏輯機制 | 對「上游開關/低表達基因」的影響 |
| --- | --- | --- |
| 經驗貝葉斯方差壓縮<br>
<br>(Empirical Bayes Shrinkage) | 利用全基因組平均表達量與方差的趨勢(Mean-Variance Relationship),將個別基因的離散度向整體壓縮(`limma-voom`, `DESeq2`)。 | **提高檢定力**:防止低表達基因因單次採樣誤差導致方差被高估,讓穩定微幅變化的上游基因有機會達到顯著。 |
| 對數倍數變化縮減<br>
<br>(LFC Shrinkage, e.g., `apeglm`) | 利用 Prior Distribution 懲罰低表達基因因抽樣波動產生的虛假高 Fold Change(如從 1 變 10 的 10 倍)。 | **抑制偽陽性**:讓真正具備高表達量且穩定的基因排在 DEG 前列,但同時也進一步壓低了低表達基因的 LFC 排序。 |
---
### 4. 統計思維的演進:從「單基因檢定」到「系統/網路推算」
正因為傳統生物統計在單基因層面(Single-gene level)難以克服「高表達 vs 低表達」的權重錯位,現代生物資訊統計學逐漸將重點轉向**系統級別的推算**:
1. **靶基因集整合推算(Regulon-based Target Inference):**
* *代表工具:* VIPER, Decoupler (ULM/MLM), DoRothEA / CollecTRI。
* *統計邏輯:* 不直接去檢定低表達 TF 本身的 mRNA 變化(因為噪訊太高),而是**觀察該 TF 下遊數十個 Known Targets(多為中高表達基因)的集體表達趨勢**。透過統計上游標的基因集的對齊程度,反向推算低表達 TF 的活性(Activity Score)。
2. **網路拓撲中心性(Network Centrality):**
* *代表工具:* WGCNA, Single-cell GRN (SCENIC)。
* *統計邏輯:* 放棄以「表達量」或「微分表達顯著性」為唯一指標,改用**共表達網路中的連通度(Degree / Hubness)**。上游驅動基因(Hub Genes)即使絕對表達量低,但因其與大量下游模組基因保持高度相關性,在統計拓撲上會被識別為核心節點。



import os

import sys

import json

import logging

import textwrap

import warnings

from dataclasses import dataclass, field, asdict

from pathlib import Path

from typing import Tuple, Optional, List, Dict, Any

from datetime import datetime


import numpy as np

import pandas as pd

from scipy.stats import spearmanr

import matplotlib.pyplot as plt

import matplotlib.path as mpath

import matplotlib.patches as mpatches

import seaborn as sns

from sklearn.decomposition import PCA

from sklearn.preprocessing import StandardScaler


try:

    import gseapy as gp

    GSEAPY_AVAILABLE = True

except ImportError:

    gp = None

    GSEAPY_AVAILABLE = False

    warnings.warn("gseapy 未安裝,富集分析將使用內建 mock 結果。")


try:

    from adjustText import adjust_text

    ADJUSTTEXT_AVAILABLE = True

except ImportError:

    ADJUSTTEXT_AVAILABLE = False



def setup_logging(level=logging.INFO, log_file=None, format_str=None) -> logging.Logger:

    if format_str is None:

        format_str = "%(asctime)s | %(levelname)-8s | %(name)s | %(message)s"

    logger = logging.getLogger("OvaryAnalysis")

    logger.setLevel(level)

    logger.handlers.clear()

    formatter = logging.Formatter(format_str, datefmt="%Y-%m-%d %H:%M:%S")

    console_handler = logging.StreamHandler(sys.stdout)

    console_handler.setLevel(level)

    console_handler.setFormatter(formatter)

    logger.addHandler(console_handler)

    if log_file:

        file_handler = logging.FileHandler(log_file, mode='w', encoding='utf-8')

        file_handler.setLevel(logging.DEBUG)

        file_handler.setFormatter(formatter)

        logger.addHandler(file_handler)

    return logger



@dataclass(frozen=True)

class AnalysisConfig:

    sample_ids: Tuple[str, ...] = ('1440MT', '2003MT', '2613MT', '1440AT', '2003AT', '2613AT')

    individual: Tuple[int, ...] = (0, 1, 2, 0, 1, 2)

    is_treated: Tuple[int, ...] = (0, 0, 0, 1, 1, 1)

    phenotype_order: Tuple[str, ...] = ('1440', '2613', '2003')

    phenotype_ranks: Tuple[int, ...] = (3, 2, 1)

    default_gmt_path: str = field(

        default_factory=lambda: str(Path.home() / "Documents/Python/c5.go.bp.v2023.2.Hs.symbols.gmt")

    )

    # asinh_lower / asinh_upper 的意義取決於 asinh_use_quantile:

    #   asinh_use_quantile=False(預設關閉時):兩者是絕對的 asinh 表現值門檻。

    #   asinh_use_quantile=True:兩者是「分位數」(0~1 之間的分數),程式會依

    #       實際資料分布動態換算成對應的 asinh 門檻值,例如 0.5 代表中位數、

    #       1.0 代表最大值。這樣門檻能隨資料自動校準,不需要每次重新猜測絕對值。

    asinh_lower: Optional[float] = 0.75

    asinh_upper: Optional[float] = 0.95

    asinh_filter_mode: str = "any"

    asinh_use_quantile: bool = True

    min_delta_range: float = 0.1

    min_rho: float = 0.8

    top_n_features: Optional[int] = None

    output_dir: str = "./ovary_analysis_output"

    dpi: int = 300

    fig_format: str = "png"


    def __post_init__(self):

        valid_modes = {"mean", "all", "any"}

        if self.asinh_filter_mode not in valid_modes:

            raise ValueError(f"asinh_filter_mode 必須是 {valid_modes} 其中之一,收到: {self.asinh_filter_mode!r}")

        if self.fig_format not in {"png", "pdf", "svg"}:

            raise ValueError(f"fig_format 必須是 png/pdf/svg 其中之一,收到: {self.fig_format!r}")

        if self.asinh_use_quantile:

            for name, val in (("asinh_lower", self.asinh_lower), ("asinh_upper", self.asinh_upper)):

                if val is not None and not (0.0 <= val <= 1.0):

                    raise ValueError(

                        f"asinh_use_quantile=True 時,{name} 必須是 0~1 之間的分位數分數,收到: {val!r}"

                    )


    def to_dict(self) -> Dict[str, Any]:

        return asdict(self)


    def save(self, path: str) -> None:

        with open(path, 'w', encoding='utf-8') as f:

            json.dump(self.to_dict(), f, indent=2, ensure_ascii=False)


    @classmethod

    def load(cls, path: str) -> "AnalysisConfig":

        with open(path, 'r', encoding='utf-8') as f:

            data = json.load(f)

        for tuple_field in ("sample_ids", "individual", "is_treated", "phenotype_order", "phenotype_ranks"):

            if tuple_field in data and data[tuple_field] is not None:

                data[tuple_field] = tuple(data[tuple_field])

        return cls(**data)


    @property

    def mt_samples(self) -> List[str]:

        return [sid for sid, t in zip(self.sample_ids, self.is_treated) if t == 0]


    @property

    def at_samples(self) -> List[str]:

        return [sid for sid, t in zip(self.sample_ids, self.is_treated) if t == 1]


    @property

    def delta_columns(self) -> List[str]:

        return [f"Delta_{p}" for p in self.phenotype_order]



@dataclass

class FeatureSelectionResult:

    positive_genes: pd.DataFrame

    negative_genes: pd.DataFrame

    full_report: pd.DataFrame


    def summary(self) -> Dict[str, Any]:

        return {

            "n_positive": len(self.positive_genes),

            "n_negative": len(self.negative_genes),

            "n_total_tested": len(self.full_report),

            "top_positive": self.positive_genes.head(5).index.tolist() if not self.positive_genes.empty else [],

            "top_negative": self.negative_genes.head(5).index.tolist() if not self.negative_genes.empty else [],

        }



@dataclass

class EnrichmentResult:

    dataframe: Optional[pd.DataFrame]

    gene_list: List[str]

    group_name: str

    success: bool = field(init=False)


    def __post_init__(self):

        self.success = self.dataframe is not None and not self.dataframe.empty

        if self.success:

            # 依顯著性排序一次,讓後續所有取用(top_terms、繪圖、匯出)都是

            # 真正依統計顯著性排序,而非 GMT/Enrichr 回傳的原始(通常是字母)順序。

            sort_col = None

            for candidate in ("Adjusted P-value", "P-value"):

                if candidate in self.dataframe.columns:

                    sort_col = candidate

                    break

            if sort_col is not None:

                self.dataframe = self.dataframe.sort_values(sort_col).reset_index(drop=True)


    def top_terms(self, n: int = 10) -> pd.DataFrame:

        """取得依顯著性排序後的前 N 個條目(最顯著者優先)。"""

        if not self.success:

            return pd.DataFrame()

        return self.dataframe.head(n)



class DataLoader:

    EXPECTED_COLS = "A,B,F,G,I,J,L,M"

    SHEET_NAME = "data"


    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log


    def load(self, file_path: Optional[str] = None) -> pd.DataFrame:

        if file_path is None:

            file_path = self._find_data_file()

        if file_path and os.path.exists(file_path):

            expr_matrix = self._load_and_process_data(file_path)

        else:

            self.logger.warning("未找到有效數據文件,生成模擬數據。")

            expr_matrix = self._generate_mock_data()

        self._qc_check(expr_matrix)

        return expr_matrix


    def _find_data_file(self) -> Optional[str]:

        candidates = [

            "./FC_GSEA.xlsx",

            "../FC_GSEA.xlsx",

            os.path.expanduser("~/Documents/Python/FC_GSEA.xlsx"),

        ]

        for path in candidates:

            if os.path.exists(path):

                self.logger.info(f"找到數據文件: {path}")

                return path

        return None


    def _load_and_process_data(self, file_path: str) -> pd.DataFrame:

        self.logger.info(f"正在載入數據: {file_path}")

        try:

            df = pd.read_excel(file_path, sheet_name=self.SHEET_NAME, engine='openpyxl', usecols=self.EXPECTED_COLS)

        except Exception as e:

            self.logger.error(f"讀取 Excel 失敗: {e}")

            raise


        df.columns = df.columns.str.strip()

        group_col = df.columns[1]


        raw_samples = list(self.config.sample_ids)

        missing_cols = [c for c in raw_samples if c not in df.columns]

        if missing_cols:

            raise ValueError(

                f"Excel 檔案缺少必要的樣本欄位: {missing_cols}。"

                f"現有欄位: {list(df.columns)}"

            )


        for col in raw_samples:

            df[col] = pd.to_numeric(df[col], errors='coerce')


        df_indexed = df.set_index(group_col)


        n_dup = df_indexed.index.duplicated().sum()

        if n_dup > 0:

            self.logger.warning(f"發現 {n_dup} 個重複的基因名稱索引,將保留第一筆出現的紀錄")

            df_indexed = df_indexed[~df_indexed.index.duplicated(keep='first')]


        expr_matrix = df_indexed[raw_samples].copy().fillna(0.0)

        self.logger.info(f"原始數據: {expr_matrix.shape[0]} 基因 x {expr_matrix.shape[1]} 樣本")


        if (expr_matrix < 0).any().any():

            n_neg = (expr_matrix < 0).sum().sum()

            self.logger.warning(f"發現 {n_neg} 個負值,asinh 轉換仍可處理,但請確認數據來源是否正確")


        expr_transformed = np.arcsinh(expr_matrix)

        self.logger.info(f"asinh 轉換完成: {expr_transformed.shape[0]} 基因 x {expr_transformed.shape[1]} 樣本")

        return expr_transformed


    def _generate_mock_data(self) -> pd.DataFrame:

        rng = np.random.default_rng(42)

        n_random_genes = 300

        pos_trend_genes = ["CYP19A1", "STAR", "LHCGR", "MAPK1", "CDK1"]

        neg_trend_genes = ["CCNB1", "TGFBR1", "CASP3", "FOXO1"]

        gene_names = pos_trend_genes + neg_trend_genes + [f"GENE_{i:04d}" for i in range(n_random_genes)]

        cols = list(self.config.sample_ids)


        data = np.abs(rng.normal(loc=5.0, scale=0.4, size=(len(gene_names), len(cols))))

        df = pd.DataFrame(data, index=gene_names, columns=cols)


        for g in pos_trend_genes:

            df.loc[g, self.config.mt_samples] = 4.0

            df.loc[g, "1440AT"] = 4.0 + 6.0

            df.loc[g, "2613AT"] = 4.0 + 3.5

            df.loc[g, "2003AT"] = 4.0 + 1.0


        for g in neg_trend_genes:

            df.loc[g, self.config.mt_samples] = 4.0

            df.loc[g, "1440AT"] = 4.0 - 6.0

            df.loc[g, "2613AT"] = 4.0 - 3.5

            df.loc[g, "2003AT"] = 4.0 - 1.0


        df = df.clip(lower=0.01)

        self.logger.info(f"模擬數據生成: {df.shape[0]} 基因 x {df.shape[1]} 樣本")

        return np.arcsinh(df)


    def _qc_check(self, expr_matrix: pd.DataFrame) -> None:

        missing_samples = set(self.config.sample_ids) - set(expr_matrix.columns)

        if missing_samples:

            self.logger.warning(f"缺少樣本: {missing_samples}")

        if expr_matrix.empty:

            self.logger.warning("表達矩陣為空")

            return

        min_val = expr_matrix.min().min()

        max_val = expr_matrix.max().max()

        self.logger.info(f"asinh 值域: [{min_val:.4f}, {max_val:.4f}]")

        zero_genes = (expr_matrix.sum(axis=1) == 0).sum()

        if zero_genes > 0:

            self.logger.warning(f"發現 {zero_genes} 個全零表達基因")



class AsinhFilter:

    """asinh 數值區間篩選器。


    此為全流程唯一負責 asinh 區間篩選的類別,篩選後的結果應直接往下游傳遞,

    不應在其他地方(例如 PhenotypeFeatureSelector)重複套用。

    """


    def __init__(self, log: logging.Logger, use_quantile: bool = False):

        self.logger = log

        self.use_quantile = use_quantile


    def _resolve_thresholds(

        self, data_df: pd.DataFrame, lower: Optional[float], upper: Optional[float], mode: str

    ) -> Tuple[float, float]:

        """將 lower/upper 轉換成實際用來比較的 asinh 數值門檻。


        若 use_quantile=False:lower/upper 本身就是絕對 asinh 值,直接使用。

        若 use_quantile=True:lower/upper 是 0~1 之間的分位數分數,會依資料

        實際分布(mode="mean" 時用逐基因平均值的分布;"all"/"any" 時用整個

        矩陣攤平後的分布)換算成對應的 asinh 數值,讓門檻隨資料自動校準。

        """

        if not self.use_quantile:

            lo = -np.inf if lower is None else lower

            hi = np.inf if upper is None else upper

            return lo, hi


        if mode == "mean":

            values = data_df.mean(axis=1).to_numpy()

        else:

            values = data_df.to_numpy().flatten()

        values = values[~np.isnan(values)]


        lo = -np.inf if lower is None else float(np.quantile(values, lower))

        hi = np.inf if upper is None else float(np.quantile(values, upper))

        self.logger.info(

            f"分位篩選: 下分位數 {lower} -> asinh {lo:.4f}, 上分位數 {upper} -> asinh {hi:.4f}"

        )

        return lo, hi


    def filter(self, data_df: pd.DataFrame, lower=None, upper=None, mode="any") -> pd.DataFrame:

        if lower is None and upper is None:

            self.logger.info("asinh 篩選: 未設置上下限,返回原始數據")

            return data_df


        lo, hi = self._resolve_thresholds(data_df, lower, upper, mode)


        if mode == "mean":

            row_stat = data_df.mean(axis=1)

            mask = (row_stat >= lo) & (row_stat <= hi)

        elif mode in ("all", "any"):

            in_range = (data_df >= lo) & (data_df <= hi)

            mask = in_range.all(axis=1) if mode == "all" else in_range.any(axis=1)

        else:

            raise ValueError(f"mode 必須是 'all'、'any' 或 'mean',收到: {mode!r}")


        result = data_df[mask]

        n_before = data_df.shape[0]

        retention = f"{result.shape[0] / n_before * 100:.1f}%" if n_before else "N/A"

        label = f"分位[{lower}, {upper}] (asinh[{lo:.3f}, {hi:.3f}])" if self.use_quantile else f"[{lower}, {upper}]"

        self.logger.info(

            f"asinh 篩選 {label} mode={mode}: "

            f"{n_before} -> {result.shape[0]} 基因 (保留率: {retention})"

        )

        return result


    def compare_modes(self, expr_matrix: pd.DataFrame, special_genes: List[str], lower, upper) -> pd.DataFrame:

        self.logger.info("展示 asinh 三種篩選模式差異...")

        records = []

        n_total = expr_matrix.shape[0]

        for mode in ("mean", "all", "any"):

            filtered = self.filter(expr_matrix, lower=lower, upper=upper, mode=mode)

            kept_special = [g for g in special_genes if g in filtered.index]

            records.append({

                "mode": mode,

                "n_genes": filtered.shape[0],

                "retention_rate": f"{filtered.shape[0] / n_total * 100:.1f}%" if n_total else "N/A",

                "kept_special_genes": ", ".join(kept_special) if kept_special else "None"

            })

        return pd.DataFrame(records)



class PhenotypeFeatureSelector:

    """表型引導特徵篩選器。


    只負責「表型趨勢」篩選(Delta / Spearman rho),不重複套用 asinh 絕對數值

    區間篩選 —— 該篩選已由 AsinhFilter 在 pipeline 更早的步驟完成一次。

    """


    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log


    def select(self, data_df: pd.DataFrame, min_delta_range=None, min_rho=None, top_n=None) -> FeatureSelectionResult:

        min_delta_range = self.config.min_delta_range if min_delta_range is None else min_delta_range

        min_rho = self.config.min_rho if min_rho is None else min_rho

        top_n = self.config.top_n_features if top_n is None else top_n


        self.logger.info("開始表型引導特徵篩選...")

        deltas = self._compute_deltas(data_df)

        results_df = self._compute_correlations(deltas)

        pos_candidates, neg_candidates = self._select_candidates(results_df, min_rho, min_delta_range, top_n)


        self.logger.info(f"篩選完成: {len(pos_candidates)} 正相關, {len(neg_candidates)} 負相關")

        return FeatureSelectionResult(positive_genes=pos_candidates, negative_genes=neg_candidates, full_report=results_df)


    def _compute_deltas(self, data_df: pd.DataFrame) -> pd.DataFrame:

        deltas = pd.DataFrame(index=data_df.index)

        for pheno in self.config.phenotype_order:

            at_col, mt_col = f"{pheno}AT", f"{pheno}MT"

            if at_col in data_df.columns and mt_col in data_df.columns:

                deltas[f"Delta_{pheno}"] = data_df[at_col] - data_df[mt_col]

            else:

                self.logger.warning(f"缺少 {at_col}/{mt_col},無法計算 Delta_{pheno}")

        return deltas


    def _compute_correlations(self, deltas: pd.DataFrame) -> pd.DataFrame:

        delta_cols = [c for c in self.config.delta_columns if c in deltas.columns]

        target_ranks = list(self.config.phenotype_ranks)[:len(delta_cols)]


        if len(delta_cols) != len(self.config.phenotype_ranks):

            self.logger.warning("Delta 欄位數與表型等級數不匹配,已依可用欄位對齊等級")


        if not delta_cols:

            self.logger.error("沒有可用的 Delta 欄位,無法計算相關性")

            return pd.DataFrame(columns=["Spearman_rho", "Delta_Span"]).rename_axis("Feature")


        records = []

        for idx, row in deltas.iterrows():

            sample_deltas = [row[c] for c in delta_cols]

            if len(set(sample_deltas)) == 1:

                rho = 0.0

            else:

                rho, _ = spearmanr(sample_deltas, target_ranks)

                if np.isnan(rho):

                    rho = 0.0

            delta_span = max(sample_deltas) - min(sample_deltas)

            record = {"Feature": idx, "Spearman_rho": rho, "Delta_Span": delta_span}

            record.update(dict(zip(delta_cols, sample_deltas)))

            records.append(record)


        return pd.DataFrame(records).set_index("Feature")


    def _select_candidates(self, res_df, min_rho, min_delta_range, top_n) -> Tuple[pd.DataFrame, pd.DataFrame]:

        if res_df.empty:

            return res_df.copy(), res_df.copy()


        pos = res_df[

            (res_df["Spearman_rho"] >= min_rho) & (res_df["Delta_Span"] >= min_delta_range)

        ].sort_values("Delta_Span", ascending=False)


        neg = res_df[

            (res_df["Spearman_rho"] <= -min_rho) & (res_df["Delta_Span"] >= min_delta_range)

        ].sort_values("Delta_Span", ascending=False)


        if top_n is not None:

            pos = pos.head(top_n)

            neg = neg.head(top_n)

        return pos, neg



class EnrichmentEngine:

    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log


    def analyze(self, gene_list: List[str], group_name: str = "Unnamed") -> EnrichmentResult:

        cleaned = self._clean_gene_list(gene_list)

        if not cleaned:

            self.logger.error(f"[{group_name}] 基因清單為空")

            return EnrichmentResult(None, [], group_name)


        self.logger.info(f"[{group_name}] 富集分析: {len(cleaned)} 基因")


        df = self._try_local_gmt(cleaned)

        if df is not None:

            return EnrichmentResult(df, cleaned, group_name)


        df = self._try_online(cleaned)

        if df is not None:

            return EnrichmentResult(df, cleaned, group_name)


        self.logger.info(f"[{group_name}] 使用 mock 結果演示")

        df = self._mock_results(cleaned)

        return EnrichmentResult(df, cleaned, group_name)


    def _clean_gene_list(self, gene_list: List[str]) -> List[str]:

        cleaned = []

        for g in gene_list:

            if pd.notna(g):

                g_str = str(g).strip()

                if g_str and g_str.lower() != 'nan':

                    cleaned.append(g_str)

        return sorted(set(cleaned))


    def _try_local_gmt(self, gene_list: List[str]) -> Optional[pd.DataFrame]:

        if not GSEAPY_AVAILABLE:

            return None

        gmt_path = self.config.default_gmt_path

        if not gmt_path or not os.path.exists(gmt_path):

            return None

        try:

            enr = gp.enrichr(gene_list=gene_list, gene_sets=gmt_path, outdir=None)

            if enr is not None and enr.results is not None and not enr.results.empty:

                self.logger.info("本地 GMT 富集分析成功")

                return enr.results

        except Exception as e:

            self.logger.warning(f"本地 GMT 失敗: {e}")

        return None


    def _try_online(self, gene_list: List[str]) -> Optional[pd.DataFrame]:

        if not GSEAPY_AVAILABLE:

            return None

        try:

            enr = gp.enrichr(

                gene_list=gene_list,

                gene_sets=['GO_Biological_Process_2023', 'KEGG_2021_Human'],

                outdir=None

            )

            if enr is not None and enr.results is not None and not enr.results.empty:

                self.logger.info("線上 Enrichr 分析成功")

                return enr.results

        except Exception as e:

            self.logger.warning(f"線上 API 失敗: {e}")

        return None




class VisualizationSystem:

    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log

        self._setup_style()


    def _setup_style(self):

        sns.set_theme(style="white", font_scale=1.0)

        plt.rcParams['font.sans-serif'] = ['DejaVu Sans', 'Arial', 'PingFang TC', 'Microsoft JhengHei']

        plt.rcParams['axes.unicode_minus'] = False

        plt.rcParams['figure.dpi'] = 100


    def _save_fig(self, fig: plt.Figure, filename: str) -> str:

        os.makedirs(self.config.output_dir, exist_ok=True)

        path = os.path.join(self.config.output_dir, f"{filename}.{self.config.fig_format}")

        fig.savefig(path, dpi=self.config.dpi, bbox_inches='tight', pad_inches=0.3)

        self.logger.info(f"圖片已儲存: {path}")

        plt.show(fig)

        return path


    def plot_chord_cnet(self, enrich_df, gene_directions: Dict[str, str], title_prefix="Combined",

                         top_n=5, max_genes=30, filename="cnetplot_chord") -> Optional[str]:

        """繪製 Chord Cnetplot(通路-基因網路圖)。


        Args:

            max_genes: 弧形上最多顯示的基因數。基因太多會讓標籤重疊、難以閱讀,

                因此當候選基因數超過此值時,只保留「連結最多通路」的前

                max_genes 個基因(即最具代表性的樞紐基因),其餘基因仍計入

                通路統計但不畫在圖上。

        """

        if enrich_df is None or enrich_df.empty:

            self.logger.warning(f"[{title_prefix}] 富集結果為空,跳過 Chord 圖")

            return None


        top_pathways = enrich_df.head(top_n).copy()

        pathway_genes_map: Dict[str, List[str]] = {}

        gene_degree: Dict[str, int] = {}


        for _, row in top_pathways.iterrows():

            term = row['Term']

            genes = [g.strip() for g in str(row['Genes']).split(';') if g.strip()]

            pathway_genes_map[term] = genes

            for g in genes:

                gene_degree[g] = gene_degree.get(g, 0) + 1


        pathways = list(pathway_genes_map.keys())


        if not pathways or not gene_degree:

            self.logger.warning("未解析出有效關係")

            return None


        n_all_genes = len(gene_degree)

        if n_all_genes > max_genes:

            # 依「連結的通路數」由多到少排序,保留最具代表性(樞紐)的基因;

            # 通路數相同則依字母排序,確保結果穩定可重現。

            ranked_genes = sorted(gene_degree.items(), key=lambda kv: (-kv[1], kv[0]))

            genes_list = sorted(g for g, _ in ranked_genes[:max_genes])

            self.logger.info(

                f"[{title_prefix}] 基因數 {n_all_genes} 超過顯示上限 {max_genes},"

                f"僅保留連結最多通路的 {len(genes_list)} 個基因"

            )

            # 只保留有畫出來的基因的連線,避免通路節點連到不存在的基因節點

            kept_set = set(genes_list)

            pathway_genes_map = {

                term: [g for g in genes if g in kept_set]

                for term, genes in pathway_genes_map.items()

            }

        else:

            genes_list = sorted(gene_degree.keys())


        fig, ax = plt.subplots(figsize=(14, 14))

        ax.set_aspect("equal")

        ax.axis("off")

        ax.set_xlim(-1.4, 1.4)

        ax.set_ylim(-1.4, 1.55)


        radius = 1.0

        n_pathways, n_genes = len(pathways), len(genes_list)


        # 基因數愈少,標籤與節點可以愈大,維持圖面可讀性

        gene_marker_size = 160 if n_genes <= 20 else max(70, 160 - (n_genes - 20) * 3)

        gene_font_size = 10 if n_genes <= 20 else max(6.5, 10 - (n_genes - 20) * 0.12)


        pathway_angles = np.linspace(0.20 * np.pi, 0.80 * np.pi, n_pathways) if n_pathways > 1 else [np.pi / 2]

        gene_angles = np.linspace(1.20 * np.pi, 1.80 * np.pi, n_genes) if n_genes > 1 else [1.5 * np.pi]


        node_coords: Dict[str, Tuple[float, float, float]] = {}

        for i, term in enumerate(pathways):

            angle = pathway_angles[i]

            node_coords[term] = (radius * np.cos(angle), radius * np.sin(angle), angle)

        for i, gene in enumerate(genes_list):

            angle = gene_angles[i]

            node_coords[gene] = (radius * np.cos(angle), radius * np.sin(angle), angle)


        # 使用局部別名,避免與 pathlib.Path 混淆

        MPath = mpath.Path

        for term, genes in pathway_genes_map.items():

            px, py, _ = node_coords[term]

            for gene in genes:

                if gene not in node_coords:

                    continue

                gx, gy, _ = node_coords[gene]

                direction = gene_directions.get(gene, "Positive")

                chord_color = "#E64B35" if direction == "Positive" else "#4DBBD5"


                path_data = [

                    (MPath.MOVETO, (px, py)),

                    (MPath.CURVE3, (0.0, 0.0)),

                    (MPath.CURVE3, (gx, gy))

                ]

                codes, verts = zip(*path_data)

                patch = mpatches.PathPatch(

                    MPath(verts, codes), facecolor="none",

                    edgecolor=chord_color, alpha=0.30, linewidth=1.4

                )

                ax.add_patch(patch)


        def _label(x, y, angle, text, **kwargs):

            deg = np.degrees(angle) % 360

            ha = "right" if 90 < deg < 270 else "left"

            rot_deg = deg if ha == "left" else deg - 180

            ax.text(x, y, text, ha=ha, va="center", rotation=rot_deg, rotation_mode="anchor", **kwargs)


        for term in pathways:

            x, y, angle = node_coords[term]

            ax.scatter(x, y, s=380, color="#00A087", zorder=4, edgecolors="black", linewidth=1.2)

            wrapped = "\n".join(textwrap.wrap(term.replace('_', ' '), width=26))

            _label(x * 1.18, y * 1.18, angle, wrapped, fontsize=9.5, fontweight="bold")


        for gene in genes_list:

            x, y, angle = node_coords[gene]

            direction = gene_directions.get(gene, "Positive")

            color = "#E64B35" if direction == "Positive" else "#4DBBD5"

            ax.scatter(x, y, s=gene_marker_size, color=color, zorder=4, edgecolors="black", linewidth=0.8)

            _label(x * 1.12, y * 1.12, angle, gene, fontsize=gene_font_size, fontweight="bold", color=color)


        handles = [mpatches.Patch(color="#00A087", label="Enriched Pathway / GO Term")]

        if any(gene_directions.get(g) == "Positive" for g in genes_list):

            handles.append(mpatches.Patch(color="#E64B35", label="Positive (Pro-development)"))

        if any(gene_directions.get(g) == "Negative" for g in genes_list):

            handles.append(mpatches.Patch(color="#4DBBD5", label="Negative (Inhibitor)"))


        ax.legend(handles=handles, loc="lower center", bbox_to_anchor=(0.5, -0.12),

                  ncol=len(handles), frameon=True, fontsize=10, title="Category & Regulation")

        ax.set_title(f"{title_prefix} Pathways-Genes Network", fontsize=18, fontweight="bold", pad=40, y=1.15)


        return self._save_fig(fig, filename)


    def plot_volcano(self, feature_result: FeatureSelectionResult, filename="volcano_plot") -> Optional[str]:

        df = feature_result.full_report.copy()

        if df.empty:

            self.logger.warning("特徵報告為空,跳過火山圖")

            return None


        fig, ax = plt.subplots(figsize=(10, 8))


        pos_idx = set(feature_result.positive_genes.index)

        neg_idx = set(feature_result.negative_genes.index)

        is_pos = df.index.to_series().isin(pos_idx)

        is_neg = df.index.to_series().isin(neg_idx)

        is_other = ~(is_pos | is_neg)


        ax.scatter(df.loc[is_other, "Spearman_rho"], df.loc[is_other, "Delta_Span"],

                   c="gray", alpha=0.3, s=20, label="Non-significant")

        ax.scatter(df.loc[is_pos, "Spearman_rho"], df.loc[is_pos, "Delta_Span"],

                   c="#E64B35", alpha=0.7, s=60, label="Positive trend")

        ax.scatter(df.loc[is_neg, "Spearman_rho"], df.loc[is_neg, "Delta_Span"],

                   c="#4DBBD5", alpha=0.7, s=60, label="Negative trend")


        ax.axvline(x=self.config.min_rho, color='red', linestyle='--', alpha=0.5)

        ax.axvline(x=-self.config.min_rho, color='blue', linestyle='--', alpha=0.5)

        ax.axhline(y=self.config.min_delta_range, color='green', linestyle='--', alpha=0.5)


        ax.set_xlabel("Spearman rho", fontsize=12)

        ax.set_ylabel("Delta Span", fontsize=12)

        ax.set_title("Phenotype-guided Feature Selection", fontsize=14, fontweight='bold')

        ax.legend(loc='upper left')

        ax.grid(True, alpha=0.3)


        # 只標註實際入選的正/負候選基因,並依 Delta_Span 取前 20 名,避免圖面過度擁擠

        labeled = df.loc[is_pos | is_neg].sort_values("Delta_Span", ascending=False).head(20)

        texts = [ax.text(row["Spearman_rho"], row["Delta_Span"], idx, fontsize=8, alpha=0.8)

                 for idx, row in labeled.iterrows()]


        if ADJUSTTEXT_AVAILABLE and texts:

            adjust_text(texts, arrowprops=dict(arrowstyle='->', color='gray', alpha=0.5))


        return self._save_fig(fig, filename)


    def plot_regulation_barplot(

        self,

        feature_result: FeatureSelectionResult,

        direction: str,

        top_n: int = 20,

        filename: str = "regulation_barplot"

    ) -> Optional[str]:

        """繪製「上調」或「下調」候選基因的長條圖(只畫單一方向,兩方向各自獨立成圖)。


        Args:

            direction: "positive"(上調,Delta AT>MT 且與表型等級正相關)

                       或 "negative"(下調,負相關)。

            top_n: 依 Delta_Span 由大到小,最多顯示幾個基因。

        """

        if direction not in ("positive", "negative"):

            raise ValueError(f"direction 必須是 'positive' 或 'negative',收到: {direction!r}")


        candidates = feature_result.positive_genes if direction == "positive" else feature_result.negative_genes

        label = "Up-regulated" if direction == "positive" else "Down-regulated"

        color = "#E64B35" if direction == "positive" else "#4DBBD5"


        if candidates.empty:

            self.logger.warning(f"[{label}] 無候選基因,跳過長條圖")

            return None


        plot_df = candidates.sort_values("Delta_Span", ascending=False).head(top_n)


        fig, ax = plt.subplots(figsize=(10, max(6, len(plot_df) * 0.35)))

        bars = ax.barh(range(len(plot_df)), plot_df["Delta_Span"], color=color, alpha=0.85,

                        edgecolor="black", linewidth=0.5)

        ax.set_yticks(range(len(plot_df)))

        ax.set_yticklabels(plot_df.index, fontsize=9)

        ax.set_xlabel("Delta Span (asinh scale, |AT - MT| range across phenotypes)")

        ax.set_title(

            f"{label} Candidate Genes (Top {len(plot_df)} by Delta Span)",

            fontsize=14, fontweight='bold'

        )

        ax.invert_yaxis()

        ax.grid(True, axis='x', alpha=0.3)


        for i, val in enumerate(plot_df["Delta_Span"]):

            ax.text(val, i, f" {val:.2f}", va='center', fontsize=8)


        return self._save_fig(fig, filename)


    def plot_heatmap(self, expr_matrix: pd.DataFrame, feature_result: FeatureSelectionResult, filename="heatmap") -> Optional[str]:

        selected_genes = list(feature_result.positive_genes.index) + list(feature_result.negative_genes.index)

        if not selected_genes:

            self.logger.warning("無候選基因,跳過熱力圖")

            return None


        if len(selected_genes) > 50:

            selected_genes = selected_genes[:50]

            self.logger.info("熱力圖限制顯示前 50 個基因")


        missing = [g for g in selected_genes if g not in expr_matrix.index]

        if missing:

            self.logger.warning(f"{len(missing)} 個候選基因不在表達矩陣中(可能已被 asinh 篩選排除),將略過")

            selected_genes = [g for g in selected_genes if g in expr_matrix.index]

        if not selected_genes:

            self.logger.warning("篩選後無可繪製之候選基因,跳過熱力圖")

            return None


        plot_data = expr_matrix.loc[selected_genes]


        row_colors = ["#E64B35" if g in feature_result.positive_genes.index else "#4DBBD5"

                      for g in selected_genes]


        fig, ax = plt.subplots(figsize=(10, max(6, len(selected_genes) * 0.3)))

        sns.heatmap(plot_data, cmap="RdBu_r", center=plot_data.values.mean(),

                    annot=False, fmt=".2f", linewidths=0.5,

                    cbar_kws={"label": "asinh(expression)"}, ax=ax)


        for i, color in enumerate(row_colors):

            ax.add_patch(plt.Rectangle((-0.5, i), 0.1, 1, facecolor=color,

                         transform=ax.get_yaxis_transform(), clip_on=False))


        ax.set_title("Candidate Gene Expression Heatmap", fontsize=14, fontweight='bold')

        ax.set_xlabel("Samples")

        ax.set_ylabel("Genes")


        legend_elements = [

            mpatches.Patch(facecolor="#E64B35", label="Positive trend"),

            mpatches.Patch(facecolor="#4DBBD5", label="Negative trend")

        ]

        ax.legend(handles=legend_elements, loc='upper left', bbox_to_anchor=(1.15, 1))


        return self._save_fig(fig, filename)


    def plot_pca(self, expr_matrix: pd.DataFrame, filename="pca_plot") -> Optional[str]:

        if expr_matrix.shape[1] < 2:

            self.logger.warning("樣本數不足以計算 PCA (需要至少 2 個樣本),跳過 PCA 圖")

            return None


        data = expr_matrix.T

        data.columns = data.columns.astype(str)


        scaler = StandardScaler()

        data_scaled = scaler.fit_transform(data)


        n_components = min(2, data_scaled.shape[0] - 1, data_scaled.shape[1])

        if n_components < 2:

            self.logger.warning("樣本或特徵數不足以計算 2 個主成分,跳過 PCA 圖")

            return None


        pca = PCA(n_components=2)

        pcs = pca.fit_transform(data_scaled)


        sample_ids = self.config.sample_ids

        treatments = ['MT' if t == 0 else 'AT' for t in self.config.is_treated]


        fig, ax = plt.subplots(figsize=(10, 8))

        colors = {"MT": "#4DBBD5", "AT": "#E64B35"}

        markers = {0: 'o', 1: 's', 2: '^'}


        for i, (pc1, pc2) in enumerate(pcs):

            marker = markers.get(self.config.individual[i], 'D')

            ax.scatter(pc1, pc2, c=colors[treatments[i]], marker=marker,

                      s=200, alpha=0.7, edgecolors='black')

            ax.annotate(sample_ids[i], (pc1, pc2), xytext=(5, 5),

                       textcoords='offset points', fontsize=9)


        ax.set_xlabel(f"PC1 ({pca.explained_variance_ratio_[0]*100:.1f}%)")

        ax.set_ylabel(f"PC2 ({pca.explained_variance_ratio_[1]*100:.1f}%)")

        ax.set_title("PCA of Samples", fontsize=14, fontweight='bold')

        ax.grid(True, alpha=0.3)

        ax.axhline(y=0, color='gray', linestyle='-', alpha=0.3)

        ax.axvline(x=0, color='gray', linestyle='-', alpha=0.3)


        from matplotlib.lines import Line2D

        legend_elements = [

            Line2D([0], [0], marker='o', color='w', markerfacecolor="#4DBBD5", markersize=10, label='MT (Untreated)'),

            Line2D([0], [0], marker='o', color='w', markerfacecolor="#E64B35", markersize=10, label='AT (Treated)')

        ]

        ax.legend(handles=legend_elements, loc='best')


        return self._save_fig(fig, filename)


    def plot_enrichment_bar(self, enrich_result: EnrichmentResult, filename="enrichment_bar") -> Optional[str]:

        if not enrich_result.success:

            return None


        df = enrich_result.top_terms(15).copy()

        fig, ax = plt.subplots(figsize=(10, max(6, len(df) * 0.4)))


        safe_p = df['P-value'].clip(lower=1e-300)

        df['neg_log_p'] = -np.log10(safe_p)

        colors = ["#E64B35" if 'GO' in gs else "#00A087" for gs in df['Gene_set']]


        bars = ax.barh(range(len(df)), df['neg_log_p'], color=colors, alpha=0.8)

        ax.set_yticks(range(len(df)))

        ax.set_yticklabels(df['Term'], fontsize=9)

        ax.set_xlabel("-log10(P-value)")

        ax.set_title(f"{enrich_result.group_name} - Top Enriched Terms", fontsize=14, fontweight='bold')

        ax.axvline(x=-np.log10(0.05), color='red', linestyle='--', alpha=0.5, label='p=0.05')

        ax.legend()

        ax.invert_yaxis()


        for i, (bar, val) in enumerate(zip(bars, df['neg_log_p'])):

            ax.text(val + 0.1, i, f"{val:.2f}", va='center', fontsize=8)


        return self._save_fig(fig, filename)



class ReportGenerator:

    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log


    def generate(self, feature_result: FeatureSelectionResult, enrichment_results: Dict[str, EnrichmentResult], image_paths: Dict[str, str]) -> str:

        os.makedirs(self.config.output_dir, exist_ok=True)

        report_path = os.path.join(self.config.output_dir, "analysis_report.html")

        html = self._build_html(feature_result, enrichment_results, image_paths)

        with open(report_path, 'w', encoding='utf-8') as f:

            f.write(html)

        self.logger.info(f"報告已生成: {report_path}")

        return report_path


    def _build_html(self, feature_result, enrichment_results, image_paths) -> str:

        summary = feature_result.summary()


        img_tags = ""

        for name, path in image_paths.items():

            if path and os.path.exists(path):

                rel_path = os.path.relpath(path, self.config.output_dir)

                img_tags += f'<h3>{name}</h3><img src="{rel_path}" style="max-width:100%;"/><br><br>'


        enrich_tables = ""

        for group_name, result in enrichment_results.items():

            if result.success:

                table_html = result.dataframe.head(10).to_html(classes='table table-striped', index=False)

                enrich_tables += f"<h3>{group_name} Enrichment</h3>{table_html}<br>"


        html = f"""

<!DOCTYPE html>

<html>

<head>

    <meta charset="UTF-8">

    <title>Ovary Analysis Report</title>

    <style>

        body {{ font-family: Arial, sans-serif; margin: 40px; background: #f5f5f5; }}

        .container {{ max-width: 1200px; margin: 0 auto; background: white; padding: 30px; box-shadow: 0 0 10px rgba(0,0,0,0.1); }}

        h1 {{ color: #2c3e50; border-bottom: 3px solid #3498db; padding-bottom: 10px; }}

        h2 {{ color: #34495e; margin-top: 30px; }}

        h3 {{ color: #7f8c8d; }}

        .summary {{ background: #ecf0f1; padding: 20px; border-radius: 8px; margin: 20px 0; }}

        .summary-item {{ margin: 8px 0; }}

        table {{ border-collapse: collapse; width: 100%; margin: 15px 0; }}

        th, td {{ border: 1px solid #ddd; padding: 8px; text-align: left; }}

        th {{ background: #3498db; color: white; }}

        tr:nth-child(even) {{ background: #f9f9f9; }}

        img {{ border: 1px solid #ddd; border-radius: 4px; margin: 10px 0; }}

        .footer {{ margin-top: 40px; color: #95a5a6; font-size: 0.9em; text-align: center; }}

    </style>

</head>

<body>

    <div class="container">

        <h1>Ovary Tissue Transcriptome Analysis Report</h1>

        <p><strong>Generated:</strong> {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}</p>


        <h2>Analysis Summary</h2>

        <div class="summary">

            <div class="summary-item"><strong>Positive Candidates:</strong> {summary['n_positive']}</div>

            <div class="summary-item"><strong>Negative Candidates:</strong> {summary['n_negative']}</div>

            <div class="summary-item"><strong>Total Genes Tested:</strong> {summary['n_total_tested']}</div>

            <div class="summary-item"><strong>Top Positive:</strong> {', '.join(summary['top_positive'][:5])}</div>

            <div class="summary-item"><strong>Top Negative:</strong> {', '.join(summary['top_negative'][:5])}</div>

        </div>


        <h2>Visualizations</h2>

        {img_tags}


        <h2>Enrichment Analysis</h2>

        {enrich_tables}


        <div class="footer">

            <p>Generated by OvaryAnalysis Pipeline v3 | Configuration: {self.config.phenotype_order}</p>

        </div>

    </div>

</body>

</html>

        """

        return html



class ResultExporter:

    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log


    def export_feature_result(self, result: FeatureSelectionResult, prefix="features") -> Dict[str, str]:

        os.makedirs(self.config.output_dir, exist_ok=True)

        paths = {}


        excel_path = os.path.join(self.config.output_dir, f"{prefix}.xlsx")

        with pd.ExcelWriter(excel_path, engine='openpyxl') as writer:

            result.positive_genes.to_excel(writer, sheet_name='Positive')

            result.negative_genes.to_excel(writer, sheet_name='Negative')

            result.full_report.to_excel(writer, sheet_name='Full_Report')

        paths['excel'] = excel_path

        self.logger.info(f"特徵結果已匯出: {excel_path}")


        for name, df in [('positive', result.positive_genes), ('negative', result.negative_genes)]:

            if not df.empty:

                csv_path = os.path.join(self.config.output_dir, f"{prefix}_{name}.csv")

                df.to_csv(csv_path)

                paths[f'csv_{name}'] = csv_path


        return paths


    def export_enrichment_result(self, result: EnrichmentResult, prefix="enrichment") -> Optional[str]:

        if not result.success:

            return None

        os.makedirs(self.config.output_dir, exist_ok=True)

        path = os.path.join(self.config.output_dir, f"{prefix}_{result.group_name}.csv")

        result.dataframe.to_csv(path, index=False)

        self.logger.info(f"富集結果已匯出: {path}")

        return path



class AnalysisPipeline:

    def __init__(self, config: AnalysisConfig, log: logging.Logger):

        self.config = config

        self.logger = log

        self.data_loader = DataLoader(config, log)

        self.asinh_filter = AsinhFilter(log, use_quantile=config.asinh_use_quantile)

        self.feature_selector = PhenotypeFeatureSelector(config, log)

        self.enrichment_engine = EnrichmentEngine(config, log)

        self.visualizer = VisualizationSystem(config, log)

        self.report_generator = ReportGenerator(config, log)

        self.exporter = ResultExporter(config, log)

        self.results: Dict[str, Any] = {}

        self.image_paths: Dict[str, str] = {}


    def run(self, data_path: Optional[str] = None) -> Dict[str, Any]:

        self.logger.info("=" * 60)

        self.logger.info("開始卵巢組織轉錄組分析流程 v3")

        self.logger.info("=" * 60)


        try:

            expr_matrix = self._step_load_data(data_path)

            expr_filtered = self._step_asinh_filter(expr_matrix)

            self._step_global_enrichment(expr_filtered)

            feature_result = self._step_feature_selection(expr_filtered)

            self._step_visualization(expr_matrix, expr_filtered, feature_result)

            enrichment_results = self._step_group_enrichment(feature_result)

            self._step_export(feature_result, enrichment_results)

            report_path = self._step_generate_report(feature_result, enrichment_results)


            self.logger.info("=" * 60)

            self.logger.info("分析流程完成")

            self.logger.info(f"輸出目錄: {self.config.output_dir}")

            self.logger.info("=" * 60)


            return {

                "success": True,

                "output_dir": self.config.output_dir,

                "report_path": report_path,

                "feature_result": feature_result,

                "enrichment_results": enrichment_results,

                "image_paths": self.image_paths,

            }

        except Exception as e:

            self.logger.error(f"分析流程失敗: {e}", exc_info=True)

            return {"success": False, "error": str(e)}


    def _step_load_data(self, data_path: Optional[str]) -> pd.DataFrame:

        self.logger.info("[Step 1/8] 載入數據...")

        expr_matrix = self.data_loader.load(data_path)

        self.logger.info(f"數據載入完成: {expr_matrix.shape}")

        return expr_matrix


    def _step_asinh_filter(self, expr_matrix: pd.DataFrame) -> pd.DataFrame:

        self.logger.info("[Step 2/8] asinh 數值區間篩選...")

        special_genes = ["CYP19A1", "STAR", "LHCGR", "MAPK1", "CDK1", "CCNB1", "TGFBR1", "CASP3", "FOXO1"]

        comparison = self.asinh_filter.compare_modes(

            expr_matrix, special_genes, self.config.asinh_lower, self.config.asinh_upper

        )

        print("\n========== asinh 篩選模式比較 ==========")

        print(comparison.to_string(index=False))

        print("=========================================\n")


        expr_filtered = self.asinh_filter.filter(

            expr_matrix, lower=self.config.asinh_lower,

            upper=self.config.asinh_upper, mode=self.config.asinh_filter_mode

        )

        self.results['asinh_comparison'] = comparison

        return expr_filtered


    def _step_global_enrichment(self, expr_filtered: pd.DataFrame) -> None:

        self.logger.info("[Step 3/8] 全域富集分析...")

        gene_list = expr_filtered.index.tolist()

        result = self.enrichment_engine.analyze(gene_list, "Global")

        self.results['global_enrichment'] = result

        if result.success:

            print("\n========== 全域富集分析結果 ==========")

            print(result.dataframe.head(10).to_string())

            print("=======================================\n")


    def _step_feature_selection(self, expr_filtered: pd.DataFrame) -> FeatureSelectionResult:

        """步驟 4: 表型引導特徵篩選。


        注意:此步驟只在 asinh 篩選後的資料 (expr_filtered) 上計算 Delta 與

        Spearman rho,不再依平均表現值做二次篩選,避免與 Step 2 的 asinh

        篩選重複套用不同標準而造成篩選結果難以解讀。

        """

        self.logger.info("[Step 4/8] 表型引導特徵篩選...")


        result = self.feature_selector.select(

            expr_filtered,

            min_delta_range=self.config.min_delta_range,

            min_rho=self.config.min_rho,

            top_n=self.config.top_n_features,

        )


        print("\n========== 特徵篩選結果 ==========")

        print(f"正相關候選: {len(result.positive_genes)}")

        print(f"負相關候選: {len(result.negative_genes)}")

        if not result.positive_genes.empty:

            print("\nTop 5 正相關基因:")

            print(result.positive_genes.head(5).to_string())

        if not result.negative_genes.empty:

            print("\nTop 5 負相關基因:")

            print(result.negative_genes.head(5).to_string())

        print("===================================\n")


        return result


    def _step_visualization(self, expr_matrix, expr_filtered, feature_result) -> None:

        self.logger.info("[Step 5/8] 生成可視化圖表...")


        path = self.visualizer.plot_pca(expr_matrix, "pca_plot")

        if path:

            self.image_paths['PCA'] = path


        path = self.visualizer.plot_volcano(feature_result, "volcano_plot")

        if path:

            self.image_paths['Volcano'] = path


        path = self.visualizer.plot_heatmap(expr_filtered, feature_result, "heatmap")

        if path:

            self.image_paths['Heatmap'] = path


        # 上調 / 下調候選基因分別各自畫一張長條圖,方便單獨檢視每個方向的結果

        path = self.visualizer.plot_regulation_barplot(feature_result, "positive", filename="barplot_upregulated")

        if path:

            self.image_paths['Barplot_Upregulated'] = path


        path = self.visualizer.plot_regulation_barplot(feature_result, "negative", filename="barplot_downregulated")

        if path:

            self.image_paths['Barplot_Downregulated'] = path


    def _step_group_enrichment(self, feature_result: FeatureSelectionResult) -> Dict[str, EnrichmentResult]:

        self.logger.info("[Step 6/8] 分組富集分析...")


        pos_list = feature_result.positive_genes.index.tolist()

        neg_list = feature_result.negative_genes.index.tolist()


        gene_directions = {g: "Positive" for g in pos_list}

        gene_directions.update({g: "Negative" for g in neg_list})


        enrichment_results = {}

        groups = [

            ("Up-regulated", pos_list, "cnetplot_up"),

            ("Down-regulated", neg_list, "cnetplot_down"),

        ]


        for group_name, genes, filename in groups:

            if not genes:

                self.logger.warning(f"[{group_name}] 無基因,跳過")

                continue


            result = self.enrichment_engine.analyze(genes, group_name)

            enrichment_results[group_name] = result


            if result.success:

                print(f"\n========== {group_name} 富集分析結果 ==========")

                print(result.dataframe.head(10).to_string())

                print("============================================\n")


                path = self.visualizer.plot_chord_cnet(

                    result.dataframe, gene_directions, group_name, top_n=10, filename=filename

                )

                if path:

                    self.image_paths[f'Chord_{group_name}'] = path


                path = self.visualizer.plot_enrichment_bar(

                    result, f"enrichment_bar_{group_name.lower().replace('-', '_')}"

                )

                if path:

                    self.image_paths[f'Bar_{group_name}'] = path


        combined_genes = list(set(pos_list + neg_list))

        if combined_genes:

            combined_result = self.enrichment_engine.analyze(combined_genes, "Combined")

            enrichment_results["Combined"] = combined_result


            if combined_result.success:

                path = self.visualizer.plot_chord_cnet(

                    combined_result.dataframe, gene_directions, "Combined", top_n=10, filename="cnetplot_combined"

                )

                if path:

                    self.image_paths['Chord_Combined'] = path


        return enrichment_results


    def _step_export(self, feature_result: FeatureSelectionResult, enrichment_results: Dict[str, EnrichmentResult]) -> None:

        self.logger.info("[Step 7/8] 匯出結果...")

        self.exporter.export_feature_result(feature_result, "features")

        for group_name, result in enrichment_results.items():

            self.exporter.export_enrichment_result(result, "enrichment")


    def _step_generate_report(self, feature_result, enrichment_results) -> Optional[str]:

        self.logger.info("[Step 8/8] 生成 HTML 報告...")

        return self.report_generator.generate(feature_result, enrichment_results, self.image_paths)



def main():

    logger = setup_logging(level=logging.INFO)

    config = AnalysisConfig()


    os.makedirs(config.output_dir, exist_ok=True)

    config.save(os.path.join(config.output_dir, "config.json"))


    pipeline = AnalysisPipeline(config, logger)

    results = pipeline.run()


    if results["success"]:

        print("\n" + "=" * 60)

        print("分析成功完成!")

        print(f"輸出目錄: {results['output_dir']}")

        print(f"報告檔案: {results['report_path']}")

        print("=" * 60)

    else:

        print(f"\n分析失敗: {results.get('error', 'Unknown error')}")

        sys.exit(1)



if __name__ == "__main__":

    main()