fix: F Tab expected dynamic + AGENTS.md encoding + dev-spec sync
This commit is contained in:
@@ -67,7 +67,7 @@ AgentsMeeting/
|
||||
│ ├── scripts/ # 运行时脚本
|
||||
│ │ ├── chat_bridge.py # xxm LLM 桥接(SessionBridge)
|
||||
│ │ ├── session_router.py # 消息路由
|
||||
│ │ ├── wechat_agent.py # ⛔ 已停用(微信桥接,由 Linux Docker bot 替代)
|
||||
│ │ ├── wechat_agent.py # ⛔ 已停用(微信桥接,由 Linux Docker wechatbot-webhook 替代)
|
||||
│ │ │ # 保留代码,不再运行
|
||||
│ │ ├── article_processor.py# 文章处理服务 (5810) ← 保持运行!
|
||||
│ │ │ # 抓微信链接全文、OCR,WeChat bot 和 莫荷 都依赖它
|
||||
|
||||
+83
-83
@@ -1,137 +1,137 @@
|
||||
# AGENTS.md 鈥?鍐崇瓥鑷鍗忚
|
||||
# AGENTS.md — 决策自检协议
|
||||
|
||||
鍩轰簬 Andrej Karpathy 鍥涘ぇ鏍稿績鍘熷垯 + 璇楀厜杩涘寲鐗?+ 瀵规姉鎬у鏌ユ満鍒躲€?
|
||||
鏈枃妗f槸鎵€鏈?Agent锛堣帿鑽枫€亁xm/绗戠瑧銆佺煡寰€佸皬鏋滐級鐨勫叡浜喅绛栫邯寰嬪崗璁€?
|
||||
基于 Andrej Karpathy 四大核心原则 + 诗光进化版 + 对抗性审查机制。
|
||||
本文档是所有 Agent(莫荷、xxm/笑笑、知微、小果)的共享决策纪律协议。
|
||||
|
||||
## 1. 鏍稿績瑙勫垯
|
||||
## 1. 核心规则
|
||||
|
||||
### 1.1 鎬濊€冨厛浜庣紪鐮?
|
||||
### 1.1 思考先于编码
|
||||
|
||||
鎷垮埌闇€姹傚悗锛岀姝㈢洿鎺ュ紑濮嬬紪鐮?璁捐銆傚繀椤讳緷娆℃墽琛岋細
|
||||
拿到需求后,禁止直接开始编码/设计。必须依次执行:
|
||||
|
||||
0. **璇诲搴旀ā鍧楃殑 搂 spec**锛氬鏋滄秹鍙婂凡鏈夊姛鑳芥ā鍧楋紝鍏堣 `gateway/scripts/specs/{module}.json` 鐨?`ai_spec` 浜嗚В瀹為檯鏋舵瀯銆佹帴鍙c€佺害鏉熷拰渚濊禆銆俿pec 鍦?`docs/dev-spec.md` 鐨勬ā鍧楁竻鍗曚腑鍙煡
|
||||
1. **澶嶈堪闇€姹?*锛氱敤鑷繁鐨勮瘽澶嶈堪瀵归渶姹傜殑鐞嗚В锛屽垪鍑轰笉纭畾鐐?
|
||||
2. **鍒楀嚭鍊欓€夋柟妗?*锛氬瓨鍦ㄥ绉嶅悎鐞嗘灦鏋勬柟妗堟椂锛屽垪鍑鸿嚦灏?2 绉嶆柟妗堝苟瀵规瘮浼樺姡
|
||||
3. **閫夋嫨鏈€绠€鏂规**锛氬熀浜庡鏉傚害銆佸伐浣滈噺銆佸彲缁存姢鎬ч€夋嫨鏈€绠€鏂规
|
||||
4. **绛夊緟纭**锛氬湪寰楀埌鐢ㄦ埛/Coordinator 纭鍓嶏紝绂佹寮€濮嬬紪鐮?
|
||||
0. **读对应模块的 § spec**:如果涉及已有功能模块,先读 `gateway/scripts/specs/{module}.json` 的 `ai_spec` 了解实际架构、接口、约束和依赖。spec 在 `docs/dev-spec.md` 的模块清单中可查
|
||||
1. **复述需求**:用自己的话复述对需求的理解,列出不确定点
|
||||
2. **列出候选方案**:存在多种合理架构方案时,列出至少 2 种方案并对比优劣
|
||||
3. **选择最简方案**:基于复杂度、工作量、可维护性选择最简方案
|
||||
4. **等待确认**:在得到用户/Coordinator 确认前,禁止开始编码
|
||||
|
||||
### 1.2 鏋佺畝浼樺厛
|
||||
### 1.2 极简优先
|
||||
|
||||
姣忎竴娆″紩鍏ユ柊鎶借薄/鏂颁腑闂村眰/鏂板伐鍏峰墠锛屽厛闂細"涓嶇敤瀹冭涓嶈锛?
|
||||
每一次引入新抽象/新中间层/新工具前,先问:"不用它行不行?"
|
||||
|
||||
- 鑳界敤鍑芥暟瑙e喅鐨勯棶棰橈紝涓嶇敤绫?
|
||||
- 鑳界敤绫昏В鍐崇殑闂锛屼笉鐢ㄦ鏋?
|
||||
- 涓嶄负"鏈潵鍙兘闇€瑕佺殑鐏垫椿鎬?鍐欎竴琛屼唬鐮?
|
||||
- 涓€娆℃€ц剼鏈笉鍋氬鏉傚皝瑁?
|
||||
- 涓嶄负绾亣璁剧殑鏋佺鍦烘櫙娣诲姞澶嶆潅鍏滃簳
|
||||
- 能用函数解决的问题,不用类
|
||||
- 能用类解决的问题,不用框架
|
||||
- 不为"未来可能需要的灵活性"写一行代码
|
||||
- 一次性脚本不做复杂封装
|
||||
- 不为纯假设的极端场景添加复杂兜底
|
||||
|
||||
### 1.3 绮惧噯淇敼
|
||||
### 1.3 精准修改
|
||||
|
||||
- 淇敼鍓嶅厛璇诲搴旀ā鍧楃殑 搂 spec锛坄gateway/scripts/specs/{module}.json`锛夛紝纭鏋舵瀯鐞嗚В鍜岀害鏉?
|
||||
- 淇敼鍓嶈緭鍑烘墽琛岃鍒掞紙鏀瑰摢鍑犱釜鏂囦欢銆佹敼浠€涔堛€佷负浠€涔堟敼锛?
|
||||
- 鍙敼鍔ㄩ渶姹傜洿鎺ユ秹鍙婄殑鏂囦欢锛屼笉椤烘墜"浼樺寲"鏃犲叧浠g爜
|
||||
- 涓€娆″彧鏀逛竴涓€昏緫鍗曞厓锛岄獙璇侀€氳繃鍐嶆敼涓嬩竴涓?
|
||||
- 涓嶉噸鏋勬病鏈夐棶棰樼殑浠g爜
|
||||
- 鍒犻櫎鏈淇敼瀵艰嚧涓嶅啀浣跨敤鐨勪唬鐮侊紝浣嗕笉鍒犻櫎淇敼鍓嶅凡瀛樺湪鐨勬浠g爜锛堥櫎闈炵敤鎴锋槑纭姹傦級
|
||||
- **鏀瑰畬鍚屾 spec**锛氫慨鏀瑰畬鎴愬悗锛屽皢瀵瑰簲妯″潡鐨?`specs/{module}.json` 鏇存柊涓轰笌瀹為檯瀹炵幇涓€鑷?
|
||||
- 修改前先读对应模块的 § spec(`gateway/scripts/specs/{module}.json`),确认架构理解和约束
|
||||
- 修改前输出执行计划(改哪几个文件、改什么、为什么改)
|
||||
- 只改动需求直接涉及的文件,不顺手"优化"无关代码
|
||||
- 一次只改一个逻辑单元,验证通过再改下一个
|
||||
- 不重构没有问题的代码
|
||||
- 删除本次修改导致不再使用的代码,但不删除修改前已存在的死代码(除非用户明确要求)
|
||||
- **改完同步 spec**:修改完成后,将对应模块的 `specs/{module}.json` 更新为与实际实现一致
|
||||
|
||||
### 1.4 鐩爣瀵煎悜浜や粯
|
||||
### 1.4 目标导向交付
|
||||
|
||||
- 澶氭楠や换鍔″厛鍐欏叆 task 娓呭崟锛屾瘡瀹屾垚涓€姝ユ爣娉ㄨ繘搴?
|
||||
- 鏀瑰姩闄勫甫绠€瑕佽鏄庯紙鏀逛簡浠€涔堛€佷负浠€涔堟敼锛?
|
||||
- 浠诲姟瀹屾垚鍚庣敤 3-5 鍙ヨ瘽鎬荤粨锛氬仛浜嗕粈涔堛€佸鍒颁簡浠€涔堛€佷笅娆℃€庝箞鍋氫笉鍚?
|
||||
- 琚籂姝e悗锛屽湪 docs/learned.md 涓拷鍔犵粡楠屾暀璁褰?
|
||||
- 多步骤任务先写入 task 清单,每完成一步标注进度
|
||||
- 改动附带简要说明(改了什么、为什么改)
|
||||
- 任务完成后用 3-5 句话总结:做了什么、学到了什么、下次怎么做不同
|
||||
- 被纠正后,在 docs/learned.md 中追加经验教训记录
|
||||
|
||||
## 2. 瀵规姉鎬у鏌ユ祦绋?
|
||||
## 2. 对抗性审查流程
|
||||
|
||||
褰?Agent 鍑嗗鎵ц涓€涓秹鍙婃灦鏋勯€夊瀷鐨勪换鍔℃椂锛屽己鍒惰蛋浠ヤ笅妫€鏌ョ偣锛?
|
||||
当 Agent 准备执行一个涉及架构选型的任务时,强制走以下检查点:
|
||||
|
||||
**Step 1**: 鍐欐柟妗堣崏妗堬紙鏈€鐩磋鐨勬柟妗堬級
|
||||
**Step 1**: 写方案草案(最直觉的方案)
|
||||
|
||||
**Step 2**: 瀵规姉鎬у鏌ワ紙鑷棶锛夛細
|
||||
- 杩欎釜鏂规鏄笉鏄垜涔犳儻鎬ч€夋嫨鐨勶紵
|
||||
- 鏈夋病鏈夋洿绠€鍗曠殑鏂规琚垜璺宠繃浜嗭紵
|
||||
- 姣忎竴灞傛娊璞$湡鐨勫繀瑕佸悧锛熷幓鎺変細鎬庢牱锛?
|
||||
- 濡傛灉鏄庡ぉ灏辫浜ゆ帴缁欏叾浠栦汉锛屼粬浼氳寰楄繖涓璁℃槸蹇呰鐨勫悧锛?
|
||||
**Step 2**: 对抗性审查(自问):
|
||||
- 这个方案是不是我习惯性选择的?
|
||||
- 有没有更简单的方案被我跳过了?
|
||||
- 每一层抽象真的必要吗?去掉会怎样?
|
||||
- 如果明天就要交接给其他人,他会觉得这个设计是必要的吗?
|
||||
|
||||
**Step 3**: 鍒楀嚭瀵规瘮鏂规锛堣嚦灏?2 涓級锛屾爣娉ㄥ鏉傚害銆佸伐浣滈噺銆佸彲缁存姢鎬у樊寮?
|
||||
**Step 3**: 列出对比方案(至少 2 个),标注复杂度、工作量、可维护性差异
|
||||
|
||||
**Step 4**: 鎻愪氦缁欑敤鎴?Coordinator 鍐崇瓥
|
||||
**Step 4**: 提交给用户/Coordinator 决策
|
||||
|
||||
**Step 5**: 纭鍚庢墽琛?
|
||||
**Step 5**: 确认后执行
|
||||
|
||||
### 2.1 绗竴鎬у師鐞?+ 濂ュ崱濮嗗墐鍒€
|
||||
### 2.1 第一性原理 + 奥卡姆剃刀
|
||||
|
||||
姣忔鍐崇瓥杩介棶锛?
|
||||
- "杩欎釜澶嶆潅搴︽槸蹇呰鐨勫悧锛?
|
||||
- "鐮嶆帀杩欏眰鎶借薄浼氭€庢牱锛?
|
||||
- "褰撳墠鐨勬柟妗堣В鍐充簡浠€涔堥棶棰橈紵鏄惁鍒涢€犱簡鏂伴棶棰橈紵"
|
||||
每次决策追问:
|
||||
- "这个复杂度是必要的吗?"
|
||||
- "砍掉这层抽象会怎样?"
|
||||
- "当前的方案解决了什么问题?是否创造了新问题?"
|
||||
|
||||
鐢ㄥ墐鍒€鐮嶆帀鎵€鏈変笉蹇呰鐨勬娊璞°€?
|
||||
用剃刀砍掉所有不必要的抽象。
|
||||
|
||||
## 3. 鍙岃褰曟満鍒?
|
||||
## 3. 双记录机制
|
||||
|
||||
### 3.1 鍐崇瓥鏃ュ織 (docs/decisions/)
|
||||
### 3.1 决策日志 (docs/decisions/)
|
||||
|
||||
姣忔鏋舵瀯鍐崇瓥锛堝紩鍏ユ柊缁勪欢銆佸鍔犳娊璞″眰銆佸彉鏇存帴鍙o級璁板綍鍦ㄩ」鐩?`docs/decisions/` 涓嬶細
|
||||
每次架构决策(引入新组件、增加抽象层、变更接口)记录在项目 `docs/decisions/` 下:
|
||||
|
||||
```
|
||||
docs/decisions/YYYY-MM-DD-绠€鐭弿杩?md
|
||||
docs/decisions/YYYY-MM-DD-简短描述.md
|
||||
```
|
||||
|
||||
鏍煎紡锛?
|
||||
格式:
|
||||
```markdown
|
||||
# 鍐崇瓥: [鏍囬]
|
||||
# 决策: [标题]
|
||||
|
||||
## Context
|
||||
涓轰粈涔堥渶瑕佸仛杩欎釜鍐崇瓥锛熷綋鍓嶇姸鎬佹槸浠€涔堬紵
|
||||
为什么需要做这个决策?当前状态是什么?
|
||||
|
||||
## Decision
|
||||
鍋氫簡浠€涔堥€夋嫨锛?
|
||||
做了什么选择?
|
||||
|
||||
## Consequences
|
||||
杩欎釜閫夋嫨鐨勫奖鍝嶅拰鍚庢灉鏄粈涔堬紵
|
||||
这个选择的影响和后果是什么?
|
||||
|
||||
## Alternatives Considered
|
||||
鑰冭檻浜嗗摢浜涙浛浠f柟妗堬紵涓轰粈涔堟病閫夛紵
|
||||
考虑了哪些替代方案?为什么没选?
|
||||
```
|
||||
|
||||
### 3.2 缁忛獙鏁欒 (docs/learned.md)
|
||||
### 3.2 经验教训 (docs/learned.md)
|
||||
|
||||
琚籂姝e悗杩藉姞涓€鏉¤褰曪細
|
||||
被纠正后追加一条记录:
|
||||
|
||||
```markdown
|
||||
- [YYYY-MM-DD] 闂: xxx | 鏍瑰洜: xxx | 姝g‘鍋氭硶: xxx
|
||||
- [YYYY-MM-DD] 问题: xxx | 根因: xxx | 正确做法: xxx
|
||||
```
|
||||
|
||||
**姣忔鏂颁换鍔″墠锛屽繀椤诲厛鎵竴閬?docs/learned.md銆?*
|
||||
**每次新任务前,必须先扫一遍 docs/learned.md。**
|
||||
|
||||
## 4. /rethink 鎸囦护锛堜笂甯濇寜閽級
|
||||
## 4. /rethink 指令(上帝按钮)
|
||||
|
||||
褰撶敤鎴锋垨 Coordinator 鍙戠幇 Agent 鏂瑰悜鍋忎簡鏃讹紝鍙戦€?`/rethink` 鎸囦护銆侫gent 蹇呴』鍦ㄦ敹鍒板悗锛?
|
||||
当用户或 Coordinator 发现 Agent 方向偏了时,发送 `/rethink` 指令。Agent 必须在收到后:
|
||||
|
||||
1. **绔嬪嵆鍋滄鎵€鏈夋搷浣?*
|
||||
2. **杈撳嚭褰撳墠鐞嗚В**锛?鎴戜互涓虹洰鏍囨槸 X锛屾鍦ㄥ仛 Y"
|
||||
3. **杈撳嚭鑷缁撴灉**锛?
|
||||
- 杩欎釜闂鎴戦亣鍒颁簡浠€涔堝洶鎯戯紵
|
||||
- 鎴戠殑鏂规鏄粈涔堬紵
|
||||
- 鏄惁瀛樺湪鏇寸畝鍗曟柟妗堬紵
|
||||
- 鏄惁鏈夐獙璇佽繃姣忎釜鍋囪锛?
|
||||
4. **绛夊緟閲嶆柊纭**锛氬湪寰楀埌鏄庣‘鐨勬柟鍚戞寚绀哄墠锛屼笉鍐嶇户缁墽琛?
|
||||
1. **立即停止所有操作**
|
||||
2. **输出当前理解**:"我以为目标是 X,正在做 Y"
|
||||
3. **输出自检结果**:
|
||||
- 这个问题我遇到了什么困惑?
|
||||
- 我的方案是什么?
|
||||
- 是否存在更简单方案?
|
||||
- 是否有验证过每个假设?
|
||||
4. **等待重新确认**:在得到明确的方向指示前,不再继续执行
|
||||
|
||||
## 5. 瑙勫垯杩涘寲
|
||||
## 5. 规则进化
|
||||
|
||||
- 鏈枃妗f斁鍦ㄩ」鐩牴鐩綍涓嬶紝鎻愪氦鍒?git
|
||||
- 姣忔浠庣籂姝d腑瀛﹀埌鏂扮粡楠岋紝鏇存柊鍒版湰鏂囨。涓?
|
||||
- 鎵€鏈?Agent 鍏变韩姝ゆ枃妗o紝璺ㄩ」鐩彲澶嶇敤
|
||||
- 本文档放在项目根目录下,提交到 git
|
||||
- 每次从纠正中学到新经验,更新到本文档中
|
||||
- 所有 Agent 共享此文档,跨项目可复用
|
||||
|
||||
## 闄勫綍: 鍒ゆ柇鍑嗗垯鏄惁鐢熸晥
|
||||
## 附录: 判断准则是否生效
|
||||
|
||||
- Diff 骞插噣涓旀渶灏忓寲锛屽彧鏈夎姹傛墍闇€鐨勬敼鍔?
|
||||
- 涓嶅啀鍥犺繃搴﹀鏉傚寲鑰岄渶瑕侀噸鍐?
|
||||
- 婢勬竻闂鍑虹幇鍦ㄥ疄鐜颁箣鍓嶏紝鑰屼笉鏄姱閿欎箣鍚?
|
||||
- 姣忔鏋舵瀯鍐崇瓥鍓嶉兘缁忚繃浜嗗鎶楁€у鏌?
|
||||
- 缁忛獙鏁欒鍦ㄥ悗缁换鍔′腑琚鐢紙涓嶅啀韪╁悓涓€鍧戯級
|
||||
- Diff 干净且最小化,只有请求所需的改动
|
||||
- 不再因过度复杂化而需要重写
|
||||
- 澄清问题出现在实现之前,而不是犯错之后
|
||||
- 每次架构决策前都经过了对抗性审查
|
||||
- 经验教训在后续任务中被复用(不再踩同一坑)
|
||||
|
||||
---
|
||||
鐗堟湰: 1.0 | 2026-07-08 | 鎻愭 #23
|
||||
版本: 1.0 | 2026-07-08 | 提案 #23
|
||||
+40
-8
@@ -69,11 +69,28 @@ specs/usage_monitor.json
|
||||
|
||||
| 模块 | spec 路径 | § 入口 | 状态 |
|
||||
|------|----------|--------|------|
|
||||
| usage_monitor | `gateway/scripts/specs/usage_monitor.json` | Dashboard I Tab → OpenCode Go Usage | ✅ |
|
||||
| agents | `gateway/scripts/specs/agents.json` | Dashboard A Tab → Agent 卡片 | ✅ |
|
||||
| api_proxy | `gateway/scripts/specs/api_proxy.json` | Dashboard I Tab → API Proxy (:8787) | ✅ |
|
||||
| article_processor | `gateway/scripts/specs/article_processor.json` | Dashboard F/I Tab → 文章抓取服务 | ✅ |
|
||||
| chat_bridge | `gateway/scripts/specs/chat_bridge.json` | Dashboard A Tab → SessionBridge | ✅ |
|
||||
| dashboard | `gateway/scripts/specs/dashboard.json` | Dashboard 全局 | ✅ |
|
||||
| dev_spec | `gateway/scripts/specs/dev_spec.json` | Dashboard G Tab | ✅ |
|
||||
| easytier | `gateway/scripts/specs/easytier.json` | Dashboard I Tab → EasyTier | ✅ |
|
||||
| ejabberd | `gateway/scripts/specs/ejabberd.json` | Dashboard F Tab → XMPP 服务器 | ✅ |
|
||||
| health | `gateway/scripts/specs/health.json` | Dashboard F Tab → 健康检查 | ✅ |
|
||||
| health_service | `gateway/scripts/specs/health_service.json` | Dashboard F Tab → 健康服务 | ✅ |
|
||||
| infra | `gateway/scripts/specs/infra.json` | Dashboard I Tab | ✅ |
|
||||
| kanban | `gateway/scripts/specs/kanban.json` | Dashboard K Tab | ✅ |
|
||||
| prd | `gateway/scripts/specs/prd.json` | Dashboard H Tab | ✅ |
|
||||
| rdp | `gateway/scripts/specs/rdp.json` | Dashboard I Tab → RDP | ✅ |
|
||||
| session_router | `gateway/scripts/specs/session_router.json` | Dashboard A Tab → SessionRouter | ✅ |
|
||||
| tests | `gateway/scripts/specs/tests.json` | Dashboard K Tab | ✅ |
|
||||
| usage_collector | `gateway/scripts/specs/usage_collector.json` | Dashboard I Tab → Usage 采集 | ✅ |
|
||||
| usage_monitor | `gateway/scripts/specs/usage_monitor.json` | Dashboard I Tab → OpenCode Go Usage | ✅ |
|
||||
| wechat_bridge | `gateway/scripts/specs/wechat_bridge.json` | Dashboard F/I Tab → 微信桥接 | ✅ |
|
||||
| xmpp_bot | `gateway/scripts/specs/xmpp_bot.json` | Dashboard A/F Tab → XMPP Bot | ✅ |
|
||||
| xmpp_watchdog | `gateway/scripts/specs/xmpp_watchdog.json` | Dashboard F Tab → 看门狗 | ✅ |
|
||||
| 开发规范本身 | `docs/dev-spec.md` | Dashboard G Tab | ✅ |
|
||||
| (新增模块) | `gateway/scripts/specs/{module}.json` | 待注册 | |
|
||||
|
||||
---
|
||||
|
||||
@@ -84,13 +101,26 @@ specs/usage_monitor.json
|
||||
│ G: 规范体系 │────→│ K: 测试 │────→│ F: 健康 │
|
||||
│ 定义期望 │ │ 验证实现 │ │ 持续监控 │
|
||||
└────────────┘ └──────────┘ └──────────┘
|
||||
↑ │
|
||||
└────────────────────────────────┘
|
||||
发现偏差 → 更新 Spec
|
||||
↑ ↑ │
|
||||
└────────────────┼────────────────┘
|
||||
│
|
||||
┌────────┴────────┐
|
||||
│ F 异常 → 触发 K │
|
||||
│ K 失败 → 更新 G │
|
||||
└─────────────────┘
|
||||
|
||||
H: 需求文档 — 双轨体系覆盖不到的架构级/跨模块需求(性能、安全、可用性等)
|
||||
```
|
||||
|
||||
### 核心反馈链路
|
||||
|
||||
| 方向 | 触发条件 | 动作 |
|
||||
|------|---------|------|
|
||||
| G → K | 新增/修改 spec | K Tab 对应测试 ID 必须新增/更新 |
|
||||
| K → F | 测试全部通过 | F Tab 组件标记为已验证 |
|
||||
| **F → K** | **F Tab 发现异常** | **应触发 K Tab 对应测试重跑,确认是服务故障还是测试过期** |
|
||||
| **F → G** | **F Tab 持续异常但测试通过** | **说明期望矩阵或 spec 过时,应更新 G 和对应的 spec** |
|
||||
|
||||
### G — 开发规范(Dashboard G Tab)
|
||||
|
||||
- 本文档,展示在 Dashboard G Tab
|
||||
@@ -99,16 +129,18 @@ H: 需求文档 — 双轨体系覆盖不到的架构级/跨模块需求(性
|
||||
|
||||
### K — 自动测试(Dashboard K Tab)
|
||||
|
||||
- `tests_api.py` 自动执行 20+ 项系统检测
|
||||
- `gateway/scripts/tests_api.py` 自动执行系统检测(Dashboard 运行时从 `gateway/scripts/tests_api` import)
|
||||
- Dashboard K Tab 实时展示 PASS/FAIL/EXPECTED
|
||||
- **部署后必做**:打开 K Tab 确认全部通过或已知失败原因
|
||||
- 新增模块时应在 `ai_spec.tests` 中添加对应的测试标识
|
||||
- **测试追溯链**:`tests_api.py` 中的每个测试用例(A1/A2/B1/B2...)应在其测试逻辑的注释中注明所验证的 `ai_spec.tests` ID(如 `# UM01`、`# XB03`)。Dashboard K Tab 前端应尽量在测试名称列展示对应的 spec 测试 ID,用于快速定位 spec 来源
|
||||
- 新增模块时应在 `ai_spec.tests` 中添加对应的测试标识,并在 `tests_api.py` 中实现
|
||||
|
||||
### F — 系统健康度(Dashboard F Tab)
|
||||
|
||||
- **期望矩阵**:应该运行的服务 vs 实际状态(Dashboard F Tab)
|
||||
- **期望矩阵**:应该运行的服务 vs 实际状态(Dashboard F Tab),从 `agents.yaml` + `PLATFORM_SERVICES` 动态生成
|
||||
- **监控数据**:Tier1(5min)/ Tier2(日报)作为实时状态输入
|
||||
- **服务拓扑**:所有服务的健康、端口、看门狗状态
|
||||
- **跨平台检测**:远程服务(非本机)标记为 `remote(见平台Tab)`,避免误报
|
||||
|
||||
### H — 需求文档(Dashboard H Tab)
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@ Flask app on :5803. Monitors agents across platforms via:
|
||||
|
||||
Auto-recovery: restarts local Windows agents after 3 consecutive offline checks.
|
||||
"""
|
||||
import os, sys, re, json, time, subprocess, logging, urllib.request, sqlite3, shutil, threading
|
||||
import os, sys, re, json, time, socket, subprocess, logging, urllib.request, sqlite3, shutil, threading
|
||||
from pathlib import Path
|
||||
from datetime import datetime, timedelta
|
||||
from flask import Flask, jsonify, request, send_from_directory
|
||||
@@ -562,9 +562,9 @@ def api_agent_restart(agent_id):
|
||||
|
||||
|
||||
PLATFORM_SERVICES = [
|
||||
{"id": "wechat_bridge", "name": "莫荷微信 (Linux)", "type": "ChannelBridge",
|
||||
"desc": "246 Docker wechatbot-webhook → webhook → Hermes Gateway",
|
||||
"host": "192.168.1.246", "port": 3001},
|
||||
{"id": "wechat_bridge", "name": "莫荷微信 (Linux :3001)", "type": "ChannelBridge",
|
||||
"desc": "246 Docker wechatbot-webhook → 微信收发 + webhook 触发 → Hermes Gateway",
|
||||
"health_url": "http://192.168.1.246:3001/", "host": "192.168.1.246", "port": 3001},
|
||||
{"id": "article_processor", "name": "文章抓取服务 (5810)", "type": "wechat-fetch",
|
||||
"desc": "fetches wechat article content + OCR (DrissionPage)",
|
||||
"host": "192.168.1.16", "health_url": "http://192.168.1.16:5810/health"},
|
||||
@@ -946,82 +946,96 @@ def api_todos():
|
||||
# ════════════════════════════════════════════════════════════
|
||||
@app.route("/api/expected")
|
||||
def api_expected():
|
||||
"""期望状态矩阵:应该运行的 vs 实际运行的"""
|
||||
expected = [
|
||||
{"name": "xmpp_bot", "port": 5802, "expected": "running", "critical": True},
|
||||
{"name": "article_processor", "port": 5810, "expected": "running", "critical": True},
|
||||
{"name": "dashboard", "port": 5803, "expected": "running", "critical": True},
|
||||
{"name": "watchdog", "expected": "running", "critical": True},
|
||||
{"name": "agents-health-check", "expected": "scheduled (5min)", "critical": False},
|
||||
{"name": "agents-daily-health", "expected": "scheduled (daily 08:00)", "critical": False},
|
||||
{"name": "agents-todo-executor", "expected": "scheduled (10min)", "critical": False},
|
||||
]
|
||||
"""动态期望矩阵:从 agents.yaml + PLATFORM_SERVICES 生成,跨平台区分检测"""
|
||||
expected = []
|
||||
# 1. 从 agents.yaml 生成各 Agent 服务的期望
|
||||
agents = load_agents_config()
|
||||
for agent in agents:
|
||||
display = agent.get("display_name", agent["name"])
|
||||
host = agent.get("host", "127.0.0.1")
|
||||
for svc in agent.get("services", []):
|
||||
svc_type = svc["type"]
|
||||
port = svc.get("port")
|
||||
entry = {"name": f"{display}({host}):{svc_type}", "host": host, "critical": True}
|
||||
if port:
|
||||
entry["port"] = port
|
||||
entry["check"] = "tcp"
|
||||
elif svc_type == "xmpp_bot":
|
||||
entry["check"] = "xmpp" # xmpp bot without port means ejabberd-connected
|
||||
else:
|
||||
entry["check"] = "agent_online"
|
||||
expected.append(entry)
|
||||
|
||||
# 2. 从 PLATFORM_SERVICES 生成平台服务的期望
|
||||
for ps in PLATFORM_SERVICES:
|
||||
entry = {"name": f"{ps['name']} ({ps['id']})", "host": ps.get("host", "127.0.0.1"), "critical": True}
|
||||
if ps.get("health_url"):
|
||||
entry["health_url"] = ps["health_url"]
|
||||
entry["check"] = "health_url"
|
||||
elif ps.get("port"):
|
||||
entry["port"] = ps["port"]
|
||||
entry["check"] = "tcp"
|
||||
else:
|
||||
continue
|
||||
expected.append(entry)
|
||||
|
||||
# 3. 固定定时任务期望
|
||||
expected.append({"name": "agents-health-check", "check": "scheduled", "critical": False})
|
||||
expected.append({"name": "agents-daily-health", "check": "scheduled", "critical": False})
|
||||
|
||||
# ── 实际状态检测 ──
|
||||
actual = {}
|
||||
for e in expected:
|
||||
name = e["name"]
|
||||
port = e.get("port")
|
||||
if port:
|
||||
actual[name] = "running" if port_open(port) else "stopped"
|
||||
elif name == "watchdog":
|
||||
if sys.platform == "win32":
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["tasklist", "/FI", "IMAGENAME eq python.exe", "/NH"],
|
||||
capture_output=True, text=True, timeout=5,
|
||||
creationflags=subprocess.CREATE_NO_WINDOW,
|
||||
)
|
||||
actual[name] = "running" if "python" in r.stdout else "stopped"
|
||||
except:
|
||||
actual[name] = "unknown"
|
||||
else:
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["pgrep", "-f", "xmpp_watchdog|health_check"],
|
||||
capture_output=True, text=True, timeout=5)
|
||||
actual[name] = "running" if r.stdout.strip() else "stopped"
|
||||
except:
|
||||
# pgrep not available
|
||||
r = subprocess.run(
|
||||
["ps", "aux"], capture_output=True, text=True, timeout=5)
|
||||
has = "watchdog" in r.stdout or "health_check" in r.stdout
|
||||
actual[name] = "running" if has else "stopped"
|
||||
elif "scheduled" in e["expected"]:
|
||||
host = e.get("host", "127.0.0.1")
|
||||
check = e.get("check", "unknown")
|
||||
|
||||
if check == "scheduled":
|
||||
task_name = name
|
||||
if sys.platform == "win32":
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["schtasks", "/Query", "/TN", task_name, "/FO", "CSV", "/NH"],
|
||||
capture_output=True, text=True, timeout=5,
|
||||
creationflags=subprocess.CREATE_NO_WINDOW,
|
||||
)
|
||||
# 中文"就绪"、英文"Ready"
|
||||
r = subprocess.run(["schtasks", "/Query", "/TN", task_name, "/FO", "CSV", "/NH"],
|
||||
capture_output=True, text=True, timeout=5, creationflags=subprocess.CREATE_NO_WINDOW)
|
||||
ready = "Ready" in r.stdout or "\u5c31\u7eea" in r.stdout
|
||||
actual[name] = "scheduled" if ready else "missing"
|
||||
except:
|
||||
actual[name] = "error"
|
||||
else:
|
||||
# Linux: 检查 systemd timer 或 crontab
|
||||
try:
|
||||
r = subprocess.run(["systemctl", "list-timers", "--all", "--no-pager"],
|
||||
capture_output=True, text=True, timeout=5)
|
||||
if task_name in r.stdout:
|
||||
actual[name] = "timer_ok"
|
||||
else:
|
||||
r2 = subprocess.run(["crontab", "-l"], capture_output=True,
|
||||
text=True, timeout=5)
|
||||
r2 = subprocess.run(["crontab", "-l"], capture_output=True, text=True, timeout=5)
|
||||
actual[name] = "cron_ok" if task_name in r2.stdout else "not_deployed"
|
||||
except:
|
||||
actual[name] = "not_deployed"
|
||||
|
||||
elif check == "health_url":
|
||||
actual[name] = "running" if _health_status(e["health_url"])["ok"] else "stopped"
|
||||
|
||||
elif check == "tcp":
|
||||
# 判断目标是否是本机
|
||||
is_local = host in ("127.0.0.1", "localhost", "::1", socket.gethostname())
|
||||
if is_local:
|
||||
actual[name] = "running" if port_open(e["port"], host) else "stopped"
|
||||
else:
|
||||
# 远程服务:状态由 /api/platform 精确检测,这里标记为远程引用
|
||||
actual[name] = "remote (见平台Tab)"
|
||||
|
||||
else:
|
||||
actual[name] = "unknown"
|
||||
|
||||
return jsonify({"expected": expected, "actual": actual})
|
||||
|
||||
|
||||
def port_open(port):
|
||||
"""Check if a port is listening. Works on both Linux and Windows."""
|
||||
def port_open(port, host="127.0.0.1"):
|
||||
"""Check if a port is listening on a specific host."""
|
||||
try:
|
||||
import socket
|
||||
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||
s.settimeout(2)
|
||||
result = s.connect_ex(("127.0.0.1", port))
|
||||
result = s.connect_ex((host, port))
|
||||
s.close()
|
||||
return result == 0
|
||||
except:
|
||||
|
||||
Reference in New Issue
Block a user