Files
AgentsMeeting/gateway/scripts/specs/article_processor.json
T

55 lines
2.6 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"module": "article_processor",
"version": "1.0",
"purpose": "文章抓取服务运行在 Windows (192.168.1.16:5810),负责抓取微信公众号全文链接并转换为纯文本,支持 OCR 图片识别。",
"ui_location": "Infrastructure tab → Platform → 文章抓取服务 (5810)",
"human_help": {
"title": "文章抓取服务 (:5810)",
"description": [
"这是整个系统的'阅读器'。当莫荷或小小莫需要阅读一个微信公众号文章链接时,由这个服务负责:",
"① 用 Chrome CDP 打开链接",
"② 等页面加载完成",
"③ 提取全文内容转为 Markdown",
"④ 如果包含图片,调用 GLM-OCR 识别图片文字",
"⑤ 返回结构化内容给请求方",
"依赖微信的 CDP session:如果微信未登录,抓取会失败('等待登录超时')。"
],
"usage": [
"监控状态:Infrastructure → Platform 下查看状态和最近抓取信息",
"如果 health_data 中有 error:表示最近一次抓取失败,查看 error 内容",
"点击服务旁的 § 查看 AI 接口详情"
],
"troubleshooting": [
"如果状态 stopped:检查 Windows 上 article_processor.py 进程",
"如果抓取失败返回 '等待登录超时':微信 CDP session 已过期,需要重新登录",
"如果图片 OCR 失败:检查 GLM-OCR 服务是否可用",
"日志位置:gateway/logs/article_processor.log"
],
"related": "WeChat Bridge(依赖文章抓取服务处理微信中的链接)"
},
"ai_spec": {
"apis": [
{"method": "GET", "path": "/health", "returns": "{ok, service, port, ocr_model, status}", "desc": "服务状态 + OCR 模型信息"},
{"method": "GET", "path": "/logs?lines=N", "returns": "{ok, lines[]}", "desc": "最近 N 行日志"},
{"method": "POST", "path": "/fetch", "body": "{\"url\":\"...\"}", "returns": "{ok, title, content, images[]}"}
],
"dependencies": [
"依赖 Chrome CDP session(微信登录态)",
"依赖 GLM-OCR 服务识别图片文字",
"依赖网络连通性访问微信公众号"
],
"constraints": [
"CDP session 过期后需要重新登录微信才能恢复",
"OCR 使用 GLM-OCR-8bit 模型",
"Linux 上访问 Windows 的此服务时需要 EasyTier VPN 连通"
],
"related_files": [
"gateway/scripts/article_processor.py — 实际运行脚本",
"gateway/scripts/templates/dashboard.html — fI() 状态展示(health_data",
"gateway/scripts/specs/article_processor.json — 本 spec 文件"
]
}
}