55 lines
2.6 KiB
JSON
55 lines
2.6 KiB
JSON
{
|
||
"module": "article_processor",
|
||
"version": "1.0",
|
||
"purpose": "文章抓取服务运行在 Windows (192.168.1.16:5810),负责抓取微信公众号全文链接并转换为纯文本,支持 OCR 图片识别。",
|
||
"ui_location": "Infrastructure tab → Platform → 文章抓取服务 (5810)",
|
||
|
||
"human_help": {
|
||
"title": "文章抓取服务 (:5810)",
|
||
"description": [
|
||
"这是整个系统的'阅读器'。当莫荷或小小莫需要阅读一个微信公众号文章链接时,由这个服务负责:",
|
||
"① 用 Chrome CDP 打开链接",
|
||
"② 等页面加载完成",
|
||
"③ 提取全文内容转为 Markdown",
|
||
"④ 如果包含图片,调用 GLM-OCR 识别图片文字",
|
||
"⑤ 返回结构化内容给请求方",
|
||
"依赖微信的 CDP session:如果微信未登录,抓取会失败('等待登录超时')。"
|
||
],
|
||
"usage": [
|
||
"监控状态:Infrastructure → Platform 下查看状态和最近抓取信息",
|
||
"如果 health_data 中有 error:表示最近一次抓取失败,查看 error 内容",
|
||
"点击服务旁的 § 查看 AI 接口详情"
|
||
],
|
||
"troubleshooting": [
|
||
"如果状态 stopped:检查 Windows 上 article_processor.py 进程",
|
||
"如果抓取失败返回 '等待登录超时':微信 CDP session 已过期,需要重新登录",
|
||
"如果图片 OCR 失败:检查 GLM-OCR 服务是否可用",
|
||
"日志位置:gateway/logs/article_processor.log"
|
||
],
|
||
"related": "WeChat Bridge(依赖文章抓取服务处理微信中的链接)"
|
||
},
|
||
|
||
"ai_spec": {
|
||
"apis": [
|
||
{"method": "GET", "path": "/health", "returns": "{ok, service, port, ocr_model, status}", "desc": "服务状态 + OCR 模型信息"},
|
||
{"method": "GET", "path": "/logs?lines=N", "returns": "{ok, lines[]}", "desc": "最近 N 行日志"},
|
||
{"method": "POST", "path": "/fetch", "body": "{\"url\":\"...\"}", "returns": "{ok, title, content, images[]}"}
|
||
],
|
||
"dependencies": [
|
||
"依赖 Chrome CDP session(微信登录态)",
|
||
"依赖 GLM-OCR 服务识别图片文字",
|
||
"依赖网络连通性访问微信公众号"
|
||
],
|
||
"constraints": [
|
||
"CDP session 过期后需要重新登录微信才能恢复",
|
||
"OCR 使用 GLM-OCR-8bit 模型",
|
||
"Linux 上访问 Windows 的此服务时需要 EasyTier VPN 连通"
|
||
],
|
||
"related_files": [
|
||
"gateway/scripts/article_processor.py — 实际运行脚本",
|
||
"gateway/scripts/templates/dashboard.html — fI() 状态展示(health_data)",
|
||
"gateway/scripts/specs/article_processor.json — 本 spec 文件"
|
||
]
|
||
}
|
||
}
|