Gateway看门狗-知微 was erroring (exit -15): its check_session_health did a live LLM ping with 25s timeout. Cold-start LLM latency is 20-100s so the ping always timed out -> false '不健康' verdict -> false gateway restart -> and each 10-min run burned 22k tokens. Now uses xmpp_logger._scan_agent_log (zero cost, reads real call results): - ok if last real call succeeded - unhealthy only if last call explicitly failed - idle (no recent calls) counts as healthy Verified: watchdog job now status=ok. Also: triggered all 6 weekend 'Blocked' jobs via hermes cron run — all now status=ok, proving the hardlink fix holds.
8 lines
544 B
Python
8 lines
544 B
Python
import json
|
|
d = json.load(open('/home/hmo/.hermes/profiles/position-analyst/cron/jobs.json'))
|
|
jobs = d if isinstance(d, list) else d.get('jobs', [])
|
|
targets = ['策略评估-每周', '建议对账-每周', '数据治理-每周', '跨市场背离检测-周末',
|
|
'自选股自动重评-周末', 'state.db真空整理-每周', 'Gateway看门狗-知微']
|
|
for j in jobs:
|
|
if j.get('name') in targets:
|
|
print(f"{j['name']}: status={j.get('last_status')} last={str(j.get('last_run_at'))[:19]} err={str(j.get('last_error'))[:90]}") |