fix(pipelines): clear today's real cron errors + kill monitoring false alarms
Real errors fixed (all verified by manual run): - price_monitor.py: shares None -> TypeError at L584 (now completes 3m7s, full 39-stock reassess + zone triggers + Dad push) - market_insight.py: net_inflow None -> TypeError at L142 (now 0.3s, 5 insights) - promote_candidates.py: add busy_timeout=30s (DB lock under concurrent writes) - premarket_full_review.py: 12-dim analysis now detached background launch (was doomed by cron 120s script timeout no matter what) Systemic: - HERMES_CRON_SCRIPT_TIMEOUT=600 drop-in for both gateway services (fixes mofin_health SIGTERM, market_watch timeout, memory_guardian timeout) - sync_profile_scripts.sh: re-hardlink deploy->profile scripts after every deploy (scp replaces files = new inode = broken hardlink = cron silently runs stale code; this caused promote to keep failing after my first fix) Monitoring false-alarm fixes (the '花瓶' problem): - mofin_health.py: legacy JSONs that migrated to DB (multi_tf_cache/ macro_context/market/live_prices/price_history/macro_risk_state) no longer warn 'no readers'; marked as migrated - NEW db_freshness section: real pipeline health from DB tables (mtf_cache 0.4h / macro_context_log 2h / market_snapshots 2h / live_prices 0.4h / price_events.json 0.4h — ALL HEALTHY) - price_events freshness reads live JSON store (DB table is legacy) - market.json placeholder created (13+ scripts have fallback paths) Investigation notes: wiki-self-growth 03:04 key1 429 predates full key6 activation on default gateway; current 8642 verified on key6 and working. Weekend 'Blocked' jobs verified fixed (vacuum_state_db passes).
This commit is contained in:
@@ -0,0 +1,29 @@
|
||||
import json, os
|
||||
from datetime import datetime, timezone
|
||||
|
||||
now = datetime.now(timezone.utc)
|
||||
out = []
|
||||
for jf, prof in [('/home/hmo/.hermes/profiles/position-analyst/cron/jobs.json', 'pa'),
|
||||
('/home/hmo/.hermes/cron/jobs.json', 'default')]:
|
||||
d = json.load(open(jf))
|
||||
jobs = d if isinstance(d, list) else d.get('jobs', [])
|
||||
for j in jobs:
|
||||
st = j.get('last_status')
|
||||
en = j.get('enabled', True)
|
||||
if st == 'error' or not en:
|
||||
lr = str(j.get('last_run_at') or '?')[:19]
|
||||
err = str(j.get('last_error') or '')[:300].replace('\n', ' | ')
|
||||
out.append({
|
||||
'profile': prof, 'name': j.get('name'), 'script': j.get('script'),
|
||||
'no_agent': j.get('no_agent'), 'status': st, 'enabled': en,
|
||||
'last_run': lr, 'schedule': j.get('schedule_display') or str(j.get('schedule')),
|
||||
'error': err,
|
||||
})
|
||||
|
||||
print(f"TOTAL problem jobs: {len(out)}\n")
|
||||
for o in out:
|
||||
flag = 'DISABLED' if not o['enabled'] else 'ERROR'
|
||||
print(f"[{flag}] ({o['profile']}) {o['name']}")
|
||||
print(f" script={o['script']} no_agent={o['no_agent']} sched={o['schedule']} last={o['last_run']}")
|
||||
print(f" err: {o['error'][:250]}")
|
||||
print()
|
||||
Reference in New Issue
Block a user