Commit Graph
65 Commits
Author SHA1 Message Date
hmo 3ea7c52112 feat(L2): 卫生审计自动收尸——scripts/影子副本+零引用孤儿(mtime>7天)自动git mv归档并提交; dev-spec新增文件居住宪法(红线#9) 2026-07-22 08:41:06 +08:00
hmo 375f6d469a chore: sync脚本补price_monitor三处副本硬链 2026-07-22 08:28:39 +08:00
hmo fbad1938f1 feat: 失败二轮重试(休息60s后失败股整体重跑) + sync脚本补7个分叉副本硬链 2026-07-22 08:26:32 +08:00
hmo 04284e5187 fix: 推送质量门禁+损盈一致性门禁(根治垃圾信号)
1. XMPP推送门禁(_validate_buy_alert): 实时价>0(live_prices)、
   区间有效、现价不超上沿5%、损<下沿且在(0.5x~1.0x)现价内、
   盈>上沿>损 —— 任一不过不推只记日志
2. DB损盈一致性门禁: 损必须在下沿之下、盈必须在上沿之上且损<盈,
   不一致字段跳过写入保留原值(浩辰 区2~3损25.11 类污染根治)
3. 推送价格源改 live_prices(不再用 holding_strategies.price 的0值)
4. 浩辰脏行已从快照恢复(区25.93~27.84 损25.11 盈31.38)
2026-07-22 08:16:40 +08:00
hmo c27388681c fix: FALLBACK_MODEL 顶层导入(NameError 曾中断批量重评) 2026-07-22 01:31:56 +08:00
hmo 6cdcd3f2e5 feat(L1): 新增reassess_daily检查——持仓当日重评覆盖率+分析存在率+自选24h补评率(修复'每日重评做没做'无监控的空白) 2026-07-22 01:03:43 +08:00
hmo 36f01dc8f0 fix: save_result区间写入门禁——下沿<上沿<下沿x3,异常整体跳过(防214.68~2.52解析污染) 2026-07-22 00:39:30 +08:00
hmo d92d6ba54b fix: 截断保护——输出<1500字且无信号时升级pro重试(flash懒答/截断导致信号丢失) 2026-07-22 00:36:53 +08:00
hmo 83b0770db0 feat(llm): flash空输出自动升级pro重试(688617实测flash三连空/pro正常输出) 2026-07-22 00:30:33 +08:00
hmo 40852de280 fix: LLM空输出视为失败 + save_result拒绝写入空分析(防止假时间戳跳过机制失效) 2026-07-22 00:24:47 +08:00
hmo 52f7df7bef feat: 提交白名单配套——guard带GIT_ALLOW_COMMIT令牌 + hygiene新增钩子存在性检查 2026-07-21 23:45:26 +08:00
hmo 9914ff9457 feat(盯盘): 推荐操作区域置顶独立 + tag与XMPP动作级信号自动同步
- mofin_db: write_holding_strategy 内置tag同步语义——动作级信号
  (买入/可买入/可加仓/卖出/止盈)→current_recommend; 信号降级→
  清除current_recommend; active_manual人工标记永不被自动流覆盖/清除
- 新增 sync_recommend_tag() 供裸SQL调用方
- batch_reassess.save_result / per_stock stage-2 调用同步
- 盯盘Tab: '重点推荐'更名'推荐操作', 区域独立琥珀色视觉, 置顶
同步语义: XMPP买入信号(ACTION级告警)的个股=推荐操作区域个股,
信号消失(重评降级)时区域同步消失
2026-07-21 23:31:57 +08:00
hmo 05a60cf0d4 revert(34337fc5): 恢复被知微二次stale提交覆盖的昨晚重构——batch/per_stock/stale_detector/fix_gateway_port/candidate_filter 回滚至1e71a2d8版本;她提交中的运行时文件(db-shm/db-wal/price_history/market_scan_summary)移出跟踪;保留其morning_health_check小果清理 2026-07-21 23:24:45 +08:00
知微 a245c8b4dd fix: price_monitor crash when shares is string + type guard in write_holding_strategy
Root cause: holding_strategies.shares field got set to literal string
'write_holding_strategy' for 6 records, causing TypeError at
price_monitor.py line 611 ('>' not supported between str and int).

Fixes:
1. price_monitor.py: Replace set comprehension with safe loop that
   checks isinstance before comparison. Non-numeric shares treated
   as holdings for safety.
2. mofin_db.py write_holding_strategy: Add type guard that resets
   non-numeric shares to 0 with warning.
3. Data fix: Updated 5 corrupted holding_strategies records from
   holdings table (300035/300308/300750/518880/00700).
   Set 002594 (watchlist) shares=0.
2026-07-21 13:47:52 +08:00
知微 91885f7623 Merge branch 'session-work' 2026-07-21 13:47:31 +08:00
hmo 5422b0a11d feat(db): 每日DB在线备份(07:50,留14天) + 管道审计malformed重试一次再告警(瞬态WAL损坏防误报) 2026-07-21 13:47:10 +08:00
知微 34337fc5b7 clean: 移除早上健康检查/系统审计中的小果残留引用(二次提交—deploy guard恢复旧版后重新应用) 2026-07-21 12:05:43 +08:00
知微 9d3cab2306 clean: 移除盘中自检/早上健康检查/系统审计中的小果残留引用 2026-07-21 12:03:21 +08:00
知微 c180129a1a fix: add SIGALRM 120s hard timeout to price_monitor lock
price_monitor.py 进程锁无硬超时,LLM策略重评耗时9-10分钟时
锁文件阻止后续*/2 cron执行,造成10分钟监控盲区。

新增:
- _handle_sigalrm: SIGALRM处理器,自动清理锁后退出
- signal.alarm(120): 获锁后设120s硬上限
- signal.alarm(0): 正常完成时取消定时器
2026-07-21 11:29:06 +08:00
知微 f6d10a41af Merge branch 'session-work'
# Conflicts:
#	deploy/profile-scripts/intraday_health_check.py
2026-07-21 10:43:41 +08:00
hmo 29e0e6c8b4 fix(688775事件): bot不再静默吞消息 + 价格推送分级 + 小果假警报清理
1. bot call_hermes 失败: 10s重试一次, 仍失败回兜底消息(网关暂时不可用)
   ——此前连接被拒时用户消息被静默吞掉(688775事件根因)
2. price_monitor 推送分级: 破止损/重评确认信号/持仓急跌=ACTION直通;
   未确认进区提示=聚合成一条摘要走INFO(30min限1条+8行截断)
   ——今早一秒一条十几连发+思特威英诺特双发
3. alert_helper ACTION 增加5min内容去重(防竞态双发)
4. intraday_health_check 删除小果全部残留检查(已退役, 假警报
   导致知微旧自愈执行器烧150k+token/session去调查并直接改代码)
2026-07-21 10:41:32 +08:00
知微 84d72dfd85 fix: intraday_health_check移除小果残留——check_xiaoguo/8645端口/小果bot/小果信号堆积
小果已全线归档(2026-07-20 Xiaoxiao重构)。盘中自检脚本仍残留:
- check_xiaoguo()函数
- 小果XMPP Bot离线检测
- 8645端口检测
- xiaoguo信号堆积检查(含import socket)
- main()中check_xiaoguo()调用

全部移除,匹配已修复的运行版本。
2026-07-21 09:46:43 +08:00
hmo 1e71a2d852 feat(alerts): alert_helper 统一告警网关——信噪比控制
两级通道:
- ACTION(买入信号/重点推荐/部署验证失败): 直通不限速, 🚨醒目前缀
- INFO(部署/卫生/修复/螺旋报备): 同类30min限1条 + ≤8行截断 +
  24h内容去重(同一问题不重复轰炸) + 压制计数透明披露

6个调用点全部迁移: deploy_guard/hygiene/spiral/self_repair(INFO)
+ batch_reassess/per_stock_reassess 买入信号(ACTION)
解决: 真正有意义的信息(重点推荐操作)不被纯通知淹没
2026-07-21 02:43:57 +08:00
hmo 1806b6a174 fix(SSOT): 4个库文件副本移出git跟踪(红线6:只允许硬链接,不允许独立副本)——硬链与git跟踪本质冲突:checkout重建文件断链→内容与提交版不同→guard判漂移回滚→sync重链→再dirty→merge被阻。移出跟踪后:根目录mofin_db.py/mo_data.py为唯一tracked canonical,副本由sync脚本以硬链重建 2026-07-21 02:09:25 +08:00
hmo 609ef06e6e fix(guard): CODE_PATHS补scripts/整目录+mofin_db/mo_data根文件——覆盖库文件硬链造成的合法dirty 2026-07-21 02:08:23 +08:00
hmo beee453bf3 fix: 同步4个库文件副本的提交内容与根canonical一致——消除deploy_guard回滚↔sync重链拉锯(guard实测抓到漂移并回滚,机制自证有效,但提交内容必须同源) 2026-07-21 02:06:55 +08:00
hmo 3e84cbc17e ux: 系统告警统一加'📟【MoFin系统·XX】(非知微本人)'前缀——告警与知微本人消息可一眼区分 2026-07-21 02:04:55 +08:00
hmo 6d1d4a5164 fix(SSOT): 库文件统一硬链到MoFin根目录canonical——今晚三处改造差点跑在陈旧副本上
- mo_data.py 提升为根目录canonical(含tag修复), deploy/profile-scripts/
  scripts/ profile脚本目录 全部硬链到根
- mofin_db.py 同理(含strategy_history/tag迁移), 四处硬链统一
- sync_profile_scripts.sh 增加库文件链接步骤, merge后自动恢复
- 根因: deploy/profile-scripts/mofin_db.py 是7-20陈旧副本(无snapshot),
  profile脚本目录mo_data.py无tag——hygiene分叉副本检查正确报警
2026-07-21 02:00:26 +08:00
hmo 2054c6e6f9 fix: state.db 时间戳单位自适应(秒/毫秒混存)+ default profile 无sessions表静默跳过 2026-07-21 01:54:59 +08:00
hmo 839c6fc2ff feat(self-heal): 三盲区系统性补丁——自愈体系覆盖今晚三类故障
1. agent_spiral_watchdog.py (新增,10min cron): state.db 检测运行>15min
   且消息>80条的 api session(螺旋特征), XMPP告警+去重。补 603288 事件
   '无watcher看agent会话本身'盲区
2. deploy_guard: 自动merge后自动跑 verify_deployment.py, 失败项立即
   XMPP告警。补'提交级回归无监控'盲区(知微stale提交事件)
3. system_hygiene_audit: 新增第7项检查'指令冻结session'——常驻session
   启动时间早于SOUL.md mtime且6h内仍活跃 → 告警需bump/重启。
   补'system_prompt冻结'盲区; 6h活跃度过滤防误报已遗弃session
2026-07-21 01:53:11 +08:00
hmo ae0c7d1ca3 fix(L3): self_repair 迁移 llm_client(OCG直连无session) 废弃常驻'self-repair' session(指令冻结+上下文累积) 2026-07-21 01:39:54 +08:00
hmo 95d07f6a11 fix(guard): porcelain路径解析健壮化(line[2:].strip + rename格式处理) 2026-07-21 01:19:53 +08:00
hmo a69b246c57 feat(guard): 部署一致性守卫 deploy_guard.py + 知微运维纪律 + 健康JSON移出跟踪
- deploy_guard.py (15min cron): 代码漂移自动回滚(仅未提交改动)+
  session-work可快进时自动merge部署+幂等重链+cron引用完整性,
  状态落盘JSON, 有动作即XMPP报备
- docs/zhiwei-ops-discipline.md: 知微纪律——禁止直接编辑被跟踪代码/
  禁stale提交/禁直调gateway批量LLM; 自愈白名单(rerun/restart/
  sync/switch_key)本不改代码, 与纪律不矛盾; 紧急热修走git流程,
  提交到session-work后guard 15min自动部署
- dev-spec 红线#6 补充部署守卫机制
- static/mofin_health.json 移出git跟踪(运行时产物, 常驻dirty
  会废掉漂移检测)
2026-07-21 01:17:01 +08:00
hmo cd530c2463 fix(llm): 重评直连OCG上游绕过hermes agent运行时 + 全线切flash + prompt输出纪律
事故根因(2026-07-21): hermes gateway /v1/chat/completions 非透传,
每个请求创建带工具的agent会话。一次603288重评螺旋35分钟/44次
terminal调用/输入153k token, 客户端超时后服务端空转, 重试叠加
新会话自我DDoS。

- llm_client 重写: OCG直连为主(key运行时从hermes config.yaml
  ocg-key6读取, 不落盘), gateway兜底(agent模式仅应急)
- REASSESS_MODEL: pro -> flash (A/B实测新prompt下质量差距微弱,
  flash快40%)
- prompt输出纪律: 【建议仓位】不可省略(非买入写'不新建仓'),
  禁止structured_data/XML/JSON块, 禁止寒暄开场白
- batch main 双通道预检(全挂才退出)
2026-07-21 01:00:47 +08:00
hmo 9a8e3344ac restore: candidate_filter busy_timeout + premarket Step1.5 摘要字段(知微17:31 stale提交回滚恢复至7b26c373已验证版本) 2026-07-20 23:53:03 +08:00
hmo de9927a627 重评核心重构:ds-v4-pro + 原策略全文 + strategy_history + 前端三改造
后端(重评管线):
- 新增 llm_client.py 共享客户端: REASSESS_MODEL=deepseek-v4-pro 单点,
  gateway预检(fail-fast), 150s超时+1次重试, 永不抛异常
- batch_reassess/per_stock_reassess: curl/urllib -> call_llm,
  prompt传入原策略全文+当前参数+最近3条变更, 输出 维持/修改判断+
  修改点理由+最终新策略, max_tokens 4096
- mofin_db: 新增 strategy_history 表 + snapshot_strategy_history(),
  write_holding_strategy 覆写前自动快照(保留20条/code)
- mofin_db: holding_strategies 补 tag 列迁移 + 写入保留
  (tag缺席=保留旧值, 显式传''=允许清除), 修复推荐标签被静默丢弃
- mo_data.read_decisions: SELECT 补 tag
- stale_detector/promote_candidates: 子进程超时 240/60 -> 480s

前端:
- 移除 报告Tab -> mofin_health 全部流程/Cron 表加 最后十次 列
  (modal列表->详情), /api/reports 支持 cron+script 多路匹配
  (jobs.json name->id 解析 + 文件名/标题子串兜底)
- 移除 决策库Tab
- 盯盘Tab 重构: 全部持仓+自选, sort_group 分组(推荐/持仓/自选),
  推荐行琥珀高亮+🔥badge+行内策略, 新增 操作策略 列查看
  最近3次完整策略(/api/strategy_history/<code>, 表缺失时降级当前行)
- 提示词Tab: registry.py 数据路径改回 /home/hmo/MoFin/data/prompts
  (红线: 数据只在规范数据根), 空态提示初始化命令
2026-07-20 23:51:24 +08:00
知微 90f07b4eba chore: sync 2026-07-20 22:05:03 +08:00
知微 3eb214f3a6 chore: sync 2026-07-20 21:43:03 +08:00
知微 ad6a416ef4 merge: watchdog agent.log check 2026-07-20 20:56:05 +08:00
知微 c971d6bced chore: watchdog fix 2026-07-20 20:55:34 +08:00
hmo 10a37f10f9 fix(watchdog): gateway session check now uses agent.log scan, not live LLM ping
Gateway看门狗-知微 was erroring (exit -15): its check_session_health did a
live LLM ping with 25s timeout. Cold-start LLM latency is 20-100s so the
ping always timed out -> false '不健康' verdict -> false gateway restart
-> and each 10-min run burned 22k tokens.

Now uses xmpp_logger._scan_agent_log (zero cost, reads real call results):
- ok if last real call succeeded
- unhealthy only if last call explicitly failed
- idle (no recent calls) counts as healthy
Verified: watchdog job now status=ok.

Also: triggered all 6 weekend 'Blocked' jobs via hermes cron run — all
now status=ok, proving the hardlink fix holds.
2026-07-20 20:55:30 +08:00
知微 dae60bb92f chore: sync before briefing merge 2026-07-20 20:37:20 +08:00
知微 9f7198dc9c chore: deploy pipeline auto-sync 2026-07-20 20:27:15 +08:00
hmo 9a359f49bd feat(deploy): automatic hardlink repair built into deployment pipeline
User insight: hardlink breakage only happens at deploy time (scp file
replacement / git checkout-merge), so detection must be welded INTO the
deploy pipeline, not left to daily audit.

Three automatic layers, no reliance on discipline:
1. systemd path watcher (profile-scripts-sync.path): watches
   deploy/profile-scripts/ directory, auto-fires sync_profile_scripts.sh
   on any change. Verified: fires within 4s of file replacement, logs to
   gateway/logs/link_sync.log (runs as hmo user)
2. git hooks (.git/hooks/post-merge + post-checkout on 246 repo):
   auto re-link after git operations
3. Manual fallback: sync_profile_scripts.sh (now self-logging)

dev-spec red line #6 updated: SSOT rule now documents the three layers
and states breakage only happens at deploy time.
2026-07-20 20:27:11 +08:00
知微 e69109fde9 merge: L0-L4 self-check architecture 2026-07-20 19:40:24 +08:00
知微 135bfced5a chore: deployed L0-L4 self-check system 2026-07-20 19:40:23 +08:00
hmo 08eef1e181 feat(self-check): L0-L4 layered self-check architecture with LLM auto-repair
User directive: daily not weekly; clear responsibilities per layer with no
overlap; functional criteria (does the function WORK) not process liveness;
problems get FIXED via LLM with file-and-report discipline (act first,
report after); plus a meta-layer watching the watchers; deeply integrated
into F健康.

Architecture (responsibility matrix in dev-spec.md):
- L0 agents_health_check (5min): port/HTTP/DB liveness + auto_heal executor
- L1 functional_health_check (15min trading): per-module FUNCTIONAL
  criteria — output freshness/validity per REGISTRY (live_prices/market_
  snapshots/mtf_cache/macro_context/bot/LLM/cron engine), not process alive
- L2 system_hygiene_audit (daily 08:20, was weekly): divergence/hardlink/
  zombie/orphan/dead-cron/db-freshness
- L3 self_repair (30min): reads L1/L2 failures -> LLM diagnoses -> executes
  WHITELISTED repair actions directly (rerun_script/restart_service/
  sync_links/switch_llm_key/none) -> repair_log.jsonl + XMPP report.
  Max 2 repairs/module/day anti-loop. LLM unavailable -> rule fallback.
- L4 meta_watchdog (hourly): checks L0-L3 output freshness + L3 cron
  registration + XMPP bridge; direct XMPP alert as last resort

Retired (overlap): Cron监护-高频 (cron_watchdog -> L3), 全局cron健康监控
(cron_health_monitor -> L1).

Dashboard: mofin_health.py now emits self_check section (functional/meta/
hygiene/recent_repairs); mofin_health.html new '🩺 自检体系' tab rendering
L4 layers, L1 module checks, L2 issues, L3 repair history.

E2E verified: stopped xmpp bot -> L1 flagged fail -> systemd recovered ->
L3 LLM correctly diagnosed 'none needed' and logged; rerun_script whitelist
path executes real scripts successfully; meta_watchdog all-green after fix.
2026-07-20 19:39:58 +08:00
知微 54c48dc5d7 chore: deployed cleanup + hygiene system 2026-07-20 19:05:03 +08:00
hmo 4f83ee8a01 feat(hygiene): anti-redundancy enforcement — spec rules + weekly audit
Root cause analysis of the 2026-07-20 redundancy incident:
1. No single-source-of-truth rule -> same file legitimately lived in 4+
   locations, diverging silently
2. Relative path resolution (Path(__file__).parent/'data') -> each
   hardlinked copy of mofin_db.py pointed to a DIFFERENT database
3. 'Backup habit' left .bak/legacy files in production dirs, which
   monitoring then scanned and reported as false alarms
4. Half-done migrations: DB tables created but old JSON writers/readers
   stayed (price_events), old files stayed
5. Dead modules never got buried: xiaoguo 'dead' but bot ran 8 days
   as root eating 2.5GB
6. Monitoring checked 'does it exist' not 'is it alive' -> stale file
   mtime reported as 'pipeline stalled 14 days' (false alarm)
7. No 'system hygiene' as a check category at all

Prevention implemented:
- dev-spec.md v2.0: 五条红线 -> 十条红线
  #6 single source of truth (hardlink only, no independent copies)
  #7 absolute data paths only (no __file__-relative data resolution)
  #8 no backups/legacy in production data dirs (archive immediately)
  #9 dead module burial checklist (6 mandatory steps)
  #10 monitor liveness (DB table freshness) not existence
- File Location Constitution: canonical location per content type
- NEW system_hygiene_audit.py: weekly Monday 07:30 cron checking
  diverged copies / broken hardlinks / zombie processes / orphan data
  files / dead cron scripts / DB freshness -> hygiene_report.json + XMPP
- specs/hygiene.json: module spec per red line #1
- Verified: audit found 5 real issues on first run, all fixed, re-run clean
2026-07-20 19:04:05 +08:00
知微 7f3ff66be4 merge: retire price_events.json 2026-07-20 18:10:40 +08:00