User insight: hardlink breakage only happens at deploy time (scp file
replacement / git checkout-merge), so detection must be welded INTO the
deploy pipeline, not left to daily audit.
Three automatic layers, no reliance on discipline:
1. systemd path watcher (profile-scripts-sync.path): watches
deploy/profile-scripts/ directory, auto-fires sync_profile_scripts.sh
on any change. Verified: fires within 4s of file replacement, logs to
gateway/logs/link_sync.log (runs as hmo user)
2. git hooks (.git/hooks/post-merge + post-checkout on 246 repo):
auto re-link after git operations
3. Manual fallback: sync_profile_scripts.sh (now self-logging)
dev-spec red line #6 updated: SSOT rule now documents the three layers
and states breakage only happens at deploy time.
Root cause analysis of the 2026-07-20 redundancy incident:
1. No single-source-of-truth rule -> same file legitimately lived in 4+
locations, diverging silently
2. Relative path resolution (Path(__file__).parent/'data') -> each
hardlinked copy of mofin_db.py pointed to a DIFFERENT database
3. 'Backup habit' left .bak/legacy files in production dirs, which
monitoring then scanned and reported as false alarms
4. Half-done migrations: DB tables created but old JSON writers/readers
stayed (price_events), old files stayed
5. Dead modules never got buried: xiaoguo 'dead' but bot ran 8 days
as root eating 2.5GB
6. Monitoring checked 'does it exist' not 'is it alive' -> stale file
mtime reported as 'pipeline stalled 14 days' (false alarm)
7. No 'system hygiene' as a check category at all
Prevention implemented:
- dev-spec.md v2.0: 五条红线 -> 十条红线
#6 single source of truth (hardlink only, no independent copies)
#7 absolute data paths only (no __file__-relative data resolution)
#8 no backups/legacy in production data dirs (archive immediately)
#9 dead module burial checklist (6 mandatory steps)
#10 monitor liveness (DB table freshness) not existence
- File Location Constitution: canonical location per content type
- NEW system_hygiene_audit.py: weekly Monday 07:30 cron checking
diverged copies / broken hardlinks / zombie processes / orphan data
files / dead cron scripts / DB freshness -> hygiene_report.json + XMPP
- specs/hygiene.json: module spec per red line #1
- Verified: audit found 5 real issues on first run, all fixed, re-run clean