Files
MoFin/scripts/test_ocr_pipeline_live.py
T
hmo 299ddc1796 fix(ocr+bot): image download race, SenseNova context, log path, encoding
Root causes of the screenshot 404 incident:
1. RACE: client uploads image AND sends message concurrently; bot received
   the message before the upload finished writing, so its GET hit a 404
   error page (<100B treated as failure). FIX: _download_image now retries
   3x with 2s backoff.
2. Zhiwei mentioned tesseract/小果 because the failure text never told her
   the pipeline IS SenseNova. FIX: failure messages now name SenseNova
   explicitly and ask for resend.
3. log_xmpp never worked for the bot: sys.path used relative '../..' from
   a symlinked __file__ which resolved to '/' instead of MoFin root. This
   is why the '最近对话' panel never had bot chat data (only cron script
   entries). FIX: absolute path per red line #7. Verified: test message
   now lands in xmpp_messages.jsonl.
4. My PowerShell -replace corrupted the file encoding (UnicodeDecodeError
   crash loop on restart). Restored from git HEAD and re-applied edits with
   the edit tool. Lesson: never use PowerShell string replace on UTF-8
   source files with Chinese content.
5. functional_health: new sense_ocr module (OCR config presence +
   SenseNova API TCP reachability), no token cost.
2026-07-20 21:42:59 +08:00

17 lines
649 B
Python

import sys
sys.argv = ['test', '--agent', 'zhiwei']
src = open('/home/hmo/MoFin/deploy/bot/xmpp_agent_core.py').read()
cut = src.find('if __name__ ==')
if cut > 0:
src = src[:cut]
ns = {}
exec(compile(src, 'xmpp_agent_core.py', 'exec'), ns)
url = 'https://upload.yoin.fun/upload/d22cef590582300cc5580c722a19280054bf0d6a/aAeuEbcXa1NNmU2b7c0Q0D75H0anzmGiB86SFxla/9a649bdc-eb7c-44b0-93eb-53ce64c446f7.png'
print('is_image_url:', ns['_is_image_url'](url))
img = ns['_download_image'](url)
print('download:', len(img) if img else None, 'bytes')
if img:
ok, text = ns['_ocr_image'](img)
print('ocr ok:', ok)
print('ocr text:', text[:400])