Repository navigation
[Bug] memos-local-plugin: capture summarizer writes Chinese summaries for English conversations (prompt's language rule + 用户说了 example) #2469
Description
Activity
- addedai:taskDispatched to AI coding agent | 已派发给 AI 编码任务Dispatched to AI coding agent | 已派发给 AI 编码任务area:pluginOpenClaw & HermesOpenClaw & Hermesstatus:in-progressSomeone or AI is working on it | 人工或 AI 正在处理Someone or AI is working on it | 人工或 AI 正在处理types:bugSomething isn't working | 功能异常Something isn't working | 功能异常
on Oct 8, 2026 🤖 AutoDev has picked up this issue and started working on it.
Task ID:
d4c48f981934baef
Working branch:bugfix/autodev-2469-20261008201723155
Target branch:latest dev* branch
Workflow: opsp (analysis → coding → testing → PR)I will post the PR link here once done. If I need more information, I will ask in the comments.
- added a commit that references this issue
on Oct 8, 2026 ✅ AutoDev task
d4c48f981934baefcompleted.Summary: Fixed issue #2469: the memos-local-plugin capture summarizer was emitting Chinese summaries for English conversations on some models (e.g. openai/gpt-6-luna at temp 0, measured ~89% CJK on tool-call summaries and ~66% on conversational summaries). Root cause: the SYSTEM_PROMPT in apps/memos-local-plugin/core/capture/summarizer.ts anchored its language rule on an ambiguous referent ("in the user's original language") and carried a lone CJK example ("用户说了") that primed models toward Chinese under the 100-character cap. gemini-2.5-flash-lite was 0/750, confirming the leak only triggered on certain models.
Applied the reporter-validated diff: replaced the language rule with "written in the same language as the USER text (English text gets an English summary)" and removed the Chinese example from the "Do NOT prefix" rule. Added a regression unit test at tests/unit/capture/summarizer-prompt.test.ts that spies on the system message the summarizer sends and asserts (a) no Han characters leak into the prompt and (b) the USER-text anchoring is preserved — guarding both failure modes against future drift.
Verification: ran the regression tests against the pre-fix prompt first (both failed as expected per TDD), then applied the fix and re-ran. Full plugin unit suite is green — 182 test files / 1592 passed / 1 skipped. TypeScript check (tsc --noEmit -p tsconfig.json) exits 0. Committed on bugfix/autodev-2469-20261008201723155 and pushed to origin; opsp task file archived to the sibling specs repo.
Base branch:
main
Branch:bugfix/autodev-2469-20261008201723155
Commit:50b1e5d
PR: #2471
Assigned to: @syzsunshine219
Reviewers: @whipser030, @hijzy- addedai:testingAI agent is running tests | AI 正在运行测试AI agent is running tests | AI 正在运行测试and removedai:codingAI agent is coding | AI 正在编码AI agent is coding | AI 正在编码
on Oct 8, 2026 - added a commit that references this issue
on Oct 8, 2026 - addedai:failedAI task failed | AI 任务失败AI task failed | AI 任务失败and removedai:testingAI agent is running tests | AI 正在运行测试AI agent is running tests | AI 正在运行测试ai:taskDispatched to AI coding agent | 已派发给 AI 编码任务Dispatched to AI coding agent | 已派发给 AI 编码任务
on Oct 8, 2026
Summary
With some models, the capture summarizer writes Chinese summaries for English-only conversations. The cause is in the system prompt in
apps/memos-local-plugin/core/capture/summarizer.ts(SYSTEM_PROMPT, upstreammain):"The user's original language" is ambiguous: the model has to work out which text is "the user's". The only non-English text in the prompt is the
"用户说了"example, and it primes the model toward Chinese. A model that is also packing output into the 100-character cap seems to pick the denser script.Evidence
Model:
openai/gpt-6-lunathrough OpenRouter, temperature 0, reasoning off. The inputs were real English exchanges with no CJK characters in them.capture.summarize, inputs with tool callscapture.summarize, conversational inputsExample: the user turn "Hey -- can you kick off a heartbeat…" was summarized as
{"summary":"用户请求手动触发一次 heartbeat,以便测试。"}.gemini-2.5-flash-liteon the same prompt produced 0 / 750, so the problem depends on the model. The prompt still leaves the language choice open.Suggested fix (tested)
Name the source of the language explicitly, and drop the Chinese example:
On the same model and the same English exchanges, this gave 0 / 10 CJK on a first run and 0 / 30 on a wider sample. Chinese input still gets a Chinese summary, because the rule follows the USER text.
Impact
The summary feeds
traces_ftsand is the snippet text injected at turn start. A wrong-language summary weakens keyword recall against English queries, and it shows up in injected context and the viewer.