Repository navigation
Improve memory eval source reconciliation and depth reporting - #256
Conversation
Co-authored-by: Medulla <medulla@tinyhumans.ai>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Tiny Sweeper reviewTiny Sweeper reviewed this change across 5 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Incomplete Review snapshot
Completeness: Incomplete FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. FindingsNo active actionable findings. Could not review: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs Before merge
Agent review detailscritique
security
tests
commits
description
Evidence and run details
|
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs.
$0.0070 · 112,796 in / 4,016 out · 7,306 cached (6%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0041 · 52,645 in / 1,650 out · 5,002 cached (10%) · gpt-5.6-luna
security: $0.0027 · 35,311 in / 913 out · 2,304 cached (7%) · gpt-5.6-luna
tests: $0.0001 · 10,096 in / 280 out · 0 cached (0%) · glm-5.3-flash
description: $0.0001 · 10,164 in / 167 out · 0 cached (0%) · glm-5.3-flash
Summary
Improve the
memory_eval --llmanswer instruction so it checks all memory sections and distinguishes an unresolved conflict with an undated source from an explicit dated update. Record the prompt version in JSON reports. Add--brain-limitto measure retrieval depth without changing the production six-document policy, and document the accuracy follow-up after #255.The CortexDB scale paraphrase still misses. Max Recall flags and a 24-document brain section did not improve its 1/2 pack-hit score; the wider section raised mean pack size from 130 to 512 tokens. This PR changes benchmark interpretation and measurement, not core retrieval or OpenHuman's production answer prompt.
Related issue
Continues #251 and #255.
API or behavior changes
No public library API change. The eval CLI accepts
--brain-limit <n>; its JSON report now includesbrain_limitandprobe_answer_prompt. The--llmanswer instruction changes to reconcile sources and explicit updates.Benchmarks
Fresh CortexDB v0.10.4/OpenRouter runs, OpenHuman v3 profile, team limit 3, five-second deadline:
The remaining live synthesis miss correctly names the Team plan but also mentions the superseded Enterprise plan, which the strict grade rejects. Mock extractive answers stayed 34/58. The mock uses keyword-based model doubles; its pack misses are not an estimate of live semantic retrieval. Fresh runs do not establish a confidence interval.
A separate 100-document middle-needle sweep (seed 251, ranked-readiness wait 30 seconds) scored 1/2 pack hits under each of: baseline with six brain documents, Max Recall server flags with six, and baseline with 24. The paraphrase missed in all three; no probe timed out. This result is why the default limit remains six.
Validation
cargo fmt --all -- --checkcargo clippy --all-targets --all-features -- -D warningscargo build --all-targets --all-featurescargo test --all-featurescargo test -p tinymemory-integrations --features full --example memory_eval brain_limit_override_is_parsed_without_changing_the_defaultTests
Added a parser regression test for the brain-limit override and its default. Model prompt behavior depends on an external model, so it was checked by replaying saved packs and by the fresh benchmark runs above.
Documentation
Updated
docs/evals/README.mdanddocs/evals/openhuman-host.mdwith the new flag, method, results, and limitations.Checklist
Summary by CodeRabbit
--brain-limitoption to evaluation runs, with the selected limit included in JSON reports.