Skip to content

Improve memory eval source reconciliation and depth reporting - #256

Merged
senamakel merged 1 commit into
tinyhumansai:mainfrom
senamakel:accuracy-paraphrase-retrieval
Oct 11, 2026
Merged

senamakel merged 1 commit into
tinyhumansai:mainfrom
senamakel:accuracy-paraphrase-retrieval

Conversation

@senamakel

@senamakel senamakel commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Summary

Improve the memory_eval --llm answer instruction so it checks all memory sections and distinguishes an unresolved conflict with an undated source from an explicit dated update. Record the prompt version in JSON reports. Add --brain-limit to measure retrieval depth without changing the production six-document policy, and document the accuracy follow-up after #255.

The CortexDB scale paraphrase still misses. Max Recall flags and a 24-document brain section did not improve its 1/2 pack-hit score; the wider section raised mean pack size from 130 to 512 tokens. This PR changes benchmark interpretation and measurement, not core retrieval or OpenHuman's production answer prompt.

Related issue

Continues #251 and #255.

API or behavior changes

No public library API change. The eval CLI accepts --brain-limit <n>; its JSON report now includes brain_limit and probe_answer_prompt. The --llm answer instruction changes to reconcile sources and explicit updates.

Benchmarks

Fresh CortexDB v0.10.4/OpenRouter runs, OpenHuman v3 profile, team limit 3, five-second deadline:

Run Pack hits Model answers, recall Model answers, synthesis Timeouts
Seven-question live slice, original prompt 7/7 5/7 5/7 0
Seven-question live slice, final prompt 7/7 7/7 6/7 0
Full mock, first conflict prompt 51/58 43/58 44/58 0
Full mock, final prompt 51/58 50/58 50/58 0

The remaining live synthesis miss correctly names the Team plan but also mentions the superseded Enterprise plan, which the strict grade rejects. Mock extractive answers stayed 34/58. The mock uses keyword-based model doubles; its pack misses are not an estimate of live semantic retrieval. Fresh runs do not establish a confidence interval.

A separate 100-document middle-needle sweep (seed 251, ranked-readiness wait 30 seconds) scored 1/2 pack hits under each of: baseline with six brain documents, Max Recall server flags with six, and baseline with 24. The paraphrase missed in all three; no probe timed out. This result is why the default limit remains six.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo build --all-targets --all-features
  • cargo test --all-features
  • cargo test -p tinymemory-integrations --features full --example memory_eval brain_limit_override_is_parsed_without_changing_the_default
  • Live and full mock benchmark runs above

Tests

Added a parser regression test for the brain-limit override and its default. Model prompt behavior depends on an external model, so it was checked by replaying saved packs and by the fresh benchmark runs above.

Documentation

Updated docs/evals/README.md and docs/evals/openhuman-host.md with the new flag, method, results, and limitations.

Checklist

  • Focused change
  • No relaxed lints or ignored tests
  • No secrets or credentials in the diff

Summary by CodeRabbit

  • New Features
    • Added a --brain-limit option to evaluation runs, with the selected limit included in JSON reports.
    • Updated evaluation answers to preserve both values when conflicting sources are undated, and to favor a newer value when it explicitly replaces an older dated one.
  • Documentation
    • Documented the retrieval-depth option and added evaluation results for retrieval settings and conflicting memory values.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@tinysweeper

tinysweeper Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 5 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Incomplete
Priority: none
Reviewed head: 6ee20d120f50
Updated: 2026-10-10T21:48:27Z

Review snapshot

Change surface Files Review signal Count
Production 2 Active findings 0
Tests 1 Noted findings 0
Documentation 2 Resolved findings 0
Configuration 0 Pending checks/questions 2

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

No active actionable findings.

Could not review: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs

Before merge

  • Complete the critique review for crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs.
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: incomplete; unanswered: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs
  • Lane summary: Reviewed 3 files; 0 findings. 2 files could not be reviewed: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs. _The code index for this repository is cold, so this review saw the diff alone._ _3 memory call(s) failed (model: cortex: v1/answer: timed out after 20s), so this review saw part of what the engine holds._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 3 files; 0 findings. 2 files were not security-reviewed: docs/evals/README.md (prose or tabular data), docs/evals/openhuman-host.md (prose or tabular data). _The code index for this repository is cold, so this review saw the diff alone._ _3 memory call(s) failed (model: cortex: v1/answer: timed out after 20s), so this review saw part of what the engine holds._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This change adds a `--brain-limit` flag and rewrites the eval answer prompt. The flag parsing is covered by a new deterministic test that checks the default, the parsed value, and the error path, and the prompt is a versioned constant whose change is exactly the kind docs record rather than something needing a test. The change looks sound. _The code index for this repository is cold, so this review saw the diff alone._ _3 memory call(s) failed (model: cortex: v1/answer: timed out after 20s), so this review saw part of what the engine holds._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The description accurately matches the diff: the eval answer prompt changes with a recorded version, `--brain-limit` is added with a parser test and JSON report fields, and the docs describe the same benchmarks. I found no reportable defects introduced by this change; it looks safe to merge. _The code index for this repository is cold, so this review saw the diff alone._ _3 memory call(s) failed (model: cortex: v1/answer: timed out after 20s), so this review saw part of what the engine holds._
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.007025
  • Tokens: 112796 input · 4016 output · 7306 cached · 0 embedding
Head State Pass summary
6ee20d120f50 incomplete 0 active finding(s), 0 resolved finding(s) (at 2026-10-10T21:48:27Z)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 89c21cb7-addc-4707-a72d-30a1b79aede4

📥 Commits

Reviewing files that changed from the base of the PR and between 86efe3a and 6ee20d1.


📒 Files selected for processing (5)
  • crates/tinymemory-integrations/examples/memory_eval/llm.rs
  • crates/tinymemory-integrations/examples/memory_eval/main.rs
  • crates/tinymemory-integrations/examples/memory_eval/main_tests.rs
  • docs/evals/README.md
  • docs/evals/openhuman-host.md

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.



📝 Walkthrough

Walkthrough

The evaluation example adds source-reconciliation instructions to its answer prompt and adds a CLI option to set the brain-document limit. The JSON report records the prompt version and effective limit. Evaluation documentation reports results for both changes.

Changes

Answer reconciliation

Layer / File(s) Summary
Reconciliation prompt and evaluation results
crates/tinymemory-integrations/examples/memory_eval/llm.rs, docs/evals/openhuman-host.md
The answer prompt distinguishes unresolved conflicts from explicit updates to older dated values and identifies the prompt as source-reconciliation-v2. The documentation reports live and mock benchmark results.

Brain-document limit

Layer / File(s) Summary
Parse and apply the brain-document limit
crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs, docs/evals/README.md, docs/evals/openhuman-host.md
The CLI accepts --brain-limit <n> and applies the value to RecallPolicy.brain_limit. The JSON report records the effective limit. Tests cover default, numeric, and invalid input. The documentation describes the option and retrieval-depth results.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~12 minutes

Change: Feature

Suggested reviewers: codeghost21


Merge Risk: ⚪ Minimal · up to 6ee20

The evaluation prompt handles unresolved conflicts and dated replacements, and retrieval-depth measurements can use an override without changing the default. No supported current-head issue remains that blocks merging.

Pre-merge checks | Passed 4 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 60.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 3 files. (2 skipped: 2… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title accurately summarizes the two main changes: source reconciliation in the memory evaluation prompt and retrieval-depth reporting through the brain limit option and JSON output.
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.

Full details: Docstring Coverage

Explanation

Docstring coverage is 60.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 3 files. (2 skipped: 2 unsupported.)


  • Fix all pre-merge checks with AI
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the memory trail,
Two dated notes can tell a tale.
A limit sets the depth to scan,
The report records the chosen span.
Then carrots celebrate the plan.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinymemory-integrations/examples/memory_eval/main.rs, crates/tinymemory-integrations/examples/memory_eval/main_tests.rs.

             $0.0070 · 112,796 in / 4,016 out · 7,306 cached (6%)  · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0041 · 52,645 in  / 1,650 out · 5,002 cached (10%) · gpt-5.6-luna
security:    $0.0027 · 35,311 in  / 913 out   · 2,304 cached (7%)  · gpt-5.6-luna
tests:       $0.0001 · 10,096 in  / 280 out   · 0 cached (0%)      · glm-5.3-flash
description: $0.0001 · 10,164 in  / 167 out   · 0 cached (0%)      · glm-5.3-flash

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Oct 10, 2026
@senamakel
senamakel merged commit b936f0a into tinyhumansai:main Oct 11, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant