Skip to content

Align OpenHuman memory eval with 10-second deadline - #257

Merged
senamakel merged 1 commit into
tinyhumansai:mainfrom
senamakel:cortex-needle-config
Oct 11, 2026
Merged

senamakel merged 1 commit into
tinyhumansai:mainfrom
senamakel:cortex-needle-config

Conversation

@senamakel

@senamakel senamakel commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Align the OpenHuman benchmark mirror with the proposed 10-second pre-turn memory deadline and record a fresh 1,000-document needle run. Historical 5-second measurements remain labeled as such.

The simple lexical owner question found the needle at rank 1 in 1.769 s. The semantic paraphrase missed in 1.767 s. Neither timed out. Direct CortexDB diagnostics attributed 1.067 s of a 1.101 s fresh simple recall to query embedding and 16 ms to vector search. Asking for 12, 24, or 100 events took 1.48, 1.49, and 1.43 s with fresh lexical phrasings. At this corpus size the external embedding call dominates; the longer deadline does not fix semantic ranking.

Related issue

Follow-up to #251. OpenHuman default change: tinyhumansai/openhuman#7368

API or behavior changes

No public API change. The benchmark's OpenHuman mirror waits up to 10 seconds for a pre-turn pack. Production OpenHuman changes in the related PR.

Validation

  • cargo fmt --all -- --check — passed.
  • cargo clippy --all-targets --all-features -- -D warnings — passed.
  • cargo build --all-targets --all-features — passed.
  • cargo test --all-features — passed.
  • Live CortexDB v0.10.4/OpenRouter 1,000-document OpenHuman-host eval — 1/2 pack hits, 0/2 timeouts.

Tests

The ten-second mirror assertion failed against the old five-second value, then passed after the change. The run used the existing deterministic needle fixture. Model answers were not requested; this comparison measures pack retrieval and timing.

Documentation

Updated docs/evals/openhuman-host.md with the deadline distinction, 1,000-document results and CortexDB timing breakdown.

Checklist

  • The change is focused on one logical change.
  • No new #[allow(...)], #[ignore], or relaxed lints.
  • No secrets, tokens, or .env contents in the diff or the description.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@coderabbitai

coderabbitai Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 04a897fa-c37a-479c-829f-87357350889b


📥 Commits

Reviewing files that changed from the base of the PR and between b936f0a and 094df91.



📒 Files selected for processing (4)
  • crates/tinymemory-integrations/examples/memory_eval/agent.rs
  • crates/tinymemory-integrations/examples/memory_eval/agent_tests.rs
  • crates/tinymemory-integrations/examples/memory_eval/main.rs
  • docs/evals/openhuman-host.md


Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.




📝 Walkthrough
📝 Walkthrough

Walkthrough

The OpenHuman pre-turn timeout changes from 5 to 10 seconds. The regression test and example documentation reflect the new value. The host profile also adds measurements from a 1,000-document benchmark.

Changes

OpenHuman evaluation updates

Layer / File(s) Summary
Update the pre-turn timeout
crates/tinymemory-integrations/examples/memory_eval/agent.rs, crates/tinymemory-integrations/examples/memory_eval/agent_tests.rs, crates/tinymemory-integrations/examples/memory_eval/main.rs, docs/evals/openhuman-host.md
The example timeout changes to 10 seconds. The regression test and example and host profile documentation reflect the updated deadline.
Record the 1,000-document benchmark
docs/evals/openhuman-host.md
The host profile reports probe and recall timings, diagnostic measurements, and model-call costs for the benchmark.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix



Merge Risk: ⚪ Minimal · up to 094df

The evaluation now allows ten seconds for a pre-turn memory pack. Late packs remain excluded, and no actionable merge risk is established.

Pre-merge checks | Passed 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 1 functions across 3 files. (1 skipped: 1 …
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly and concisely describes the main change: updating the OpenHuman memory evaluation to use a 10-second deadline.

  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the timer's run,
Ten seconds now, the work gets done.
A thousand notes pass through the test,
Recall times join the measured quest.
The rabbit hops, then takes a rest.

Comment @coderabbitai help to get the list of available commands.

@senamakel
senamakel merged commit c764dc3 into tinyhumansai:main Oct 11, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant