Skip to content

Establish a golden evaluation for memory usefulness #4

Description

@miniLV

This was generated by AI during triage.

What to build

Establish a small deterministic golden evaluation that proves the memory system helps answer future engineering questions, rather than only proving that each pipeline step runs.

The evaluation should exercise a complete path from representative local session fixtures through the bounded evidence and Wiki surfaces to the read-only query contract. It should make regressions visible without requiring live personal session logs or a model-as-judge.

Acceptance criteria

  • The corpus contains 10–20 representative questions covering recent state, exact tickets and paths, historical decisions, root causes, superseded conclusions, reusable guidance, and no-match cases.
  • At least one long multi-workstream session fixture verifies whether important earlier outcomes survive the bounded Evidence Snapshot.
  • Expected facts, expected evidence pages or cards, and expected NO_MATCH behavior are explicit and reviewable.
  • Evaluation output reports fact hits, citation correctness, false reuse, and no-match correctness per case and in summary.
  • The first version is deterministic and local-substitutable; it does not require an external model judge.
  • Fixtures contain no private user data and are safe to keep in the repository.
  • Existing tests and strict Wiki lint remain green.

Blocked by

None — can start immediately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestready-for-agentScoped and ready for an agent to implement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions