Skip to content

Methodology review requested: regulated-memory evaluation #40

Description

@ebeirne

We are preparing a preprint around a narrow question: can an agent memory layer prove what the agent knew at a specific time?

The evaluation covers five invariants:

  1. Stale-value suppression after correction
  2. Point-in-time reconstruction
  3. Verifiable erasure
  4. Lookahead-bias prevention
  5. Auditable memory state

The draft paper and artifact structure are in #26. A fresh Lians 0.4.1 local run is archived as schema-validated JSON under paper/regulated-memory-eval/evidence/.

We are specifically asking reviewers to challenge:

  • Whether each invariant is defined narrowly enough to reproduce
  • Whether pass and fail criteria are implementation-neutral
  • Whether default configurations are being represented fairly
  • Which exact Mem0 OSS and Graphiti OSS versions and public APIs should be used for the reruns
  • Whether any claim goes beyond the archived evidence

We will not reconstruct missing raw outputs from prose. Competitor rows will not be treated as publication-ready until fresh runs and exact JSON artifacts are archived. Material vendor corrections will be incorporated before submission.

If you review only one thing, please inspect the invariant definitions and identify a counterexample that our harness would score incorrectly.

Disclosure: this evaluation is maintained by the Lians authors, and Lians is one of the evaluated systems. Runnable adapters, raw outputs, and explicit limitations are the intended safeguards against favorable interpretation.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions