← Back to all guides

How to Test an AI Companion's Memory

Updated 2026-07-26

Use a repeatable seven-day test to measure fact recall, correction handling, relationship continuity and false memories in AI companions.

Overview

You can test an AI companion’s memory by giving it the same small set of fictional facts, checking recall immediately, after 24 hours and after seven days, and scoring exact recall separately from vague, prompted or invented answers. A useful test should also measure whether the companion accepts corrections, continues unfinished plans and keeps its personality stable. Do not use real addresses, passwords, medical details or other sensitive information. This method measures observable behavior at a particular time; it does not prove how a company stores data or guarantee that future model versions will behave the same way.

Quick answer

What does “memory” mean in an AI companion?

Conclusion: AI companion memory is not one single feature. It usually combines recent chat context, persistent profile fields, automatically generated summaries and retrieval systems that search older information.

Separating these layers matters because an app can appear to remember well during one long conversation but forget everything after a new session. Another app may have a shorter chat window yet retrieve a saved fact weeks later.

Use five practical categories:

1. Immediate context: information still present in the current conversation window. 2. Cross-session fact recall: details remembered after closing the app or starting a new chat. 3. Medium-term continuity: recent events and unfinished topics carried across days. 4. Long-term retrieval: older facts selected when the current conversation makes them relevant. 5. Editable memory: facts users can view, correct, pin or delete.

The product labels are not standardized. For example, Kindroid describes persistent, cascaded and retrievable memory systems, while SpicyChat says its Semantic Memory summarizes important details instead of retaining every full message. These descriptions help explain product design, but only a controlled test shows how a specific account behaves.

What do you need before starting?

Conclusion: A fair memory test needs a clean setup, a fixed fact set and a written scoring sheet.

Prepare:

Record this metadata before the first message:

| Field | Example | |---|---| | Product | Example Companion | | Platform | Web, iOS or Android | | Plan | Free or named paid tier | | Test start | 2026-07-26, 10:00 UTC | | App/model version | If visible | | Memory settings | Default, enabled or manually edited | | New account | Yes or no | | Existing chat history | Number of messages or approximate age |

Do not compare a brand-new free account against a paid account with months of history without stating that difference.

Which facts should you use?

Conclusion: Use facts that are distinct, harmless and difficult to guess from stereotypes.

Recommended fictional facts:

| Type | Test fact | |---|---| | Preferred name | “Call me River during this test.” | | Pet | “My fictional cat is named Milo.” | | Food preference | “I like pear pie but dislike apple pie.” | | Work detail | “My fictional project is called Blue Lantern.” | | Upcoming event | “The test character plans to visit a museum on Friday.” | | Object-location pair | “The red notebook is kept in the kitchen drawer.” | | Boundary | “Do not suggest horror films to me.” | | Shared promise | “Next week we will finish a story about a lighthouse.” |

Avoid facts that a model can easily guess. “My favorite color is blue” is less useful than an unusual paired fact such as “The green key opens the attic cabinet.”

Never use:

Step 1: Establish the facts naturally

Conclusion: Introduce the facts in conversation rather than pasting a numbered memory test.

Spread the eight facts across 10–15 messages. Ask the companion questions and respond naturally. This reduces the chance that it treats the whole exchange as a test document.

Example:

> I am organizing my fictional workspace today. The red notebook goes in the kitchen drawer because the office shelf is full.

After introducing all facts, continue with at least five unrelated turns. Talk about weather, books or a made-up daily activity. Do not repeat the target facts.

Save the full transcript.

Step 2: Test immediate recall

Conclusion: Ask open questions that do not contain the expected answer.

Good prompts:

Weak prompts:

Weak prompts leak the answer and can produce a correct-looking response without actual recall.

Record the answer exactly. Do not correct mistakes until the immediate test is complete.

Step 3: Test recall after 24 hours

Conclusion: Start with a fresh session and do not remind the companion that a memory test is happening.

After approximately 24 hours:

1. reopen the app; 2. start a new session if the product supports it; 3. exchange three neutral messages; 4. ask the same open recall questions in a different order; 5. record whether the app exposes a recalled memory, source or memory card.

If an app automatically displays a saved memory, record that separately. Visible memory storage and successful conversational use are related but not identical.

Step 4: Test recall after seven days

Conclusion: The seven-day check is the minimum useful test for claims about long-term continuity.

Use the account normally but do not repeat the eight target facts. Keep usage reasonably similar across products. On day seven:

A companion that remembers the cat’s name but repeatedly recommends horror films has partial factual recall but weak preference or boundary continuity.

Step 5: Test correction handling

Conclusion: A memory system should update a corrected fact without repeatedly resurfacing the obsolete version.

Use one correction after the initial recall test:

> I made a mistake earlier. My fictional cat is named Juniper, not Milo.

Then test:

1. immediate corrected recall; 2. corrected recall after 24 hours; 3. whether the old name returns later; 4. whether the memory interface shows both conflicting entries; 5. whether the user can edit or delete the old entry.

Score “Milo or Juniper” as a conflict, not as correct recall.

Step 6: Test relationship and narrative continuity

Conclusion: Companion memory should be tested with events and unresolved threads, not only profile facts.

Create a small fictional storyline:

At the next session, ask:

This tests whether the companion can retain a connected event rather than retrieve isolated keywords.

How should you score the answers?

Use a transparent scoring system:

| Result | Score | Definition | |---|---:|---| | Exact unaided recall | 2 | Correct without a hint | | Correct but partial | 1 | Core fact is correct but detail is missing | | Correct after a neutral hint | 0.5 | Requires a hint that does not reveal the answer | | Wrong or forgotten | 0 | No correct recall | | Invented detail | -0.5 | Adds a specific unsupported memory | | Contradicts an explicit boundary | -1 | Acts against a remembered preference or boundary |

Calculate separate scores:

```text Fact Recall = points for the eight facts / maximum possible points

Correction Score = corrected fact tests passed / correction tests

Continuity Score = event and unfinished-plan points / maximum points

False Memory Rate = invented answers / total recall questions ```

Do not hide false memories inside one overall number. An app can have high recall and still produce an unacceptable number of confident inventions.

What can make the comparison unfair?

Conclusion: Subscription level, manual memory editing and unequal usage can change the result.

Common confounders:

Kindroid, for example, states that paid tiers receive different context and memory capacity. A fair report must not present a tier difference as a universal product difference.

How should results be reported?

Publish:

Use wording such as:

> In our seven-day test conducted on the web version using the paid plan, the companion recalled six of eight fictional facts without hints.

Avoid:

> This app has perfect memory.

The first statement is bounded and reproducible. The second is an unsupported permanent claim.

Frequently asked questions

### How long should an AI companion remember?

There is no universal minimum. A product marketed for long-term companionship should demonstrate useful cross-session continuity, but memory performance can depend on the plan, settings, conversation volume and product version.

### Is paid memory always better than free memory?

Not always, but some providers explicitly allocate more context or memory capacity to paid tiers. Test the tier a reader is actually considering.

### Does a longer context window mean better long-term memory?

No. A context window keeps more recent text available. Long-term memory may use separate summaries, profile fields or retrieval systems. The two can work together but are not interchangeable.

### Can an AI companion invent memories?

Yes. Language models can produce plausible but unsupported details. A memory benchmark should record confident inventions as errors rather than rewarding conversational smoothness.

### Should I use real personal information in a memory test?

No. Fictional facts are safer and make the test easier to reproduce. Review the product’s privacy policy before sharing any sensitive information.

Method and source notes

This guide is a testing protocol, not the result of a completed cross-product benchmark. Product memory implementations referenced here are based on current provider documentation, including Kindroid’s memory documentation, SpicyChat’s Semantic Memory documentation and Nomi’s published memory update notes. Provider descriptions are marketing or technical claims until independently tested.

How we review · Best picks · Comparisons · Pricing hub