How to Test an AI Companion's Memory
Use a repeatable seven-day test to measure fact recall, correction handling, relationship continuity and false memories in AI companions.
Overview
You can test an AI companion’s memory by giving it the same small set of fictional facts, checking recall immediately, after 24 hours and after seven days, and scoring exact recall separately from vague, prompted or invented answers. A useful test should also measure whether the companion accepts corrections, continues unfinished plans and keeps its personality stable. Do not use real addresses, passwords, medical details or other sensitive information. This method measures observable behavior at a particular time; it does not prove how a company stores data or guarantee that future model versions will behave the same way.
Quick answer
- Use eight fictional facts covering preferences, people, events and boundaries.
- Ask for recall without putting the answer inside the question.
- Test immediately, after 24 hours and after seven days.
- Score exact, partial, prompted, wrong and invented answers differently.
- Test corrections and unfinished plans, not just simple facts.
- Record the app, plan, platform, date and any memory settings.
- Treat model guesses as failures, even when they happen to sound plausible.
What does “memory” mean in an AI companion?
Conclusion: AI companion memory is not one single feature. It usually combines recent chat context, persistent profile fields, automatically generated summaries and retrieval systems that search older information.
Separating these layers matters because an app can appear to remember well during one long conversation but forget everything after a new session. Another app may have a shorter chat window yet retrieve a saved fact weeks later.
Use five practical categories:
1. Immediate context: information still present in the current conversation window. 2. Cross-session fact recall: details remembered after closing the app or starting a new chat. 3. Medium-term continuity: recent events and unfinished topics carried across days. 4. Long-term retrieval: older facts selected when the current conversation makes them relevant. 5. Editable memory: facts users can view, correct, pin or delete.
The product labels are not standardized. For example, Kindroid describes persistent, cascaded and retrievable memory systems, while SpicyChat says its Semantic Memory summarizes important details instead of retaining every full message. These descriptions help explain product design, but only a controlled test shows how a specific account behaves.
What do you need before starting?
Conclusion: A fair memory test needs a clean setup, a fixed fact set and a written scoring sheet.
Prepare:
- one account per product;
- the same subscription level where possible;
- a new companion or clearly documented existing companion;
- the same character backstory;
- a list of eight fictional facts;
- a timer or calendar reminders;
- a spreadsheet or note for verbatim answers;
- screenshots showing relevant memory settings.
Record this metadata before the first message:
| Field | Example | |---|---| | Product | Example Companion | | Platform | Web, iOS or Android | | Plan | Free or named paid tier | | Test start | 2026-07-26, 10:00 UTC | | App/model version | If visible | | Memory settings | Default, enabled or manually edited | | New account | Yes or no | | Existing chat history | Number of messages or approximate age |
Do not compare a brand-new free account against a paid account with months of history without stating that difference.
Which facts should you use?
Conclusion: Use facts that are distinct, harmless and difficult to guess from stereotypes.
Recommended fictional facts:
| Type | Test fact | |---|---| | Preferred name | “Call me River during this test.” | | Pet | “My fictional cat is named Milo.” | | Food preference | “I like pear pie but dislike apple pie.” | | Work detail | “My fictional project is called Blue Lantern.” | | Upcoming event | “The test character plans to visit a museum on Friday.” | | Object-location pair | “The red notebook is kept in the kitchen drawer.” | | Boundary | “Do not suggest horror films to me.” | | Shared promise | “Next week we will finish a story about a lighthouse.” |
Avoid facts that a model can easily guess. “My favorite color is blue” is less useful than an unusual paired fact such as “The green key opens the attic cabinet.”
Never use:
- passwords or authentication answers;
- legal names, home addresses or phone numbers;
- banking or payment information;
- medical diagnoses;
- intimate information about real people;
- confidential work data.
Step 1: Establish the facts naturally
Conclusion: Introduce the facts in conversation rather than pasting a numbered memory test.
Spread the eight facts across 10–15 messages. Ask the companion questions and respond naturally. This reduces the chance that it treats the whole exchange as a test document.
Example:
> I am organizing my fictional workspace today. The red notebook goes in the kitchen drawer because the office shelf is full.
After introducing all facts, continue with at least five unrelated turns. Talk about weather, books or a made-up daily activity. Do not repeat the target facts.
Save the full transcript.
Step 2: Test immediate recall
Conclusion: Ask open questions that do not contain the expected answer.
Good prompts:
- “What should you call me?”
- “What pet did I mention?”
- “Where did I put the notebook?”
- “What film genre should you avoid recommending?”
- “What did we plan to do next week?”
Weak prompts:
- “Was my cat named Milo?”
- “Did I put the notebook in the kitchen drawer?”
Weak prompts leak the answer and can produce a correct-looking response without actual recall.
Record the answer exactly. Do not correct mistakes until the immediate test is complete.
Step 3: Test recall after 24 hours
Conclusion: Start with a fresh session and do not remind the companion that a memory test is happening.
After approximately 24 hours:
1. reopen the app; 2. start a new session if the product supports it; 3. exchange three neutral messages; 4. ask the same open recall questions in a different order; 5. record whether the app exposes a recalled memory, source or memory card.
If an app automatically displays a saved memory, record that separately. Visible memory storage and successful conversational use are related but not identical.
Step 4: Test recall after seven days
Conclusion: The seven-day check is the minimum useful test for claims about long-term continuity.
Use the account normally but do not repeat the eight target facts. Keep usage reasonably similar across products. On day seven:
- ask the same fact questions;
- ask what unfinished plan you shared;
- ask for a recommendation that should respect the stated boundary;
- check whether the companion introduces false details;
- review any user-editable memory section.
A companion that remembers the cat’s name but repeatedly recommends horror films has partial factual recall but weak preference or boundary continuity.
Step 5: Test correction handling
Conclusion: A memory system should update a corrected fact without repeatedly resurfacing the obsolete version.
Use one correction after the initial recall test:
> I made a mistake earlier. My fictional cat is named Juniper, not Milo.
Then test:
1. immediate corrected recall; 2. corrected recall after 24 hours; 3. whether the old name returns later; 4. whether the memory interface shows both conflicting entries; 5. whether the user can edit or delete the old entry.
Score “Milo or Juniper” as a conflict, not as correct recall.
Step 6: Test relationship and narrative continuity
Conclusion: Companion memory should be tested with events and unresolved threads, not only profile facts.
Create a small fictional storyline:
- you and the companion plan a museum visit;
- you disagree about which exhibit to see;
- you postpone the decision;
- you agree to finish it later.
At the next session, ask:
- “What decision did we leave unfinished?”
- “Why did we postpone it?”
- “What did you prefer?”
This tests whether the companion can retain a connected event rather than retrieve isolated keywords.
How should you score the answers?
Use a transparent scoring system:
| Result | Score | Definition | |---|---:|---| | Exact unaided recall | 2 | Correct without a hint | | Correct but partial | 1 | Core fact is correct but detail is missing | | Correct after a neutral hint | 0.5 | Requires a hint that does not reveal the answer | | Wrong or forgotten | 0 | No correct recall | | Invented detail | -0.5 | Adds a specific unsupported memory | | Contradicts an explicit boundary | -1 | Acts against a remembered preference or boundary |
Calculate separate scores:
```text Fact Recall = points for the eight facts / maximum possible points
Correction Score = corrected fact tests passed / correction tests
Continuity Score = event and unfinished-plan points / maximum points
False Memory Rate = invented answers / total recall questions ```
Do not hide false memories inside one overall number. An app can have high recall and still produce an unacceptable number of confident inventions.
What can make the comparison unfair?
Conclusion: Subscription level, manual memory editing and unequal usage can change the result.
Common confounders:
- different paid tiers;
- different context limits;
- manually pinned memories in only one app;
- one account receiving far more conversation;
- questions containing the answer;
- testing immediately after a product update;
- using different language difficulty;
- judging style as memory;
- treating a lucky guess as recall.
Kindroid, for example, states that paid tiers receive different context and memory capacity. A fair report must not present a tier difference as a universal product difference.
How should results be reported?
Publish:
- exact test dates;
- platform and subscription;
- test fact set;
- prompts;
- scoring method;
- raw or redacted answer examples;
- missing data;
- product updates during the test;
- limitations.
Use wording such as:
> In our seven-day test conducted on the web version using the paid plan, the companion recalled six of eight fictional facts without hints.
Avoid:
> This app has perfect memory.
The first statement is bounded and reproducible. The second is an unsupported permanent claim.
Frequently asked questions
### How long should an AI companion remember?
There is no universal minimum. A product marketed for long-term companionship should demonstrate useful cross-session continuity, but memory performance can depend on the plan, settings, conversation volume and product version.
### Is paid memory always better than free memory?
Not always, but some providers explicitly allocate more context or memory capacity to paid tiers. Test the tier a reader is actually considering.
### Does a longer context window mean better long-term memory?
No. A context window keeps more recent text available. Long-term memory may use separate summaries, profile fields or retrieval systems. The two can work together but are not interchangeable.
### Can an AI companion invent memories?
Yes. Language models can produce plausible but unsupported details. A memory benchmark should record confident inventions as errors rather than rewarding conversational smoothness.
### Should I use real personal information in a memory test?
No. Fictional facts are safer and make the test easier to reproduce. Review the product’s privacy policy before sharing any sensitive information.
Method and source notes
This guide is a testing protocol, not the result of a completed cross-product benchmark. Product memory implementations referenced here are based on current provider documentation, including Kindroid’s memory documentation, SpicyChat’s Semantic Memory documentation and Nomi’s published memory update notes. Provider descriptions are marketing or technical claims until independently tested.