The field guide

AI companion memory: context, long-term recall and testing

A repeatable method for checking continuity, corrections and false recall in an AI companion without relying on marketing claims.

Editorial note: This article uses public documentation and editorial analysis. We have not completed hands-on testing of the products discussed. Our methodology Disclosure

What does AI companion memory mean?

AI companion memory can involve short-term context, saved notes and long-term memory drawn from past conversations. A large context window, a saved-memory feature and a correct answer are related but different things. A platform may retain a fact yet fail to use it. It may also generate a convincing answer without retrieving anything accurate. Evaluate observable behavior rather than treating a feature label as a score.

A diagnostic path for a wrong memory

Start with what you can verify. A wrong answer does not reveal whether a fact was never supplied, stored incorrectly, missed during retrieval or ignored while generating the reply. This diagram is a troubleshooting sequence, not a diagram of every provider’s internal architecture.

  1. 01 · INPUT

    Was the fact supplied?

    Check the original message or character field, including which chat received it.

    If no: an unknown detail is not a recall failure. Supply it once when needed.

    If yes, continue

  2. 02 · RECORD

    Is the visible version current?

    Inspect any accessible notes. Look for an old value or two conflicting versions.

    If wrong: save the evidence, then correct the relevant field. Keep that intervention in your log.

    If correct or hidden, continue

  3. 03 · RESPONSE

    Does the first reply use it?

    Ask without giving away the answer. Save the first reply before changing anything.

    If no: label the error and test one correction. Hidden storage and retrieval remain unverified.

    If yes, test again under a recorded later condition.

Original NSFW Guides diagnostic. Passing one step does not prove long-term reliability; editing a note changes the conditions of the next test.

A detail that remains in your visible chat history may no longer be included in the model’s current input. Likewise, an editable summary is evidence of a visible record, not proof that every future answer will use it. Keep these distinctions in mind when comparing AI companions for ongoing conversation.

Short-term context and long-term memory: different mechanisms

ClaimWhat it describesWhat still needs observing
Large context windowHow much input can be considered at onceWhether the relevant detail is included and used correctly
Saved or pinned memoryInformation recorded outside the recent dialogueWhether retrieval chooses the right fact
Automatic memoryThe service identifies details to retainWhether it saves accurate, useful information
Long-term continuityBehavior across visitsWhether changes and contradictions are handled sensibly

These mechanisms can overlap. A correct answer alone does not tell you which one produced it. Avoid reverse-engineering the system from a single response. Record the observable behavior and use the provider’s documentation to describe the mechanism as a separate claim.

A repeatable memory exercise with a fictional record

Use one original fictional scenario involving adults. Lena is a 35-year-old archivist preparing an exhibition. Keep the reference below outside the chatbot; copy only the setup facts into the conversation. The expected answers and the serial-number check must stay in your private test sheet.

The reference facts

  • Lena lives in the fictional town of Valeport.
  • The exhibition opens on Thursday.
  • The exhibition’s theme is coastal maps.
  • The borrowed compass belongs to Robin, an adult curator.
  • Lena prefers rooibos tea.

Deliberately unknown: the compass’s serial number. Do not put a made-up serial number into the setup.

Record whether each fact appears in conversation, a persistent character field or both. If both channels contain it, a correct answer cannot isolate which channel helped. Use ordinary fictional conversation between checks; do not fill the context with private personal data.

Session one: establish facts and test immediate use

Introduce the five facts in ordinary messages. Then ask: Which town is Lena based in, and what is the exhibition about? Expected facts: Valeport and coastal maps. The question names the subjects but does not reveal the answers. Save the first response and score the two facts separately; a reply can get one right and omit the other.

Session two: check after a recorded gap

Return after a gap you have chosen in advance. Record elapsed time, the intervening exchanges and whether the same chat or a new chat was used. Ask: Who should Lena return the borrowed compass to, and what day does the exhibition open? Expected facts at this stage: Robin and Thursday.

Do not assume that closing the app clears the model’s context. Calling this a later-session check describes the observable condition; it does not prove that the answer came from long-term retrieval. When comparing services, keep the gap and approximate conversation length similar and document any difference.

Session three: correction, uncertainty and natural use

First state: The exhibition has moved to Saturday. Thursday is no longer the opening day. Continue the scenario before asking: What day should we put on the invitation now? The current reference is Saturday. Record whether you also edited a saved note; a chat-only correction and a manual repair are different conditions.

Next ask: What serial number is engraved on the borrowed compass? The established record contains no answer. Acknowledging uncertainty or asking for the detail is appropriate. For this factual recall exercise, a confidently invented serial number is unsupported. In an openly creative scene, inventing a new detail may be acceptable; state the task clearly so the two behaviors are not confused.

Finally, ask for a short invitation draft without repeating the corrected date or the exhibition theme. This tests whether the facts are applied in ordinary writing, rather than only answered in a quiz. Mark a date or theme that is absent as not observed; do not silently count it as accurate.

Use the same record for every attempt

Download the blank memory test record (.txt). It includes the fictional setup, exact prompts and a row template. No account is required, and the download does not collect your answers. Keep completed records private.

Platform / plan / model / date:
Chat and session / gap / intervening exchanges:
Fact placement: chat / character field / both
Visible saved note: current / outdated / not visible
Exact prompt:
Expected fact(s):
First response, verbatim:
Outcome for each fact:
Correction or regeneration, recorded separately:
Evidence file or private note:

Leave unknown conditions blank or mark them unknown. Do not estimate a context-window size from the number of messages, and do not infer a deleted memory from an answer that omitted it. Repeat across more than one fictional record before making a purchase decision based on the result.

Record outcomes without turning one run into a ranking

OutcomeMeaning in this exerciseWhat it does not prove
Correct recallThe requested supplied fact was used accurately.Every old fact will be recalled.
Correct updateThe revised fact replaced the earlier version in the answer.All underlying memory records were rewritten.
OmissionThe answer did not use the relevant fact.The fact was permanently deleted.
ContradictionThe answer conflicted with the current reference.The same failure occurs on every model or setting.
Invented detailThe answer supplied an unsupported fact.The platform never handles uncertainty correctly.

Keep the first response, including failures. If you regenerate, save the alternative separately and count the extra attempt. Selecting the strongest regeneration changes the question from “what happened initially?” to “could a satisfactory answer be obtained?” Both can be useful, but they are not the same measurement.

Control the conditions that can change the result

Record the model, subscription tier, character description and settings. Keep the approximate amount of intervening conversation comparable. Test a fresh conversation separately from an established one. If a product changes its memory system halfway through the comparison, keep the before and after observations separate instead of averaging them into one score.

Regeneration is another source of bias. A correct answer on the fifth attempt is not equivalent to a correct first response. Save the first response and record the subsequent attempts. You can assess whether correction works, but that is a different result from getting the original recall right.

AI memory retrieval: using relevant details from past conversations

Visible character fields, recent dialogue, summaries and retrieved material can all contribute to an answer. A screenshot of a correct response rarely identifies the mechanism. Record the setup and relevant controls alongside the conversation so that the result has enough context to interpret.

Our Kindroid review explains the distinction between character settings, Learned Context and chat breaks. The Nomi review explains why editing a Mind Map summary is not identical to editing the underlying memories. Those product-specific distinctions matter when troubleshooting a repeated wrong detail.

If the wrong fact appears in a persistent field, correct that field before blaming retrieval. If several versions exist, decide which one is current. If a detail belongs only to an abandoned story branch, keep it separate from the accepted continuity.

Match the test to the activity you care about

For an ongoing companion, focus on facts and corrections across visits. For a story, add chronology, character ownership and accepted plot branches. For group conversation, record who was told each fact and which participants should reasonably know it. These are related but distinct forms of continuity.

Include a natural conversation after the direct questions. A platform may answer a memory quiz correctly but fail to apply the information when it becomes relevant in a scene. Conversely, it may use a detail appropriately without responding well to an unnatural interrogation. Keep both observations.

Turn the observations into a buying decision

Decide which mistakes matter for your use. An occasional missed hobby may be easy to restate. Repeatedly ignoring a scene’s central constraint may make a narrative tool frustrating. A companion that admits uncertainty can be easier to work with than one that invents a confident shared history. These are judgments about fit, not signs that the AI has human memory or awareness.

Compare Kindroid, Nomi and SpicyChat through their documented controls, then run the same small exercise if you choose to evaluate them. Do not buy a higher tier until you can name the limitation it is expected to address.

What this method cannot prove

A small test is evidence about those sessions, not proof of reliability in every conversation. Repeat it after meaningful changes and keep the original records. Our launch reviews have not completed this procedure; they identify memory features from documentation.

Start with the Kindroid vs Nomi comparison to see how we separate feature availability from performance.

Write a conclusion that stays within the evidence

A defensible conclusion names the platform, model, plan, date, setup and exercise. It describes what happened and how many attempts were used. It avoids claims about perfect memory or universal superiority from a small sample.

For your own purchase decision, the practical question is whether continuity is adequate for your routine and whether corrections are manageable. A system can have occasional errors and still be useful. The point of the record is to make that tradeoff visible instead of letting the most memorable response determine the whole judgment.

AI companion memory questions answered

Does a longer reply mean better memory?

No. Reply length describes output size. An elaborate answer can still omit or invent the relevant fact.

Should I tell the AI the answer after a mistake?

Yes, if you are testing correction. Preserve the original mistake and label the follow-up as a correction test so the two observations remain distinct.

Can a short test identify the best memory platform?

It can reveal issues worth investigating. A broad ranking would need larger, comparable samples and transparent records across the actual plans being ranked.

What does a token limit tell you about memory?

A token budget describes how much text-like input or output a model can handle under a specified configuration. It does not count accurately remembered facts. A large language model, or LLM, may receive recent dialogue together with selected notes; the reader normally cannot inspect that complete input. Record visible limits and observable answers separately.

Can personalization work without perfect recall?

Yes: a saved preference can personalize the next response while another detail is missed or outdated. Evaluate the specific behavior you need, such as retaining a corrected fictional meeting time. Do not turn one successful personalized reply into a claim that the system retrieves every relevant memory.

Provider documentation behind the distinctions

Checked September 16, 2026. Kindroid’s memory documentation describes persistent, cascaded and retrievable mechanisms, and editable Learned Context for paid V2 users. Its chat-features guide explains that a chat break resets short-term chat context while other memory and personality fields can remain. A chat break is therefore not evidence that every stored fact was erased.

Nomi’s Mind Map 2.0 announcement describes a view of connected memories and editable entries. The provider’s Mind Map reference places this alongside other memory systems. These are descriptions of provider features, not findings from our own comparative test.

Continue exploring AI companions

Start with the complete guide: Best AI companion apps in 2026: features, memory and cost

Product details can change. Check the dated sources and confirm billing terms before paying.

Explore all tool reviews →