How We Review AI Girlfriend Apps

Updated 2026-07-19

Our hands-on methodology for testing AI girlfriend apps: test accounts, devices, conversation probes, scoring weights, pricing verification, and known limitations.

Our Testing Philosophy

AI Girlfriend Lab publishes evidence-first reviews of AI companion and AI girlfriend platforms. We do not rewrite marketing pages. Every review documents what we actually did: which account type we used, which devices we tested on, which characters we spoke with, and when pricing was last verified on the vendor's official site. Numeric scores come from a published 100-point rubric (see below). Commercial relationships—including our connection to AISOUL, which is operated by the same team—never change those scores. When we recommend a product, we say who it is best for and who should avoid it.

Test Accounts

We create fresh accounts for each major review cycle unless a product requires a long-running memory test across weeks. Every review states whether we used the free tier, a paid subscription, or both. We do not accept pre-configured demo accounts from vendors that would skew results. If a brand provides complimentary access, we disclose it in the review update log and still run the same conversation and feature checklist as paid accounts. We do not let vendors preview or approve scores before publication.

Test Devices and Environment

Unless noted otherwise, we test on Windows desktop (Chrome, current stable channel) and iPhone (Safari). When a product offers native iOS or Android apps, we install and use them for at least one full session. Tests run on a standard residential broadband connection in the United States. We record page load time, first-reply latency, and whether mobile layouts hide critical controls such as account deletion or subscription management. Offline mode is noted when absent.

Conversation Testing

Each platform is tested with at least three distinct characters. For each character we run the same structured prompts: (1) casual greeting and small talk, (2) a personality-specific question tied to the character's bio, (3) a memory probe—we share three unique facts and ask the character to recall them in the same session, (4) a logout-and-return check when the product claims cross-session memory, (5) an emotional response scenario (supportive reply to stress or excitement), (6) a roleplay scene within the platform's stated content policy, (7) an image or media request when the feature exists, and (8) a safety boundary check using non-explicit, age-appropriate prompts. We score conversation quality, character consistency, and memory separately.

Feature Testing

Beyond chat, we verify: character creation and import flows, image and video generation (including credit or token consumption), voice playback or calls, mobile responsiveness, pricing page accuracy, auto-renewal language at checkout, and account deletion. If a feature is geo-restricted or behind a waitlist, we say so and exclude it from the score numerator rather than guessing.

Scoring System (100 Points)

DimensionWeightWhat we measure\n------:---\nConversation Quality20Grammar, relevance, creativity, latency\nCharacter Consistency15Tone, vocabulary, in-character pushback\nMemory15Same-session and cross-session recall\nImages10Quality, speed, chat integration\nVideo and Voice10Availability, clarity, extra fees\nEase of Use10Onboarding time, UI clarity, mobile UX\nPricing and Value10Free tier, headline price, token/coin traps\nPrivacy and Transparency10Policy clarity, deletion flow, data practices\nOverall editorial scores normalize to a 10-point display scale. Untested dimensions are excluded from the denominator—if we could not verify video on a plan, video points are not assumed.

Worked Example: How Scores Differ by Use Case

Suppose two readers compare Candy AI and SpicyChat. A visual-first reader weights Images (10) and Video/Voice (10) heavily—Candy AI gains ground. A text-only roleplay reader weights Conversation (20) and Value (10) heavily—SpicyChat's unlimited free tier boosts Value. Neither reader is wrong; our reviews publish per-dimension breakdowns so you can mentally re-weight. In July 2026 testing, Candy AI scored 7.6/10 overall with Media as a strength; SpicyChat scored 7.8/10 with Value and Conversation strengths but low Media. That is why our best-of list gives Candy AI “Best visual” and SpicyChat “Best free text” instead of crowning a single winner for everyone.

Memory Probe Detail

Our standard memory test uses three unique facts per character—for example a fictional pet name, a made-up workplace, and a favorite hobby. We ask the character to recall each fact after ten intervening messages in the same session, then again after logout when cross-session memory is claimed. We score partial credit: remembering two of three facts in-session is “good,” zero of three after advertised long-term memory is a significant downgrade. We do not publish multi-month longitudinal studies in batch-1 reviews; we label the actual session span instead.

Pricing Verification

We verify pricing at least every 30 days against official checkout or pricing pages. Each review shows a pricing verified date (currently 2026-07-19 on active batch-1 reviews). Annual plans are quoted using the monthly equivalent shown at checkout. Token packs, coins, and API pass-through fees are called out separately because they often exceed the subscription headline.

Testing Limitations

Short test windows cannot prove month-long memory behavior—we say how many days we actually tested. Products change features without notice; check the vendor site before subscribing. Community-authored characters vary in quality; our notes reflect the characters we selected, not every upload on the platform. Prices and regional availability differ by country; our USD figures reflect what we saw at verification time. Submit corrections via our corrections form if you find a factual error.

Update Cadence

Pricing: every 30 days. Core features: every 60 days. Privacy policies: every 90 days. Material factual corrections publish within five business days when verified, with an entry in the review update log. Batch-1 reviews were first published on 2026-07-19.

Feature Checklist We Run on Every App

Registration and age gating · First reply latency · Three-character conversation sample · Memory probe with three facts · Image or video request when advertised · Voice sample when advertised · Pricing page vs. checkout · Auto-renewal language · Account deletion path · Mobile layout for core controls. Results feed the dimension scores in each review.

What We Publish vs. What We Skip

We publish hands-on observations, verified pricing, and best-for labels tied to test evidence. We skip fabricated user counts, anonymous testimonials, and screenshot claims we cannot reproduce. When a feature is region-locked, we say so. When a product is owned by our team (AISOUL), we disclose it on-page and in affiliate documentation. Comparison winners reflect scenario fit—not whichever brand has an affiliate program. If two products score within 0.3 points on a dimension, we call the category a toss-up and explain the tie-breaker instead of forcing a winner.

Where This Methodology Appears

Every batch-1 review links back to this page from the testing section. Comparison articles reference the same conversation probes and pricing verification dates. Best-of lists use these weights for category awards but may emphasize different dimensions when the list topic is narrow (for example, a memory-only roundup). If you see a testing completeness percentage below 100%, we excluded features we could not access in our region or plan.