Rank 23
EVA AI
56.9C/O Score
Illustrative preview: in this staging dataset EVA AI reads strongest on emotional response and weakest on narrative initiative under methodology v0.1.
Two material weaknesses
- Illustrative record: narrative initiative scored 50/100 in the staging scenario set.
- Illustrative record: delayed recall scored 51/100 in the staging scenario set.
Who it fits
- Best fit
- Readers weighting emotional response above the other criteria.
- Poor fit
- Readers who need verified production evidence rather than a staging preview.
Test dimensions
| Dimension | Result |
|---|---|
| Conversation quality | 64.3 |
| Local language and culture | 56 |
| Continuity and memory | 55.6 |
| Price and value | 58.8 |
| Product usability and reliability | 52 |
| Character and roleplay | 53.5 |
Category scores
| Category | C/O Score |
|---|---|
| Overall | 56.9 |
| Relationship | 61.9 |
| Roleplay | 53.6 |
| Memory | 55.7 |
| Value | 58.5 |
Memory timeline
- continuity 53/1002026-08-15
- contradiction-resistance 58/1002026-08-15
- delayed-recall 51/1002026-08-15
- memory-controls 55/1002026-08-15
- relationship-continuity 61/1002026-08-15
Local-language findings
- local-language-performance: 56/100 (Illustrative staging scenario result recorded on the published 0-100 scale.)
Pricing and paywall observations
| Tested tier | paid-standard |
|---|---|
| Observed price | $17.99 |
| Free limit | 12 messages per day |
| Observed on | 2026-08-12 |
Three alternatives
- Nomi 83.3
- Character.AI 79.2
- Kindroid 78.5
Methodology and dates
Calculated under methodology v0.1 on 2026-08-24.
Corrections
- 2026-08-24 app: apps — A destination may only be published once it has been resolved, not assumed. (Edition-specific official destinations assumed → Each edition points at the same verified official entry point until Task 7 verifies edition-specific destinations; no score change)
- 2026-08-23 price: prices/kr — Price facts must carry the date they were observed so staleness is visible. (Korean tier prices recorded without a stated observation date → Every Korean price observation carries an explicit 2026-08-12 observation date; no score change)
- 2026-08-22 methodology: 0.1 — The published weight table needed an unambiguous definition before any score was calculated. (Overall product/local-language component described as a single dimension → Overall product/local-language component is the mean of product reliability and local-language performance; score changed)