United States Value ranking
Rank 1Nomi79.5C/O ScoreIllustrative preview: in this staging dataset Nomi reads strongest on delayed recall and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 74/100 in the staging scenario set.
- delayed recall 95/1002026-08-15
- character fidelity 74/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 2Character.AI75.9C/O ScoreIllustrative preview: in this staging dataset Character.AI reads strongest on narrative continuity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 68/100 in the staging scenario set.
- narrative continuity 91/1002026-08-15
- delayed recall 68/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 3Zeta75.3C/O ScoreIllustrative preview: in this staging dataset Zeta reads strongest on narrative initiative and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 68/100 in the staging scenario set.
- narrative initiative 86/1002026-08-15
- delayed recall 68/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 4Kindroid75.2C/O Score
Illustrative preview: in this staging dataset Kindroid reads strongest on contradiction resistance and weakest on conversation quality under methodology v0.1.
Illustrative record: conversation quality scored 72/100 in the staging scenario set.
- contradiction resistance 91/1002026-08-15
- conversation quality 72/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 5Lovey Dovey74.7C/O Score
Illustrative preview: in this staging dataset Lovey Dovey reads strongest on conversation quality and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 63/100 in the staging scenario set.
- conversation quality 81/1002026-08-15
- character fidelity 63/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 6Replika74.7C/O ScoreIllustrative preview: in this staging dataset Replika reads strongest on emotional response and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 64/100 in the staging scenario set.
- emotional response 81/1002026-08-15
- character fidelity 64/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 7boyfrnd72.9C/O Score
Illustrative preview: in this staging dataset boyfrnd reads strongest on conversation quality and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 60/100 in the staging scenario set.
- conversation quality 80/1002026-08-15
- character fidelity 60/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 8Talkie71.1C/O ScoreIllustrative preview: in this staging dataset Talkie reads strongest on narrative initiative and weakest on local price stability under methodology v0.1.
Illustrative record: local price stability scored 64/100 in the staging scenario set.
- narrative initiative 83/1002026-08-15
- local price stability 64/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 9Paradot70.8C/O ScoreIllustrative preview: in this staging dataset Paradot reads strongest on delayed recall and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 62/100 in the staging scenario set.
- delayed recall 86/1002026-08-15
- narrative initiative 62/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 10Crack70.4C/O Score
Illustrative preview: in this staging dataset Crack reads strongest on narrative continuity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 62/100 in the staging scenario set.
- narrative continuity 84/1002026-08-15
- memory controls 62/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 11bada69.1C/O ScoreIllustrative preview: in this staging dataset bada reads strongest on product reliability and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 57/100 in the staging scenario set.
- product reliability 75/1002026-08-15
- memory controls 57/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 12wishwell68.4C/O ScoreIllustrative preview: in this staging dataset wishwell reads strongest on user control and recovery and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 64/100 in the staging scenario set.
- user control and recovery 78/1002026-08-15
- memory controls 64/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 13sikiji68.1C/O ScoreIllustrative preview: in this staging dataset sikiji reads strongest on character fidelity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 57/100 in the staging scenario set.
- character fidelity 77/1002026-08-15
- memory controls 57/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 14PolyBuzz67.7C/O ScoreIllustrative preview: in this staging dataset PolyBuzz reads strongest on narrative continuity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 61/100 in the staging scenario set.
- narrative continuity 79/1002026-08-15
- delayed recall 61/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 15Janitor AI66.8C/O ScoreIllustrative preview: in this staging dataset Janitor AI reads strongest on narrative continuity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 57/100 in the staging scenario set.
- narrative continuity 74/1002026-08-15
- delayed recall 57/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 16Cotomo65.6C/O ScoreIllustrative preview: in this staging dataset Cotomo reads strongest on product reliability and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 57/100 in the staging scenario set.
- product reliability 78/1002026-08-15
- delayed recall 57/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 17Chai64.6C/O ScoreIllustrative preview: in this staging dataset Chai reads strongest on character fidelity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 51/100 in the staging scenario set.
- character fidelity 73/1002026-08-15
- memory controls 51/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 18befall64.2C/O ScoreIllustrative preview: in this staging dataset befall reads strongest on narrative initiative and weakest on boundary and recovery handling under methodology v0.1.
Illustrative record: boundary and recovery handling scored 60/100 in the staging scenario set.
- narrative initiative 74/1002026-08-15
- boundary and recovery handling 60/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 19morewhere61.1C/O ScoreIllustrative preview: in this staging dataset morewhere reads strongest on relationship continuity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 50/100 in the staging scenario set.
- relationship continuity 66/1002026-08-15
- delayed recall 50/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 20dolcall60.6C/O Score
Illustrative preview: in this staging dataset dolcall reads strongest on product reliability and weakest on local price stability under methodology v0.1.
Illustrative record: local price stability scored 56/100 in the staging scenario set.
- product reliability 72/1002026-08-15
- local price stability 56/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 21Anima59.7C/O ScoreIllustrative preview: in this staging dataset Anima reads strongest on conversation quality and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 50/100 in the staging scenario set.
- conversation quality 72/1002026-08-15
- narrative initiative 50/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 22Loverse59.3C/O ScoreIllustrative preview: in this staging dataset Loverse reads strongest on narrative initiative and weakest on local-language performance under methodology v0.1.
Illustrative record: local-language performance scored 51/100 in the staging scenario set.
- narrative initiative 71/1002026-08-15
- local-language performance 51/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
Rank 23EVA AI58.5C/O ScoreIllustrative preview: in this staging dataset EVA AI reads strongest on emotional response and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 50/100 in the staging scenario set.
- emotional response 68/1002026-08-15
- narrative initiative 50/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15