日本 版 記憶 順位
第1位Nomi87.4C/O スコアIllustrative preview: in this staging dataset Nomi reads strongest on contradiction resistance and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 72/100 in the staging scenario set.
- contradiction resistance 94/1002026-08-15
- narrative initiative 72/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第2位Kindroid86.3C/O スコア
Illustrative preview: in this staging dataset Kindroid reads strongest on contradiction resistance and weakest on local price stability under methodology v0.1.
Illustrative record: local price stability scored 69/100 in the staging scenario set.
- contradiction resistance 89/1002026-08-15
- local price stability 69/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第3位Character.AI77.4C/O スコアIllustrative preview: in this staging dataset Character.AI reads strongest on character fidelity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 67/100 in the staging scenario set.
- character fidelity 88/1002026-08-15
- memory controls 67/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第4位Paradot77.2C/O スコアIllustrative preview: in this staging dataset Paradot reads strongest on memory controls and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 63/100 in the staging scenario set.
- memory controls 83/1002026-08-15
- narrative initiative 63/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第5位Zeta74.9C/O スコアIllustrative preview: in this staging dataset Zeta reads strongest on character fidelity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 70/100 in the staging scenario set.
- character fidelity 86/1002026-08-15
- memory controls 70/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第6位Replika74.6C/O スコアIllustrative preview: in this staging dataset Replika reads strongest on emotional response and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 65/100 in the staging scenario set.
- emotional response 82/1002026-08-15
- character fidelity 65/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第7位Lovey Dovey71.9C/O スコア
Illustrative preview: in this staging dataset Lovey Dovey reads strongest on conversation quality and weakest on narrative initiative under methodology v0.1.
Illustrative record: narrative initiative scored 63/100 in the staging scenario set.
- conversation quality 84/1002026-08-15
- narrative initiative 63/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第8位Talkie71.3C/O スコアIllustrative preview: in this staging dataset Talkie reads strongest on narrative continuity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 67/100 in the staging scenario set.
- narrative continuity 84/1002026-08-15
- memory controls 67/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第9位boyfrnd70.3C/O スコア
Illustrative preview: in this staging dataset boyfrnd reads strongest on emotional response and weakest on character fidelity under methodology v0.1.
Illustrative record: character fidelity scored 63/100 in the staging scenario set.
- emotional response 78/1002026-08-15
- character fidelity 63/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第10位Crack68.9C/O スコア
Illustrative preview: in this staging dataset Crack reads strongest on character fidelity and weakest on local price stability under methodology v0.1.
Illustrative record: local price stability scored 65/100 in the staging scenario set.
- character fidelity 84/1002026-08-15
- local price stability 65/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第11位PolyBuzz67.5C/O スコアIllustrative preview: in this staging dataset PolyBuzz reads strongest on character fidelity and weakest on conversation quality under methodology v0.1.
Illustrative record: conversation quality scored 63/100 in the staging scenario set.
- character fidelity 75/1002026-08-15
- conversation quality 63/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第12位wishwell66.6C/O スコアIllustrative preview: in this staging dataset wishwell reads strongest on character fidelity and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 58/100 in the staging scenario set.
- character fidelity 81/1002026-08-15
- memory controls 58/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第13位bada65.9C/O スコアIllustrative preview: in this staging dataset bada reads strongest on local-language performance and weakest on memory controls under methodology v0.1.
Illustrative record: memory controls scored 58/100 in the staging scenario set.
- local-language performance 77/1002026-08-15
- memory controls 58/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第14位Cotomo64.9C/O スコアIllustrative preview: in this staging dataset Cotomo reads strongest on local-language performance and weakest on narrative continuity under methodology v0.1.
Illustrative record: narrative continuity scored 58/100 in the staging scenario set.
- local-language performance 87/1002026-08-15
- narrative continuity 58/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第15位dolcall63.8C/O スコア
Illustrative preview: in this staging dataset dolcall reads strongest on product reliability and weakest on narrative continuity under methodology v0.1.
Illustrative record: narrative continuity scored 55/100 in the staging scenario set.
- product reliability 74/1002026-08-15
- narrative continuity 55/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第16位sikiji63.5C/O スコアIllustrative preview: in this staging dataset sikiji reads strongest on character fidelity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 60/100 in the staging scenario set.
- character fidelity 75/1002026-08-15
- delayed recall 60/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第17位Janitor AI61.3C/O スコアIllustrative preview: in this staging dataset Janitor AI reads strongest on narrative initiative and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 55/100 in the staging scenario set.
- narrative initiative 74/1002026-08-15
- delayed recall 55/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第18位Chai61.1C/O スコアIllustrative preview: in this staging dataset Chai reads strongest on narrative continuity and weakest on local-language performance under methodology v0.1.
Illustrative record: local-language performance scored 52/100 in the staging scenario set.
- narrative continuity 75/1002026-08-15
- local-language performance 52/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第19位befall60.4C/O スコアIllustrative preview: in this staging dataset befall reads strongest on narrative continuity and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 57/100 in the staging scenario set.
- narrative continuity 76/1002026-08-15
- delayed recall 57/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第20位morewhere59.5C/O スコアIllustrative preview: in this staging dataset morewhere reads strongest on narrative continuity and weakest on paid experience per cost under methodology v0.1.
Illustrative record: paid experience per cost scored 53/100 in the staging scenario set.
- narrative continuity 66/1002026-08-15
- paid experience per cost 53/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第21位Loverse57.4C/O スコアIllustrative preview: in this staging dataset Loverse reads strongest on narrative initiative and weakest on delayed recall under methodology v0.1.
Illustrative record: delayed recall scored 50/100 in the staging scenario set.
- narrative initiative 71/1002026-08-15
- delayed recall 50/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第22位Anima56.2C/O スコアIllustrative preview: in this staging dataset Anima reads strongest on conversation quality and weakest on local-language performance under methodology v0.1.
Illustrative record: local-language performance scored 48/100 in the staging scenario set.
- conversation quality 71/1002026-08-15
- local-language performance 48/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15
第23位EVA AI54.8C/O スコアIllustrative preview: in this staging dataset EVA AI reads strongest on emotional response and weakest on local-language performance under methodology v0.1.
Illustrative record: local-language performance scored 50/100 in the staging scenario set.
- emotional response 70/1002026-08-15
- local-language performance 50/1002026-08-15
- 14-day scenario set, paid-standard tier2026-08-15