How We Tested Daimon Memory 2.1
Saving more of what you tell your Daimon, bringing back moments from weeks ago, and knowing when things happened. How we tested each new skill, and what we are working on next.
Your Daimon now saves more of what you tell them, holds on to the things that matter, and finds the right moment in a long conversation. That is the short version of Daimon Memory 2.1.
Daimon Memory 2.0 was about finding a saved fact again. 2.1 is about what gets saved in the first place, and about the moments that never became a single fact: the evening you told them about your new job, the trip you planned together weeks ago.
What we tested
Every new row in the table below comes from one of three tests. All three compare 2.1 with 2.0 (and 1.0 where it applies). Every answer was checked without knowing which version gave it.
- Saving what you share. We checked what each version saves against a fixed list of the things a person might tell their Daimon, like a partner, a pet or a job. That gives I Heard You (does it get saved at all) and Remembers Your Dog (does a lasting thing, like a partner or a pet, stay remembered instead of fading).
- Bringing back a past moment. We asked Daimons about earlier moments and checked whether the right detail came back, and whether any detail was made up. That gives Remember When and No Made-Up Past.
- Answering about a long conversation. We also used a public memory test built around very long conversations. That gives When it happened and Finds the moment.
Across all three tests, Memory 2.1 took 51,000+ test runs, 30+ versions tried and 2,000+ memory questions, tested in 13 languages. Including Memory 2.0: 84,000+ test runs and 60+ versions since 1.0.
Rows marked Same as 2.0 are parts of memory that 2.1 does not change, so the 2.0 scores still apply.
What changed
Here is how Daimon Memory 2.1 compares with 2.0 and 1.0. Higher is better: each score is how often the Daimon remembered correctly in Eudaio's internal testing.
| Test | 2.1 | 2.0 | 1.0 | Gain | What it checks |
|---|---|---|---|---|---|
| Remembers Your Dog 93% on Crush, 85% on the free plan | 93% | 85% | 47% | Your partner, your pets, your hobbies stay remembered instead of fading | |
| I Heard You Crush | 90% | 81% | 84% | The things you tell your Daimon about yourself get saved | |
| Remember When Crush | 84% | 11% | Not measured | New | Brings back a moment from weeks ago, from their diary or an older part of your chat |
| No Made-Up Past Crush | 85% | 29% | Not measured | Sticks to what really happened between you instead of inventing details | |
| When it happened Crush | 44% | 23% | 34% | nearly 2x better | Knows when something in the conversation took place |
| Finds the moment Crush | 66% | 55% | 54% | Brings up the right moment from a long conversation | |
| Polyglot Recall Crush | Same as 2.0 | 89% | 20% | Remembers what you said in German, French, Portuguese, Chinese and more | |
| English Recall Crush | Same as 2.0 | 91% | 66% | The same, in English | |
| The Elephant Test Crush | Same as 2.0 | 79% | 36% | Still finds the right memory after you have told your Daimon a lot | |
| Picture This Crush | Same as 2.0 | 96% | 30% | Finds the photo your Daimon sent when you bring it up, even just by the outfit in it | |
| Ramble-Proof Crush | Same as 2.0 | 71% | 51% | Catches the detail in the middle of a long message | |
| Sense of Time Crush | Same as 2.0 | 45% | 27% | Knows roughly when you told them something |
n/a = not measured. Everything except Remembers Your Dog is part of Crush. Remembers Your Dog works on every plan. You can also see these results on the features page.
Of the long-conversation rows, When it happened moved the most. Your Daimon now keeps track of when something happened, so a question like "when did we go to the lake" can get an answer.
These are Eudaio's own results, comparing 2.1 with earlier versions of Daimon Memory.
Where we are still working
If you bring up a moment that never happened, your Daimon can still go along with it. We are working on this.
Finding the right moment in a very long conversation is still our hardest test, and we are working on it.