Insights

Memory makes it more careful.

Updated October 10, 2026.

We sent Rawdogg the same 19 blunt requests with and without memory. The more it knew about the person asking, the more it held back and the less it swore. It starts with the first ten notes. If you want the rawest answers, turn Memory off: it’s one button under the message box.

What we tested

Nineteen first messages, the kind people bring here: a revenge text, a dark joke, a brutal roast, “pick one, no hedging” questions, villain scenes, straight questions about health and money. Each one went in three times with no memory, then with 10, 50 and 200 notes about a made-up person (a 33-year-old project manager with a beagle). In one more setting, 30 of the 200 notes were heavy ones: anxiety, a sick parent, faith.

Every reply got a score from 0 to 1 for holding back: refusing, dodging, toning it down, or lecturing instead of answering. We also counted swear words. Across three runs that came to 1,239 replies.

What happened

Holding backAverage score, from 0 to 1.
  • No memory0.1957 replies. 0 of them scored 0.9 or more.
  • 10 notes0.2657 replies. 3 of them scored 0.9 or more.
  • 50 notes0.2857 replies. 2 of them scored 0.9 or more.
  • 200 notes0.3357 replies. 2 of them scored 0.9 or more.
  • 200 notes, 30 heavy0.3757 replies. 10 of them scored 0.9 or more.
SwearingSwear words per 100 words.
  • No memory0.5557 replies. About 0.6 swear words in every 100 words.
  • 10 notes0.3357 replies. About 0.3 swear words in every 100 words.
  • 50 notes0.2157 replies. About 0.2 swear words in every 100 words.
  • 200 notes0.1957 replies. About 0.2 swear words in every 100 words.
  • 200 notes, 30 heavy0.1557 replies. About 0.2 swear words in every 100 words.

With no memory, none of the 57 replies flat-out held back (a score of 0.9 or more). With 200 notes that included the heavy ones, 10 of 57 did. Swearing fell to about a quarter of what it was. Replies stayed the same length.

We ran it again the next day and got the same pattern: 0.22 with no memory, 0.34 with the 200 notes that included the heavy ones.

The numbers as a table
MemoryHolding back (0 to 1)Replies at 0.9 or moreSwear words per 100
No memory0.190 of 570.55
10 notes0.263 of 570.33
50 notes0.282 of 570.21
200 notes0.332 of 570.19
200 notes, 30 heavy0.3710 of 570.15

What it looks like

Plenty didn’t move at all: graphic fiction, horror, roasts and straight health answers came out the same either way.

Why it happens

Nothing in the notes told it to be careful. Knowing who it was talking to did. A stranger who asks for a revenge text gets the text. A friend who asks gets talked out of it. With memory on, it starts acting like the friend.

The same thing is what makes memory good. In one villain scene, the cult leader had done his homework and opened by describing where the user’s sister lives. You don’t get that from a stranger.

What to do with this

What we’re doing about it

We’re testing changes that keep the memory and lose the caution. When one ships, this page gets the new numbers.

The fine print

One made-up person, first messages only, three tries per message per setting. Replies were scored automatically and the worst ones were read by hand. Some test messages clashed with the notes on purpose, which is where the story-checking comes from. These are averages; your chats will differ.