what sign up
nixm nixm

← learn

Signal vs thread — we measured it

September 28, 2026

We compared a normal chat thread, which sends the whole transcript to the model every turn, with a nixm signal, which sends a fixed intent, a running log of decisions and the last five exchanges. Same model on both sides, 30 scripted conversations, design registered before any run. A blind AI judge preferred the signal's replies 28 to 12, with 5 ties, winning where plans changed and early constraints mattered. The thread was better at exact recall: 82% correct against 60%. After about turn 8 the signal's context stays flat while the thread's keeps growing until near the end, but today the signal costs more per turn because of a large fixed rulebook.

Why nixm works this way

nixm doesn't keep a transcript. A distilled log of the session stays with the signal for seven days, then dissolves, along with any images you uploaded. That was a privacy decision, and it has a consequence: with no thread to reread, the context has to travel with every message.

So each nixm conversation is a signal with two parts:

  • The intent: what the conversation is for, set once and fixed.
  • The running log: what has been decided so far, rewritten as the conversation moves.

The model sees the intent, the log, the last five exchanges and the new message. It never sees the full history. The obvious question is whether that makes the AI worse at keeping track. We tested it instead of assuming.

The test

The design, the scenarios and the answer keys were written and committed before any run: the pre-registration.

  • Five scripted conversations of 22 to 25 user turns each: a trip plan that keeps changing, a café-name brainstorm with rejects, a long tangent followed by "where are we?", an early constraint that matters late, and planning a scene.
  • Two conditions, same model on both sides: A is a plain chat thread with the full transcript every turn; B is live nixm.
  • Three runs of each: 30 conversations and 708 turns in total.
  • Grading: check questions with answer keys written in advance, a blind pairwise judge (a separate, different AI model, shown the two replies in random order), and tokens and latency measured per turn.

The results

MeasureThread (A)nixm signal (B)
Exact-recall checks: correct82%60%
…partly correct9%33%
…wrong9%7%
"Where are we" and secondary checks: correct1 of 125 of 12
Blind judge preference, final three turns1228 (5 ties)
Context tokens at turn 5 / 10 / 20 / last1,174 / 2,519 / 4,312 / 5,1211,630 / 1,913 / 1,936 / 2,307
Fixed system prompt, tokens1125,895
Mean reply time5.3 s8.4 s
Line chart of context tokens per turn: the thread's context grows steadily, while the signal's stays roughly flat after about turn 8.
Context sent to the model per turn. The lines cross around turn 8; after that the signal stays flat and the thread keeps growing until near the end. Measured to turn 25 only.

Where the signal won

The blind judge preferred the signal's replies about two to one, and none of its stated reasons were about style. The wins clustered where conversations change under you:

  • The changing trip plan: 8 to 0 for the signal.
  • The early constraint that matters late: 8 to 1.
  • Scene planning: 7 to 0.

A running log of decisions keeps the current state of a plan in front of the model. A long transcript keeps every superseded version of it there too. The "where are we?" checks point the same way: the signal got 5 of 12 right, the thread 1 of 12.

Why the goal holds

A thread has no fixed goal. What the conversation is for is simply its first message, and by turn 20 that message is one of dozens, weighed against every tangent and every plan that has since been dropped. A signal's intent is set once and travels with every message, so on each turn the model is reminded what the conversation is for, not only what was said most recently.

That is where the signal's clearest wins came. When a constraint set early mattered late, the judge preferred the signal's replies 8 to 1. Recall alone does not explain it: on the direct checks for that constraint, both sides scored 3 of 9. The difference was in how the replies acted on it. On the "where are we?" checks, the signal was right 5 times out of 12 and the thread once.

In the judge's own words: "a creamy pasta that is dairy-free, uses one pan and is spicy, so it respects all the constraints" and "It stays dairy-free by using coconut cream, cooks in one pan, and adds chili for heat."

It is not a guarantee against drift. In the long-tangent conversation the thread won 5 to 2, so a fixed intent does not, by itself, pull a derailed conversation back on course.

Where the thread won

Exact recall. When a check asked for a specific detail from earlier, the thread was right 82% of the time and the signal 60%, with most of the signal's misses partly right rather than wrong.

The café brainstorm exposed why. The thread got its recall checks 9 of 9; the signal got 0 of 9. On the very first message, nixm wrote its own opinions into the log as if the user had decided them, noting names as cut that the user had never rejected. Once that exchange scrolled out of the last five, the log was the only record, and it was wrong: names the user still wanted were gone. The thread also won that scenario on the judge's preference, 6 to 3, and won the long tangent 5 to 2.

What we changed

Two earlier attempts at a fix did not hold. The one that worked is simpler. The log now holds only what the user said, including every item they list, and nixm's own opinions stay in its reply instead of being written down as decisions.

In a targeted rerun of the brainstorm, the first message kept all ten names in 5 of 5 tries, and every name the user cut was recorded correctly through turn 15 in both 15-turn runs.

One known limit remains: when a user cuts a name, the log drops it rather than recording why it was cut.

What it costs

Plainly: the signal is more expensive per turn today. It carries a fixed rulebook of about 26,000 tokens, which the plain thread does not, and its replies took 8.4 seconds on average against 5.3. The part that stays small is the conversation's own context: around 2,000 tokens from turn 10 to the end, against a thread that peaked at about 6,000 tokens near the end. We only measured to turn 25, so we are not claiming where the curves go after that.

The limits of this test

  • Five scenarios and three runs each is enough to see a pattern, not to settle a question.
  • The judge is a blind AI model, not people.
  • Both sides used one model; another model might behave differently.
  • Conversations ran to 25 turns at most.

The pre-registration, with its scripts and answer keys, is linked above with the date it was first committed, so the design can be checked against the results.

Why this matters beyond nixm

Most AI chat keeps everything and rereads it every turn, because that is the easy way to keep context. This test suggests you do not need the whole transcript to keep a conversation on track, and that a decision log can do better when plans change. It is weaker at word-for-word recall, and it has a cost you can see. For a tool whose point is that there is no transcript to keep, that is the trade we would rather make in the open. How nixm's storage works is in what ephemeral AI chat is, and how it compares with other assistants is in what each AI chatbot keeps.

frequently asked

Does an AI chat need the whole transcript to keep track?

Not in this test. A fixed intent plus a running log of decisions was preferred by a blind AI judge 28 to 12, especially when plans changed. The full transcript was better at exact recall of details.

What is a nixm signal?

A conversation's fixed intent plus a running log of what has been decided. The model sees those, the last five exchanges and the new message, never the full history. The signal dissolves seven days after it is created.

Where did the full transcript do better?

Exact recall: 82% correct against 60%. It was far better in the brainstorm test, because the signal's log recorded nixm's own opinions as if the user had decided them. The fix keeps only what the user said in the log.

Is the signal approach cheaper?

Not today. Its context stays around 2,000 tokens after turn 10, but it carries a fixed rulebook of about 26,000 tokens, so each turn costs more and replies were slower: 8.4 seconds against 5.3.

Was the test fair?

The design and answer keys were committed before any run, both sides used the same model, and the judge saw replies blind and in random order. It is still a small test: five scenarios, three runs each, and an AI judge rather than people.

Does nixm keep a transcript?

No. A distilled log of the session stays with the signal for seven days, then dissolves, along with any images you uploaded.

Can I try nixm without an account?

Yes. Open nixm.ai and seed a signal, with no sign-up and no transcript kept.

try it now — no account, no transcript kept. even the signal dissolves in seven days.

seed a signal →

Enter your email and receive a link to sign in. No passwords, just a secure signal to connect.

why connect

  • 20 free credits to start
  • your signals, remembered
  • buy more credits anytime
  • image & video generation features
or
+