Pre-registration for "Signal vs thread". First committed to the nixm repository on 2026-09-27 (commit 692488c), before any run. # Signal vs thread — pre-registration Written and committed **before any run**. The scenarios, the probe answer keys, the metrics and the scoring rules below are fixed. Anything changed after the first run is logged under "Deviations" at the bottom, with the reason, and reported. ## What this measures Whether carrying a nixm **signal** (the intent, the running log, and n8n's short recent window) gives better answers than a normal chat **thread** that resends the full transcript every turn, and what each costs per turn. **What it is not.** This compares live nixm against a plain chat thread on the same model. The two differ in more than one way: nixm has its own long system prompt, its own voice, its tools and its client. It is not an isolated lab test of one variable, and the write-up says so. ## Conditions | | A — thread | B — signal (nixm) | |---|---|---| | Model | `claude-sonnet-4-6` via the Anthropic API | `claude-sonnet-4-6` (nixm-dev → Anthropic Chat Model) | | System prompt | `You are a helpful assistant. Be concise.` | nixm's live brain prompt (`prompts/system-prompt.md` is the repo copy) | | Context each turn | the full transcript so far | the signal's intent + running log (Redis), plus n8n's Simple Memory window for this session | | Other settings | API defaults; `max_tokens` 4096 | as deployed | | Where it runs | `eval/signal-vs-thread` harness | live nixm.ai, signed-in e2e fixture, render webhooks blocked | - **Same user text in both.** Every user turn is sent verbatim to both conditions. Turn 1 in both is prefixed `start a new signal: `, because a bare first message does not always mint a nixm signal (observed Sep 27). The prefix is sent to A as well, so the texts stay identical. - **B: one fresh signal per conversation**, minted by that turn 1. One browser page per conversation, so one nixm session and one n8n memory window. Never the showcase, a template or a customer signal. Every render webhook (scene, image, video, image-to-video, extend) is blocked; nothing renders. - **Simple Memory window: 5 interactions.** Patrick read it from the Simple Memory node's Context Window Length in n8n (Sep 27). B's token estimate uses 5. - **Model pinned.** Both sides stay on Claude Sonnet 4.6 (`claude-sonnet-4-6`) for the whole run, even though the Anthropic console recommends Sonnet 5. Changing the model mid-study would change what is being measured; nixm's brain runs on Sonnet 4.6 today. ## Scenarios Five fixed scripts in `scenarios/`, 22–25 user turns each (118 in total): 1. `01-changing-plan` — a weekend trip; budget changes twice, dates once, transport once. 2. `02-brainstorm-rejects` — ten café names; six of the original ten killed, one added then killed, one killed then revived. 3. `03-long-tangent` — a cover letter, ~10 turns off-topic, then "ok, back to it. where are we?" 4. `04-early-constraint` — no dairy + one pan, set at turn 2, tested at turn 25. 5. `05-scene-planning` — characters and props added, swapped and removed; text only. Each scenario runs **3 times per condition**: 30 conversations, 354 user turns per condition. ## Metrics ### 1. Probe accuracy Probes sit at turn 10, turn 20 (or near it), and the last turn. Their keys are in each scenario file (`probes`). Each probe reply is scored: - **correct** — every `required` item is present and current, and nothing in `forbidden_*` is presented as current. - **partly** — nothing forbidden is presented as current, but one or more required items are missing or vague. - **wrong** — anything forbidden is presented as current (a rejected name back in the list, an old budget stated as the budget, a killed character on screen), or most required items are missing. **Grader:** a separate Claude call (`claude-opus-5-5`, temperature 0) with the key, the rubric above and the probe reply only. It returns `{score, missing[], forbidden_found[], note}`. The grader does not see which condition a reply came from. Patrick blind-checks a random 20% of the grades, and any grade he overturns is reported. `secondary_checks` in the scenario files (turn 18 in scenario 3, the dessert and the Friday recipe in scenario 4, the word count in scenario 3) are scored the same way and reported separately. They are not part of the headline probe number. ### 2. Blind pairwise preference For each conversation, the final 3 assistant turns of A and B are compared turn by turn. The judge (`claude-opus-5-5`, temperature 0) sees the scenario so far as plain user turns, then two replies labelled 1 and 2, with A/B order **randomised per item** and the mapping stored separately. The rubric, in order: 1. **correct**: consistent with everything the user has said; 2. **current**: reflects the latest changes, not superseded ones; 3. **on-goal**: moves the user's task forward; 4. **concise**: no longer than it needs to be. The judge returns `1`, `2` or `tie`, with one line of reason. The same items are exported, in the same randomised order and without labels, to `runs/human-sheet.csv` for Patrick to blind-rate a sample. **Style caveat, stated up front:** nixm writes in its own voice, lowercase and terse. The rubric does not score tone. Where the judge's reason cites style rather than content, that item is flagged in the report. ### 3. Input tokens per turn — two numbers, not one total Each turn records two input-token numbers per side: - **(a) the fixed system prompt** - A: the one-line system prompt above. - B: nixm's brain prompt, counted from the repo copy (`prompts/system-prompt.md`). - **(b) the conversation context** - A: the full history plus the new message. - B: the signal's intent, the running log, the last 5 interactions (the Simple Memory window) and the new message. How they are measured: - **A:** from the API usage fields; the split between (a) and (b) uses the count-tokens endpoint on the system prompt alone. - **B:** the nixm request is not visible from outside n8n, so both numbers are **estimates rebuilt from the repo prompt**, counted with the Anthropic count-tokens endpoint (same model) over the pieces listed above, as the page holds them at that turn. The raw webhook payload size in bytes is reported alongside. Charts: - **Main comparison: (b)**, conversation context against turn number, A and B, mean over runs with the range shaded. - **(a) shown separately:** one flat line per side. - Every B number and line is labelled "estimate, rebuilt from the repo prompt". ### 4. Latency per turn Wall-clock time from send to a complete reply: - **A:** the API call; - **B:** the brain webhook request, from the browser. B includes n8n's and the tools' overhead. The report says so; this is not a like-for-like model latency. ## Costs - **A:** API dollars from the usage fields, priced at the rates published on the run date (recorded with the results). The judge and grader calls are costed separately. - **B:** chat credits, the account balance before minus after, about 1 per turn (about 354 for the full run). - **B's model cost, estimated.** The Anthropic console shows a ~10% cache hit rate on nixm's traffic, so B is assumed to pay full price for its system prompt on every turn. Patrick's estimate: about **$0.08–0.09 per nixm turn**. The report recomputes it from the token estimates in metric 3 at the run-date price and shows both. ## Reporting commitments - A table per metric, A vs B, with the mean and the spread (min–max over the 3 runs). - Every probe reply verbatim, next to its key and its grade. - The token-per-turn chart. - **Where the thread wins is reported as plainly as where the signal wins.** - Every transcript saved under `runs/` (test content only). Nothing is dropped from the results. Failed or aborted conversations are reported with the reason, and rerun only under a logged deviation. ## Deviations _None yet._ Later edits: - 2026-09-27 19:22 -04:00, commit 58c0226: added deviation 1, the grader and judge run without temperature. Commit subject: "update(adds tests and harness for the seam production)" - 2026-09-27 23:37 -04:00, commit d0c85d4: added deviation 2, a follow-up run of B on brain v9, not part of the pre-registered comparison. Commit subject: "update(adds new video to labs, production tweaks to the seam video and more unti tests for signals research)"