Last verified: September 10, 2026
What You Can Now Do
Context Memo now streams ad brief results in the simulator: each configured engine's answer paints in as it generates, rather than the whole run sitting dark until the slowest model returns. The run's coverage calculation also begins earlier, so the share of queried prompts where your brand appears accumulates live instead of arriving as a terminal number attached to a finished brief. What changes for you is the decision point. You no longer have to complete a run to learn that the prompt phrasing or the persona was wrong. You see it forming, stop the run, fix the input, resubmit.
Where It Is in Context Memo
This lives in the ad brief simulator. Submit a brief exactly as before, with your prompt set and persona selection. The difference is on the results screen: engine answers fill in progressively as tokens arrive, and the coverage figure updates alongside them instead of appearing once at the end. Streaming applies to new runs.
How to Use It
- Open the simulator and submit a brief with your prompt set and target persona. The run begins and the first engine responses start appearing immediately.
- Read the earliest two or three answers for the thing most runs are actually checking: does your brand appear when a buyer asks the category question, and which domains get cited instead of you.
- Track the live coverage figure as responses land. Coverage is the share of queried prompts where the brand shows up in the answer, and now you can see its trajectory while the run is still executing.
- Kill the run if the early signal is clearly bad. A run that looks wrong at 20% completion usually is, and the cost of stopping it is now close to zero.
- Rewrite the prompt phrasing or swap the persona, then resubmit. This is where the time savings compound: iteration cycles that used to fit a few into an hour now fit a few into ten minutes.
- Wait for the completed brief before you record a finding, brief a writer, or prioritize a memo. Triage on the stream, decide on the finished result set.
Why We Built It
Nobody submits a brief and stares at a spinner for the fun of it. Under batch execution, the answer you needed was almost always visible in the first two or three engine responses, but the wait was set by the slowest model in the run rather than the average one, so a single lagging engine dictated your whole feedback loop. Coverage made it worse by arriving only with the finished brief, which meant a badly phrased prompt set burned a full run before anyone could see it had failed.
What It Does Not Do Yet
Streaming applies to new runs, so anything submitted before this change still behaves as a batch job.
The tradeoff deserves stating plainly: streamed partial output is genuinely partial. A brand mention in the first sentence of a streaming answer can be qualified, hedged, or flatly contradicted three sentences later. Reading a half-generated response and drawing a conclusion from it is a real failure mode, and it is the one way this feature can make your analysis worse instead of faster. Use the stream to decide whether a run is worth finishing. Use the completed brief to decide what's true.
Final coverage is unchanged in how it's computed. The live figure is the same metric arriving earlier, not a new estimate, so a mid-run number will move as slower engines report and should not be quoted as a result on its own.