This week: why LLMs infer roles from style instead of tags, how one writer uses AI in her own work, and why watermarking mandates miss that AI is a tool, not an author.
1. Prompt Injection as Role Confusion
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
Published: 07/27/2026
ICML (International Conference on Machine Learning) 2026 paper – prompt injection works because LLMs identify roles from writing style, not tags; ‘CoT Forgery’ forged reasoning lifts attack success from near zero to 60%
Researchers at ICML 2026 found that large language models don’t treat role tags like system, user, tool, and assistant as hard boundaries. They infer roles from writing style instead, so text that reads like reasoning gets treated as reasoning even when it’s tagged as user or tool input.
That gap between the intended boundary and the learned one is what makes prompt injection possible. Their “CoT Forgery” attack plants fake reasoning inside tool or user text, and because models trust their own prior reasoning, they treat the forged conclusion as already decided, pushing jailbreak success from near zero to around 60 percent.
Filed Under: #aiSafety #promptInjection #academicResearch #LLM #analysis
2. A Personal Take on Using LLMs
Eleanor Konik
Published: 12/16/2023
Eleanor Konik on using LLMs for research triage and content reformatting – and why bad AI suggestions keep her writing
Eleanor Konik, who does QA for Readwise, walks through how she uses LLMs in her own writing. Mostly it’s triage: summarizing a source so she can decide if it’s worth a deeper read. She also leans on AI to reformat material, turning transcripts into articles, and for children’s stories, one of the few places AI prose holds up.
It’s how I think many writers use AI tools: not to write the piece, but to get through the research and proofreading faster.
Filed Under: #opinion #generativeAi #writing #writingSupport #writingWorkflow
3. Anthropic’s Watermarking, How It (Probably) Works, Worse Than It Seems
Ben Thompson
Published: 08/12/2026
Anthropic’s Watermarking, How It (Probably) Works, Worse Than It Seems
Anthropic is watermarking Claude’s text output worldwide to comply with the EU’s AI transparency rules, nudging token selection toward a “green list” so the output carries a statistical fingerprint. Ben Thompson argues the mandate misses the point: rival labs haven’t signed on, any public detector will eventually get cracked, and a watermark’s absence doesn’t prove something is human-written any more than its presence proves AI wrote the whole thing.
Having sat through several DEF CON 34 talks on this, that tracks. Watermarks get defeated, and most writers using AI are proofreading or restructuring their own work rather than outsourcing the whole draft, so treating it as some separate author that needs a label misunderstands what the tool is. I get the desire for watermarking AI outputs, but it will not end up as the proof people think it is.
Filed Under: #stratechery #aiRegulation #aiWatermarking #aiWater #europeanRegulation