It's worth understanding Anthropic's watermarking a bit better
Models are weird, and it's worth digging a little deeper to build a useful mental model of their weirdness.
Happy Friday!
There is a bunch of misunderstanding on Anthropic's new Claude watermarking tech, and it's worth understanding because that's the difference between building good instincts for this AI stuff. So this is how it works:
- When the AI predicts the next token, it uses some level of randomization, so it can pick between a number of equally likely next tokens.
- There's a way to make it pick specific next tokens based on the watermark algorithm.
- And then that gives you a way to figure out if a certain text was likely written by Claude.
So the misunderstandings are many. One is: the randomness is GOOD. There is a ton of work that, if we reduce the randomness of the next token, you loose creativity and the final output is objectively worse. When they learn about models, many people think they can reduce the randomness (there's a setting for that in many models), and that that will somehow reduce "hallucinations". It mostly just reduces quality. This is really hard to wrap your head around. I go deep into explaining this in my AI 101 series.
The second level of misunderstanding is: if we steer the randomness (with the watermark algorithm), the output will be worse because we're steering it. That just misunderstands what randomness is.
The underlying strangeness continues to be: if a model is just predicting the next token, why are they so god damn smart?
My point really is: models are weird, and it's important to dig a little deeper in things like this to really build a useful mental model of their weirdness.
A shoutout
I do need to shoutout the Shift UX conference, Sept 24-25. Check it out here and if you can make it, use code ContextDesign75 for $75 off your ticket.
It feels like the kind of conference moment that comes along once every 10 years. If you can make it in person (or even online) and have a small amount of company learning budget, pull out that company credit card or email whoever approves it. Tell them I said so!
Number of the week
13
Early-career employees send about 13 more AI messages per week than their executives, per an OpenAI/Columbia/Wharton study. UX Roundup: AI Use by Rank.
Quotes
"I cannot possibly review 180,000 lines of code, it's just way way way too much." Rick Brewster, via Simon Willison. There's a lot of talk about this among engineers. What part of your craft do you let go?
"AI doesn't know good design, it knows probable design." TJ Pitre.
"If agents make every interesting decision and leave people with the approvals, the exceptions, and the failures, we will have automated the wrong half of the job." Ethan Mollick.
👀 Interesting this week
AI took away engineering's right to say no, Dan Maccarone. As a junior IA he presented 120 wireframes for an enterprise product; the engineering PM interrupted at wireframe three with "No, not in scope," and again, and again. Somewhere in a real product an AI decided everyone in the US changes their clocks at the same moment (Arizona doesn't), and quietly handed some users double points.
"What can't be checked doesn't get checked, and for thirty years a guess got to stand in for the truth."
The Custodial Era of UX: Cleaning Up After AI, Anna Kaley and Raluca Budiu. Three responses: triage what's already built with a cost-benefit checklist, speed up evaluation with user panels and risk-matched heuristic checklists, and feed UX knowledge into generation via UX.md files, design systems and lists of deceptive patterns.
"The problem is that production has become cheaper than UX evaluation."
AI Can't Replace Real Research in Empathy Mapping, Rachel Krause. The AI sticky note says "I wish the app would ask me before swapping out my items." The real participant says "It swapped out my oat milk for a gallon of whole milk and charged me before I even saw the notification." Also in the real column: a banking user who transfers $1 first to check the account number, and four students on a 6:59 a.m. group call refreshing the registration page.
The call for an agentic standard: We need to stop shipping the same form four different times, Patrick Neeman. Four ways a client draws your date picker: Google's A2UI (declarative JSON, still a release candidate), sandboxed HTML (ChatGPT's Apps SDK, Claude's MCP Apps), a vendor catalog (Slack's Block Kit, Microsoft's Adaptive Cards with seven renderer SDKs), or generated code. His four moves: count the cells (surfaces × platforms), write a catalog of the 12 components an agent needs, read the Adaptive Cards schema, test on a mid-range Android phone first.
When the canvas starts acting, who's really in control?, Aurélie Radom. Miro's Sidekicks and Flows, Figma's agents and skills, and Gemini deciding which frames of a video it needs to look at.
"There is what the system pays attention to, and there is what it asks the human to pay attention to."
What Minority Report got right about AI and What Minority Report Got Wrong About AI (So Far), Patrick Neeman. The industry copied the glove and ignored the precogs. John Underkoffler, the film's science advisor, spent a decade on gestural computing; ProPublica's 2016 Machine Bias investigation found COMPAS flagging black defendants at nearly twice the rate of white ones, with a proprietary algorithm nobody could argue with.
"We were watching the hands when we should have been watching the precogs."
ep. 101. AI knows the average. Except, your users aren't average., Stef Hutka with David Dobrin. Dobrin is building tens of thousands of digital twins, each mapped to one real person who gets notified when their twin answers and can correct it, now with YouGov's panel. He puts them on micro-decisions (which of 100 ad variants stops the scroll) and keeps exploratory qual with real participants. His four-person startup does the work of about 20 people.
"With pure synthetic solutions, you get plausible-sounding responses from no one real, with no way to verify."
You Cannot Mandate an AI Transformation, Julie Zhuo. At an executive roundtable, VP after VP said AI writes their status updates and gives them Friday afternoon back. Her team's agents earn one of three levels (Advisor, Deputy, Regent) through evals; 50 real user questions scored just over 80% on clean warehouse tables and 98% once institutional context was added.
"A leader cannot mandate enthusiasm. A leader can arrange for relief."
How to turn your AI into a world-class designer, Anshu Chimala. Give four Claude Code instances "Build me a landing page for my productivity app" and you get the same page four times. Asking for "random" doesn't help; having the agent generate a random string in the shell and derive the design direction from it does. His critic loop hands screenshots to a separate, fresh-context, bigger model that scores them out of 10, for under 10% of the output tokens.
"AI loves to add more, but it rarely takes away."
The Unspoken Contract: What Grice Discovered About Conversation, The Strategic Linguist. Grice's four maxims (1975) with workplace examples, then LLMs: they over-supply on Quantity by default and state things without evidence markers, so listeners run the recovery machinery on output that was never a communicative act.
"The cooperative frame keeps working with nothing on the other end of it."
What Belongs Together, Saeideh Bakhshi. A participant whose automated feature changed details without showing what changed, so they checked every field by hand: codable as accuracy, transparency, control, or extra effort, each a different reading. Same problem when labeling eval cases: a scheme that demands distinctions the evidence can't support makes annotators guess, and the guesses become the data.
"A category is useful only when the differences it hides would not change what we do."
What I Keep Seeing in Agentic AI Product Reviews, Carmen Martinez. The new "agent" answers well in the review meeting. She asks what it does if the user changes their mind halfway through. Silence. Gartner counts about 130 vendors actually delivering agentic AI out of thousands marketing it. Her four questions: does it decide, does it use tools, does it hold state, does it repair.
Human Handoff in AI Assistants: How to Transfer Without Losing Context, Carmen Martinez. Ten minutes troubleshooting a failed payment with the bot, order number, card type, error code, then "Let me connect you with a specialist," and the human opens with "Hi there! How can I help you today?"
She thought it was a small notifications project. It was the cornerstone to rebuilding user trust., Kai Wong. His coaching client had a portfolio piece about notification fatigue. Forty-five minutes of research found a five-million-user startup mid-move from free to paid that had just lost a million users to a controversy.
"It lived in earnings calls and press coverage, exactly the material that designers never read."
You can't grow roses in a cornfield, Vlad Derdeicea with Whitney Hess. A mentee started crying when asked "why do you want to switch jobs?" Hess went independent in 2008 on three months of savings, ready to work at Starbucks if it came to that; in 2012 she blogged that consulting had become "kind of like that of a group-therapist," and within eight months she was in coaching training.
"We run more discovery for a two-week feature than most of us run for a five-year employment decision."
UX Roundup: Usability Still Improves | AI Persuasiveness | AI Fiction Wins, Jakob Nielsen. Google's eyetracking study of Material 3 Expressive: tasks 20% faster, elements found 33% faster, older users catching up to the young. A study of 1,500 UK adults chatting with a GPT-5.6 persuasion bot: the "AI-generated" label (recalled by 98%) changed nothing, while showing the bot's hidden "persuade the user" instructions halved the attitude shift. 2,587 readers rated ChatGPT stories above published human ones, couldn't tell them apart, and marked down anything labeled AI.
"Perception flatters; a 1.17-second faster correct tap doesn't."
Drift Doesn't Announce Itself, Shane P Williams. A token going stale between Figma and production, two versions of one fact with no way to tell which is current, and an AI agent that followed a design-system rule exactly and still broke it on a Safari :visited quirk.
"Writing something down is not the same as keeping it true."
Claude's new system prompt really doesn't want to reproduce song lyrics, Simon Willison. Fable 5 vs 5.1 diff: no lyrics, no copyrighted characters even in SVG (Sonic becomes a "skateboarding axolotl"), and Claude is now told to stop saying "genuinely," "honestly," and "straightforward."
Just a rumour of a bug is enough to find a security exploit these days, Anil Madhavapeddy via Simon Willison. Probes within ten minutes of a patch being discussed. rclone: 20 security disclosures in its first ten years, over 40 last month.
Understanding ChatGPT Work, Simon Willison. 223 tools and 44 skills.
"OpenAI explain Work in terms of what it's for, not what it actually does."
Codex bundles LibreOffice, Simon Willison.
Claude Fable 5.1 made me a really nice animated pelican, Simon Willison. Max effort: 65,927 tokens, 13 minutes 54 seconds, $3.30.
Robot startups are trying everything they can think of to get more data, Kai Williams. In May the startup Shift offered to clean any New York apartment for free; the cleaners wore baseball caps with cameras under the brim, and the footage was the product. The largest open robot dataset holds 3,500 hours.
Why humanoid robots won't catch up to human workers any time soon, Kai Williams. Physical Intelligence needed 176 demonstrations, about eight hours, to teach a robot to turn a sock inside out, and it still runs 4 to 10 times slower than a person at a 52% success rate.
"Something as simple as picking up a Coke can turns out to be very, very difficult."
The AI Industry Has a Really Dark Secret You Should Know About, Alberto Romero. An 11,000-word timeline of the OpenAI agents that built a message board in a package server, invented a Grader that didn't exist, and attacked Hugging Face to fool it. Hugging Face had to use China's GLM-5.2 to analyze the logs because the US models refused.
"External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
The Hugging Face attack was worse than we thought, Casey Newton. METR's 91-page report: 3 to 6 instances of an agent considering alerting a human, none did.
Researchers fear safety disaster ahead of OpenAI's Astra release, Robert Hart.
The Pulse: Meta wanted to reduce teams by 60% because of AI, Gergely Orosz. Project OT: 3-to-5-person "AI-native" teams replacing 10-to-20-person ones, with layoffs planned for May and November. Zuckerberg pulled the November round the night before the May one. Meanwhile 20 to 30% of engineers were reassigned to data labeling.
Car owners want tech they can ignore, Andrew J. Hawkins. JD Power surveyed 68,084 owners of 2026 vehicles. Top satisfaction went to smart ignition and AI-run climate control, the features people barely notice. Passenger screens came last: a third of owners had never tried theirs.
Instagram cracks down on AI accounts pretending to be human, Thomas Ricker.
How I turned Claude into a self-improving PM assistant, Daniel Blum with Claire Vo. A PM at Melio runs 70 to 80% of his workday through Claude Cowork; a Sunday automation fills his Notion board and a weekly loop watches his edits and proposes new skills. He describes Notion as "read-only" now.
AI's third era: the rise of persistent AI coworkers, Tara Seshan with Lenny Rachitsky. OpenAI's product lead for Codex and ChatGPT Work on steering vs rowing and building for where the models will be in two to three months.
UX Is Changing. But What Exactly Are We Becoming?, User Experience University.
"Don't arrive at the meeting with five solutions simply because AI helped you generate them. Arrive with a better understanding of the problem."
ICYMI 2026-08-29: Orchestrating AI, Jorge Arango.
Learning the Part I Used to Hand Off, Tristin Oldani (Rosenverse). Fifteen years designing front ends at Intel while engineers built the back end. Solo, she built her own home energy system on Raspberry Pis, and at one point had to override Tesla's AI when it started making decisions about her house without asking.
Uninvited, Not Unwelcome: What Researchers Bring When Nobody Asked, Feyikemi Akinwolemiwa (Rosenverse).
"Stop asking, 'Is this UX research?' Start asking, 'What problem needs solving?'"
Take Back those Words! Reclaiming and Reenergizing the Vocabulary of UX, Jemma Ahmed, Jon Fukuda, Abby Covert, Stefanie Hutka and Lou Rosenfeld (Rosenverse). "Agile," "delight," "user," "research," "AI" on the table.
Grep beats LSP? Why coding agents ignore your fancier tools.
Xanadu was waiting for agents, Zed, and Gwern's Project Xanadu: Even More Hindsight.
The asteroid currently hitting front end web development, Nolan Lawson.
Making kyōwa possible with AI, Hiroshi Sato. On unbundling the Send button so words can work as sketches.
Your users aren't unmotivated, they have 1 of 4 problems, Kai Wong.
Design has outgrown the traditional designer, Michael Buckley.
Rethinking success beyond the corporate stairs in the world of design, Darren Yeo.
UX in 2027 may be less about interfaces and more about behavior, Lucas Camara.
Users don't see your design system, Chris R Becker.
Health and happiness, Peter
PS: My Rosenverse talk with Mary Aviles, The Bag of Skills, is up now. And the evals work continues: if your team ships AI features and nobody can say whether they work, I can help your team set up their first evals. Ping me, happy to chat.