§ 01TL;DR
Crypto jargon from AI agents dropped from 7 in 9 test runs with no guidance to 1 in 9 with EthUX.design's, including our contributions.
§ 02The invitation
“Bad UX” has dogged web3 for so long it's a cliché, and words are often where it starts. A new user meets “gas” (the fee for using the network), “slippage” and “approve” before they've done anything at all.
Increasingly, AI coding agents write these words as they build wallets, websites and apps from a developer's instructions. An agent uses whatever words it's given or trained on, so if you fix the guidance it reads, the fix reaches every screen it builds.
That's the idea behind EthUX.design. Jakub Konopka, a designer at the Ethereum Foundation who we'd worked with on ethereum.org, started it as his own project. It draws on more than 32,000 reports from Ethereum users about where they got stuck, and turns them into fixes published as skills: instruction files an AI agent reads before it writes code.
The project was already live when Jakub invited us into the small private group contributing to it. We chose to work on its plain-language guidance.
I want it to write copy a first-time user understands
so that people don't give up at the first unfamiliar word
I want to understand what each screen is asking me to do
so that I can act with confidence and avoid costly mistakes
§ 03The work
1. What was there
EthUX.design's onboarding skill already had a word table the agent could use: 13 crypto terms on the left, a plain word for each on the right. “Gas fee” becomes “Network fee”, “Revoke” becomes “Remove permission”.
However, some common terms were missing, like allowance (how much an app may spend from your wallet), testnet and transaction hash. On top of that, we felt some translations might mislead users: the table suggested using “coin” in place of “token”, but NFTs are tokens and not coins.
2. What we proposed
A fuller table, grouped by what the agent should do with each word: replace it with a plain one, keep it and explain it the first time it appears, or avoid it altogether when writing for beginners.
Our goal was to get agents to write copy a first-time user understands, without confusing the people who already know the words.
3. Why we started small
EthUX.design was already live, and builders could install its skills into their own projects. Every renamed word costs something for everyone who learned the old one. So we split the work into two parts.
The first part added 8 words we felt were missing but were least likely to be contested. For example, we suggested swapping allowance for “Spending limit” and testnet for “Practice network”. Jakub accepted them in 4 days.
The second took on the contested words, like “token” and “bridge”, and split the table into three groups: replace with a plain-language alternative, keep and explain, or avoid the term altogether for beginners. We spent time navigating these changes with feedback from the group, including UX researcher Sasha, and from Jakub's own review. He asked for three changes, and we made all three.
Direction is right and the split is worth keeping. Jakub Konopka, EthUX.design, review of pull request #10
| Crypto term | Before our contributions | After our contributions |
|---|---|---|
| Replace with a plain word | ||
| Allowance | Not in the table | Spending limit |
| Tx hash | Not in the table | Transaction ID |
| Mainnet | Not in the table | Main network |
| Testnet | Not in the table | Practice network |
| Layer 2 | Not in the table | The network's name, e.g. Base |
| ENS name | Not in the table | Username / .eth name |
| dApp | Not in the table | App |
| Secret Recovery Phrase | Only “Seed phrase / Mnemonic” listed | Recovery phrase / Backup words |
| Smart contract | App / Service | App |
| Keep the word, explain it on first use | ||
| Token | Asset / Coin | Token: “an item you own, like a coin or a collectible” |
| Bridge | Not in the table | Bridge: “moves your tokens from one network to another” |
| Leave out for beginners | ||
| Nonce, Gwei, Wei, ERC-20 | Nonce left out; Gwei and Wei banned by a separate rule; ERC-20 not covered | One rule for all four: shown only in advanced mode |
4. Test it before it merges
Before he accepted the second set of larger changes, Jakub asked for proof they wouldn't make agents write worse copy. So we built a small blind test:
- Three beginner scenarios: letting an app spend their USDC (a dollar-backed token) and later taking that permission away, a swap that fails because the price moved, and moving funds to a cheaper network.
- Three AI models: Claude Haiku 4.5, Claude Opus 5.5 and OpenAI's GPT-5.3 Codex.
- Three setups: no guidance, EthUX.design's guidance before our change, and after it.
That's 3 scenarios, 3 models and 3 setups: 27 test runs, each run once. A separate model, Claude Sonnet 5, graded every result without knowing which model or setup wrote it. It gave one point for each rule the copy passed:
- No bare jargon like “gas”, “nonce” or “seed phrase”.
- Every crypto term is replaced or explained where it first appears.
- It uses the plain word, like “Network fee” for gas fee.
- It names the app (say, “Uniswap”) instead of saying “smart contract”.
- Any technical term appears only in an advanced section.
§ 04What we found
Test runs that slipped into jargon went from 7 in 9 with no guidance, to 3 before our change, to 1 after it.
With guidance, the models scored 39 out of 45, before and after our change. With none, 35.
| Model | No guidance | Before our change | After our change |
|---|---|---|---|
| Claude Haiku 4.5 | 11 | 11 | 11 |
| Claude Opus 5.5 | 12 | 15 | 14 |
| GPT-5.3 Codex | 12 | 13 | 14 |
| All models, out of 45 | 35 | 39 | 39 |
| Runs with a plain-word miss | 7/9 | 3/9 | 1/9 |
What this doesn't show. While jargon successfully fell from 3 to 1, our change didn't raise the overall score across the 5 checks: it stayed at 39 out of 45. Claude Haiku slipped into jargon on one scenario, and the grader failed Claude Opus for a sentence it had passed, word for word, under the old guidance.
This shows the limit of running each combination just once: an AI grader isn't perfectly consistent. Given the chance to run these tests again we would:
- conduct several runs per combination
- grade each result a few times
- check the word rules with a fixed list instead of an AI
We sent Jakub every result without a rerun, and flagged the grading slip. He merged on 2 October.
§ 05Impact
Both contributions now ship in EthUX.design's guidance, for any web3 builder who points their agent at it.
Thanks to Jakub, for the invitation and rounds of reviews that made our work more rigorous. And to Sasha, whose research-backed feedback helped decide what went in first.
Pull request #9 → Pull request #10 →
Working on agent experience (AX)? Let's compare notes.
We're working on AXBeat. AX is how well AI agents can find, read and act on your product, its site and its docs. If you're building for or with agents, on Ethereum or anywhere else, email us. We'd love to talk.
hey@mooch.agency