Mooch  /  Work  /  EthUX.design
2026
Case Study

Contributing to Ethereum UX for humans and AI agents

We contributed plain-language guidance to EthUX.design, an open project that teaches AI agents UX best practice for web3.

Project
EthUX.design
Sector
Web3 / AI
Year
2026
Services
Content design, AI testing
Deliverables
2 merged contributions, blind AI test
Discipline
UX + AX

§ 01TL;DR

Crypto jargon from AI agents dropped from 7 in 9 test runs with no guidance to 1 in 9 with EthUX.design's, including our contributions.

§ 02The invitation

“Bad UX” has dogged web3 for so long it's a cliché, and words are often where it starts. A new user meets “gas” (the fee for using the network), “slippage” and “approve” before they've done anything at all.

Increasingly, AI coding agents write these words as they build wallets, websites and apps from a developer's instructions. An agent uses whatever words it's given or trained on, so if you fix the guidance it reads, the fix reaches every screen it builds.

That's the idea behind EthUX.design. Jakub Konopka, a designer at the Ethereum Foundation who we'd worked with on ethereum.org, started it as his own project. It draws on more than 32,000 reports from Ethereum users about where they got stuck, and turns them into fixes published as skills: instruction files an AI agent reads before it writes code.

The project was already live when Jakub invited us into the small private group contributing to it. We chose to work on its plain-language guidance.

As a builder using an AI agent on an Ethereum app
I want it to write copy a first-time user understands
so that people don't give up at the first unfamiliar word
As a first-time user of a crypto wallet or app
I want to understand what each screen is asking me to do
so that I can act with confidence and avoid costly mistakes

§ 03The work

1. What was there

EthUX.design's onboarding skill already had a word table the agent could use: 13 crypto terms on the left, a plain word for each on the right. “Gas fee” becomes “Network fee”, “Revoke” becomes “Remove permission”.

However, some common terms were missing, like allowance (how much an app may spend from your wallet), testnet and transaction hash. On top of that, we felt some translations might mislead users: the table suggested using “coin” in place of “token”, but NFTs are tokens and not coins.

2. What we proposed

A fuller table, grouped by what the agent should do with each word: replace it with a plain one, keep it and explain it the first time it appears, or avoid it altogether when writing for beginners.

Our goal was to get agents to write copy a first-time user understands, without confusing the people who already know the words.

3. Why we started small

EthUX.design was already live, and builders could install its skills into their own projects. Every renamed word costs something for everyone who learned the old one. So we split the work into two parts.

The first part added 8 words we felt were missing but were least likely to be contested. For example, we suggested swapping allowance for “Spending limit” and testnet for “Practice network”. Jakub accepted them in 4 days.

The second took on the contested words, like “token” and “bridge”, and split the table into three groups: replace with a plain-language alternative, keep and explain, or avoid the term altogether for beginners. We spent time navigating these changes with feedback from the group, including UX researcher Sasha, and from Jakub's own review. He asked for three changes, and we made all three.

Direction is right and the split is worth keeping. Jakub Konopka, EthUX.design, review of pull request #10
The word table before and after our two contributions. “After” quotes the merged text.
Crypto termBefore our contributionsAfter our contributions
Replace with a plain word
AllowanceNot in the tableSpending limit
Tx hashNot in the tableTransaction ID
MainnetNot in the tableMain network
TestnetNot in the tablePractice network
Layer 2Not in the tableThe network's name, e.g. Base
ENS nameNot in the tableUsername / .eth name
dAppNot in the tableApp
Secret Recovery PhraseOnly “Seed phrase / Mnemonic” listedRecovery phrase / Backup words
Smart contractApp / ServiceApp
Keep the word, explain it on first use
TokenAsset / CoinToken: “an item you own, like a coin or a collectible”
BridgeNot in the tableBridge: “moves your tokens from one network to another”
Leave out for beginners
Nonce, Gwei, Wei, ERC-20Nonce left out; Gwei and Wei banned by a separate rule; ERC-20 not coveredOne rule for all four: shown only in advanced mode

4. Test it before it merges

Before he accepted the second set of larger changes, Jakub asked for proof they wouldn't make agents write worse copy. So we built a small blind test:

  • Three beginner scenarios: letting an app spend their USDC (a dollar-backed token) and later taking that permission away, a swap that fails because the price moved, and moving funds to a cheaper network.
  • Three AI models: Claude Haiku 4.5, Claude Opus 5.5 and OpenAI's GPT-5.3 Codex.
  • Three setups: no guidance, EthUX.design's guidance before our change, and after it.

That's 3 scenarios, 3 models and 3 setups: 27 test runs, each run once. A separate model, Claude Sonnet 5, graded every result without knowing which model or setup wrote it. It gave one point for each rule the copy passed:

  1. No bare jargon like “gas”, “nonce” or “seed phrase”.
  2. Every crypto term is replaced or explained where it first appears.
  3. It uses the plain word, like “Network fee” for gas fee.
  4. It names the app (say, “Uniswap”) instead of saying “smart contract”.
  5. Any technical term appears only in an advanced section.

§ 04What we found

Test runs that slipped into jargon went from 7 in 9 with no guidance, to 3 before our change, to 1 after it.

With guidance, the models scored 39 out of 45, before and after our change. With none, 35.

Points out of 15 per model (3 scenarios × 5 rules), 45 per column. A plain-word miss is a run that failed rule 3, like “Revoke” instead of “Remove permission”.
ModelNo guidanceBefore our changeAfter our change
Claude Haiku 4.5111111
Claude Opus 5.5121514
GPT-5.3 Codex121314
All models, out of 45353939
Runs with a plain-word miss7/93/91/9

What this doesn't show. While jargon successfully fell from 3 to 1, our change didn't raise the overall score across the 5 checks: it stayed at 39 out of 45. Claude Haiku slipped into jargon on one scenario, and the grader failed Claude Opus for a sentence it had passed, word for word, under the old guidance.

This shows the limit of running each combination just once: an AI grader isn't perfectly consistent. Given the chance to run these tests again we would:

  • conduct several runs per combination
  • grade each result a few times
  • check the word rules with a fixed list instead of an AI

We sent Jakub every result without a rerun, and flagged the grading slip. He merged on 2 October.

§ 05Impact

1/9Plain-word misses, down from 7/9
2/2Contributions merged
27Blind-graded test runs

Both contributions now ship in EthUX.design's guidance, for any web3 builder who points their agent at it.

Thanks to Jakub, for the invitation and rounds of reviews that made our work more rigorous. And to Sasha, whose research-backed feedback helped decide what went in first.

Working on agent experience (AX)? Let's compare notes.

We're working on AXBeat. AX is how well AI agents can find, read and act on your product, its site and its docs. If you're building for or with agents, on Ethereum or anywhere else, email us. We'd love to talk.

hey@mooch.agency

§ 06Related work