User testing for legal documents: the evidence layer

Does a real person, doing a real task, reach the right understanding?

We design and run user testing for legal documents and AI pipelines globally — measuring comprehension, usability and sensitivities, and benchmarking every result against our international portfolio.

Why testing is the part you can't skip

Every redesigned contract looks clearer to the team that redesigned it. The only way to know whether readers genuinely understand — and whether an AI answering from the document leads them right — is to test with real people. Since 2018, testing has been the "meat" of our practice: it's what the FT highlighted when HSBC selected Inkling to examine how neurodiverse users read the bank's terms, and it's the data foundation our award-recognised AI tools are built on. Design without testing is a hypothesis. We don't ship hypotheses.

What we test

Documents.

Comprehension, findability, task success and sensitivities on contracts, terms, policies and disclosures — before-and-after where you're redesigning, standalone where you need a baseline.

AI pipelines.

The full path: your document, your AI system, your user. We measure whether real people asking realistic questions through your AI reach correct decisions — and where retrieval, accuracy, comprehension or interaction failures mislead them. As far as we know, nobody else tests this end to end.

Audiences.

We recruit and test across jurisdictions, cultures and cognitive profiles — our cognitive-diversity research means results reflect your whole audience, not a convenient sample.

What you receive

A benchmarked report: comprehension and task-success scores against our global portfolio, error patterns ranked by risk, verbatim reader evidence, and specific redesign recommendations. Numbers you can put in front of a board, a regulator or a product team.

Legal design services →
How testing fits the full pipeline →

FAQ

How many participants does a legal document test need?

Fewer than most expect. Comprehension failures are systematic, not random — a well-designed study with a modest, carefully recruited cohort per audience segment reliably surfaces the patterns that matter. We scope cohort size to your audiences and risk.

Can you test documents in other languages and markets?

Yes. We design and conduct user testing globally, across jurisdictions and cultures, and our benchmarks reflect that international portfolio.

What does "benchmarked" mean?

Your results are compared against our portfolio of prior studies on similar document types and audiences — so you learn not just your score, but whether it's good.

Can you test our AI assistant, not just our documents?

Yes — that's our signature study: real users, realistic tasks, your actual AI, your actual documents, measured end to end from source to decision.

Claims are cheap. Deltas aren't.

Design a study with us →