Back to work
Case study · Safe Intelligence

Spec27: an engineer-built AI-validation platform, made clear.

Safe Intelligence built Spec27 to validate AI agents without teams building their own test infrastructure. The engineering was real; the interface fought its users. At North, I led the redesign across research, strategy, branding, every screen, and the dev-ready build engineering shipped from.

Live in production. The relaunch drew 500+ sign-ups in its first cohort.

From empty to alive: the redesigned console, live in production
01 · Problem
01

The problem, and what we found

Before designing a single screen, I mapped the territory: how analogous consoles organise themselves, what the product actually was underneath, and what its users kept tripping on. The client's own material (repo threads, Looms) was mined with AI for speed and crossed with my audit. Everything landed in one feedback document that aligned client and team on what was broken.

competitive scan 6 consoles product audit annotated sample concept map the real model feedback document problems, prioritised github repo looms /ai-digest CLIENT UX green light prototype client material processed with AI for speed: the document aligned client and team before a single screen was designed
Render Colour that guides choices; component treatment patterns OpenAI Architecture, recognisable on a deliberately generic DS Latitude How web and product split; their onboarding flows Cala The organic, human LLM look, studied, then set aside Anthropic Confirmed the organic direction wasn't ours Galtea Closest competitor: components, colour, IA, forms, modals the Spec27 console language dashed: direction studied and discarded · violet: the closest competitor

Five findings

01
Navigation didn't match the model.

The structure fought how the product actually thinks: projects, agents, specs and evals were scattered rather than related.

02
A rudimentary UI.

Inconsistent, dated and poorly accessible, with components that behaved differently on every page.

03
A brand that couldn't be worked with.

Light and dark modes disagreed; the logo had no application for either. Nothing to build screens on. (It even had a different name: DeepScan.)

04
No empty states.

New users met blank tables: the product at its emptiest exactly when it most needed to explain itself.

05
Complexity without guidance.

Deeply technical language, no onboarding, no visible path to a first result.

02 · Phases
02

Two deadlines, two phases

Two hard dates shaped the whole engagement: a client event weeks away, and a relaunch after it. So the work split in two: a surgical quick win first, the real redesign second.

the event: presentation + user tests relaunch · 500+ sign-ups phase 1 · quick win phase 2 · full redesign pause next cycle MAR APR MAY JUN JUL AUG SEP feedback rounds dates approximate, reconstructed from project artifacts

Phase 1 · the quick win

Three touchpoints restructured without touching the brand: the home, each project's overview, and a first pass of navigation and topbar. Enough clarity, fast enough, for the client to present at the event and run user tests on something that held together. The outcome was qualitative, and it bought trust for phase 2.

Phase 1: the same product, restructured The original product, before the redesign BeforePhase 1

Drag to compare the same product weeks apart. Phase 1 stayed inside the old brand on purpose.

03 · Redesign
03

The redesign

Phase 2 was the full rebuild, run on an agentic workflow with humans at every gate. Plans became design.md and roadmap.md; prototype and critique agents produced and challenged the work; the client and their engineers reviewed each step. Critique lived in chat, implementation and audit in a code agent, exploration in a design agent.

Agentic execution · plan Agentic execution · build ux audit client feedback concept map /plan design.md roadmap.md tasks.md /ux-prototype /ux-critique /batch-review prototypehtml dev-ready kithtml · css · js ENG implementation 2 weeks relaunch Human review UX CLIENT UX CLIENT UX ENG UX: me · CLIENT: the product owner · ENG: the client's engineering team critique in chat · implementation & audit in a code agent · exploration in a design agent · nothing ships without human review

A brand that could work

Two weeks of light branding before the screens: a style guide, a motion language, and the product shell wearing the new identity, with light and dark finally agreeing and a logo that works in both.

The product shell wearing the new Spec27 identity

The corrected model

The navigation only made sense once the model did. A project holds three first-class objects: agents, specifications, and the evals that pair them. Datasets and judges became plumbing: configured when needed, never navigated.

organisation registry agents built-ins · published copy in Project agents what you validate specifications versionable contract eval agent + spec runs → results Three first-class objects.The user never meetsthe plumbing. datasets · judges: configured behind the scenes, never navigated

The pivot: from wizard to loop

Halfway in, spec creation changed nature. The client's own prototyping showed users don't configure once and move on; they iterate between an agent and its spec until it behaves. The sequential wizard I was designing died that week. In its place: a micro-wizard for the essentials, then a guided loop inside the spec editor where the user, not the form, decides when the spec is good enough.

The plan · sequential The redesign · micro-wizard + loop create project add agent create spec create eval run one direction · the form decides when you're done “Users don’t set uponce — they iteratebetween agent andspec until it works.” the client feedback loop micro-wizardproject · essentials spec editor guided · onboarded draft spec run vs agent refine dummy agentalways available eval → run good enough? the user decides 0→1 in 7 steps spec creation is now a loop the user leaves when the spec is good enough

Designed by states

With the 0→1 defined, every core surface was designed three times: empty, first project, and full of data. State 0 walks you straight to your first specification, with a built-in test agent always available, so integration friction never blocks the way.

State 0: the product guides you to create your first specification State 1: first project configured, everything covered State 30: the console full of live runs and data

State 0, nothing yet. The overview becomes an onboarding: create your first specification.

Results: one atom, three lenses

Results used to live scattered across pages that mirrored the resource pages: two “Evals” pages, robustness split by type. The redesign named the valuable datum, the spec run, and made every view a lens over it: group by eval, by agent or by spec; rows expand in place; everything crosslinks.

Before · grouped by page After · one atom, three lenses results pages resource pages eval results spec robustness agent robustness evals specifications agents Every object type got a results pageand a resource page: two “Evals”pages, and the data split by type. where does a result live? it depends spec run clean → robust · the atom the data that matters, every other view rolls it up Results: one page every spec run · read-only by eval by agent by spec rag agent · non-stream multi-turn agent google vertex rag agent view spec run #493 → same atoms, regrouped: never duplicated · crosslinks everywhere the fix wasn't more pages: it was naming the valuable datum (the spec run) and letting the user change the lens
The consolidated Results page: spec runs grouped by eval, agent or spec

The consolidated Results page: one place, grouped three ways, every row crosslinked.

Delivered as a dev-ready kit: HTML, CSS and JS. Engineering shipped the rebuild in two weeks, in time for the relaunch, which the conventional route wouldn't have allowed.

04 · Outcome
04

Outcome

When the model is right, the product explains itself. The client wrapped the cycle happy and booked the next one.

0steps
0→1 set-up, down from ~30–40
0weeks
dev-ready kit to shipped rebuild
0+
sign-ups in the relaunch cohort
0states
every core surface: empty, first, full

First real cohort: sign-ups, not yet activation. The next cycle starts in September.

Interactive

Don't take my word for it. Open it.

The full redesign runs in your browser: sign in, switch between the empty, first and full states, and click through the spec loop and results.

Open the live prototype
Spec27 live prototype preview
More work