Summary
Neverfail Continuity Engine keeps critical applications available regardless of what threatens them: an application-aware core, a warm stand-by server architecture, deep VMware vSphere integration, multi-server failover and software-based WAN optimisation. I designed the web app used to manage it, the Engine Management Service.
The engine was strong. The interface was the reason people weren’t using it. Features were hidden, the visual language was dated enough to cost the product trust, and the flows were cumbersome.
7
Target users in moderated testing
5
Flows tested, across five days of sessions
120+
Wireframes in the interactive prototype
NOTE:Draft pending Paul’s review. Everything below comes from the project artefacts; the round-two outcome numbers are the gap.
The problem
The brief named four problems with the old system: low business adoption caused by hidden features and poor usability, eroded trust because the interface looked its age, cumbersome flows with hard-to-reach features, and no room to add what was coming next.
The reasons for acting were commercial as much as they were usability ones. The product needed to stay relevant in its market and get ahead of competitors, and the interface was what stood in the way.
My role
I ran the UX end to end: heuristic evaluation, competitive analysis, personas, top task analysis, user interviews, information architecture, user flows, wireframes, visual design and usability testing.
The hard parts were not the deliverables. Finding and scheduling real users was a constant constraint, the system itself carried genuinely complex flows and features that had to be restructured rather than merely restyled, and the feedback had to be gathered and analysed in a form the team could act on.
Restructuring before restyling
Top task analysis established what the product actually had to be good at, and the old user flow was mapped in full before anything was redrawn. Seeing the existing flow at full size is what made the case for restructuring rather than resurfacing: the map is dense enough that its complexity is the finding.
What testing showed
Seven target users, five flows, sessions run across five days. Each task was scored on the same three-point scale (0 for not completed, 1 for completed with difficulty or help, 2 for completed easily), timed, and recorded alongside verbatim comments and observed barriers. The tasks were the real ones: add a server, create a disaster recovery pair, create a high availability pair, stop replications.
01
The tree lost people after every action
Completing a task returned users to the top of the tree rather than to the item they had just been working on, which meant re-navigating to continue.
02
Operators chose the edit screen over inline edit
Inline editing tested as pleasant but risky. Participants preferred a dedicated edit screen for values they change rarely.
03
A confirmation message was actively wrong
The description shown for the default distribution did not match what the system did, which is worse than no description at all.
04
Ambiguous button naming
A button labelled “add” replaced existing values. Participants suggested renaming it, or splitting it into separate add and replace actions.
“It’s annoying not to be redirected to the tree in the right place. You always get sent back to the first item.”
Design response
Rebuilt navigation around top tasks
Top task analysis, not the system's internal structure, decided what got surfaced. Hidden features were the single biggest adoption complaint.
Returned users to where they were working
Participants lost their place in the tree after every completed action, which turned a one-step task into a re-navigation exercise.
Kept a dedicated edit screen
Participants said plainly that inline editing can introduce errors on values they change rarely. Convenience was not worth the risk to them.
Made system status legible at a glance
Status summaries and dashboard-level indicators, so an operator arriving mid-incident can read the state before acting.
