• Case study • Design leadership •

Running design across a brand portfolio

Role

Head of UI/UX

Date

2024 – 2025

Responsibilities

Design leadership
Operating model
Team development

Timeline

2024 – 2025

Context

Seven designers. Twenty-two properties across eleven markets. This was performance marketing and lead generation at portfolio scale, which is what explains why a portfolio of that shape existed. My job was not to design the sites. It was to make the design function legible: to say what good meant, to show whether we were hitting it, and to back both with evidence rather than opinion.

01

What I inherited

The work was getting done. Nobody could say whether it was any good. Twenty-two properties, each with its own history, its own shortcuts and its own idea of quality. Ask whether one brand was in better shape than another and the honest answer was a shrug and a hunch.

That is a harder problem than it sounds. You cannot fix what you cannot compare, and you cannot compare twenty-two sites by looking at them. Every review turned into an argument about taste, because taste was the only instrument anyone had.

So the first months were not about redesigning anything. They were about building instruments that would let us argue from evidence.

02

Making quality countable

Two things had to be true before an audit meant anything: the same questions had to be asked of every brand, and no single person could be the judge.

1.

1.

One instrument, applied to every property

I built a heuristic instrument of thirteen categories and ran it across twenty-two properties in eleven markets. The categories were fixed, so a finding on one brand meant the same thing on another. Comparability is what turns a pile of observations into a portfolio view. We completed thirty-two site audits in Q1 2025.

2.

2.

Five evaluators, so a score was not one person’s taste

One reviewer produces one opinion with a number attached to it. Five evaluators working the same instrument produce something that survives a room full of stakeholders who each have their own view of what good looks like. The disagreements were useful in themselves: where evaluators diverged, the category was usually ambiguous and needed rewriting.

03

Measuring people without pretending

Auditing sites is one thing. Measuring designers is where this kind of work usually goes wrong, so I want to be precise about what I built and about what I deliberately left out.

The system covered seven designers, scored quarterly, drawn from eight stakeholder responses and seven internal survey responses. Every metric was bound to a named evidence source, so any score could be traced back to something real. If I could not point at the evidence, the metric did not belong on the scorecard.

The most important decision was a removal. The first version scored each designer on Outcome. I took it out. Revenue is not attributable to one person: a designer can do excellent work on a page that a pricing change then buries. Scoring someone on a result they do not control teaches them to chase attribution instead of doing the work. What remained measures contribution, not weather.

04

What it produced

These are adoption and delivery numbers. They say the operating model was working, and that is all they say.

7

designers on the system

22

properties audited

11

markets

13

audit categories

81.5%

of tasks on time

The delivery changes behind that number were unglamorous: the team moved to Kanban, and I added a Design QA stage before anything reached stakeholder review, so we stopped discovering problems in the room where they were most expensive.

One product outcome sits alongside these, and it needs its attribution stated plainly. Click-through on a rebuilt paid-acquisition landing template went from 39% to 71% year on year. Design, marketing, product and engineering delivered that together, and these are early results. The scoring system did not move that number, and I am not going to pretend it did.

I am often asked for the development-time saving. We never measured it, so I do not have one. The mechanism is real, a shared component library removed a rework loop, but I will not put a percentage on something nobody instrumented.

A score you can’t trace back to evidence is an opinion with a number on it.

05

What I’d say about it now

I built the instruments before I had the team’s trust in them, and that cost me a quarter. Scoring arrived as something being done to people rather than something they had helped shape. If I ran it again I would draft the categories with the designers first, even though the instrument would come out much the same, because adoption is a function of authorship.

I would also cut the scorecard down. Quarterly scoring across seven people, bound to named evidence, is a lot of ceremony for a team that size, and some quarters the reporting cost more than the signal was worth.

The audit instrument was the part that earned its keep. It outlived my involvement, which is the only real test.

What I would keep without hesitation is the evidence rule. Once every metric had to name its source, the arguments changed character. They stopped being about whose judgement was better and started being about whether the evidence supported the claim.

That is the habit I took with me, and it is the one I would install first anywhere else.

Hiring a design leader, or need one for a while?

Let’s talk.

Marko Rosić PR Beyond Clicks Studio · MB 67353170 · PIB 114143089
Cara Lazara 26/21, 34000 Kragujevac, Serbia

© 2026 rosic.net