adrian_zbercea

10+YEARS
6SECTORS
30+ENGINEERS LED
4DISCIPLINES
17CERTIFICATIONS

> where_i_add_value

What a company gets

Four problems I’m brought in to solve, and what changes when they’re solved.

01

Reliability that compounds

Incident response is the easy half — every company has it. I build the problem management underneath: severity classification that produces usable data, root cause analysis held to a real quality bar, and a known error database that works as organisational memory instead of a graveyard.

→ The same failure stops costing you twice.

02

Delivery you can plan against

Roadmap and quarterly planning alongside Product, release coordination across team boundaries, and forecasting that survives contact with reality. Scrum, Kanban and SAFe used as tools rather than as identity, and budget held as part of the same job.

→ Dates the rest of the business can commit to without hedging.

03

An organisation that outgrows any one person

Team structure and ownership boundaries, hiring and onboarding that hold up under growth, career paths people can actually see, and a second layer of leadership able to decide without me in the room.

→ Capacity that scales past the bottleneck at the top.

04

AI adoption that survives the pilot

The governance questions answered first — what data may leave, what output needs human review, who is accountable when it’s wrong. Then rollout into the workflows people already resent, where the person doing the work can judge the output immediately.

→ A measured change in the work, not a usage dashboard.

> tools

Three things I actually use

Not a portfolio — working versions of the thinking I bring to a team. Have a go.

TOOL 01

Incident severity matrix

Pick an impact and an urgency.

Low
Medium
High
Critical
All customers
Business unit
Single team
One user

CLASSIFICATION

Select a cell

Severity is the first thing most organisations get wrong. If it is assigned by whoever shouts loudest, the incident data is noise and no amount of trend analysis will rescue it.

TOOL 02

What changes as a team grows

Drag to set headcount.

15ENGINEERS
2TEAMS AT ~8
105COMMUNICATION PATHS

WHAT BREAKS HERE

Stretching

You are starting to hear about problems rather than see them. This is the point to name owners, before your instinct starts firing with the same confidence on much worse data.

TOOL 03

Is that actually a root cause analysis?

Tick what your last one contained.

VERDICT · 0/7

Nothing ticked yet

Tick the boxes above. The bar is not a score for its own sake — each line is something I would send an analysis back for.

> perspective

How I think about the work

Four positions I hold strongly enough to argue for. Open one.

Reliability is a memory problem

Most organisations are good at incident management and almost nobody is good at problem management. That difference is why some teams get quieter and others just get busier.

+

Incident management restores service. Problem management makes sure the same class of failure costs less every time it appears. A team that only does the first is very busy and never gets quiet.

The mechanism is boring and administrative, which is exactly why it gets skipped. Classification is what lets you see patterns instead of anecdotes: if severity is assigned by whoever shouts loudest, your incident data is noise.

The known error database is the organisation’s memory. When it is maintained, an engineer meets a failure a colleague already solved and closes it in ten minutes. When it isn’t, they rediscover it, and the company pays twice for the same lesson.

“Human error” is never a root cause. It is a description of a system that permitted the error and didn’t catch it.

The RCA quality bar is where I am least willing to compromise. An analysis that terminates at a person has failed, and I’ll send it back. The useful question is never who typed the command — it is why the command was available, why nothing flagged it, and why recovery took as long as it did.

Two things make it hold in practice. RCA needs a protected turnaround time, because an analysis delivered weeks late is a document rather than a control. And blamelessness has to be real rather than declared — everyone can tell the difference between a company that has written the word on a wiki and one where naming your own mistake has visibly never harmed anyone. Only the second kind gets told the truth.

The payoff isn’t a dashboard. It’s that the same incident stops arriving.

What breaks when a team becomes an organisation

Fifteen engineers and thirty engineers are not the same job scaled up. They are different jobs, and the common failure is doing the second with the habits of the first.

+

At fifteen you can hold every thread yourself. You know roughly what each person is working on, and your instinct is usually right because it rests on direct observation. That competence is exactly what betrays you at thirty: your information becomes second-hand, and the instinct keeps firing with the same confidence on much worse data.

So the first real decision at that size is what you are going to be deliberately ignorant of. Not what you’ll delegate, since delegation still implies you’re tracking it, but which categories of detail you will genuinely stop holding. A leader who keeps a hand in everything at thirty people becomes the queue the whole organisation waits in.

The second change is where predictability comes from. On a single team it is mostly the team’s own discipline. Across four teams, almost all the variance lives in the connections between them — the dependency raised too late, the shared service nobody owns, the release that needs three teams to align on one afternoon. Working harder inside each team does very little for that.

You stop optimising decisions and start optimising who gets to make them.

Making the seams explicit does almost all of the work, which is why cross-team release coordination stops being an administrative chore and becomes a first-class responsibility with a name attached to it.

The third change is the uncomfortable one. You now lead through people whose judgment you have to trust before you have much evidence for it, and the only way to get evidence is to let them decide things while the stakes are still survivable. Managers who can’t tolerate that never build a second layer — and an organisation without a second layer is just a large team with a bottleneck at the top.

Adopting AI without a mandate

AI adoption fails when it is announced. A target arrives from above, tools get bought, dashboards start counting usage — and engineers become fluent at generating whatever the dashboard measures.

+

It works when it starts from a task somebody already resents. Not the most impressive use case: the most annoying one. Release notes assembled by hand. First-pass triage on a noisy alert stream. Turning an incident timeline into a draft analysis that a human then corrects.

These are unglamorous, low-stakes and high-frequency, and they have the property that matters most — the person doing them can immediately tell whether the output is any good. That feedback loop turns curiosity into habit, and habit is the only thing that survives the enthusiasm phase.

Measure the work, not the tool. Usage statistics tell you how compliant people are, not whether anything got better.

Before any of it reaches a team, the governance questions need real answers, because engineers will ask and a vague answer stops adoption dead. What data is allowed to leave the building. What output requires human review before it reaches a customer, a repository or a regulator. Who is accountable when it’s wrong — and the answer is always a person, never the tool.

The last thing I’d say is that this is mostly a permission problem wearing a technology costume. Most engineers have already tried these tools privately. What they lack is a clear signal about what is sanctioned, what is off limits, and whether using them reads as competence or as cutting corners. Answer that plainly and adoption largely handles itself.

What I look for, and what I owe in return

Hiring and growing people are the same activity separated by a start date, so it’s worth being consistent about what you’re selecting for.

+

I screen for three things. How someone reasons when the information is incomplete — not whether they reach my answer, but whether they can say what they’d need to know and how they’d find it. Whether they can disagree well, holding a position under pressure without turning it into a contest. And whether they have ever changed their mind about something technical, and can explain what changed it.

That last one is the most predictive question I know, and it’s remarkable how many strong applications have no answer to it. What I’m not selecting for is recall: trivia interviews measure preparation and confidence, both distributed for reasons that have nothing to do with how good someone will be a year in.

Feedback delivered at review time isn’t feedback. It’s a verdict on a period the person can no longer influence.

The other half is what I owe. Clarity about what “good” looks like at their level, specific enough that they could assess themselves against it without me. One-to-ones that aren’t status meetings. Feedback early enough to act on, including the uncomfortable kind. And honesty about what I can’t promise — I would rather say plainly that the promotion isn’t available than manage someone with an implication I can’t back.

I treat attrition as feedback on management before anything else. People rarely leave a role they’re growing in, for a manager who is straight with them, on a team that ships. When someone leaves, at least one of those three was missing, and working out which one is more useful than concluding that the market is hot.

> certifications

Formal grounding

Seventeen, across service management, agile delivery and AI.

ITIL® Foundation in IT Service ManagementAXELOS
ITIL® Continual Service ImprovementGlobal Knowledge
Certified Agile Service Manager (CASM)DevOps Institute
ISO/IEC 27001 Information Security AssociateSkillFront
ISO 9001 Quality Management Systems AssociateSkillFront
Certified SAFe® 6 Product Owner / Product ManagerScaled Agile, Inc.
Certified SAFe® 6 Scrum MasterScaled Agile, Inc.
Professional Scrum Product Owner I (PSPO I)Scrum.org
Professional Scrum Master I (PSM I)Scrum.org
Certified Kanban AssociateInternational Scrum Institute
Certified Scrum Product Owner (CSPO)International Scrum Institute
Scrum Master Certified (SMC)International Scrum Institute
AI in Agile DeliveryProject Management Institute
Generative AI Overview for Project ManagersProject Management Institute
Generative AI Learning Plan for Decision MakersAmazon Web Services
AI for Product ManagementPendo.io
Snowflake Sales Professional AccreditationSnowflake

showing all 17

BSc Marketing, West University of Timișoara · Romanian (native), English (full professional), German & Spanish (elementary)

> contact

Get in touch

Always glad to talk about engineering organisations, service reliability, or the practical side of AI adoption — whether or not there’s a role attached. A message on LinkedIn is the surest way to reach me.

LINKEDINlinkedin.com/in/adrianzbercea
BASED INTimișoara, Romania · CET · remote across Europe