Applied AI · Quality · Delivery

01 / 10

Evidence before adjectives.

I build AI systems, make their results visible, and help teams adopt them without losing human ownership.

Paul Leung 16 years in quality & delivery 3 years building with applied AI Hong Kong
Open with the position, not a list of technologies. I am not presenting myself as a model researcher. My value is turning AI capability into controlled, measurable work that teams can adopt.

Why my background fits AI

AI made quality engineering more important.

02 / 10

An AI can suggest, plan, and act. It should not be the only party deciding that its own work is correct.

Foundation QA engineering across web, mobile, API, payment, gaming and media
Leadership QA management, release delivery, coaching and cross-team execution
Applied AI Agentic support PoC and AI-driven test automation at work
Platform depth Self-built multi-agent platform with tools, policies and operations
My QA background is not something I am trying to escape. It is the reason I naturally ask for expected outcomes, failure evidence, release gates and ownership. Those disciplines are essential when system behaviour is probabilistic.

Three pieces of evidence

Built, transferred, and operated.

03 / 10
Work · Active use

AI test automation

Natural-language test scenarios become browser-executed steps with bounded retry, evidence capture and reporting.

Proof: adopted on QA suites; regression scope expanded from about 10% to 80%.
Work · Handoff

Agentic IT support PoC

Built a bot-first support journey, then coached the delivery team on prompts, tools, escalation and productionisation.

Proof: moved from one-person PoC into a separate team's delivery path.
Personal · Engineering lab

Multi-agent platform

Built an enterprise-pattern platform to learn orchestration, MCP tools, approvals, model routing, governance and operations together.

Boundary: serious personal engineering, not presented as enterprise production scale.
These three examples prove different things. Test automation proves adoption and measurable work impact. The support bot proves use-case shaping and handoff. The homelab proves hands-on platform depth. I do not combine them into one inflated claim.

Case study · AI test automation

More regression, without turning QA into programmers.

04 / 10
10% 80%
Approximate expansion of repeatable regression scope.
Before Manual capacity limited regression breadth
After QA authors intent in plain-language test files
  • AI works within the current test step rather than freely exploring the whole system.
  • Browser actions must leave evidence before a step can be accepted.
  • Failure context is carried into bounded retry instead of blindly repeating.
  • The wider automation scope created room for more system-level testing.
Honest boundary: the claim is expanded regression scope and team adoption—not that AI removed QA, guaranteed defect reduction, or replaced engineering ownership.
The 10-to-80 figure is useful because it describes a concrete operational change. I should explain the denominator if asked. I should not turn it into a code-coverage or defect-reduction claim unless I have that separate evidence.

What the system actually does

Intent enters. Evidence comes out.

05 / 10
The prompt guides.

Code still owns permissions, limits and hard stops.

The agent acts.

It cannot declare success without supporting evidence.

The team owns.

QA keeps the test meaning and final release responsibility.

If the interviewer uses “agent harness,” this is my concrete example. The harness is everything around the model that limits the task, gives tools, preserves context, captures evidence and decides whether the result is acceptable.

Personal engineering laboratory

I built the platform to understand the whole chain.

06 / 10

Not a tutorial chatbot: a working environment for learning where agents, tools, policy, data and operations meet.

16MCP services exposing domain tools
5+model families routed through one platform
Boundary: this proves personal technical depth and operating judgment. It does not claim enterprise traffic, revenue, or a staffed production organisation.
ExperienceChat, approvals, pending work and user-facing run status
Agent runtimeDAG workflows, per-step instructions, skills and cross-session memory
Tool layerMCP services, tool discovery and bounded execution
Control layerRBAC, scoped policy, audit trails and encrypted secrets
OperationsPostgreSQL, Redis, object storage, CI/CD and Kubernetes
The point of this slide is not technology counting. It shows I have touched the boundaries that cause real agent failures: tool contracts, pending state, permissions, model choice, auditability, deployment and recovery.

How I would scale beyond one builder

A core AI team, with shared ownership.

07 / 10
Domain teams

Own the meaning

They know the workflow, exceptions and business consequence.

  • Use-case definition
  • Domain tools and knowledge
  • Acceptance criteria
  • Final outcome sign-off
AI platform & enablement

Build the shared road

A small engineering-led team creates reusable capability and helps lighthouse teams deliver.

  • Agent foundation and tool standards
  • Evaluation and observability
  • Common controls and cost management
  • Coaching, patterns and adoption
Independent gates

Protect the company

SRE and Security do not need to sit inside the AI organisation.

  • Reliability requirements
  • Security and data policy
  • Production readiness
  • Incident and audit standards
I would not duplicate Security or SRE inside the AI team. I would involve them early, build the controls they require, and keep their approval independent. The central AI team owns the reusable road; domain teams still own destination and business correctness.

Adoption without theatre

Autonomy should be earned.

08 / 10
Stage 1

Assist

AI reviews code, tests or delivery information. A human decides what to use.

Gate: usefulness and review burden
Stage 2

Prepare

AI proposes a plan or change. The engineer approves before execution.

Gate: plan quality and safe tool scope
Stage 3

Execute

The agent performs a bounded task. Evidence and human sign-off close the work.

Gate: repeatable success and recovery
Stage 4

Automate

Only stable, reversible and observable tasks move to controlled automation.

Gate: agreed risk, monitoring and rollback
This is how I would introduce coding agents into an organisation. I do not begin by asking the agent to own delivery. I begin where output is reviewable, build the evaluation data, and only increase autonomy after the system proves stable.

What I would measure

Usage is a signal. Outcome is the result.

09 / 10
CompletionTask success

Did the requested outcome actually happen—not merely return HTTP 200?

Human costReview and rework

How much correction was required before the output became usable?

QualityDefect escape

Did speed improve without moving defects into later stages?

ControlSafe execution

Were permissions, policies and approval boundaries respected?

FlowEnd-to-end time

Did the complete workflow become faster, including review and recovery?

EconomicsCost per success

Model and platform cost divided by accepted, useful outcomes.

One overall percentage is not enough. Break results down by task type, domain, risk and failure severity.
I use percentages, but I care about what sits under them. An 80 percent average can hide a complete failure on a high-risk workflow. The management outcome is faster useful delivery after including review, rework, defects and cost.

The role I am ready to play

Build the evidence. Build the team. Earn the trust.

10 / 10

I sit between hands-on AI engineering, quality discipline, delivery leadership and team adoption.

People Manager and coach

Hiring, team growth, delivery rhythm and cross-functional alignment.

Product Applied AI builder

Use-case shaping, agent workflows, tools, platform decisions and handoff.

Production Quality-led operator

Evidence, failure analysis, controls, adoption and measurable closure.

[email protected] linkedin.com/in/paul-leung-39479a51 Hong Kong · Cantonese · Mandarin · English
Close by connecting the evidence to the role. I am strongest where an organisation needs someone who can prove the technical path, structure delivery, build team capability and turn an AI experiment into an adopted operating system.
Speaker note

Back to site