Why Competency-Based Training Can't Scale Without Intelligence

Contact Our Team

For more information about how Halldale can add value to your marketing and promotional campaigns or to discuss event exhibitor and sponsorship opportunities, contact our team to find out more

 

The Americas -
holly.foster@halldale.com

Rest of World -
jeremy@halldale.com



Image credit: Amris Aviation

From training outcomes to operational outcomes — across MPL, EBT and AQP

Andy O'Shea, FRAeS — Executive Chairman, Amris Aviation

Cedric Paillard — Chief Executive Officer, Amris Aviation


Executive summary

The most valuable training programme is not the one that delivers the most hours. It is the one that makes the operation perform better. Safer line flying, more resilient crews, lower cost. That is the shift running through APATS 2026, from training outcomes to operational outcomes, and it is what CBTA and EBT were built to deliver.

But the link between what happens in the simulator and what happens on the line only holds if an airline can measure competency and feed it back into the operation. Almost every operator has adopted Competency-Based Training and Assessment (CBTA) and Evidence-Based Training (EBT); the harder, shared task is making them work the way the founding policies intended. Drawing on real implementations we have worked on — anonymised to protect the operators — this paper looks honestly at a challenge the whole industry is navigating: a programme can be fully compliant on paper and still leave a training team short of evidence it can act on. That is not a mark against anyone. It is the nature of the transition, and it is exactly where the right support makes the difference.

We look at four challenges that turn up again and again. Grades can be recorded faster than the evidence behind them. The grading form is rarely built to capture what an instructor actually saw, behaviour by behaviour. Assessors drift apart when no one is resourced to check. And hard-won data can end up stored rather than used. For each one we set out the practice that closes it, explain why those practices are hard to run on spreadsheets and paper, and show how AI, used well, extends what a training team can do when it implements MPL, EBT and AQP.

The point is simple. The instructor is the decisive factor, and everything built around the instructor should exist for one reason: to let a skilled professional get on with training pilots. Closing the gap between what CBTA promises and what it delivers in practice is what Amris Aviation was built to do.

The gap between a sound design and a working programme

Adopting a competency framework is now the easy part. The hard part starts when an airline tries to run it, and here the experience is remarkably common across the industry: a programme can tick every box — the right competency model, a sound three-year cycle, the correct phase sequence — and still find it hard to produce evidence worth having. This is not a failing of any one operator. It is simply where a sound design meets the realities of a live operation.

Implementation does not live in the architecture. It lives in the form the instructor grades on, in the habits that build up under time pressure, and in the culture and management around them. That is where competency-based training quietly succeeds or struggles. None of the challenges below is a sign of a poor programme. Each is where a system tends to drift unless something is designed to hold it in place — which is exactly what a well-built system is there to do.

A programme is not built when its architecture is approved. It is built when instructors grade honestly and consistently, and when the evidence they produce is used to make decisions rather than filed away.

Lesson 1 — The silent, unsubstantiated grade

The clearest of these challenges is also the easiest to measure. In one operator's evidence-based cycle, around 95% of 'meets standard' grades, across the nine competencies, carried no supporting comment and no record of the observable behaviours behind them. The grade meant to be the workhorse of the whole system had quietly become a reflex: a record that looked complete and told you very little.

It would be wrong to blame tired or careless instructors, and it is worth seeing why. Grading a full session across dozens of observable behaviours and many phases of flight, late at night after a demanding session, is genuinely hard work. When the tool allows it, 'meets standard, no comment' is the easy way out at the end of a long day. The answer is better design and honest measurement, not exhortation. Track the silent-grade rate as a number the training department watches. Set a reasonable minimum for how much instructors comment, or how many observable behaviours they record. And give them a simple template so a good comment takes seconds to write. A grade with nothing behind it is a missed opportunity to learn something for the pilots, the instructors and the operation.

Lesson 2 — You cannot record what the form cannot hold

Underneath the silent grade sits a more basic issue: the grading record itself, the fields the instructor actually fills in. In many systems the form captures a single grade per competency and a free-text box, and little else. There is no straightforward way to record, behaviour by behaviour, what has been demonstrated and what has not, so the evidence the standard asks for is hard to capture even with the best will in the world. And where one button can set every competency to a grade of 3 at a stroke, the record can look finished while saying little about observable behaviours.

The lesson is a simple one. The form is the foundation of the whole programme. It has to capture behaviour-level data before any procedure, dashboard or analysis can be built on top of it.

Lesson 3 — The evaluator variance nobody was watching

If an airline cannot show how closely its assessors agree with one another and with the standard, it is hard to call the result evidence-based, whatever the manuals say. Assessor concordance is a basic requirement of EBT, and it is often the first thing to go unmeasured — usually for want of time rather than intent. Without a regular check (more than one a year), assessors gradually drift apart over months of live delivery and nobody notices. No other safety-critical industry would leave a measuring instrument unchecked like that.

Concordance is a standing obligation, not a box ticked at launch. Once the initial standardisation is over, instructors settle back, gradually and in good faith, into their own comfortable way of grading. What keeps a cadre of instructors aligned is not the quality of the first course but a properly resourced, repeating routine: measuring how closely assessors agree with each other and with the standard, a regular grading and evidence meeting led by the Head of Training, and calibration against a common, impartial benchmark.

Lesson 4 — The open loop

Evidence-based training is meant to run as a loop: better training and safety standards, driven by a cycle of training, data capture, analysis and decisions that feed back into the next round. In practice the loop is often left open, and rarely by choice. When automation and flight-path issues that are plainly visible in an airline's own flight data barely surface in instructor comments, the system is effectively running open loop. Training data, safety data and even selection data sit in separate systems, joined only by manual effort that rarely has the time to happen.

Part of this is a genuine mismatch of language, not carelessness. Training data is organised by competency, a grade against KNO, SAW, WLM and the rest. Safety and flight data are organised by event: an unstable approach, a level bust, a go-around. Joining the two is a real problem, and an operator that has not deliberately built that link will find its best sources of evidence cannot talk to one another. A programme whose data never feeds its own decisions is not yet evidence-based. It is evidence-storing.

Airlines do not need more training history. They need competency evidence they can trust, compare and act on, produced early and by design.

Why this keeps happening

Two things sit underneath all four challenges. The first is trust. An instructor will only record honest, detailed judgements if it is safe to do so for everyone whose name is on them. Without a proper “just-culture agreement”, settled with pilot representation before any behaviour-level recording begins, guaranteeing that the data will be de-identified and used to develop people rather than to punish them, the safe choice for the instructor is the silent grade of 3 that hollows the system out. Confidentiality is not something to bolt on at the end. It sits underneath everything else.

The second is ownership. Evidence-based training is a change programme that runs for years and touches recruitment, training, standardisation and safety. A large share of operators that enter the transitional (Mixed-EBT) phase find it hard to leave — not for lack of effort, but because ownership is spread too thinly. When it fragments, the form is never improved, the timeline to full implementation is never published, and standardisation quietly slips. The fix is organisational and unglamorous. One senior person owns the whole crew life cycle, and the role is treated as safety-critical rather than a project that ends on launch day.

The cognitive limits of CBTA

There is a more fundamental reason these challenges recur, and it has nothing to do with diligence. CBTA asks a single instructor to run a demanding session and, at the same time, observe, classify and grade against nine competencies and dozens of observable behaviours, across every phase of flight. Attention is finite, and it is drawn from one pool. When the cockpit saturates, a conscientious expert does what anyone would: they focus on the aircraft and the crew, and the recording waits.

The difficult part is the timing. Instructor workload spikes hardest at exactly the moments a scenario is built around, and those are the moments that generate the richest evidence. A rejected take-off or an engine failure produces an avalanche of assessable behaviour in seconds, and it arrives precisely when the instructor has least capacity to record it.

The best evidence is created at the very moments it is most likely to be lost. That is a structural challenge, not a personnel one.

This reframes the whole problem. The bottleneck is attentional, not judgemental. Instructors are not failing to judge well; the evidence is lost before judgement can begin. You do not fix a cognitive limit with more encouragement, more forms, or a better competency model. You fix it by giving the instructor a way to capture what they saw without taking their eyes off the crew, and by keeping the human firmly in command of the grade. That is assistance that sits low on the regulators' own scale of autonomy, fully monitored and always overridable. The instructor is augmented, not replaced, and that is what the rest of this paper builds on.

Closing the gap: from evidence-storing to evidence-based

Every challenge above has an answer, and together they describe the Amris approach. Our position is deliberately narrow. We build intelligent assistance that extends what an instructor can do and never replaces their judgement, with the technology there to support mentoring rather than take it over. In practice that means:

  • A form built to capture behaviour. Every exercise is tied to its observable behaviours, so the evidence is objective rather than anecdotal, and there is no shortcut that lets a grading finish without discriminating between behaviours. This addresses Lesson 2 at the source.
  • The silent grade, measured and designed out. The silent-grade rate becomes a number the department can see, and a simple, structured template makes commenting the easy path rather than the exception.
  • Concordance as a standing practice. Assessor agreement is measured every cycle, through regular evidence meetings and joint grading of recorded sessions, so drift shows up early and standardisation can be shown to improve rather than simply claimed.
  • Governance built in. Raw observations are kept separate from validated results. The original record is preserved and never overwritten, only reviewed and signed-off evidence leaves the building, and the analysis is de-identified and used to develop people, not to discipline them. Honest recording is protected, and the evidence stands up.
  • A loop that closes. Training evidence is linked to operational and safety data and fed back in two directions: into scenario design, so clusters of weak behaviours trigger a refreshed exercise, and into selection, so recurring in-training weaknesses inform who is recruited and how.

Put together, this is what we mean by decision-grade evidence: structured, traceable, aligned to CBTA, EBT and AQP, and solid enough for a training leader to act on. It turns the competency blind spot from a permanent feature into a solved problem.

How an intelligent system makes these practices work

There is a reason these practices so rarely hold on their own. Capturing behaviour, keeping assessors aligned and closing the data loop mean recording, mapping and reconciling far more evidence than any team can manage on spreadsheets and paper. The very things that make competency-based training powerful, its nuance and its focus on the individual, are what make it hard to run consistently at scale. This is where well-governed AI earns its place: not as a stand-in for the instructor's judgement, but as the layer that makes the practices workable.

AMRIS CMS is built to support the three points where traditional systems struggle. Observations from line checks, simulator sessions and briefings that would otherwise sit in separate places are captured in one structured record. Assessments that would vary from instructor to instructor are mapped to observable behaviours and continuously calibrated. And the sheer volume of evidence that no team could process by hand becomes visible patterns and trends. The method behind it is ORCA (Observe, Record, Classify, Assess), which sets a common way to document an observation, map it to a competency and classify it as the session runs, so the record is usable by the time it is saved. Personalised reporting then builds a picture for each pilot and each instructor, and dashboards surface trends across a cohort early enough to act before a problem becomes a failure.

The same engine serves the three frameworks operators actually run:

  • iATPL / MPL. In an ab-initio programme, the evidence trail starts on day one. Observation is structured from the first exercise, the long-term record the framework expects is built as you go, and the pathway adapts to what a cadet has demonstrated. It is the personal, feedback-rich experience the next generation of pilots expects.
  • EBT. In recurrent training, the core phases (evaluation, manoeuvres training and scenario-based training) are joined up so each grade feeds the next phase and the wider programme. Concordance is checked continuously, drift and outliers are flagged, and the loop between training, safety and selection is closed.
  • AQP. In an FAA Advanced Qualification Programme, the same evidence base gives you an audit-ready spine: proficiency objectives tracked against observable behaviour, continuing-qualification cycles backed by traceable data, and documentation a regulator can inspect on demand.

None of this works without governance, and the regulators have already set the bar. EASA's Artificial Intelligence Roadmap 2.0 and the FAA's guidance both stress explainability, fairness and human oversight. Amris is built to that standard. Every recommendation can be explained, so a Head of Training can ask why it was made. All assessments stay under human review, so the instructor keeps the final say. And the system points to the inconsistencies and outliers that signal bias or drift instead of hiding them. The technology is the enabler; the instructor is still the one who decides.

From training outcomes to operational outcomes

CBTA and EBT were never really about training for its own sake. They exist to make the operation safer and more efficient, by developing the competencies that drive line performance: decision-making, threat and error management, communication and workload management. The industry now has the numbers to show that link is real, and that the route to it runs through evidence, not through more hours.

Well-run, data-driven training moves operational metrics. Lufthansa's OPS Sustainability programme, which threads efficiency through the whole training chain, reports saving more than 54,000 tonnes of fuel and avoiding 170,000 tonnes of CO2 since 2022. Work by Boeing with airline customers on fuel and operational data has delivered 1–4% fuel-burn reductions worth $2–5 million a year per carrier, and broader studies put AI-assisted flight-path optimisation at 8–12%, depending on network and routing. In an industry running on a forecast net margin below 4%, numbers like these matter.

The same logic runs through the cost of getting competency wrong. For a representative 250-aircraft fleet (assuming a legacy mix of roughly 80% narrowbody and 20% widebody), we model the value of managing the competency framework with intelligent evidence capture at more than €40 million a year. That figure builds up across roughly €19 million in incidents and attritional damage (a single tail strike can exceed €12 million, a hard landing €200,000 to €1 million), about €15 million in operational inefficiency from fuel, track deviation and SOP drift, and around €8 million in training waste. These totals scale with fleet size, mix and utilisation, so each operator should rebuild them from its own numbers.

This is where evidence changes the economics. When training is measured properly and fed back into the operation, our modelling of the Amris CMS integration points to a 40% reduction in grading variance, a 50% reduction in remedial training, and a 30% cut in instructor administrative burden. EBT proportionality then lets a proven fleet reduce recurrent training volume by 20–25%. That footprint reduction alone is a material saving, before a single incident is avoided.

None of this comes from training more. It comes from training on evidence: knowing which competencies are strong, which are drifting, and feeding that back into scenarios, selection and the line.

A note on the numbers. The operational and financial figures in this section combine published industry sources with Amris CMS modelling and implementation results. They are indicative rather than guarantees. Every airline, ATO and flight school is different, so each should recalculate them against its own fleet, operation and cost base before using them in a business case.

What to take away

A competency-based programme improves operational performance when its instructors grade honestly and consistently, when the evidence they produce is used to make decisions rather than filed away, and when one senior person owns the whole thing from selection to the line. None of that comes from a better competency model, or from more hours. It comes from the systems that support the instructor and connect training to the operation, and that is the work Amris Aviation is focused on and Amris CMS is built for.

 

Sources. Lufthansa Group OPS Sustainability programme reporting; Boeing airline fuel-efficiency and operational-data studies; IATA industry financial outlook (net margin); ICAO Doc 9995, Manual of Evidence-based Training; EASA Artificial Intelligence Roadmap 2.0 and FAA guidance on AI and human oversight. The operator findings in this paper are drawn from real, anonymised implementations Amris has worked on; the operators are not named in order to protect confidentiality, but the findings are genuine.

 


Pilot competency. Clear. Consistent. Provable.

Amris Aviation is the exclusive AI Leadership Sponsor of APATS 2026. We are meeting delegates by appointment — to arrange a conversation, contact Cedric Paillard at cpaillard@amrisaviation.com.


Related articles



More Features

More features