A new study proposes a four-level model for deciding when an automated load control decision can be trusted without human review.
Ground operations are digitalising quickly, and load control β the process that plans an aircraft's load and calculates its mass and balance β is beginning to run without a person in the loop. A new paper by Ivan JakovljeviΔ, SVP of Operations at Ink, and researchers from the University of Belgrade's Department of Air Transport and Traffic ask a question that neither regulation nor academic research has yet answered:Β
When can an automated load control decision be trusted without human review?
The systems that automate the process today approve a flight under two conditions: the aircraft is within its certified mass and balance limits, and the operator's business rules are satisfied. Both checks matter, and neither is controversial. What they leave out is the judgment an experienced load control agent supplies without being asked β noticing that a flight is not behaving the way this route normally behaves, and weighing how much room is left for whatever happens between the final calculation and departure. Automation, as currently practised, removes that judgment.
The paper's solution is to translate those human checks into measurable ones. Its four-level validation model combines statistical process control with operational risk assessment, and it clears a flight for automatic release only when four conditions hold at once β the statistics look normal, the margins are adequate, the certified limits are respected, and the business rules are met β judged flight by flight, with no exceptions.
Why does the question need answering now?
In the traditional, centralised model, every participant in the turnaround is responsible for the accuracy of the data they submit β but one person, the load control agent, is responsible for checking all of it. The agent receives figures from check-in, the ramp, cargo and fuelling, verifies them, and reacts to changes that, by the nature of the process, arrive late: close to departure, when the options for correction have narrowed and decisions are made under time pressure. Attention is spread across dozens of routine confirmations rather than concentrated on the handful of events that actually carry risk.
Meanwhile, the industry is removing the technical reasons for keeping the process manual. IATA's digital ground-operations standards β the electronic loading instruction report (eLIR), the electronic dangerous goods notification to the captain (eNOTOC), which are part of the digital loading verification (DLV) and the xAHM aircraft data exchange β close the loop between load planning and the ramp, so that every figure can be captured, transmitted and confirmed electronically. Automation of load control is not the future; it is already being deployed.
What has not kept pace is the basis for trusting it. Regulation requires that the output of computerised mass and balance systems be verified and that every departure leave an audit trail, but says little about how to do either systematically once the primary executor of the process is software. Academic research, for its part, has focused on optimising where the load should go, rather than on validating the decision to release the aircraft. The result is an industry automating a safety-critical process without an agreed answer to when the automation should be believed β the gap this model sets out to fill.
How does the model learn what counts as routine?
It begins with the discrepancy between plan and outcome, tracked across 15 mass and balance parameters: the forward and aft holds, individual compartments, baggage, cargo, dangerous goods, zero-fuel mass distribution, and unit load device discrepancies.
Some variation between plan and outcome is unavoidable, given shifting passenger and baggage numbers, late-arriving bags, cargo adjustments and fuel changes. A model built to expect an exact match with the original plan would flag almost every flight, which is another way of saying it would flag none usefully.
The alternative is a benchmark drawn from history rather than theory. Historical flight data establishes the normal range of variation for each parameter, and a new flight is judged against what is typical for that specific route, aircraft type and set of conditions, rather than against a single rule applied across the network. A discrepancy is treated as significant only when it departs from what that particular context would predict.
Why do equivalent changes carry unequal weight?
Mass and balance behave unevenly across an aircraft. Baggage totals fluctuate considerably from flight to flight, while a given cargo compartment may barely move. An identical numerical shift can therefore mean very different things depending on where it occurs.
The model corrects for this by weighting parameters that are normally stable more heavily than those known to vary. A modest deviation in an otherwise steady pattern can carry more significance than a larger shift in a naturally volatile one.
Safety criticality reinforces the weighting further. Positions that have a bigger effect on the centre-of-gravity limit, and changes there, receive closer scrutiny even on routes where that hold typically varies more than on others.
Treating every kilogram as equally significant produces two failure modes: a flood of alerts over routine baggage fluctuation, or a missed shift in a compartment where the consequences would be far more serious.
Can single-parameter checks catch everything?
Certain faults are self-evident from a single figure β a large cargo offload, an unexpected dangerous goods adjustment β and individual checks remain essential for catching them.
Others emerge only from the combination. A baggage shift here, a compartment change there, a modest zero-fuel mass difference elsewhere: each within its own threshold, together forming a pattern that none of the individual checks would flag. The model therefore analyses data on two levels β each parameter against its own history, and the combined pattern across the aircraft as a whole. A system confined to isolated checks can approve a flight in which every figure passes individually, even where the aggregate loading pattern is markedly unusual.

Is an unusual flight necessarily a risky one?
Not always, and the mismatch runs in both directions. A cargo offload can produce a sizeable statistical deviation while the aircraft remains comfortably within every structural and centre-of-gravity limits, worth investigating, but not necessarily a cause for concern. Equally, a flight that closely resembles its historical pattern can still be operating near its maximum mass or centre-of-gravity boundary.
The model resolves this by running two assessments in parallel. A statistical assessment asks whether the outcome matches the baseline; a risk assessment considers certified limits, operator rules and the margin available for further change. Automatic approval requires both to be satisfied: adequate margin does not excuse an unexplained anomaly, and a statistically normal pattern does not compensate for insufficient margin.
Why does the remaining margin matter as much as the outcome itself?
Loading can continue changing until the moment of departure. Fuel uplift may deviate from plan, baggage may be accepted late, passengers may be reseated, cargo may be repositioned β each affecting total mass, distribution, or centre of gravity.
A compliance check speaks only to the moment it was run; it says nothing about the capacity to absorb what follows. For the centre of gravity, the model applies an allowance and evaluates the aircraft against a working margin tighter than the certified envelope, so that a flight comfortably inside that margin carries less operational risk than one at its edge, even when both are technically legal. The same logic applies to maximum mass, where the model considers remaining underload, likely fuel variation and possible late payload; an aircraft can sit below its maximum weight and still lack meaningful capacity for further adjustment.

The distinction, in short, is between compliance and resilience β between where the aircraft stands and what it can still absorb.
How does the analysis resolve into a decision?
The model reduces its output to three statuses.
A green status signals that the flight matches its expected pattern, remains within every limit and retains sufficient margin, and is therefore eligible for automatic release.
A yellow status signals a condition that is unusual but not blocking. A supervisor is notified, though the flight is not necessarily at risk; these cases prove useful in retrospect, since a recurring pattern of yellow flags can reveal a baseline drifting before it becomes a genuine problem.
A red status requires human review before release, whether the trigger is a genuine statistical anomaly, a failed compliance check, or a margin that has run out.
The most serious assessment governs the outcome. A favourable statistical result cannot override a red risk finding, nor can adequate mass margin offset an abnormal dangerous goods change. Automatic approval follows only when every criterion is satisfied β a deliberately conservative rule that allows automation and human oversight to operate concurrently rather than in tension.
What would this mean for an airline's operations?
In the study, 39 of the 48 flights examined qualified for automatic release; the remaining nine were referred for reasons that held up under scrutiny β cargo offloads, dangerous goods adjustments, and baggage differences that genuinely departed from the historical record.
The figure itself matters less than what it implies for how load controllers spend their time. Rather than reviewing every flight, they would review only the minority that requires judgment β a shift from uniform oversight to targeted attention.
The scale of the study invites caution: 48 flights constitute a proof of concept rather than an industry benchmark, and each airline would need to establish its own baseline and recalibrate it as data accumulates. What the results do suggest is that greater automation can be achieved in a safe way, as long as the verification layer is in place and human attention is targeted at anomalies and risks
Does automation make load controllers redundant?
The paper argues for a redefinition of the role rather than its removal. In a conventional, centralised model, a load controller reviews every flight, including the routine cases that follow familiar patterns and remain comfortably within limits. An exception-based model instead delegates routine validation to the system and escalates only flights that display unusual behaviour, fail a compliance check, or carry reduced margin.
This leaves load controllers free to concentrate where judgment is genuinely required β cargo changes, dangerous goods adjustments, baggage discrepancies that break from pattern, flights operating close to a limit β while routine releases no longer demand manual verification.
There is also a traceability benefit, and it is not a minor one in a safety-critical process: the system can identify precisely which parameter triggered an alert, whether the concern was statistical or operational, and why human review was necessary. An airline can therefore account, with precision, for why one flight was released automatically and another was not.

When, then, can automated load control be trusted?
Trust rests on whether the system can define the limits of its own authority. A credible automated decision demonstrates that the flight matches its expected pattern, remains within every limit, and retains margin for what may follow before departure and, critically, recognises the moment at which any of those conditions no longer hold.
This yields a clear division of labour: routine flights proceed automatically, unusual patterns are flagged, thinning margins are escalated, and load controllers handle what remains. Automation becomes safer not by taking every decision, but by knowing precisely which ones should still be left to a person.
β
Source: Validation Framework for Automated Aircraft Load Control Operations by Ivan JakovljeviΔ, Olja Δokorilo, LjubiΕ‘a Vasov; Published: 3 August 2026, the Special Issue, Interdisciplinary Insights in Engineering Research 2026. [https://www.mdpi.com/2673-4117/7/8/379]

