Following-on from my previous article, Why the “Manual or Automated?” Study Is Fundamentally Flawed, I decided to deploy ChatGPT to conduct a full forensic, page-by-page critique of the 51 page Manual or Automated study, produced by Dr. Piers Robinson, in May 2026.
I would advise readers to read the study for themselves, so they can fully understand the problems which I pointed out in my previous article. The construction of the study itself was set-up to give a false, flawed binary choice, and conclusion.
This article covers a ChatGPT unbiased request to analyse the study, and report its results, which I have published here.
A Forensic, Page-by-Page Critique of the Manual or Automated? A Flight Simulation Study and Analysis of Reported Aircraft Maneuvers on September 11, 2001 The study’s central methodological problem can be stated very simply:
It attempts to determine whether Boeing 767 and 757 aircraft could have been manually flown along particular reconstructed trajectories, but the principal simulator experiment was conducted using a Boeing 737 simulator whose handling characteristics the authors themselves acknowledge may differ materially from those of the aircraft being investigated.
That problem is compounded by:
Extremely small sample sizes
Artificial starting conditions
Researcher-controlled timing cues
Changing instructions between experimental stages
Deliberate omission of parts of the manoeuvres
Simulator familiarisation effects
Reconstructed rather than directly measured flight paths
Ambiguous coding of some results
Reliance on first-attempt performance
Lack of an equivalent experimental demonstration of the proposed automated mechanism
A substantial inferential leap from “difficult manually” to “therefore automated.”
Most importantly, the study repeatedly moves from what the experiment actually measured to what the authors believe probably happened historically. Those are not the same thing.
Pages 1–2 — Authorship, credentials and framing
Pages 1–2 identify Piers Robinson as the author and research director of IC911, with a PPL/IMC rating and approximately 200 hours of flying experience. Seven pilots participated in the simulations, alongside contributors and reviewers.
Methodological issue
The author’s 200 hours of flying experience is not itself a problem. Nor is institutional affiliation evidence that the study is incorrect.
But the report subsequently makes highly specialised claims about aircraft handling, flight-path reconstruction, probability and automated guidance.
The study therefore requires particularly strong methodological safeguards.
What is missing from these opening pages is an independent statement concerning:
Independent peer review
Simulator validation against the actual aircraft
Independent reconstruction of the flight paths
Pre-registration of hypotheses
Statistical analysis plan
Or independent replication
This becomes important later because the report’s conclusions are considerably broader than its experimental design.
Page 3 — The first major logical problem
The executive summary states that direct routes to the targets were readily accomplishable, while the reported routes involved “unnecessarily demanding” manoeuvres.
This is already an important methodological distinction. The experiment demonstrates that alternative routes were easier. It does not demonstrate that the reported routes were impossible or implausible. An aircraft can follow a more difficult route for many reasons.
The study needs to establish:
Why should the existence of an easier route be evidence against manual flight?
That proposition is never experimentally established.
A pilot does not necessarily choose the geometrically simplest possible route merely because it is easier.
The study effectively introduces an unstated behavioural premise:
A manually controlled pilot would necessarily choose the easiest available route.
That is an assumption, not a measured law of aviation behaviour.
Page 4 — The conclusion appears before the evidence
Page 4 goes substantially further.
It argues that the reported flight paths show deliberate precision and that the simulations support automated control.
The crucial problem is the phrase:
“Taken together, the flight simulations and analysis of available flight path data support the hypothesis of automated control.”
The experiment does not actually test an automated-control system.
It tests human pilots flying a 737 simulator.
Thus the experiment provides data concerning the difficulty of a manually performed task. It does not provide corresponding experimental data demonstrating the operation of an automated system. That creates an asymmetrical comparison:
Manual: Experimentally tested.
Automated: Inferred.
That is not a controlled comparison between two competing hypotheses.
Pages 4–5 — The study crosses from aviation analysis into a much larger historical conclusion
The report then states that its findings strengthen the thesis that 9/11 was a false-flag event rather than a terrorist attack.
This is an enormous evidential leap.
Even if the aviation experiment were completely sound, demonstrating that a particular manoeuvre was difficult for a sample of simulator pilots would not establish:
Who controlled the aircraft
Whether remote control existed
Whether an automated guidance system existed
Whether such a system was installed
Whether it was activated
Who operated it
Or why?
The experiment therefore cannot bear the evidential weight assigned to it in the broader conclusion.
Pages 5–6 — Aircraft mismatch becomes central
The study identifies UA175 and AA11 as Boeing 767s and AA77 as a Boeing 757.
Then, on page 8, it reveals the experimental aircraft:
“a full motion (Category D equivalent) Boeing 737 flight simulator was secured for use.”
This is arguably the single most important methodological weakness.
A Boeing 737 is not a Boeing 767.
Nor is either aircraft a Boeing 757.
The relevant question is therefore not whether both are commercial Boeing jets.
The relevant question is whether the 737 simulator accurately reproduces the specific aircraft-performance characteristics that determine the manoeuvre being studied.
Those include:
roll response
pitch response
control sensitivity
inertia
aerodynamic response
high-speed handling
turn performance
energy retention
thrust response
and control-system behaviour
The report does not provide a quantitative validation demonstrating equivalence in these characteristics.
Page 8 — The “validation” of the 737 is inadequate
The authors attempt to deal with the aircraft mismatch.
Two highly experienced pilots, each with more than 10,000 hours, evaluated the simulator. They did not actually fly the 9/11 manoeuvres. They assessed handling at high and low speeds and concluded that it was a reasonable approximation of a 767 or 757.
This is not a rigorous validation.
A qualitative opinion that:
“this feels reasonably similar”
is not equivalent to establishing:
“the simulator reproduces the relevant 767/757 aerodynamic and control characteristics within a defined error margin.”
That distinction is critical.
The study needed to establish fitness for purpose, not merely general plausibility.
And later in the report, the researchers themselves admit that the 737 may have been less manoeuvrable than the actual 767 or 757.
That admission substantially weakens the earlier validation.
Page 9 — The first-attempt premise
The study decides that first attempts are particularly important because the aircraft supposedly hit the targets on their first attempts.
This is a major assumption.
The authors argue that prior simulator practice would not have provided knowledge of how the real aircraft handled.
But the relevant question is not simply:
Did the pilot practise this exact manoeuvre?
It is:
What transferable aviation skills did the original pilots possess?
A pilot does not learn every manoeuvre from scratch.
Pilots transfer skills between aircraft and situations.
The experiment’s own results demonstrate the importance of familiarity with the particular simulator.
That creates a confounding variable:
failure may reflect unfamiliarity with the simulator rather than inability to perform the manoeuvre.
Page 9 — Another problem: 15 minutes of familiarisation
The study states that the pilots received at least 15 minutes to practise handling the simulator before commencing runs.
This is problematic because the later pilot testimony explicitly identifies simulator unfamiliarity as a major factor.
Pilot 6 said that the difficulty involved his unfamiliarity with the 737 simulator’s responses and that he believed he could adapt more sharply in an aircraft he normally flew.
Pilot 3 similarly described unfamiliarity with the control inputs and differences between aircraft control systems.
This is extraordinarily important.
The study is effectively using failure caused partly by unfamiliarity with the experimental aircraft as evidence concerning the ability to fly a different aircraft.
That is precisely the confound that should have been eliminated.
Pages 10–11 — Direct routes are not a control for manual capability
The direct-route experiments show that pilots could easily hit the targets.
That is useful.
But the study interprets this as evidence that a manually controlled aircraft would have taken the direct route.
That conclusion does not follow.
The experiment demonstrates:
The direct route was easier.
It does not demonstrate:
A manually controlled aircraft would necessarily take the easiest route.
Those are completely different propositions.
This is one of the study’s recurring problems: an empirical observation is converted into a behavioural assumption.
Pages 12–13 — The experiment is not actually a complete replication
The South Tower experiment is particularly problematic.
Pilots were given a command approximately ten seconds before impact telling them when to initiate the turn.
This means the experimental pilot was given information that the original pilot supposedly did not have.
The researcher knew:
the target
the desired trajectory
the timing
the intended manoeuvre
and when the turn was supposed to begin
The original pilot, by contrast, was not operating under experimental instructions.
More importantly, the study explicitly says:
“we did not give pilots instructions to initiate the last-second pull-out and instead focused on the execution of the final turn only.”
This is a fundamental limitation.
The experiment did not reproduce the complete manoeuvre.
Yet the conclusions discuss the difficulty of the complete last-second manoeuvre.
That is a classic construct-validity problem.
Page 14 — A potentially serious heading error
The study says it identified an error in the NTSB’s conversion of true headings to magnetic headings and corrected it.
This is potentially important.
If the researchers are correcting the source data used to construct the experimental trajectory, then the study must establish:
precisely what the original error was
how it arose
how the correction was calculated
whether all subsequent coordinates were recalculated
whether the NTSB’s underlying radar data were affected
and whether the corrected path is independently validated
The report does not provide sufficient detail in the study itself to establish the uncertainty introduced by this correction.
This matters because small heading differences at high speed can produce substantial positional differences over seconds.
Pages 15–17 — The Pentagon experiment introduces additional artificial constraints
The Pentagon experiment is even more complicated.
Pilots were instructed to reproduce the 330-degree orbit and low-level approach.
The instructions were then modified.
At one point pilots were told to attempt to clip light poles; later the instructions were changed so that the poles became visual guides.
This creates a major experimental-design problem:
The experimental conditions were not constant.
If instructions change between participants or stages, then performance cannot simply be treated as though all participants were operating under identical conditions.
The experiment is no longer testing merely:
“Can pilots perform this manoeuvre?”
It is testing:
“Can pilots perform this manoeuvre under this particular set of instructions?”
Those are not equivalent questions.
Pages 18–22 — The direct-route results actually establish something useful
The direct-route experiment is arguably the strongest part of the study.
All four lower-experience pilots successfully aligned with the South Tower, and three of four successfully struck the Pentagon directly.
This demonstrates something valuable:
Hitting a large stationary target with a large aircraft in a simulator is not necessarily difficult merely because the aircraft is travelling fast.
But it does not establish the subsequent proposition:
Therefore the original pilots would necessarily have chosen the simplest trajectory.
That remains an inference.
Pages 23–24 — The most damaging experimental admission
Page 23 contains one of the most important passages in the entire study.
The authors explicitly acknowledge:
uncertainty about the correct offset
uncertainty about when to instruct the turn
possible lower manoeuvrability of the 737
graphics limitations
problems with the instructions
and the possibility that other simulators could produce higher completion rates
This is not a minor limitations paragraph.
These are precisely the variables that determine whether the experiment is measuring what it claims to measure.
The report nevertheless says the results can still support substantive conclusions.
That is where the methodological argument becomes weak.
If the apparatus itself may systematically increase the difficulty of the manoeuvre, then the failure rate cannot be treated as an unbiased estimate of real-world difficulty.
Pages 23–26 — The results reveal the simulator-learning problem
The study reports zero first-attempt completions for the South Tower indirect manoeuvre.
But the following pages explain why.
The pilots had difficulty determining:
appropriate roll angle
appropriate back pressure
correct initiation timing
and how the simulator responded
Critically, the pilots themselves describe unfamiliarity with the simulator.
Pilot 2 said:
“The roll rate is quite low.”
and explicitly referred to getting used to how the simulator performed.
Pilot 3 described unfamiliar control fidelity and stated that he was unfamiliar with the control inputs despite having thousands of hours in aircraft.
Pilot 6 was even more explicit: he said his difficulty resulted from unfamiliarity with the 737 simulator’s response and that he believed he could adapt the turn more sharply in an aircraft he normally flew.
This is perhaps the strongest evidence against treating the first-attempt failures as direct evidence of real-world impossibility.
The pilots themselves identify simulator unfamiliarity as a cause of failure.
Page 27 — The study turns pilot preference into evidence
The pilots repeatedly say that they would personally have flown a more straightforward route.
This is interesting qualitative evidence.
But it is not proof of what another pilot would have done.
Nor does it establish that a pilot choosing another route is evidence of automation.
This is essentially an appeal to expert preference:
“I would not fly it that way.”
That is not the same as:
“A competent pilot could not fly it that way.”
The distinction is crucial.
Page 28 — The critical inferential leap
The study concludes that the required precision was inconsistent with manual control by the alleged hijackers.
But the experiment has not established the probability of successful manual execution.
It has established a small experimental failure rate under specific conditions.
Those are not equivalent.
There is no statistical model that converts:
0/5 first attempts
into:
probability that the original pilot could not perform the manoeuvre.
Nor is there a Bayesian analysis incorporating:
prior pilot experience
aircraft differences
route uncertainty
actual visual conditions
possible navigation assistance
and uncertainty in the reconstructed trajectory
The phrase “inconsistent with manual control” therefore goes beyond the experimental data.
Pages 28–30 — Pentagon completion rate: 28%
The study reports 25 attempts at the Pentagon low-level manoeuvre, of which seven succeeded, giving a 28% completion rate. None succeeded on the first attempt.
At first glance, this appears powerful.
But the 25 runs are not a clean statistical sample.
They include:
different pilots
different attempts
repeated attempts by the same pilots
different starting points
evolving pilot experience
modified instructions
and different experimental conditions
Therefore the 25 runs cannot simply be treated as 25 independent observations.
This is a major statistical issue.
If Pilot 1 performs five runs, those five observations are not equivalent to five independent pilots.
The study does not provide an appropriate statistical treatment of this repeated-measures structure.
Consequently, quoting 28% completion creates an appearance of statistical precision that the experimental design does not support.
Pages 29–30 — Visual acquisition is confounded with flight skill
The study says pilots had difficulty reacquiring the Pentagon visually after the orbit.
But this is not necessarily evidence of the difficulty of manually controlling the aircraft.
It is partly a test of:
How easily can a simulator pilot visually reacquire a target after being deliberately instructed to fly away from it?
That depends upon:
simulator graphics
field of view
rendering
cockpit geometry
visual fidelity
atmospheric representation
and pilot familiarity with the simulated environment
The report itself acknowledges that graphics quality may have complicated the Pentagon runs.
Therefore visual reacquisition failures cannot simply be interpreted as evidence that the original manoeuvre was beyond manual flying capability.
Pages 30–33 — The study effectively admits that the manoeuvre was abnormal
The pilots repeatedly say the manoeuvre was unnatural, counter-intuitive and difficult.
That is not surprising.
They were explicitly instructed to reproduce a highly unusual trajectory.
Pilot 1 even described it as something outside normal piloting behaviour.
But again:
Unusual does not mean impossible.
The study repeatedly conflates:
unusual
difficult
counter-intuitive
unlikely
and impossible
These are different propositions.
A manoeuvre can be:
possible + difficult + unusual + successfully executed.
That combination is entirely compatible with manual flight.
Pages 34–36 — Reconstruction is treated as established fact
The study reconstructs UA175 and AA11 trajectories from NTSB headings, radar and film analysis.
But the reconstruction itself is a model.
The researchers are not measuring the original control inputs.
They are reconstructing a path from available observations.
That means the experiment actually contains two models:
Model 1: reconstruction of the historical flight path.
Model 2: simulation of a pilot attempting to reproduce Model 1.
Any uncertainty in Model 1 propagates into Model 2.
Yet the study’s conclusions frequently speak as though the reconstructed path were an exact record of the original aircraft’s control trajectory.
That is an important epistemological problem.
Page 37 — “Perfect perpendicularity” is treated as stronger evidence than demonstrated
The study argues that AA11 achieved an almost perfectly perpendicular impact trajectory and that this is strongly suggestive of automated control.
There are several problems.
First, the reported trajectory itself has uncertainty.
Second, an impact angle close to perpendicular does not inherently identify the control mechanism.
A manually flown aircraft can produce a straight final trajectory.
The relevant question is not:
“Could automation produce this?”
Obviously it could, assuming an appropriate system.
The relevant question is:
“What is the likelihood of obtaining this trajectory manually versus automatically, given all known information?”
The study does not actually calculate those competing probabilities.
It instead asserts that the symmetry is “strongly suggestive.”
That is evidentially weaker than the language used elsewhere in the report.
Pages 38–40 — The “symmetry” argument contains an important assumption
The report argues that because both aircraft eventually approached the towers at approximately perpendicular angles, the probability of this occurring under manual control was extremely low.
But no actual probability calculation is presented.
This is therefore a qualitative probability assertion.
To claim that the probability is “extremely low,” the researchers would need a model defining:
the distribution of possible approach angles under manual control
pilot behaviour
target geometry
available visual cues
navigation information
aircraft performance
and the probability of correcting toward a target
None of this is quantified.
Therefore:
“extremely low probability”
is asserted rather than demonstrated.
Page 40 — The terminal-guidance hypothesis is introduced without establishing the system
Page 40 is where the argument shifts decisively from observation to speculation.
The authors propose that a terminal guidance system may have “kicked in” during the final seconds.
But the study does not establish:
the existence of such a system
its architecture
its availability in 2001
whether it could control a 767/757
whether it could identify the target
whether it could generate the required bank/pitch inputs
whether it could operate at the reported speeds
or whether such a system was installed on the aircraft
Thus the proposed mechanism is hypothetical.
It cannot serve as the demonstrated explanation for the simulator results.
Pages 41–43 — The “manual versus automated” table is not actually a test
The study produces a table categorising findings as either consistent with manual or automated control.
This creates the appearance of a formal hypothesis test.
But the categories are qualitative.
There is no:
likelihood ratio
confidence interval
probability
statistical test
Bayesian posterior
error estimate
or quantitative comparison between competing hypotheses
The table therefore represents the authors’ interpretation rather than a statistical test.
Pages 43–45 — The alternative hypothesis is unfairly constructed
The study presents the manual explanation as requiring multiple unlikely propositions.
But notice how the alternative is constructed.
It assumes:
the pilots deliberately chose convoluted routes
the routes were mistakes
they placed themselves in highly difficult positions
they then successfully recovered
and they independently achieved similar impact geometry
The problem is that this is not the only manual-control hypothesis.
There are other possibilities.
For example:
The pilot deliberately flew a particular route for reasons unrelated to minimising the geometric distance to the target and then successfully controlled the aircraft during the final approach.
The study does not adequately model the full range of possible manual-control behaviours.
It constructs a relatively unattractive version of the manual hypothesis and compares it with a highly purposeful automated hypothesis.
That introduces hypothesis asymmetry.
Pages 45–46 — “Planning would have made it simpler” is not demonstrated
The study argues that if the routes had been planned and practised, the pilots would have chosen easier trajectories.
Again, this is a behavioural assumption.
Planning does not necessarily produce the mathematically simplest route.
A route can be selected for many reasons:
navigation
target orientation
concealment
timing
geographic constraints
visual acquisition
or simply pilot preference
The study does not establish that the alleged pilots were optimising route simplicity.
Pages 46–47 — The conclusion exceeds the evidence
The conclusion says the pattern of evidence fits better with automation than manual control.
“Fits better” is a legitimate scientific formulation if the competing models have actually been quantitatively compared.
Here they have not.
Instead, the conclusion rests upon a chain:
737 simulator failures, reported manoeuvres were difficult
the manoeuvres were counter-intuitive
pilots would probably have chosen easier routes
similar final geometries are suspicious
and therefore automated guidance provides a better explanation.
Each, introduces an assumption.
The study does not independently validate every link.
Pages 47–48 — A particularly revealing contradiction
The final section acknowledges that further research is needed into:
the missing FDRs
AA77’s FDR
FDR interpretation
UA93
earlier flight paths
AA77’s radar tracking
ACARS
possible guidance systems
and other evidence
This is extremely important.
The study therefore effectively says:
The evidence necessary to resolve several of the central questions is still unavailable or requires further investigation.
Yet the earlier sections present the automated-control conclusion with considerably greater confidence.
That is an internal tension.
If further investigation is required to determine whether the FDR data can distinguish manual from automated control, then the study should be cautious about claiming that its own simulator results establish automated control.
Page 49 — Source selection
The reference list includes a mixture of:
NTSB
NIST
9/11 Commission material
academic publications
documentary material
advocacy organisations
YouTube material
and sources advancing competing 9/11 interpretations
The problem is not that controversial sources are cited. The problem is that the study does not consistently distinguish between:
Primary evidence and interpretations of primary evidence.
For a study whose central purpose is to determine aircraft performance, the hierarchy should strongly favour:
original radar data
original FDR data
aircraft performance data
validated simulator models
independent aerodynamic analysis
peer-reviewed aviation research
Interpretive material should occupy a secondary position.
Page 51 — The appendix creates another reproducibility problem
The study says that videos of the simulator runs are available through hyperlinks.
That is useful, but video observation alone does not provide the underlying experimental dataset.
A reproducible experiment should ideally provide:
exact starting coordinates
altitude
airspeed
aircraft weight
centre of gravity
wind
temperature;
atmospheric model
engine settings
control inputs
bank angles
pitch
vertical acceleration
heading
simulator model/version
software version
sampling rate
and precise success/failure criteria
Without those data, independent researchers cannot fully reproduce the experiment.
The five most serious methodological flaws
After reviewing the entire report, I would rank the problems as follows.
1. Aircraft-model validity
This is the most fundamental.
The study investigates 767/757 manoeuvres using a 737 simulator.
The authors themselves acknowledge that the 737 may have been less manoeuvrable than the aircraft actually under investigation.
That means the independent experimental apparatus may systematically bias the result.
2. The experiment does not test the complete manoeuvres
The South Tower experiment explicitly omits the last-second pull-out.
Therefore the experiment cannot legitimately claim to reproduce the complete manoeuvre.
3. First-attempt failure is confounded by simulator unfamiliarity
The pilots themselves repeatedly identify unfamiliarity with the 737 simulator as a source of difficulty.
That directly undermines the use of first-attempt failures as evidence about real-world first-attempt performance in a different aircraft.
4. The study does not establish the probability of manual success
The study reports completion rates but does not establish a statistically valid probability that the original pilots could or could not have performed the manoeuvres.
The 28% Pentagon figure, for example, comes from repeated runs involving the same pilots under varying conditions.
Those are not 25 independent trials.
5. Failure of manual reproduction does not prove automation
This is the fundamental logical error.
The alternatives include:
manual control
manual control plus navigation assistance
manual control by a more capable pilot
prior knowledge or practice
differences between aircraft
errors in the reconstructed trajectory
or simply successful manual execution of a difficult manoeuvre
The study does not experimentally eliminate these alternatives. Nor does it demonstrate an actual automated system reproducing the historical flight paths.
The fundamental evidential chain is therefore incomplete
The study essentially needs to establish this:
A. The reconstructed flight path is correct.
B. The 737 simulator accurately represents the relevant 767/757 characteristics.
C. The experimental instructions faithfully reproduce the original circumstances.
D. The pilot sample represents the relevant population.
E. First-attempt simulator performance accurately predicts real-world first-attempt performance.
F. The observed failure rate establishes that manual execution was highly improbable.
G. No alternative manual explanation adequately accounts for the trajectory.
H. A specific automated system can reproduce the trajectory.
I. That system existed and was available on the aircraft.
J. Therefore automation was the most probable cause.
The study does not establish this entire chain.
At several points it effectively jumps from A–F to J.
The strongest criticism of the study
The most defensible criticism is therefore not:
“The study proves the official account.”
Nor should the criticism simply be:
“The study is wrong because it used a 737.”
That would be too simplistic.
The much stronger methodological argument is:
The study’s experimental apparatus, experimental conditions and inferential framework are insufficiently validated to support the strength of the conclusions drawn from them.
The 737/767/757 mismatch is especially important because it directly affects the quantity being measured: manoeuvrability and pilot control difficulty.
The researchers acknowledge this limitation themselves.
They also acknowledge uncertainty concerning the starting offset, timing, graphics and instructions, and concede that different simulators and instructions could produce higher completion rates.
Once those limitations are recognised, the simulator results become exploratory evidence, not decisive evidence, and that distinction is critical.
What the study actually demonstrates
A cautious interpretation would be:
The selected pilots generally found the reconstructed indirect manoeuvres difficult to reproduce under the particular experimental conditions used.
What the study claims
It moves considerably further:
The reported manoeuvres were unlikely to have been manually flown and are better explained by automated guidance.
Those propositions are not equivalent.
The first is supported by the experiment.
The second requires additional evidence that the experiment itself does not provide.
That is where I believe the study’s fundamental methodological weakness lies.
Thanks for reading and caring!
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.