Skip to content
Daniel Johnston Aerospace engineering — guidance, navigation and control

Optimisation and control

MasterGA

A genetic algorithm that tunes its own settings, grown out of a flight-control capstone.

The capstone's GA settings were tuned once by an outer search, then hard-coded. MasterGA turns that step into a tool: it tunes the settings for each problem, writes them down as a preset, and reports what each preset costs across seeds rather than on one lucky run.

A two-variable search space with three wide decoy wells and one small true well. A population of 24 points settles in the lower-right decoy; when it stalls, a scout scatters samples over the least-visited ground and moves the search to the upper-left decoy, and a second scout moves it to the true well in the upper right.
2-D decoy map — escape preset, seed 8
Fig. 1 A modular genetic-algorithm system in MATLAB: plug-in plants and fitness profiles, a preset ladder from a 100-evaluation smoke test to a calibrated multi-hour run, and a meta-GA that tunes the GA's own settings and measures what they cost.
Tighter spread across seeds
26 ×

Meta-tuned preset against the best hand-written one: IQR 3.4e-5 against 8.9e-4, 5 seeds, equal 5 000-evaluation budget

Cheapest preset on target every time
1035 evaluations

The balanced tuned preset, on target on 5 of 5 seeds of the short-period pitch problem

Escaped the decoys to the true well
28/50 seeds

The escape preset on a 2-D decoy map, against 10 / 50 for standard

Passing
399/400 tests

Fast suite, 2026-10-06. The one skip is the no-Simulink error path

Measured against limits

Evaluations spent, tightest preset 1802 evals
0 5500 evals

Every preset had the same 5 000-evaluation cap; this one stopped at its own generation limit on every seed, 36 % of what thorough spent.

Cessna trim residual, median 2.61 × 10⁻³
0 16 × 10⁻³

The accepted criterion is 1e-2. No preset reaches the tighter 1e-3 at the median; the gap is a badly scaled valley, not the tuning.

In sixty seconds

The whole system in one silent minute: the problem, a search getting out of a trap, the preset ladder, the meta-GA and the test suite. Every curve, point and number in it is drawn from the same seeded runs this page cites.

Fig. 2 MasterGA in sixty seconds. Silent; the footage is the real Studio and terminal, filmed headless, with sped-up stretches labelled.

The question

The capstone tuned PID gains and trim schedules for a Cessna's longitudinal dynamics with a genetic algorithm. The GA's own settings — population, generations, selection pressure, mutation — came from an outer tuning harness that searched over them once. The result, a population of 172 run for 28 generations, was then hard-coded into the final scripts.

That leaves the real question unanswered: which settings suit which problem, and what does each one cost? MasterGA answers it with measurements. Every preset is a point on a quality-against-cost curve, and every claim about a preset is a median over seeds with its spread beside it.

A pitch-attitude step response rising from 0 to a 5 degree command in about a quarter of a second with a small overshoot, above the elevator deflection, which saturates at minus and then plus 25 degrees before settling to zero.
Fig. 3 What is being tuned: the best pitch response at each improving generation of one tuned run, a 5° step with the elevator limited to ±25°.

One engine, plug-in problems

A problem is a plant — anything that turns a candidate into a response — plus a fitness profile that scores the response. The engine never knows whether it is tuning a pitch controller, trimming an aircraft or hunting a synthetic trap.

  • Every run writes a self-contained folder: configuration, history, the best response at up to eight improving generations, and figures.
  • A bounded digest, about 3.4 KB for a 100-generation run, states the outcome first and lists findings with stable rule ids — "reached the target after 656 of 7 208 evaluations", "diversity fell to 2.6 %".
  • Runs checkpoint and resume, and a calibrated machine profile prints a wall-clock estimate before any long run starts.
Best score against evaluations for five seeds of two presets. The five amber tuned-preset curves all finish at the same low score; the five grey standard curves stop early at scattered, higher scores.
Fig. 4 Five seeds each on the short-period pitch problem. The tuned preset ends at 0.0589 on every seed; standard ends anywhere from 0.0591 to 0.0807.

The Studio

The engine also has a window: the MasterGA Studio, a MATLAB app built in code rather than App Designer. Pick a problem and a preset, press Run, and watch the score, where the search is looking, its diversity and the best response redraw every generation; afterwards, replay the run in two or three dimensions and compare presets on the same problem.

Fig. 5 The Studio on the pitch problem, tuned preset, seed 1. Filmed headless: every frame is the real window, exported as it ran.

Tuning the tuner

The meta-GA is a GA whose candidates are GA settings. Each candidate is scored by running the inner GA over several seeds, so a setting that wins once by luck does not survive. Two targets come out of it for each problem: a balanced preset that is cheap and always on target, and a reliable one whose worst seed is still good.

On the short-period pitch problem, at an equal 5 000-evaluation budget, the reliable preset has a seed spread 26× tighter than the best hand-written preset, thorough, while spending 36 % of its evaluations. The outer level runs on 8 parallel workers, which cut one tuning stage from about 11 minutes to about 2.5.

A log-log chart of final score against evaluations used for five presets. quick sits at 100 evaluations with a poor, widely spread score; standard near 1 000 with a wide spread; the two tuned presets in amber sit at the lowest score with almost no spread; thorough reaches the same score at about 5 000 evaluations.
Fig. 6 The preset ladder as a quality-against-cost curve. Dots are medians of 5 seeds; bars span every seed. Amber marks the presets the meta-GA produced.

Traps, and getting out of them

A GA that converges fast converges into the nearest trap. To see that happen, a decoy problem hides one narrow true well among wider, shallower decoys. When a run stalls with budget left, a stall scout searches the least-visited part of the space and, if it finds something better, moves the population there.

On a fixed 2-D decoy map, the escape preset ends in the true well on 28 of 50 seeds, against 10 for standard. It is not a reliable show: a scout that finds nothing better ends the run, which happens on 22 of 50 seeds.

A heat map of a two-variable search space with four circular wells. A white path runs from the start in the lower-right decoy, across to the upper-left decoy at the first scout, then to the small true well in the upper right at the second scout.
Fig. 7 Seed 8: settled in one decoy, scouted into a better one at generation 22, and into the true well at generation 37.

No preset wins both

A second proving ground asks the GA to find two to four tones in a noisy signal, where a weak tone beside a strong one is the trap. Run against the decoy problem at three difficulties, ten seeds a cell, the presets split.

The capstone's tuned set escapes the decoys — 9 of 10 at medium difficulty — but cannot separate close tones. thorough is the reverse, finding the hard resonance case 6 times in 10. Nothing yet finds the hard decoy, which makes it the next target for the meta-GA.

Two grids of found-out-of-ten counts, one per problem, with presets across and difficulty down. The decoy grid is brightest under the capstone preset; the resonance grid is brightest under thorough; the hard decoy row is zero throughout.
Fig. 8 Runs that found the true answer, out of 10 seeds, with the median evaluations spent.

Against simpler methods

Is a genetic algorithm the right tool at all? Four methods, the same evaluation budget, 20 seeds each: MasterGA's best preset, its standard preset, random search, and Nelder–Mead — MATLAB's fminsearch — restarted from a random point each time it converges, until the budget runs out.

On the case study's own small problems, the honest answer is no. On the three-gain pitch controller, restarted Nelder–Mead reached the target on 19 of 20 seeds against 18 for the tuned GA, and sooner. On the 2-D decoy map, 3 000 random samples cover the square so densely that random search found the true well 18 times in 20; the escape preset managed 9.

Left, final scores of 20 seeds for four methods on the pitch problem: the tuned GA and Nelder–Mead cluster at the lowest score, the standard GA is spread out, random search sits highest; 18, 13, 19 and 4 of 20 reach the target. Right, a map of the 2-D decoy problem with where each run ended, and counts of 9, 4, 13 and 18 of 20 in the true well.
Fig. 9 Equal budgets, 20 seeds a method: 2 000 evaluations on the pitch problem, 3 000 on the decoy map.

Where a GA earns its keep

The same four methods on four harder problems, chosen and fixed before any of them was run: Rastrigin, Schwefel and Ackley in 10 variables, and a 5-variable decoy map, 20 000 evaluations each.

On the rugged 10-variable landscapes the GA was the only method to get close. On Rastrigin it reached the target on 11 of 20 seeds; nothing else reached it once. No method reached the Schwefel or Ackley targets, but the GA's median final scores, 574 and 0.115, compare with at best 1530 and 16.7 for the baselines.

The 5-variable decoy map went the other way. Restarted Nelder–Mead found the true well 12 times in 20 and the GA never did: its escape preset ends the run when a scout finds nothing better, after a median of 1753 of its 20 000 evaluations. A continue option for the scout now exists (below), but it has not been re-run on this map.

Four panels of best score against evaluations on log axes. On Rastrigin, Schwefel and Ackley in 10 variables the GA's line falls far below Nelder–Mead and random search, which stay near their starting scores; on Rastrigin the GA crosses the target. On the 5-variable decoy map the lines flatten together until Nelder–Mead drops to the true well at the end.
Fig. 10 Median best score over 20 seeds, band 25–75 %. Counts: seeds that reached the target, GA best preset first, then GA standard, Nelder–Mead and random search.

One problem, four answers

"Best" depends on what you ask for. The same three-gain pitch controller was tuned four times, changing only the fitness: settle fastest, never overshoot, move the elevator least, and the balanced tracking score used everywhere else. Each goal got its own answer: 0.22 s to settle, 0 % overshoot, and 0.96 rad of elevator travel against about 2 for the others. The tracking optimum turned out to be the no-overshoot design again.

Asking for two things at once gives a trade-off rather than a winner. A two-objective run returned 100 designs, none beaten on both settling time and elevator travel. One surprise: a plain grid search found a gentler corner of that trade-off (score 0.279) than any of the five GA runs aimed at it (0.981), a sliver about 1 % wide where the derivative gain is zero.

Left, four pitch step responses and their elevator traces for four fitness goals: the fastest settles in 0.22 s with a small overshoot, two others coincide with no overshoot, and the gentlest rises slowly with far less elevator movement. Right, a log-log trade-off of settling time against elevator travel, with the four designs marked on the front.
Fig. 11 shortPeriodPID, `thorough` preset, best of 5 seeds per goal. Right: seed 1's front (line) and the non-dominated union of 5 seeds (dots).

Damping you can see

Two new problems where the answer is visible. In the first, a motor turns an arm through a springy shaft, and the arm is what gets scored. With fixed gains it overshoots 70 % and rings for nearly two seconds. The GA's best plain PID settles in 0.45 s. Given an input shaper as well — the command sent in two steps, timed so the second cancels the ringing the first starts — it settles in 0.17 s, 2.6 times sooner. It did not rediscover the textbook shaper: it paired a mild shaper with much stiffer gains, which without the shaper overshoot 14 %.

Arm angle after a 1 radian command for five designs. Fixed gains overshoot to 1.7 radians and ring for two seconds; the textbook shaper on the same gains cuts that to about 15 %; the GA's PID settles smoothly in under half a second; the GA's PID plus shaper reaches the target in under 0.2 s without overshoot. In the video, two arms turn by the simulated angle.
Fig. 12 flexibleServo, 1 rad step, torque limited to ±5 N·m; GA designs are the best of 5 seeds. The video runs at a third of real time and says so.

Comfort against grip

The second is a car suspension: one wheel, a spring and a damper, driven over a 5 cm kerb. A soft suspension is comfortable but lets the tyre lose grip; a firm one keeps the tyre planted and shakes the passengers. There is no single best answer, so the GA returns the whole front: 100 designs, spanning 2.9 times in comfort and 2.6 times in grip. The comfort end sits on the softest spring the design box allows.

Left, body height over time after a 5 cm kerb for three suspensions: the softest overshoots to 8.4 cm and oscillates slowly, the firmest peaks at 7 cm and settles fastest. Right, the comfort-against-grip front with those three designs marked. In the video, three schematic cars ride over the kerb.
Fig. 13 quarterCar, `pareto` preset, seed 1. The video runs at 0.4× real time; heights are to scale on screen.

One controller for every flight condition

A controller tuned at one flight condition is optimal there and nowhere else. The pitch problem was extended to 25 conditions — 100 to 190 ft/s, sea level to 10 000 ft, light and heavy, centre of gravity fore and aft — and tuned against the worst of them, then judged on 200 conditions it never saw. Tuned at one condition, the best design overshoots up to 16 % and fails the 10 % spec at the slow, high corner. Tuned for the envelope, the worst overshoot is 0.4 % and all 20 seeds meet the target on the held-out set, against none. The price is paid at the nominal condition, where the score rises from 0.0589 to 0.1065.

Two panels of 25 pitch step responses each. Tuned at one condition, the responses fan out and four, at low speed and high altitude, overshoot the 10 % line. Tuned for the envelope, all 25 rise into a tight bundle with almost no overshoot.
Fig. 14 shortPeriodEnvelope: textbook short-period derivatives scaled from the nominal problem (representative, not identified). GA tuned preset, median of 20 seeds.

The objective did the work, not the GA

Is that a win for the genetic algorithm? The experiment was fixed in advance to find out: restarted Nelder–Mead got the same envelope objective and the same budget. It found the same controller. By the pre-registered test it even edged the GA, by 0.007 %, with 20 of 20 seeds on target for both. The gain comes from asking the right question — tune for the whole range — and any competent optimiser can then answer it on three gains. Where the GA should earn its keep is a bigger version: gains scheduled on airspeed, six to nine variables, the size at which it already pulled ahead above.

Worst held-out score for 20 seeds of six arms on a log scale. The three arms tuned at the nominal condition sit near 650 with 0 of 20 on target; the three tuned for the envelope sit at 0.112, on the target line, with 20 of 20. Nelder–Mead and the tuned GA are indistinguishable.
Fig. 15 20 seeds an arm, 2 000 evaluations each; the model, box, arms and decision rule were written down before the runs. Below each arm, its score at the nominal condition.

Can the GA beat the obvious alternative?

A stricter test, set before any run. Nelder–Mead solves the problem first; the level it reaches and the evaluations it needs become the line to beat, and a GA that needs more is a fail. That rule became the fitness for the meta-GA, which tuned the GA's settings to beat it. Only seeds the tuner never saw are quoted.

On Rastrigin in 10 variables the tuned GA reaches Nelder–Mead's final quality in 2.6 % of the evaluations Nelder–Mead needs, on every held-out seed, and 11 of 20 seeds go on to the real target. On the three-gain pitch controller it cannot: it needs 1.28 times Nelder–Mead's evaluations, with a confidence interval from 0.43 to 8.4. Adding a Nelder–Mead polish to the end of the GA did not beat it either. The GA is faster in wall-clock time on the pitch problem, 0.6 s against 9.4 s, but only because it evaluates candidates in batches, which is not better search.

A side effect: letting the scout keep going when it finds nothing better (a new continue option) took the 2-D decoy escape from 28 to 50 of 50 seeds, at the cost of far more evaluations.

Left and centre, the fraction of 20 held-out seeds that have reached Nelder–Mead's level against evaluations spent. On the pitch problem Nelder–Mead's line rises first and the tuned GA catches up only after Nelder–Mead's median; on 10-D Rastrigin every GA line reaches 20 of 20 within about 600 evaluations while Nelder–Mead reaches 8 of 20 near 20 000. Right, the meta-GA's 288 candidates against a fail line at 1 times Nelder–Mead's cost.
Fig. 16 Held-out seeds 101–120, equal budgets (2 000 and 20 000 evaluations); a seed that never reaches the level is charged twice the budget. Ratios with a 2 000-resample bootstrap interval.

What it has not solved

  • Cessna trim to 1e-3. Every seed of every preset ends in the right basin, within 0.5° and 3.3 hp of the exact trim, but no preset reaches a 1e-3 residual at the median. More generations, the stall scout and new operators did not close it; the criterion is 1e-2.
  • The hard decoy. Zero of ten for every preset.
  • Beating Nelder–Mead on small smooth problems. Even tuned specifically to beat it, the GA needs more evaluations on the three-gain pitch controller.
  • Small problems. On two or three variables, restarted Nelder–Mead or plain random search does as well as the GA or better (see above), including over a whole flight envelope. The GA is for rugged, higher-dimensional problems.
  • Trim through Simulink. The capstone's Simulink trim problem is flat for a GA, because trim() solves it from any starting elevator. What to optimise instead is open.
  • A near-bang-bang controller. Both fitness profiles choose gains that saturate the elevator about 2 % of the time. It was accepted as the design point, and the response figure shows it.