The Air Force
2011–2024 · 38.34°N 23.57°E
Four years at the Academy and thirteen in uniform. People assume an air force
engineer comes out knowing how to maintain an aircraft. I did not. I never
worked as a mechanic here.
What Tanagra taught me was how to run a large organisation under pressure: how
to deal with people, how to get a workshop of thirty to the right answer when
you are not the one holding the spanner, and how to put your name on work you
did not do with your own hands. That turned out to be the harder skill — and
the one everything since has rested on.
The work
I came out of the Air Force Academy in 2011 with a degree in aeronautical
engineering, and spent the first year at the Directorate of Airspace
Applications: CAD modelling in CATIA V5, airflow and loads in ANSYS, on the
structural integrity programme for the F-16.
Then Tanagra. Flight line maintenance officer on the Mirage 2000-5, supervising
line maintenance. Commander of the base maintenance workshops — airframe,
electrical, hydraulics, bonded stores, paint and corrosion protection. Stage A
cross-servicing supervisor for F-16 and F/RF-4, preparing them to fly.
From 2014, quality control: monitoring Mirage inspections, airworthiness
records, ground support equipment, tool calibration. From 2016, quality
assurance — accident and incident investigation, auditing, technical compliance
on the Mirage fleet and its support equipment, defective maintenance reports,
and issuing engineers' maintenance authorisations.
For four of those years I also taught at the Non-Commissioning Officers Academy:
advanced mathematics, science of materials, fluid dynamics, to first and second
grade cadet officers.
What it taught me
Command a workshop and you stop being the person who knows the answer and start
being the person who has to get thirty other people to the answer. Sign an
authorisation and you are staking your name on somebody else's competence.
I did the master's at the National Technical University of Athens while I was
serving, which the Air Force was good enough to note. The subject was
mathematical modelling. The lesson was time management.
Athens
2011–2024 · 37.98°N 23.73°E
Lived here for the long middle of the story.
Thirteen years between an Air Force base and whichever civil organisation I was
working for that year — usually more than one at a time. The career pins nearby
are what happened while this was home.
Boarding Lab
47.98°N 7.81°E
Every published comparison of aircraft boarding methods starts its clock when
the first passenger walks through the aircraft door. But somebody has to sort
180 people into that clever sequence first, and that work happens to the
passengers, standing at a gate.
So I built a simulator that starts the clock earlier — at the announcement to
prepare for boarding — and followed the same 180 people all the way to their
seats. Strict Steffen, the mathematically optimal order, fills the cabin faster
than anything else and still finishes the whole journey last, because building
its perfect queue takes thirteen minutes.
That conclusion rests almost entirely on one number I could not find published:
how long a gate agent needs to call a single passenger. So the number is a
slider. Move it below about two and a half seconds and the answer flips.
Run the boarding experiment
Quality & compliance
2019– · 37.94°N 23.70°E
Where the work moved from the aircraft to the system that keeps the aircraft
right.
Manuals, audits, transitions, and the unglamorous business of making a
requirement survive all the way to the person doing the task.
Air Power, 2019–2024
I was there from the start of the organisation. Quality Manager in the CAMO
first, then Safety and Compliance Monitoring Manager, and eventually covering
both the CAMO and operations. The company was renamed Hoper toward the end; I
stayed a few months after that and then left for TUS.
Starting with an organisation is a different job from joining one. There is no
procedure to follow because you are the one writing it, and no precedent to
point at when somebody asks why.
I led the transition to Part-CAMO and wrote the CAME and the operations manuals
that went with it. I also built the training behind them — theory and practical,
for airworthiness staff and for pilots: maintenance programmes, technical
records, pre-flight inspections, aircraft familiarisation. From 2021 I was
airworthiness review staff as well, which is the role that signs an aircraft's
airworthiness review certificate — the annual check that the paperwork and the
aeroplane still agree.
A transition is not a document exercise. It means changing procedures, manuals,
training and everyday decisions at the same time, and the only test that matters
is whether the person doing the task ends up doing it differently.
Marathon, 2022–2023
An airline mid-transition. I covered the CAMO safety and compliance role
temporarily while it moved to Part-CAMO, then worked as Deputy Safety and
Compliance Monitoring Manager.
The Part-CAMO transition and the CAME. The Management System Manual.
Management-of-change workflows built in Centrik. Registration and upkeep of the
external accreditations — Argus, WYVERN. And I led the WebManuals transition
end to end.
What mattered was a management system people could operate, not a shelf of new
documents. Workflows, manuals and records had to point at the same action. When
they did not, that was the work.
TUS Airways, 2024–
Engineering Safety and Compliance Manager. I started on the CAMO side. In
November 2025 we obtained the Part-145 approval; in 2026 I led the Part-IS
introduction — EASA's requirement for cybersecurity and digital assets — into
the engineering system.
The job is the same thread as the earlier chapters, only further down the line:
find the requirement, walk it into practice, and leave evidence that it stayed
there.
Corfu
1989 · 39.62°N 19.92°E
Born here, on the first of October 1989.
An island teaches you early that everything arrives by sea or by air. I think
that is where the rest of it started.
Unfortunately, aviation on Corfu is mostly transit — wheels up, wheels down,
and the island is already behind you. I left the other way: I stayed long enough
to learn that the interesting part is not the landing, but what you build after
you arrive.
EASA Caveman
50.95°N 6.96°E
BIG RULE. FEW WORD.
Part-M, Part-145 and Part-CAMO Section A enter cave. Legal throat-clearing gets
clubbed. Competent-authority Section B stays outside. Rule references, numbers,
units and dangerous little words such as not, only and except stay.
Rules kept best we can.
Part-M
Part-145
Part-CAMO
Hands on the fleet
2015–2024 · 37.98°N 23.73°E
In 2015 I walked into a civil hangar not knowing how to screw in a bolt.
Thirteen years of Air Force engineering had taught me to manage aircraft and the
people around them. It had not taught me to touch one. This is where that
changed, alongside the uniform rather than after it.
Interisland, 2015–2019
Mechanic on a mixed light fleet — Robinson R22, R44 and R66 helicopters, and
Maule MXT-7-180A, Robin ATL, Diamond HK36 and Van's RV-6A aeroplanes.
Within months of starting we had done a twelve-year inspection and assembled an
R66. That is the whole story of this chapter in one sentence: the distance
between not knowing which end of a spanner to hold and putting a helicopter
together is shorter than anybody admits, and it is walked by doing it beside
somebody who already can.
Everything I know about airworthiness and maintenance in practice — the records,
the plan, the directive, the reason a task exists — I learned here, on aircraft
I had my hands inside.
Jetway, 2018–2024
Robinson R44 maintenance, and certifying staff.
This is where the Part-66 licence stopped being a certificate. Certifying is a
small physical act with a long shadow: you inspect the work, you read the
record, and then you sign, and the signature is the part that matters.
Flyence, 2022–2024
Certifying staff first, then CAO Manager. The CAE was rewritten rather than
amended — the point at which a manual has so many changes stacked on it that
rebuilding it around the way the work actually runs is both faster and more
honest.
Four helicopters came from Nepal under that approval: documentation, assembly,
two twelve-year inspections, and autopilot installations. R22, R44 and R66
capability went onto the organisation with them. It was the best project of the
stretch.
The work on that fleet was not only managed. Handling, fuelling, procurement,
maintenance, airworthiness monitoring — every job. That is rare once the title
says manager. A manager who has not recently done the work forgets what the
work costs; I try not to become that person.
Alongside the helicopters: certifying staff on the Cessna F152 and 172,
maintenance programmes tightened, procedures standardised, and custom
modifications and installations under CS-STAN.
The Hangar
38.15°N 21.85°E
Four things I have built by describing what I want to a machine and arguing
until it was right. Open a bay for the short version — what it is for, and what
it taught me.
avioverse.io
The main one, and the reason the others exist. Coming soon — with real test
cases and evaluations behind it, not a landing page and a promise.
Everywhere I have worked I have tried to improve the system I found there.
Usually that meant automating whatever was actually in front of me, and for most
of my career what was in front of me was Excel: the workbook that tracks the
fleet, the workbook that tracks the training, the workbook somebody rebuilt from
scratch because the last one finally broke.
Then one job put a complete SaaS in my hands, and it impressed me. Not because
it did everything — it did not — but because it showed me the shape the answer
could take.
So I called Thanos, the best developer I know, and told him we were going to
build the next aviation SaaS: everything the existing ones were missing, plus
everything twenty-five years had taught me about what the work actually needs.
We started. And while we were building it I carried on doing the thing I have
always done — one tool for this, another tool for that, half the tasks
impossible to track across either — until the obvious finally landed.
The market does not need another piece of software for businesses. Companies
are well served. What nobody has built is something for the individual: the
auditor, the mechanic, the pilot, the post holder, the admin who holds the whole
thing together and owns none of the systems.
I built the core of avioverse.io in one night, and that is who it is for.
The agent: a colleague, not a coder
I have used the real thing. Claude Code, Codex, my own Hermes — run hard, on
work that mattered, for long enough to know exactly what they can do. They are
extraordinary.
They are also, in the hands of somebody who is not a programmer, extremely
dangerous. The whole skill is knowing when a confident answer is wrong, and that
is precisely the skill a non-programmer has not built yet. The machine does not
sound less certain when it is mistaken. I have the aviation judgement to catch
it in my own field and the scars to catch it in software; most people have one
of those at best.
So the agent inside avioverse.io is not a coding agent pointed at aviation. It
is built for the person doing the job — an expert colleague. One who has read
the regulation, knows where the evidence is meant to live, and will tell you
plainly when what you have does not support what you are about to sign. Bounded
so that a wrong answer costs you an argument, not an airworthiness decision.
The core feature will be free. The AI on top is a subscription priced like
software a person buys for themselves — not enterprise SaaS priced per seat by
somebody who has never done the job.
fleetnode.org
Built to answer one question before committing to it in the main application:
could a framework I was considering really carry what I needed?
So I gave it the hardest thing I had — nodes, and how you visualise a second
brain. Ideas connected to ideas, and a way to see the shape of them rather than
scroll a list. It is a proving flight, not a product: I wanted to know how the
airframe behaved before I trusted it with passengers.
flybycode.com
This. A personal site you fly over instead of tabbing through — the CV, the
articles and the guidance, arranged as a place rather than a page.
Built the way I build everything now: plain-language direction, review of what
comes back, and refusal of anything I do not understand. Still flying, still
being refined.
mxcourse.net
An asynchronous learning management system, currently open only to
organisations. I built and shipped it because I wanted to see how Cloudflare
actually behaves in production — not the brochure, the real thing.
Contact me if you like it.
Larnaca
2024 · 34.92°N 33.63°E
Home in Cyprus since 2024 — one house on the map for the whole island.
Engineering Safety and Compliance Manager at TUS Airways. Twenty-five years of
aviation — if you linearise the parallel work — arrives here: safety,
compliance, quality, airworthiness, training, operations, maintenance.
Part-147, Part-CAMO, Part-145, Air Operations.
Evenings are for side projects: describe what you want to a machine, argue
until it is right.
The Lighthouse
35.52°N 23.48°E
Thoughts — short pieces on aviation, safety, and building software by talking
to a machine.
Things worth saying live here when they do not need a whole carrier deck of
chapters.
Practical drift
The first thing I do in a company is read the documentation against the work.
Not for typos — for drift. People leave. Regulations change. A procedure that
was true three years ago is often still on the shelf, and the person doing the
task has quietly invented a different way.
Find the gap between what is written and what is done. Improve the documents
with focused amendments. The point is not a prettier manual; it is a manual
somebody can still follow on a Tuesday afternoon.
The broken screw
The day after my engagement, before the last screw torque on a job, I skipped
something I never skip: I did not open both the IPC and the AMM.
The result is a broken screw in my bag — and the biggest human-factors lesson I
own. Distraction does not announce itself as distraction. It looks like
confidence. It looks like you already know. It is the first and last time I
signed off without both books open.
Learn SMS from life
To learn a safety management system, think about your life — not aviation
examples. How you notice risk at home. How you decide what to fix first. How you
check that a change actually stuck.
SMS does not need a huge budget or a shelf of consultants. It needs commitment,
and enough love of the work that you would rather find the problem yourself
than have it find you.
Where I am now
2024– · 34.91°N 33.63°E
Cyprus, an airline, and a company of my own.
TUS Airways, 2024–
Engineering Safety and Compliance Manager. Engineering compliance and safety
across the CAMO, the Part-145 and the AOC.
I own the compliance-monitoring plan. I track (and issue) findings, corrective
actions and evidence. I identify and manage engineering safety risk, liaise
with the Cyprus DCA on engineering and continuing-airworthiness matters, and
support accountable and safety management on engineering topics.
The work is quiet when it is working: a requirement is found, an audit follows
it into practice, and an action closes with evidence. I keep the public account
at that level. Internal findings, people and occurrences stay private.
AvioVerse
My own company, set up in the United Kingdom: consulting, human resources and
training for aircraft operators.
Regulatory and management-system training, and CAR.66 and CAR.147 courses
delivered to a CAR.147 organisation. Where an operator needs advice, or needs
people rather than a document, AvioVerse is how I provide them.
Independent work
I have prepared organisations for the transition to Part-CAO and Part-CAMO,
reviewed and amended maintenance programmes, and written a safety manual and a
CAE for a declared training organisation. I am independent certifying staff for
piston engine aeroplanes and R22 and R44 helicopters, and a mechanic on the
R66. I deliver safety management training to maintenance organisations, CAMOs
and operators.
For Avilaw I built the continuing-airworthiness section, pulling the regulation,
the acceptable means, the guidance material and the EASA publications into one
readable place — including one of the very few guides written for the Part-CAMO
transition and Part-145 SMS implementation.
The Schoolhouse
37.20°N 18.40°E
I have never written software in my life. I have shipped it.
I am an aeronautical engineer. My hands know airworthiness and compliance, not
compilers — so I learned to direct software instead of typing it. I say what the
work needs, the machine proposes the code, and I refuse anything I do not
understand. The product is mine; the typing is not.
Everything that taught me is here: forty cards in seven decks, from what a
function actually is to how to keep an agent on a lead. Each one is short enough
to read on the walk out to the aircraft, and each one stands on its own — so
start with whichever deck sounds like your problem, not with the first card.
Teaching
2011– · 38.02°N 23.76°E
Years spent teaching the work — and finding out how much of it I actually
understood.
If you cannot explain why a limit matters, you do not understand it yet.
Teaching is where you find that out, in front of people who will ask.
Non-Commissioned Officers Academy
While I was still in uniform I taught at the Non-Commissioned Officers Academy:
advanced mathematics, science of materials, and fluid dynamics, to first and
second grade cadet officers.
Four years of standing in front of people who will ask the awkward question.
That habit is older than any civil title on this map.
Aviatec, 2020–2024
Quality Manager of a Part-147 training organisation. A training organisation has
two products: the course, and the evidence that the course is controlled.
I owned the MTOE and the audit programme, and wrote the procedures that kept us
running through the COVID restrictions. I taught as well as managed — EASA
regulations, safety management systems, human factors, maintenance programme
development and planning — and the non-EASA courses too: human factors, EWIS,
SMS, regulations.
I built the learning-management system (in Google Workspace) for developing
instructors and examiners, and evaluated the new ones, permanent and ad-hoc. We
expanded to satellite facilities in Doha and Morocco. I developed the Robinson
R66 type rating course, along with Group 3 Level 1 training.
Training quality is easy to make bureaucratic. I tried to keep it usable: clear
procedures, controlled questions, instructors who understood the rule, and a
system that still worked when the classroom moved online.
Jetstream, 2021–2024
I qualified as an R22/R44 type rating instructor and delivered one course in
2021. What stayed with me was not the certificate — it was students contacting
me long after the course ended, asking for guidance on modifications and
troubleshooting. That is the test that matters: whether people still trust you
when the classroom is over.
AvioVerse
Through AvioVerse I deliver regulatory and management-system training for
aircraft operators — and, where needed, the people who make the system run
rather than another document on the shelf. The same rule as every other deck on
this map: if you cannot explain it, you do not understand it yet.
Control Tower
40.25°N 24.05°E
About Dennis Kefalas.
Aeronautical engineer. Four years at the Hellenic Air Force Academy and thirteen
in uniform, then civil aviation from 2015 — the two ran alongside each other for
nearly a decade. Safety, compliance, quality, airworthiness, training,
operations and maintenance. EASA Part-66 licensed engineer, B1.2, B1.4 and C.
Now Engineering Safety and Compliance Manager at TUS Airways, in Cyprus.
In every organisation I have worked in I have tried to improve the system I
found there — usually by automating whatever was in front of me, which for most
of my career meant Excel. That habit is what eventually turned into building
software.
I do not write software by typing it. I direct machines in plain language,
review what they propose, and ship only what I understand. This site is one of
those products.
What I trained in
A degree in aeronautical engineering from the Air Force Academy, and a master's
in mathematical modelling from the National Technical University of Athens,
taken while I was serving.
Accident investigation at Embry-Riddle. Flight safety, quality assurance,
IOSA airline auditing with IATA, and the Inter-European Squadron Officer School.
The certificates matter less than the habit they leave: find the requirement,
follow it into practice, and check what actually happens rather than what is
written down.
Family
Home is Cyprus, with the people I love. That is the part that matters, and the
only part that needs a pin on this map. Names, ages and private detail stay off
it — not because they are an afterthought, but because they are not for
strangers. The people who matter already know where they sit in the story;
visitors get the outline, and that is enough.
The Schoolhouse library
Short lessons about building software by directing machines in plain language.
The airframe
Every application has the same four parts
Learn the shape once and every new tool becomes their version of a part you know.
An aircraft has a cabin, a bay where the systems live, records that follow it
for life, and defined joints between them. Software has the same gross anatomy,
and it barely varies.
The frontend is what a person sees and touches. It runs on their device.
The backend runs on a computer you control. The rules live there, and so do the
keys.
The API is the agreed joint between the two: which messages may pass, and in
what shape.
The database keeps the permanent records. Everything else can crash and be
restarted. This must not be lost.
Learn this once and new tools stop being alien. Whatever a framework calls
itself, it is somebody's version of one of those four parts. It also tells you
where an argument belongs: a colour change is a frontend question, a permission
is a backend question, and what must survive a restart is a database question.
Do this: Take an application you use daily and name its four parts out loud.
Common mistake: Treating the screen as the whole system, and forgetting that
rules and records have to live somewhere a user cannot reach.
The frontend is the cabin
Anything the passenger can reach, the passenger can tamper with.
The cabin is the part of the aircraft the passengers can get at. You design it
knowing that. Nothing important is left loose in there.
The frontend is the cabin. It runs on somebody else's phone or browser — a
machine you do not control and cannot inspect. Everything sent there can be
read, and everything on it can be altered by anyone patient enough. So no keys,
no prices, no permissions, no final decisions.
The hard problem a frontend solves is keeping the screen in step with the data.
When one status changes, every place showing it has to change too. Doing that by
hand is miserable, so frameworks do it for you: you declare that the screen is a
picture of the data, you change the data, and the framework redraws whatever
depended on it. That single idea is what React, Vue, Svelte and the rest are all
selling.
Do this: Name one thing on a screen that must never be decided there.
Common mistake: Hiding a button and calling that a permission check.
The backend is the systems bay
Rules and keys belong behind a panel the passenger cannot open.
Behind the cabin lining is the bay where the systems actually live. Passengers
do not go in there. It is not hidden to be mysterious; it is where things are
enforced.
The backend is that bay. It is a program sitting on a computer you control,
waiting for requests. When one arrives it checks who is asking, applies the
rules, reads or writes the records, and sends an answer back.
Three things live here and nowhere else. The rules of the work, because this is
the only place they cannot be bypassed. The keys and passwords, because this is
the only place they are not handed to a stranger. And the decisions that count,
because a decision made on a screen is a suggestion until the server agrees.
You do not need to read the code to ask which part owns a rule. If nobody can
answer clearly, the design is mixing responsibilities that should stay apart.
Do this: For one feature, write down which rule the server must enforce even
if the screen is bypassed entirely.
Common mistake: Letting the screen decide something and asking the server
only to store the result.
The API is the interface control document
Two parts built by different people fit because a document says exactly how.
When two manufacturers build parts that have to bolt together, neither one
guesses. A document specifies the joint: the exact dimensions, the fasteners,
the loads it carries. Build to the document and the parts mate on the line.
An API is that document, for software. It says which addresses exist, what you
may send to each, and what shape the answer comes back in. Send me this, and I
will reply with that.
This is why parts written by strangers, in different languages, on different
continents, fit together. The screen does not need to know how the server
stores anything. The server does not need to know what the screen looks like.
They only need to agree on the joint.
It is also why changing an API carelessly breaks things far away. Anyone built
to the old document is now built to a joint that no longer exists.
Do this: For one feature, write the joint in a sentence: what goes in, what
comes back.
Common mistake: Changing what an existing address returns, and being
surprised when something else stops working.
The database is the technical record
Screens get repainted. The record outlives all of them.
An aircraft's records outlive its paint, its interior, its avionics and usually
its owner. They are the thing you protect, because they are the only account of
what happened that survives everything else.
The database is that record. It is structured, permanent storage, and it is the
single source of truth. The screen only shows it. Closing a browser must not
erase anything, and renaming a label on a page must not quietly rewrite history.
Two families exist. Relational databases keep data in tables with strict
columns, linked by identifiers, and the database itself enforces the rules — no
work order without a valid aircraft. That is the sensible default, because most
real information is naturally relational. The other family trades those enforced
rules for flexibility, which means the rules now live in somebody's head.
SQL is the small, stable language for asking these questions. It is fifty years
old, it is not going anywhere, and it is genuinely worth learning.
Do this: List the real things one tool must remember, and treat that list as
the design.
Common mistake: Designing the pages first and bending the records to fit
them.
Two languages, and a pairing that already flies
You do not need a rating on every type. Two languages cover almost all of it.
Nobody flies every type. You get rated on what the job needs and you recognise
the rest by name.
Two languages cover nearly all of this work. JavaScript, usually written as
TypeScript, is the only language a browser runs — so anything on a screen is
that, whatever else you choose. It runs servers too. Python is the language of
automation, data and AI work. SQL asks the records questions. Everything else is
a rating you take when a project demands it, which for most people is never.
TypeScript is JavaScript with a checking step. It catches a whole class of
mistakes before the program runs — the difference between finding a defect at
inspection and finding it in flight.
Then do not assemble from the full catalogue. Fly a proven pairing: a
combination thousands of people already run, which is also the one AI assistants
know best. Pick one and stay on it six months. Depth compounds.
Do this: Choose one combination for your first project and refuse to change
it for six months.
Common mistake: A new set of tools every project, which makes you a
permanent beginner in each.
How a request travels
A request goes out, a reply comes back
Press Enter and a strict little conversation happens, always in the same order.
A radio call has a format. You transmit in that format, and what comes back
tells you how it went before it tells you anything else.
The web works the same way. You type an address and press Enter. The name is
looked up and turned into a numeric address, the way a phone book turns a name
into a number. An encrypted connection opens. Code on the server runs, usually
asking the database something. A reply travels back and the browser draws it.
Every request carries a method saying what you want: read, create, update or
delete. Every reply carries a status number saying what happened. Two hundred
means fine. Four hundred and four means there is no such thing. Four hundred and
one means you are not allowed. Five hundred means the server itself broke — that
one is yours to fix, and the others usually are not.
The common style is to make addresses the things and methods the actions.
Do this: Next time a page fails, find the status number before theorising.
Common mistake: Treating every failure as one undifferentiated broken.
JSON is how systems speak to each other
One plain text format, readable by every language, is why strangers’ parts fit.
A shipping label works because everyone reads it the same way. Labelled lines,
plain text, no interpretation required: part number here, quantity there,
destination underneath.
JSON is the shipping label of software. It is text, laid out as labels and
values, and it is what nearly every system sends when it has structured
information to pass on. A report might travel as a number labelled id, a word
labelled status, and a list labelled items. You can read it yourself without
tools, which is the point.
Its value is that every language reads and writes it. A part written in Python
and a part written in TypeScript exchange work through JSON without knowing
anything about each other. That is what makes it possible to choose the right
language for each part instead of one language for everything.
When somebody says an API returns JSON, they are saying the answers arrive as
labelled text you can look at.
Do this: Find one API's example answer online and read it. You will
understand more of it than you expect.
Common mistake: Assuming the format is technical and therefore unreadable.
Never trust the cabin
Anyone can send your server a hand-made request that never touched your screen.
This is the most common serious mistake beginners make, and it is worth its own
card.
Your screen is a convenience for honest users. It is not a gate. Anyone can send
requests to your server directly, without ever loading your page — the tools to
do it are free and take a minute to learn. So every check that matters has to be
made again on the server: is this person allowed, does this record belong to
them, is this price the real price, is this quantity sane.
The screen may still do the check, because telling somebody early is kinder than
telling them late. But the screen's check is a courtesy. The server's check is
the control.
The same logic you already use elsewhere: a placard on the outside of a panel is
information, not a lock.
Do this: For each rule in your feature, ask where it is enforced. If the
answer is only the screen, it is not enforced.
Common mistake: Believing that because your page cannot send a bad value,
nobody can.
The technical records
One table per real thing
Get the record shape right and everything is easier. Get it wrong and data piles on top of the mistake.
The data model is the one decision that gets ten times more expensive to change
later, because by then real data has accumulated on top of it.
Four rules cover most of it. One table for each real kind of thing — aircraft,
reports, people — and one row for each individual one. Every row gets its own
identifier. Rows point at other rows by storing that identifier; where two
things relate many to many, a small table in the middle holds one row per
pairing. And do not store anything you can work out — totals, counts, ages. A
stored copy drifts from the source, and then you have two versions of the truth.
The tell that you have it wrong: you are adding columns called item1, item2,
item3. That is a second table trying to escape.
Sketch the tables before any code. Describe the work, ask for a proposed
structure, then attack it with the awkward cases you know and the machine does
not.
Do this: Sketch one tool's tables and test them against your two most
awkward real cases.
Common mistake: Designing attractive pages first and forcing the records to
fit them.
A migration is the modification record
You never modify an aircraft without recording the modification.
No modification goes onto an aircraft without paperwork that says what changed,
when, and against what approval. The point is not bureaucracy. It is that the
next person needs to know what state the aircraft is actually in.
As an application grows, the shape of its records has to change: a new column, a
new table, a renamed field. A migration is a small numbered script that makes
exactly that change and can be replayed, in order, on any copy of the database.
Run them all against an empty database and you get today's structure. That is
the modification record, and it is how a change made on your machine reaches the
live one intact.
The rule that follows: never hand-edit a live database because it seems faster.
The change then exists in one place only, recorded nowhere, and the next
deployment has no idea it happened. Structural changes need a repeatable method
and a backup you have actually restored.
Do this: Before any structural change, ask how it will be applied to the
live records and by what script.
Common mistake: One quick manual fix in production that nobody can find six
months later.
A backup you have never restored is a rumour
Backups are the fire bottle. You must have one, and you must have fired one.
A fire bottle you have never checked is not a fire bottle. It is a red cylinder
you are hoping about.
Backups are the same. Almost everyone sets one up. Very few ever restore one,
which means very few know whether the file being written every night is
complete, readable, or in fact empty. The moment you find out is the worst
possible moment to find out.
So restore one. Not in theory — take last night's backup, load it into a spare
copy of the database, open it and look. Check the newest records are there, and
that the awkward parts came across intact. Write down how long the whole thing
took, because on the day it matters somebody will ask.
Do it again occasionally, because backups quietly break when the structure
changes.
Do this: Restore one backup into a scratch copy this month and look at the
newest record in it.
Common mistake: Treating the existence of a backup file as evidence that
restoring it will work.
Why a page goes slow
Almost every slow page is one of three things, and none of them is a slow computer.
Slowness almost never comes from the language or the processor. It comes from
waiting, and from repeating work.
The first cause is a missing index. An index is the index at the back of a
manual: without one the database reads every row to find a match; with one it
goes straight to the page. The difference is routinely a thousandfold. Index any
column you regularly filter, sort or join by.
The second is fetching a list, then fetching each item's related record one at a
time. Fifty rows becomes fifty-one trips to draw one page. It is invisible with
five rows of test data and brutal with five thousand, which is why it survives to
production. When you ask for a list view, say plainly: fetch the related data in
one query.
The third is repeating expensive work — an outside lookup on every page view.
Store the answer with an expiry and reuse it.
Then measure. Speeding up something that was never the bottleneck is the classic
wasted week.
Do this: When a page feels slow, check for a missing index and repeated
per-row queries before anything else.
Common mistake: Assuming a bigger machine will fix it.
Ground equipment
Version control is an audit trail that writes itself
It records every change, which is what makes bold experiments cheap.
Revision control on a manual is a discipline: every amendment recorded, every
superseded page traceable. Git does that for code, and it does it automatically.
Git saves snapshots of the project. Each one carries who made it, when, and a
message saying why — a full audit trail, without anyone maintaining it. You can
branch, meaning work on a copy without touching the good version, merge it back
when it is right, and return to any earlier point exactly as it was. GitHub is
the hosted home for these projects, and adds a review step before changes are
accepted.
The habit that matters is small and frequent snapshots, with messages that say
why rather than what. And commit before every experiment. When an assistant
mangles the code at midnight, one command throws away the mess and puts you back
where you were. I have never once regretted a snapshot taken before trying
something.
That is the whole trick: version control is what makes fearlessness cheap.
Do this: Commit before you let anything make a change you are unsure about.
Common mistake: One enormous snapshot at the end of the day, labelled
“stuff”.
Packages are the parts catalogue
Modern software is mostly assembled from published parts.
Almost nothing is made from raw material any more. You order the part, you
record which part number and which revision went on, and the record is what
makes the next build identical.
A package manager is the catalogue and the stores counter together. It fetches
published code your project depends on, and writes down the exact version of
every single one in a lockfile, so that every machine builds the same aircraft.
The JavaScript world uses npm; Python uses pip, or the newer and faster uv. The
folder of fetched parts is large, regenerable, and never stored with your own
code.
Version numbers read as three numbers: the last one is fixes, the middle one
additions, the first one breaking changes. Treat a first-number change as an
inspection task with a test flight, not a routine top-up.
And weigh each new dependency. Every one is somebody else's code you now depend
on, update with, and inherit the faults of.
Do this: Before adding a part, ask what it does that twenty lines of your own
would not.
Common mistake: Accepting a major version jump as if it were a patch.
Secrets live in a locked cabinet, never in the code
A key left in the code is found by automated scanners within minutes.
Keys are not left in the aircraft. They live in a cabinet, and taking one out is
a recorded act. Nobody thinks this is excessive.
Software keys — database passwords, service credentials, API keys — follow the
same rule and it is not optional. They never go in the code. On your own machine
they live in a separate settings file that version control is explicitly told to
ignore. In production they live in the hosting service's settings, handed to the
program when it starts.
The reason for the strictness: public code is scanned continuously by people
looking for exactly this. A key committed by accident is typically found and
used within minutes, and it is usually your card that pays for the usage.
If a key does escape, the fix is not to delete the line. It is to revoke that
key and issue a new one. The old one is in the history for ever.
Do this: Check that your settings file is on the ignore list before the first
commit, not after.
Common mistake: Pasting a key in temporarily and meaning to move it later.
Deployment is getting it airborne
Three postures, and the choice is mostly about who does the maintenance.
Running on your own machine is a ground run. Deployment is putting the thing on
a computer that is always on and reachable, which is a different set of
concerns.
Three postures. A managed platform takes your code, builds it, serves it and
handles the certificates — least work, least control, and free at small scale.
Your own rented machine is cheap and will run anything, including jobs that
never stop, but you are now the mechanic and the operations department:
updates, firewall, certificates, monitoring, restarts. A managed backend service
gives you database, logins and file storage ready-made, which is an enormous
head start on their terms.
A container packages a program with its exact surroundings so it runs the same
everywhere. It is the standard shipping container of software, and it is worth
the extra learning only once deployment actually demands it.
Then automate the release: on every push, run the tests and deploy only if they
pass. That is continuous inspection built into the process.
Do this: Choose your posture by who you want doing the maintenance at 2am.
Common mistake: Renting a machine to look after, for something a managed
platform would have carried for nothing.
Airworthiness
Every input is unverified cargo
Nothing arriving from outside is trusted until it has been checked.
Cargo does not go in the hold because somebody says it is fine. It is checked,
weighed and documented first. Anything arriving from outside your system
deserves the same suspicion.
Check every incoming value on the server: the right type, a sensible range,
present when it must be, allowed to be there at all. Libraries exist that do
this for you from a description of the expected shape, so it costs a few lines
rather than an afternoon.
The classic failure has a name. If you build a database question by gluing the
user's text into it, a user can write their own question — and read or delete
everything. The fix is free: never assemble a query by pasting text together.
Pass the values separately, so data stays data and can never become an
instruction. Any decent database library does this by default.
The same shape of attack now reaches AI systems, where hostile instructions hide
inside content the model is asked to read.
Do this: For one form, write down every value it accepts and what would
happen if each arrived as nonsense.
Common mistake: Checking the shape of data on the screen only.
Who you are is not what you may do
Checking the licence is not the same as checking the privilege.
A valid licence in somebody's pocket does not tell you they may sign for this
particular work, on this particular type. Two questions, two checks, and mixing
them up is how things get signed that should not have been.
Software has exactly the same pair. Authentication answers who are you — the
login. Authorisation answers what may you do — the roles and the ownership. The
classic hole is checking the first and forgetting the second: the user is
definitely logged in, and is definitely reading somebody else's record because
nobody asked whether it was theirs.
The check belongs on every request that touches a record, not once at the door.
And do not build the login itself. Password storage, resets, sessions, lockouts
— this is solved, dangerous-to-improvise territory with a long history of
expensive mistakes. Use an established service and spend your attention on the
part only you can build.
Do this: For one record type, write the sentence: may this user do this to
this specific record?
Common mistake: Assuming that being logged in is a permission.
Tests are the alarm, not the proof
Their job is not today. Their job is the day a distant change breaks something.
People assume tests exist to prove the thing works. That is the small part. A
test is a permanently installed warning system: it rings when a change somewhere
else breaks something over here, months after everyone forgot the two were
connected.
Tests are code that checks your code, and they rerun in seconds, for ever. Some
check one small piece in isolation. Some check that parts work together. Some
drive the whole thing the way a user would.
This matters more with an AI assistant than without one. The assistant edits
boldly and confidently across files; the tests are what catch the damage it did
not know it was doing. Without them you are reviewing every line by eye for
ever.
Two cheaper cousins do related work. Automatic style and defect checkers flag
likely mistakes without running anything. A type-checking language catches a
whole class of errors at build time — at inspection rather than in flight.
Do this: Ask for a test on the one path that must never break, before asking
for more features.
Common mistake: Testing everything shallowly and the critical path not at
all.
Logs are the flight recorder
When it misbehaves at two in the morning, logs are the only evidence you have.
After an event, you do not rely on anyone's recollection. You read the recorder.
It was running before anybody knew there would be something to investigate,
which is the entire point.
Logs are that. A program writes down what it did, in order, as it does it: this
request arrived, this user, this record, this decision, this failure. When
something goes wrong on a machine you cannot see, at an hour when nobody was
watching, the log is the only account that exists.
So log what happened and which identifiers were involved — the record number,
the user, the outcome. A line saying “error” with nothing attached is noise.
And never log secrets, or personal detail you do not need. Logs get copied,
shipped to other services and read by people who were not thinking about
privacy. Whatever is in them has effectively been published inside your
organisation.
Do this: For one important action, write down the line you would want to
read the next morning.
Common mistake: Adding logging after the incident that needed it.
Reading the code
Every value has a type, including nothing
A surprising share of all crashes is code expecting something that turned out to be nothing.
Syntax is costume. The ideas underneath are the same in every language you will
ever meet, which is why learning them once makes every tutorial, error message
and explanation readable.
All data is values, and every value has a kind. Numbers. Text, called a string,
and always in quotes. True or false, called a boolean. And nothing at all, which
different languages call null, undefined or None. A variable is simply a name
bound to a value so the code can refer to it later.
That last kind causes more crashes than anything else. Code assumes something is
there — a record, a field, an answer — and it is nothing. The message you will
meet says it cannot read a property of undefined. You now know exactly what that
means: something you expected to exist did not, and the line named in the error
is where the assumption was made, not necessarily where the mistake was.
Do this: When something crashes, ask first which value turned out to be
nothing.
Common mistake: Reading an error as gibberish when it is naming the exact
file, line and assumption that failed.
Two shapes describe nearly all data
A list of things, and a thing with labelled parts. That is most of it.
There are two containers, and between them they describe almost every piece of
information you will ever handle.
A list is an ordered many: three registrations, forty reports, the items on one
work card. Use it when you have several of something.
A labelled set — languages call it a dictionary, an object or a map — is one
thing with several named parts: registration, hours, status, owner. Use it when
one thing has properties.
Now nest them. A fleet is a list of labelled sets. A work order is a labelled
set containing a list of task sets. A month of reports is a list of those. That
is genuinely the shape of most real data, and JSON is exactly these two shapes
written down.
The moment you can look at real information — a form, a spreadsheet, a report —
and see its list-and-labels shape, you can describe it to a machine precisely.
That skill is worth more than any amount of syntax.
Do this: Take one document from your work and write it as lists and labels.
Common mistake: Describing data in prose when the shape would have said it
exactly.
A function turns inputs into a result
A part you can bench-test beats a part you can only test installed.
A function takes inputs, does work, and returns a result. That is the whole
idea, and code is mostly functions calling other functions.
Two habits make code dramatically easier to reason about, and you can ask for
both without writing a line yourself.
Return early: handle the “no” cases at the top and get out, instead of nesting
conditions five deep. Deeply nested code is where mistakes hide.
Prefer functions that touch nothing else. Give one the same inputs and it always
gives the same output, changing nothing anywhere in the system. That kind of
function can be bench-tested on its own, in seconds, in isolation. A function
that quietly modifies things elsewhere can only be tested installed on the
aircraft, with everything else running, and when it misbehaves you are searching
the whole system for the cause.
Do this: Ask for the awkward part as a separate function that only takes
inputs and returns a result.
Common mistake: Accepting one enormous function that does the fetching, the
deciding and the displaying at once.
State is what changes while it runs
Most bugs are state bugs: something changed when you did not expect it to.
State is any data that changes while the program is running: what is currently
selected, who is logged in, which items are loaded, whether the save is in
progress.
Almost all difficult bugs are state bugs. Something changed when nothing should
have. Something did not change when it should have. Two parts held different
ideas of the same fact at the same moment. Nothing is broken in isolation —
every piece works alone — and the fault only appears in a particular order of
events. That is why “it only happens sometimes” is so common, and why writing
down the exact sequence is half the fix.
Keeping a screen in step with changing data is the hardest version of this
problem, which is precisely why frontend frameworks exist.
Scope is the smaller cousin: where a name is visible. A value created inside a
function exists only there. When code cannot see something you are certain
exists, scope is usually the answer.
Do this: When a bug is intermittent, write down the exact order of actions
that triggers it.
Common mistake: Hunting for a broken part when the parts are fine and the
sequence is not.
Standard equipment in standard places
Thirty unfamiliar files, and nearly all of them are standard fit.
Open an unfamiliar aircraft and you do not know that airframe — but you know
where the manuals are, where the placards are, and what a fuse panel looks like.
A project folder is the same. Almost everything in it is standard equipment in a
standard location.
Start with the readme: what this is and how to run it. Always. Then the manifest
— the file listing the project's parts and the commands it knows how to run,
which is effectively its list of capabilities. Beside it sits the lockfile,
which is machine-maintained and never hand-edited. An example settings file
names which secrets the program expects without containing any. The ignore list
says what must never be uploaded. The code itself lives in one folder, usually
called src or app, and the tests in another.
Frameworks put files in fixed places on purpose: the location means something.
When a folder structure seems rigid, the rigidity is usually load-bearing.
Do this: In any new project, read the readme and the list of commands before
opening a single code file.
Common mistake: Browsing files alphabetically and mistaking volume for
understanding.
Trace one action from end to end
One completed trace teaches more than an hour of browsing files.
Pick one button. Follow it the whole way: the piece of screen that holds it, the
request it sends, the server route that receives it, the question that hits the
database, the answer coming back, the screen changing. One trace, all the way
through, and the shape of the whole project appears.
Doing it by reading takes a while. Asking is faster: give an assistant the
project and say walk me through this — where is the screen, where are the rules,
where are the records — then trace what happens when a user does this one thing.
It is the fastest way into unfamiliar code, and you can check the answer by
following it yourself.
A few words you will meet while tracing. A monolith is one program containing
everything, and it is the right default. Layers means code separated by job:
routes, then rules, then data. Middleware runs on every request before anything
else — the guard at the single entrance. A background job is work done outside
the request, so nobody waits for it.
Do this: Trace one action end to end before you change anything.
Common mistake: Reading everything and understanding nothing.
You are the pilot in command
Ask for the plan before the build
Fixing a plan costs one sentence. Fixing the wrong code costs an evening.
An AI assistant is a very fast, very well-read, occasionally overconfident first
officer. The arrangement that makes that safe is the one cockpits already use:
the pilot in command sets the destination, delegates the flying, and cross-
checks continuously. Never sleeps in the seat.
So for anything beyond the trivial, say: propose a plan first, do not write code
yet. Then read the plan. A wrong assumption in a plan is one sentence to
correct. The same assumption discovered after four files were written costs an
evening.
Then one task per request. Add the login page, not build the application. Small
changes can actually be reviewed. Thousand-line changes get waved through, and
waving things through is exactly how bad code gets in — not through malice, but
through volume.
The pattern is a loop you run every time: plan, small step, read what changed,
save it, check it actually works.
Do this: Say “propose a plan first, do not write code yet” and read what
comes back.
Common mistake: Asking for the whole application in one go, and getting
something too large to check.
Read what changed, not what you were told
You do not need to write code to smell trouble in it.
Every tool can show you exactly what changed: lines added, lines removed, file
by file. That view is called a diff, and reading it is a skill you already have
in another form — you have reviewed paperwork for things that should not be
there.
You are not checking whether the code is elegant. You are looking for four
smells, and all four are visible without knowing the language.
Code deleted that you did not ask about. A new outside dependency appearing from
nowhere. A key, password or address typed directly into a file. And changes in
files that have nothing to do with the task you asked for.
Any of those gets the same question: why did you change this? A good assistant
answers it plainly. An unconvincing answer is itself information.
The discipline that makes this survivable is small changes. Nobody can read a
thousand-line diff honestly, and pretending otherwise is how the smells get
through.
Do this: Read the diff before saying yes, and ask about anything surprising.
Common mistake: Approving on the strength of a confident summary of what was
done.
Done is a claim, not evidence
Run it, click it, look at it. Then it is done.
An assistant reporting that a feature is complete is making a claim about work
it cannot fully observe. Sometimes the claim is true. Often it is true of the
part it was thinking about and not of the whole.
So treat “done” the way you would treat any unwitnessed sign-off: it is the
beginning of the check, not the end of the job. Run the thing. Click the button.
Try the awkward case. Look at the record afterwards and confirm it says what it
should.
Better still, invert the order. Ask for the test first, then the code that
passes it. A test is a contract that cannot be quietly talked around, and it
turns “it works” from an opinion into something that either passes or does not.
This is not distrust, it is verification, and it is the same reason a signature
comes after an inspection rather than before it.
Do this: Before accepting any change, run the thing yourself and try one
case nobody mentioned.
Common mistake: Accepting a summary of the work as evidence of the work.
Context is the fuel
It knows only what is in front of it. Nothing carries over by itself.
Nothing carries over between flights. You check the fuel before every one,
because what was in the tanks last time is not there now. An assistant starts
each session the same way: knowing nothing about your project except what it is
handed, and everything you told it yesterday gone.
The fix is a standing orders file in the project: how to run it, the conventions
you keep, what must never be touched, and why. It loads automatically into every
session. Whenever you correct the same thing twice, the correction belongs in
that file rather than in your repetition.
Keep it lean. A bloated file buries the rules that matter, and models over-obey
shouted language: write “CRITICAL — you MUST” and you will get work contorted to
satisfy the shouting. State each rule plainly and delete anything the assistant
already gets right.
Within a session, name files exactly and paste errors verbatim. And start fresh
when the topic changes: long conversations accumulate stale material and get
worse.
Do this: Move your next repeated correction into the project's standing
orders file.
Common mistake: Writing a long instruction file and wondering why the
important rules get missed.
Where the first officer fails
Four failure modes, all four predictable, all four cheap to counter.
Assistants fail in patterns. Knowing the patterns is most of the defence.
Invented equipment. It produces a function or a setting that sounds exactly
right for a library and does not exist, because it inferred it from the
library's general feel. If code fails with “no such method”, suspect this first
and tell it to check current documentation.
Sweeping changes. Large multi-file rewrites are where errors compound silently,
because nobody can review them honestly. Break them into staged steps, saved
between each.
Judgement calls. What to build, what good enough means, which trade-off suits
your situation. It will pick one confidently and the pick is yours to make.
Force the alternatives out: give me three approaches with trade-offs before
choosing.
Agreeableness. It leans towards telling you your idea works. So do not ask
whether it works. Ask what is wrong with this approach, and argue against it.
Do this: On your next design question, ask for three approaches with
trade-offs before anything is written.
Common mistake: Reading confident agreement as a second opinion.
Writing the work card
Treat a prompt like a work card
If a colleague reading only your instruction would be confused, so is the model.
A work card does not say “look at the brakes”. It says who does it, to what, to
which limits, and what is recorded at the end. That precision is what lets the
work happen without the author standing there.
Write instructions to a model the same way. State the task, who the output is
for, how long it should be, and what shape it should take. The test is exactly
the one you would apply to a work card: if somebody competent read only this,
would they know what to do?
Two more levers cost nothing. Give it a role — “you are a careful technical
editor reviewing a procedure” — which measurably shifts rigour and vocabulary.
And keep the material separate from the instruction: wrap anything you paste in
simple tags, so pasted content cannot be mistaken for orders. When the material
is long, put it first and your question at the very bottom. Models answer
measurably better when they read the evidence before learning what to look for.
Do this: Rewrite one instruction to state task, audience, length and format
explicitly.
Common mistake: A one-line request, then blaming the answer for guessing
wrong.
Say what to do, not what not to do
“Write in flowing paragraphs” beats “do not use bullet points”, every time.
This is the smallest change with the largest effect, and it takes one rewrite to
adopt.
A positive instruction gives a target to hit. A prohibition fences off one
option and leaves everything else to guesswork. Tell a model not to be verbose
and you have ruled out one style from a hundred. Tell it to reply in one
sentence and there is nothing left to interpret.
So phrase your rules as the thing you want. Not “do not invent references” but
“cite only passages that appear in the text I gave you”. Not “avoid technical
language” but “explain it to somebody who has never opened a terminal”. Not “do
not change other files” but “change only the file I named”.
The same instinct works on people, which is why briefings say what to do. It
just matters more here, because a model has no shared context to fall back on
when your instruction only rules one thing out.
Do this: Take your last prohibition and rewrite it as the behaviour you
actually want.
Common mistake: A list of things not to do, and no description of the target.
Show it two worked examples
Two examples beat three paragraphs of description.
Describing a format is slow and lossy. Showing it is fast and exact. Two or
three worked examples — this input, that output — are the single strongest lever
you have for getting consistent results.
It works for the same reason a filled-in sample form works better than a page of
notes on how to fill in the form. The pattern is visible, including all the
small conventions you would never have thought to write down: how dates look,
what to do with a missing value, how long an entry should be.
Choose the examples carefully, because they are your specification now. Include
one straightforward case and one awkward one — the model will copy both the
shape and the judgement. If your examples all show tidy inputs, you have taught
it nothing about mess.
Where the output has to be machine-readable, examples and a defined shape work
together: examples teach the judgement, the shape enforces the structure.
Do this: Add one clean example and one awkward one to an instruction that
keeps producing the wrong shape.
Common mistake: Explaining the format at length instead of showing it once.
Let it think first, and let it say no
Thinking after the answer is worthless. Order matters.
Ask for the reasoning before the answer, in its own section. Accuracy climbs on
anything that is not trivial, and you get something you can check rather than a
conclusion you must take on faith.
The order is not a detail. A model writes one piece at a time and cannot go back
and revise what it already said, so reasoning written after an answer is
justification, not thinking. It will happily defend a conclusion it reached
carelessly.
The second half of this card is just as valuable. Say explicitly that “I do not
know, the text does not say” is an acceptable answer. Invented answers drop
sharply once refusing is permitted, because you have removed the pressure to
produce something.
For questions about a document, go further: require it to quote the relevant
passage first and answer only from the quote. Then you can check the answer
against its own evidence in seconds, without reading the document yourself.
Do this: Add “reason it through first, then answer” and “say if the text does
not cover it” to your next question.
Common mistake: Asking for an answer with an explanation, which produces a
guess with a defence attached.
An eval set is your test equipment
A prompt that worked three times is untested equipment.
You would not sign off a measurement taken with an instrument nobody had
calibrated. A prompt that seemed to work on three tries is exactly that
instrument.
Build a small set of representative inputs together with the answers you would
accept. Twenty is enough to start. Then every time you change the instruction,
run the whole set and grade the results — by exact comparison where the answer
is a fact, by reading them where it is judgement, or by a second model checking
against your criteria.
Two things happen. You stop guessing whether a change helped, because you can
see it. And you find out that the change which fixed one case broke two others,
which is otherwise invisible until a user finds it.
This is the difference between a demo and a product. Anyone can produce an
impressive demonstration in an afternoon; the distance from works once to works
reliably on real, messy input is covered by exactly this.
Do this: Write down ten real inputs and the answers you would accept, before
tuning any instruction further.
Common mistake: Tuning by impression, one example at a time.
Operating the engine
It predicts the next piece of text
One fact explains every strength it has and every way it fails.
A large language model does one thing: it predicts the next small piece of text.
Trained on an enormous amount of writing, it learned the shape of language,
code and argument well enough that this prediction produces genuinely useful
work.
The unit it predicts is called a token — roughly three quarters of a word. It
produces one, appends it, and predicts the next from everything so far.
Everything else follows from that. Answers arrive a word at a time because they
are being produced a piece at a time. It cannot go back and revise what it has
already said, only continue — which is why reasoning before an answer helps and
justifying afterwards does not. And because it is producing what is plausible,
it is superb at pattern work and unreliable on exact recall.
Understanding this one mechanism is the difference between being impressed by
the machine and being able to operate it.
Do this: Before any task, decide whether you are asking for a pattern or a
fact.
Common mistake: Treating fluency as evidence of knowledge.
The context window is working memory
It is a clipboard, not a filing cabinet.
Everything a model can see in one conversation — your instructions, the
documents you pasted, everything either of you has said — sits in one working
memory called the context window. It is measured in tokens, and it is all the
model knows about your session.
Three consequences matter.
Nothing persists. Between conversations, the clipboard is wiped. If it needs to
know something today that you told it yesterday, you send it again. That is not
forgetfulness in the human sense; there is simply nowhere for it to go.
It fills up. When it does, the oldest material effectively falls off the edge —
including, quite often, the instruction you gave at the start.
And quality sags in the middle of a very full window. Relating everything to
everything gets harder as the pile grows, so the material in the middle of an
enormous context gets the least attention.
Do this: Start a fresh conversation when the topic changes, and re-state
what matters.
Common mistake: A single enormous conversation that has quietly forgotten
its own instructions.
There is no database inside
It has no store of facts to look things up in. That is why it invents.
This is the mental model that explains every failure you will meet.
There is no filing cabinet inside a model. Its knowledge is not stored anywhere
retrievable — it is spread across billions of learned numbers as statistical
tendencies. Nothing can be looked up, only reconstructed.
So it cannot tell you where it learned something, because there is no where. Its
grasp of patterns is superb and its recall of specifics — part numbers,
citations, figures, function names — is unreliable, exactly where accuracy
matters most. Invention is intrinsic, not a bug somebody will fix. And its
knowledge stops at the date its training stopped.
The working rule that follows: pattern work, trust it — drafting, classifying,
explaining, restructuring. Specifics, make it look them up and quote what it
found.
And no, training it on your own documents is almost never the answer. That is
expensive, brittle and does not reliably add knowledge. Better instructions,
examples, and handing it the right pages solve nearly every case.
Do this: Split your task into pattern work and exact facts, and treat the two
differently.
Common mistake: Believing a confident, well-written specific because it is
confident and well-written.
A model call is just a request
You pay by the word, in and out, on every single turn.
Using a model from code is not exotic. It is the ordinary request and reply of
any other service: send the conversation so far, get the next message back. That
loop is the entire foundation of every AI product you have used, which is worth
knowing before anyone sells you a platform.
Cost follows the text. You pay for what you send and what comes back, every
call — and a long conversation re-sends everything each turn, so cost grows with
the square of a chat rather than in a line. Long documents in the instruction
are the usual bill.
Two dials worth knowing. Model size trades intelligence for price: route the
easy, high-volume work to the small cheap model and keep the expensive one for
the hard part. And the same input can produce a different output each time,
because probability is the mechanism — so for extraction and classification, ask
for the most predictable setting, and never assume two runs will match.
Do this: Before automating anything, work out the cost of one call and
multiply by the real volume.
Common mistake: Sending an entire document on every turn of a long
conversation.
Structured output makes it a component
This is the step from chatbot to software part.
A chat window is a demonstration. A part your program can rely on is a product.
The bridge between the two is telling the model the exact shape of the answer
and requiring it to match.
Instead of asking for a summary and getting prose, you hand it a defined shape —
these four labelled fields, this one from a fixed list of values — and the reply
is constrained to that. Messy human text goes in; clean, labelled data comes
out. A rambling report becomes a category, a severity, a component and an action,
which your program can then store, count and route.
That is the moment a model stops being something a person talks to and becomes a
function inside a system, sitting between two ordinary pieces of code that know
nothing about it.
It also makes checking possible. A shape can be validated automatically before
anything downstream trusts it, and anything that fails validation can be retried
or set aside for a person.
Do this: For one task, write down the exact fields you want back before you
write the instruction.
Common mistake: Parsing prose with clever text matching instead of asking for
a defined shape.
Meaning as coordinates
An embedding is a position in meaning-space
Turn meaning into coordinates and “find things like this” becomes arithmetic.
Plot two numbers and you get a point on a chart. Plot several hundred and you
get a point in a space nobody can picture, but the arithmetic works identically.
An embedding is a piece of text turned into a few hundred numbers — a position
in that space — arranged so that texts meaning similar things land near each
other. Brake wear limits sits close to minimum pad thickness before replacement,
with not one word in common. Nobody named the axes; training discovered them.
Two consequences matter. Similarity becomes a distance you can measure, so
finding things like this one is arithmetic rather than guesswork. And it is
extraordinarily cheap: a whole document collection turns into coordinates on an
ordinary laptop in minutes.
That is why the right shape is almost always coordinates find the material and
the model reads it — never the reverse.
Do this: Think of one pile of documents where “find things like this one”
would save real time.
Common mistake: Asking a model to read a whole archive when arithmetic could
have found the ten relevant pages first.
Hand it the right pages
The model is not retrained. It is handed the relevant pages at question time.
The most useful pattern in this whole field, and the answer to “how do I make it
know my documents”.
The question arrives. Your system searches your own documents, takes the best few
passages, puts them in front of the model and asks it to answer from those,
citing what it used. The model has learned nothing new — it has been handed the
right pages, rather than sent on a course.
The consequence: answer quality is dominated by the search, not the model. When
answers come back bad, debug the search first. Almost everybody blames the model,
and almost everybody is wrong.
Two things decide the search. How documents are split — one position cannot
represent a two-hundred-page manual, so split into coherent sections. And
searching by words as well as by meaning: meaning blurs exact identifiers like
part numbers and error codes, and keyword search catches precisely those. Run
both, then re-score the top candidates with a second, more careful pass.
Do this: When answers are wrong, look at which passages were retrieved
before touching the instruction.
Common mistake: Blaming the model for what the search failed to find.
Coordinates organise the haystack
Search is only the first thing coordinates are good for.
Once a pile of documents has coordinates, several things become nearly free that
would otherwise be projects.
Grouping. Ten thousand reports fall into recurring themes nobody ever tagged,
because similar ones sit together. Recurring problems are where useful ideas
live.
Near-duplicates. The same issue reported twice in different words lands in almost
the same place — deduplication without matching a single phrase.
Labelling without training. To classify a new item, find its nearest already-
labelled neighbours and take their vote. Ten lines of code, no training, works
from a handful of examples.
The odd one out. An item far from every group is unusual by definition. “Show me
this month's entries least like anything we have seen” is a genuinely powerful
review question.
Model calls are slow, expensive, and handle one item at a time. Geometry is
instant and covers the whole pile. If an idea involves understanding a whole
pile of documents, it is a coordinates problem first.
Do this: Ask what your largest pile of text would look like grouped by
similarity.
Common mistake: Reaching for a model per item when the whole corpus could be
organised at once.
Running the crew
An agent is a loop with tools
A model that can only talk becomes a model that can do.
On its own, a model produces text. Give it tools and a loop, and it acts.
The mechanism is simpler than the word suggests. You describe the tools available
— search the documents, read this file, run this command. The model replies: call
that one, with these arguments. Your code performs the action and feeds the
result back. Round it goes until the job is done. The coding assistants people
use daily are exactly this.
There is now a standard plug for the wiring: expose a tool once, in an agreed
way, and any assistant can grip it, instead of a custom fitting for each one.
The caveats arrive with the capability. Errors compound across steps. Tokens
burn fast, because every round re-sends the growing conversation. And something
that can act needs limits on what it may do unsupervised.
Start with one plain request. Add the loop only when the task genuinely needs
several steps.
Do this: For one task, write the list of tools it would need. If the list is
one item, you do not need an agent.
Common mistake: Reaching for an agent when a single well-written request
would have finished the job.
Stand on the lowest rung that works
Every rung up multiplies cost, delay and ways to fail.
There is a ladder. One request and one answer. A fixed chain of requests, each
feeding the next. One agent with tools, deciding its own steps. A crew of agents
coordinated by another. Each rung costs more, takes longer, and adds ways to go
wrong.
The distinction worth learning: a workflow is a pipeline where your code decides
the steps and the model fills them in — predictable, debuggable, cheap. An agent
is a loop where the model decides what to do next. Workflows for known processes,
agents for open-ended work. Most agent products that survive in production are,
quietly, mostly workflow.
One piece of arithmetic explains most failures of ambitious systems. A step that
succeeds ninety-five times in a hundred, run twenty times in sequence, produces a
correct result about a third of the time. Reliability does not survive length.
The answers are fewer steps and a check between stages. I have yet to meet a job
where a crew beat a chain somebody had thought about properly.
Do this: Before building a crew, write down what one careful request cannot
do.
Common mistake: Building the impressive version first, and discovering the
arithmetic afterwards.
Scope the task, name the handback
A worker given “look into this” wanders. Given “return four fields” it delivers.
Think of each worker as a specialist hired for one task. It arrives with its
instructions, does the work, hands back a result, and its memory is gone.
Anything not in the handback is lost for ever. That single fact drives most of
the design.
So write the work card properly. Scope small enough to finish in one sitting —
one file, one chapter, one question. If a worker needs its own workers, the job
was cut wrongly.
Then name the handback exactly. Between workers, prose is where information goes
to die: a paragraph of summary loses precisely the details the next step needs.
Say instead: return these fields, with these names, one entry per finding. The
work card has always specified what gets recorded at the end; this is the same
discipline.
Anything that must outlive a worker goes into a file or a database that others
can read, never into a conversation, which evaporates.
Do this: Write the exact fields you want back before dispatching any task.
Common mistake: Asking a worker to “look into” something and receiving an
essay you cannot use.
Never let the finder verify its own finding
The inspector is not the person who did the work. The rule transfers exactly.
Agents treat other agents' output as fact. That is the quiet failure mode of
every multi-step system: one invented finding upstream becomes three paragraphs
of confident analysis downstream, and nothing in between ever questions it.
Garbage does not just propagate, it gets elaborated.
The counter is independent verification. Where a wrong finding is expensive —
reviews, audits, safety claims, anything somebody will act on — send each finding
to a separate sceptic whose only instruction is to refute it against the source
text. Only findings that survive count.
The word separate is doing the work. A model asked to check its own output tends
to agree with itself, in the same way and for the same reasons that people do.
And a panel of identical checkers shares identical blind spots, so vary how they
are asked.
This is the quality inspector, in software form, and it exists for the same
reason: the person who did the work is the worst-placed person to find what they
missed.
Do this: Add one independent checking step wherever a wrong answer would be
acted on.
Common mistake: Asking the same model, in the same conversation, whether it
is sure.
A human signs the release
The system can prepare the decision. A person releases it.
Decide in advance which actions require a person: anything irreversible, anything
that leaves your organisation, anything expensive. Automation prepares the
decision — gathers the evidence, drafts the assessment, lays it out — and a
person releases it. That division is not caution. It is the arrangement that
makes the automation usable at all.
Three limits belong with it.
Every loop gets a budget: a maximum number of steps, a maximum spend, and a
defined path for giving up and reporting. Something unable to finish will happily
burn money for ever. Nothing flies without a fuel gauge.
Anything reading untrusted text can be steered by instructions hidden inside it.
Cap what it may do without confirmation, especially anything with write access.
And record every input, output and action. When a long run produces a wrong
answer, “why did it do that?” is unanswerable unless it was all logged. Install
the recorder before the incident, not after.
Do this: List the actions in your system that must never happen without a
person, and make them impossible without one.
Common mistake: Discovering which actions needed approval by watching one
happen.
Fault isolation
Reproduce it before you repair it
A fault you can trigger on demand is half fixed.
Debugging is fault isolation, and every habit that works in a hangar works here
unchanged: reproduce, isolate, one change at a time, fix the cause and not the
symptom. You are not learning a new discipline. You are applying one you have to
a new system.
Start where you always start. Can you make it happen again, on demand? If yes,
you can test whether your fix worked, which is the whole reason it matters. If
it only happens sometimes, you have not found the real trigger yet — and “it is
intermittent” is a description of your knowledge, not of the fault.
Write the exact steps down. Not roughly: exactly. Which record, which user,
which order, which browser. Half of all intermittent faults become repeatable the
moment somebody writes the steps down honestly, because the writing exposes the
step everybody was doing without noticing. I have never regretted the ten minutes
that took.
Do this: Write the reproduction steps before you touch anything.
Common mistake: Starting to fix something you cannot yet make happen, and
never knowing whether you fixed it.
Read the error, bottom-up
It usually names the file, the line and the cause. Go and look.
When something impossible happens, the code raises an error. It travels back up
through whatever called it until something handles it, or nothing does and the
program stops, printing the trail of calls that led there.
That printout looks like noise and is not. It is a breadcrumb trail, and you read
it from the bottom: the deepest line names the actual failure, and the lines above
show how execution arrived there. It usually gives you the file and the line
number. Go and look at that line before theorising about anything.
Reading a stack trace calmly is half of debugging, and it takes no programming
ability at all — only the willingness to read something that looks unfriendly.
When you ask anybody for help, human or machine, give four things: the exact
error text copied verbatim, the relevant code, what you expected against what
happened, and what you already tried. Then say: diagnose the cause before
proposing a fix. Vague descriptions get guesses; exact ones get diagnoses.
Do this: Copy the error verbatim and open the file and line it names.
Common mistake: Reporting that it is broken, and receiving a guess.
Suspect the last change first
It worked yesterday. What is different?
The most efficient question in troubleshooting is not what is wrong. It is what
changed.
Version control answers it exactly: it will show you every line that differs
from the last known-good state. Not what you remember changing — what actually
changed, including the thing an assistant altered while you were looking
elsewhere.
For something that broke further back, there is a tool that binary-searches your
history: it checks out a point halfway back, you say whether the fault is present,
and it halves again. A dozen answers finds the exact change among thousands. It is
the same method you would use on a wiring run, applied to time instead of length.
This is also the practical argument for frequent small snapshots. When each one
contains one small change, “what changed” has a short answer. When yesterday is
one enormous snapshot, the tool can only tell you that something in it did it.
Do this: Before investigating, look at what changed since it last worked.
Common mistake: Theorising about the architecture when a one-line change
three hours ago did it.
Find the fault by halving
Same logic as splitting a wiring run: check the middle, and you have halved the problem.
You know this method. You do not trace a circuit end to end; you check the
midpoint, and whichever half is wrong, you halve that. Three or four checks
isolate almost anything.
In software the equivalent is printing the suspect value at the midpoint of its
journey. Correct there? The fault is downstream. Wrong there? Upstream. Repeat.
Three checks find most faults, and each one is a single line you delete
afterwards.
Four suspects come up again and again, and it is worth checking them before
inventing anything more interesting. The data is not what you think — log the
actual value, because it is empty, or missing, or text where you expected a
number far more often than anybody believes. The outside part does not behave as
you assumed — read its documentation for the two minutes it takes. Timing —
something used a result before the slow step had finished. And drift between
machines — a different version, a missing setting, a stale part.
Do this: Print the value at the midpoint before forming any theory.
Common mistake: Reading code for an hour instead of looking at one actual
value.
Fix the cause, then trap it
One change at a time, and then a test so it can never come back unnoticed.
Two habits finish the job properly, and both come straight from the hangar.
One change at a time, retesting between changes. Shotgunning five fixes at once
means that even success teaches you nothing: you do not know which one worked,
whether two of them cancel out, or what the other four broke. It feels slower and
it is not.
And fix the cause, not the symptom. Catching an error and hiding it makes the
message go away and leaves the fault installed. Ask what allowed the bad value to
exist in the first place.
Then trap it. Add the test that would have caught this fault. It costs a few
minutes now, and that failure mode is permanently caught from here on — including
by whoever, or whatever, edits this code in a year. This is how a system gets
sturdier with age instead of more fragile: every fault that ever occurred leaves
a permanent sentry behind it.
Do this: After every fix, add the test that would have caught it.
Common mistake: Five changes at once, and no idea afterwards which one
mattered.
Learning and shipping
Build first, theory second
You learn this the way you learn an aircraft: on the floor, with the manual open.
Nobody learned an aircraft by reading the manual cover to cover. They learned it
on the hangar floor with the manual open at the right page and a job in front of
them. Software is the same, and the temptation to study first is the main thing
that wastes people's first year.
So pick a real problem from your own work. Always your own — you can judge
whether the result is right, which is the part no course can give you.
And use the machine as an instructor, not only as a mechanic. The highest-value
request in the whole toolkit is: explain what you just wrote and why, line by
line, to somebody who knows the basic concepts. Followed by: what are three ways
to do this and their trade-offs? Every task becomes a lesson, at no extra cost.
Concepts over syntax. The machine writes the syntax now. Your job is the layer it
cannot do: knowing what to build, how the parts should connect, and whether the
output is right.
Do this: After the next thing gets built for you, ask for the line-by-line
explanation.
Common mistake: Studying for months before building anything you care about.
The traps that claim beginners
Five traps, and every one of them feels like progress at the time.
Rewriting instead of debugging. When something misbehaves, the itch is to start
over. Resist it. Debugging builds the skill you actually need; rewriting reliably
reproduces the same fault in new words, three days later.
Tool tourism. A different set of technologies every project means being a
permanent beginner in each. Pick a proven combination and stay on it at least six
months. Depth compounds; novelty does not.
Tutorial paralysis. Watching courses feels productive and mostly is not. Enforce
a ratio: one hour of learning to three hours of building.
Premature scaling. But will it handle a million users? You do not have ten. The
simple version is correct at your size, and complexity bought for imaginary load
is pure cost paid today.
Polishing before validating. Weeks perfecting features nobody asked for. Find out
whether anybody wants it first.
I would add three operational sins to the list: accepting work unread, saving
your progress rarely, and keys in the code.
Do this: Name which of these five you are doing right now.
Common mistake: Believing the trap you are in is the exception.
Mine the work you already know
The best problems are boring, specific, and already annoying somebody daily.
Innovation rarely comes from better technology. It comes from insight into a real
problem multiplied by the ability to ship — and shipping just got cheap for
everybody, which leaves the insight as the scarce half.
Insight is what you already have and a general-purpose developer does not. Nobody
outside your trade knows which five-minute task everybody does eleven times a
day. Hunt workflow pain, not app ideas: the spreadsheet everybody curses, the
document nobody can find, the question asked five times a day. Boring pain,
reliably felt, beats any brainstorm.
Three patterns are worth hunting in any industry. Turning documents into answers
you can query, which is the largest open field right now. The queue behind one
scarce expert, where a machine does the first pass and the expert signs. And glue
work — anywhere a person retypes data from one system into another.
Your own trade is the first domain, not the ceiling. The instincts transfer: any
regulated, document-heavy, safety-critical industry works the same way.
Do this: Write down three pieces of pain from your own week, not three app
ideas.
Common mistake: Looking for a clever idea when a boring, daily irritation was
the better business.
Five conversations before you build
Polite enthusiasm is worthless as evidence.
Talk to five people who actually have the problem before you build anything. Not
to pitch — to find out whether the pain is real, frequent, and worth money. Those
are three separate questions and an idea can pass one and fail the others.
Ask about the last time it happened, what they did instead, and what it cost
them. Past behaviour is evidence; enthusiasm about a future product is not.
Everybody is encouraging about an idea that costs them nothing.
Then build the crude version in a weekend and put it in front of one real person
— even if that person is you. Watch what they actually do with it. Real usage
steers better than any amount of planning, and a running ugly prototype beats a
perfect plan every time.
Ask for money, or at least a concrete commitment, early. The moment you ask, the
polite answers stop and the honest ones start.
Do this: Have five conversations about the pain before writing anything.
Common mistake: Treating “that sounds great” as validation.
Messages from visitors
Hello! Nice to meet you!
— anonymousDid you visit my page?
— anonymousGreat! 👋
— anonymous