flybycode

Everything on this site, as a readable list. Fly through the world instead.

The Air Force

2011–2024 · 38.34°N 23.57°E

Four years at the Academy and thirteen in uniform. People assume an air force engineer comes out knowing how to maintain an aircraft. I did not. I never worked as a mechanic here.

What Tanagra taught me was how to run a large organisation under pressure: how to deal with people, how to get a workshop of thirty to the right answer when you are not the one holding the spanner, and how to put your name on work you did not do with your own hands. That turned out to be the harder skill — and the one everything since has rested on.

The work

I came out of the Air Force Academy in 2011 with a degree in aeronautical engineering, and spent the first year at the Directorate of Airspace Applications: CAD modelling in CATIA V5, airflow and loads in ANSYS, on the structural integrity programme for the F-16.

Then Tanagra. Flight line maintenance officer on the Mirage 2000-5, supervising line maintenance. Commander of the base maintenance workshops — airframe, electrical, hydraulics, bonded stores, paint and corrosion protection. Stage A cross-servicing supervisor for F-16 and F/RF-4, preparing them to fly.

From 2014, quality control: monitoring Mirage inspections, airworthiness records, ground support equipment, tool calibration. From 2016, quality assurance — accident and incident investigation, auditing, technical compliance on the Mirage fleet and its support equipment, defective maintenance reports, and issuing engineers' maintenance authorisations.

For four of those years I also taught at the Non-Commissioning Officers Academy: advanced mathematics, science of materials, fluid dynamics, to first and second grade cadet officers.

What it taught me

Command a workshop and you stop being the person who knows the answer and start being the person who has to get thirty other people to the answer. Sign an authorisation and you are staking your name on somebody else's competence.

I did the master's at the National Technical University of Athens while I was serving, which the Air Force was good enough to note. The subject was mathematical modelling. The lesson was time management.

Athens

2011–2024 · 37.98°N 23.73°E

Lived here for the long middle of the story.

Thirteen years between an Air Force base and whichever civil organisation I was working for that year — usually more than one at a time. The career pins nearby are what happened while this was home.

Boarding Lab

47.98°N 7.81°E

Every published comparison of aircraft boarding methods starts its clock when the first passenger walks through the aircraft door. But somebody has to sort 180 people into that clever sequence first, and that work happens to the passengers, standing at a gate.

So I built a simulator that starts the clock earlier — at the announcement to prepare for boarding — and followed the same 180 people all the way to their seats. Strict Steffen, the mathematically optimal order, fills the cabin faster than anything else and still finishes the whole journey last, because building its perfect queue takes thirteen minutes.

That conclusion rests almost entirely on one number I could not find published: how long a gate agent needs to call a single passenger. So the number is a slider. Move it below about two and a half seconds and the answer flips.

Run the boarding experiment

Quality & compliance

2019– · 37.94°N 23.70°E

Where the work moved from the aircraft to the system that keeps the aircraft right.

Manuals, audits, transitions, and the unglamorous business of making a requirement survive all the way to the person doing the task.

Air Power, 2019–2024

I was there from the start of the organisation. Quality Manager in the CAMO first, then Safety and Compliance Monitoring Manager, and eventually covering both the CAMO and operations. The company was renamed Hoper toward the end; I stayed a few months after that and then left for TUS.

Starting with an organisation is a different job from joining one. There is no procedure to follow because you are the one writing it, and no precedent to point at when somebody asks why.

I led the transition to Part-CAMO and wrote the CAME and the operations manuals that went with it. I also built the training behind them — theory and practical, for airworthiness staff and for pilots: maintenance programmes, technical records, pre-flight inspections, aircraft familiarisation. From 2021 I was airworthiness review staff as well, which is the role that signs an aircraft's airworthiness review certificate — the annual check that the paperwork and the aeroplane still agree.

A transition is not a document exercise. It means changing procedures, manuals, training and everyday decisions at the same time, and the only test that matters is whether the person doing the task ends up doing it differently.

Marathon, 2022–2023

An airline mid-transition. I covered the CAMO safety and compliance role temporarily while it moved to Part-CAMO, then worked as Deputy Safety and Compliance Monitoring Manager.

The Part-CAMO transition and the CAME. The Management System Manual. Management-of-change workflows built in Centrik. Registration and upkeep of the external accreditations — Argus, WYVERN. And I led the WebManuals transition end to end.

What mattered was a management system people could operate, not a shelf of new documents. Workflows, manuals and records had to point at the same action. When they did not, that was the work.

TUS Airways, 2024–

Engineering Safety and Compliance Manager. I started on the CAMO side. In November 2025 we obtained the Part-145 approval; in 2026 I led the Part-IS introduction — EASA's requirement for cybersecurity and digital assets — into the engineering system.

The job is the same thread as the earlier chapters, only further down the line: find the requirement, walk it into practice, and leave evidence that it stayed there.

Corfu

1989 · 39.62°N 19.92°E

Born here, on the first of October 1989.

An island teaches you early that everything arrives by sea or by air. I think that is where the rest of it started.

Unfortunately, aviation on Corfu is mostly transit — wheels up, wheels down, and the island is already behind you. I left the other way: I stayed long enough to learn that the interesting part is not the landing, but what you build after you arrive.

EASA Caveman

50.95°N 6.96°E

BIG RULE. FEW WORD.

Part-M, Part-145 and Part-CAMO Section A enter cave. Legal throat-clearing gets clubbed. Competent-authority Section B stays outside. Rule references, numbers, units and dangerous little words such as not, only and except stay.

Rules kept best we can.

Part-M

Part-145

Part-CAMO

EASA Community Brain

50.88°N 7.08°E

A map of public talks on the EASA Community. Each point is a post. Click it and you land on the original page on EASA's site.

This is people discussing the work — it is not the rulebook, and it is not binding. Treat it as conversation, not as the law.

Open the EASA Community Brain map

Hands on the fleet

2015–2024 · 37.98°N 23.73°E

In 2015 I walked into a civil hangar not knowing how to screw in a bolt.

Thirteen years of Air Force engineering had taught me to manage aircraft and the people around them. It had not taught me to touch one. This is where that changed, alongside the uniform rather than after it.

Interisland, 2015–2019

Mechanic on a mixed light fleet — Robinson R22, R44 and R66 helicopters, and Maule MXT-7-180A, Robin ATL, Diamond HK36 and Van's RV-6A aeroplanes.

Within months of starting we had done a twelve-year inspection and assembled an R66. That is the whole story of this chapter in one sentence: the distance between not knowing which end of a spanner to hold and putting a helicopter together is shorter than anybody admits, and it is walked by doing it beside somebody who already can.

Everything I know about airworthiness and maintenance in practice — the records, the plan, the directive, the reason a task exists — I learned here, on aircraft I had my hands inside.

Jetway, 2018–2024

Robinson R44 maintenance, and certifying staff.

This is where the Part-66 licence stopped being a certificate. Certifying is a small physical act with a long shadow: you inspect the work, you read the record, and then you sign, and the signature is the part that matters.

Flyence, 2022–2024

Certifying staff first, then CAO Manager. The CAE was rewritten rather than amended — the point at which a manual has so many changes stacked on it that rebuilding it around the way the work actually runs is both faster and more honest.

Four helicopters came from Nepal under that approval: documentation, assembly, two twelve-year inspections, and autopilot installations. R22, R44 and R66 capability went onto the organisation with them. It was the best project of the stretch.

The work on that fleet was not only managed. Handling, fuelling, procurement, maintenance, airworthiness monitoring — every job. That is rare once the title says manager. A manager who has not recently done the work forgets what the work costs; I try not to become that person.

Alongside the helicopters: certifying staff on the Cessna F152 and 172, maintenance programmes tightened, procedures standardised, and custom modifications and installations under CS-STAN.

The Hangar

38.15°N 21.85°E

Four things I have built by describing what I want to a machine and arguing until it was right. Open a bay for the short version — what it is for, and what it taught me.

avioverse.io

The main one, and the reason the others exist. Coming soon — with real test cases and evaluations behind it, not a landing page and a promise.

Everywhere I have worked I have tried to improve the system I found there. Usually that meant automating whatever was actually in front of me, and for most of my career what was in front of me was Excel: the workbook that tracks the fleet, the workbook that tracks the training, the workbook somebody rebuilt from scratch because the last one finally broke.

Then one job put a complete SaaS in my hands, and it impressed me. Not because it did everything — it did not — but because it showed me the shape the answer could take.

So I called Thanos, the best developer I know, and told him we were going to build the next aviation SaaS: everything the existing ones were missing, plus everything twenty-five years had taught me about what the work actually needs.

We started. And while we were building it I carried on doing the thing I have always done — one tool for this, another tool for that, half the tasks impossible to track across either — until the obvious finally landed.

The market does not need another piece of software for businesses. Companies are well served. What nobody has built is something for the individual: the auditor, the mechanic, the pilot, the post holder, the admin who holds the whole thing together and owns none of the systems.

I built the core of avioverse.io in one night, and that is who it is for.

The agent: a colleague, not a coder

I have used the real thing. Claude Code, Codex, my own Hermes — run hard, on work that mattered, for long enough to know exactly what they can do. They are extraordinary.

They are also, in the hands of somebody who is not a programmer, extremely dangerous. The whole skill is knowing when a confident answer is wrong, and that is precisely the skill a non-programmer has not built yet. The machine does not sound less certain when it is mistaken. I have the aviation judgement to catch it in my own field and the scars to catch it in software; most people have one of those at best.

So the agent inside avioverse.io is not a coding agent pointed at aviation. It is built for the person doing the job — an expert colleague. One who has read the regulation, knows where the evidence is meant to live, and will tell you plainly when what you have does not support what you are about to sign. Bounded so that a wrong answer costs you an argument, not an airworthiness decision.

The core feature will be free. The AI on top is a subscription priced like software a person buys for themselves — not enterprise SaaS priced per seat by somebody who has never done the job.

fleetnode.org

Built to answer one question before committing to it in the main application: could a framework I was considering really carry what I needed?

So I gave it the hardest thing I had — nodes, and how you visualise a second brain. Ideas connected to ideas, and a way to see the shape of them rather than scroll a list. It is a proving flight, not a product: I wanted to know how the airframe behaved before I trusted it with passengers.

flybycode.com

This. A personal site you fly over instead of tabbing through — the CV, the articles and the guidance, arranged as a place rather than a page.

Built the way I build everything now: plain-language direction, review of what comes back, and refusal of anything I do not understand. Still flying, still being refined.

mxcourse.net

An asynchronous learning management system, currently open only to organisations. I built and shipped it because I wanted to see how Cloudflare actually behaves in production — not the brochure, the real thing.

Contact me if you like it.

Larnaca

2024 · 34.92°N 33.63°E

Home in Cyprus since 2024 — one house on the map for the whole island.

Engineering Safety and Compliance Manager at TUS Airways. Twenty-five years of aviation — if you linearise the parallel work — arrives here: safety, compliance, quality, airworthiness, training, operations, maintenance. Part-147, Part-CAMO, Part-145, Air Operations.

Evenings are for side projects: describe what you want to a machine, argue until it is right.

The Lighthouse

35.52°N 23.48°E

Thoughts — short pieces on aviation, safety, and building software by talking to a machine.

Things worth saying live here when they do not need a whole carrier deck of chapters.

Practical drift

The first thing I do in a company is read the documentation against the work. Not for typos — for drift. People leave. Regulations change. A procedure that was true three years ago is often still on the shelf, and the person doing the task has quietly invented a different way.

Find the gap between what is written and what is done. Improve the documents with focused amendments. The point is not a prettier manual; it is a manual somebody can still follow on a Tuesday afternoon.

The broken screw

The day after my engagement, before the last screw torque on a job, I skipped something I never skip: I did not open both the IPC and the AMM.

The result is a broken screw in my bag — and the biggest human-factors lesson I own. Distraction does not announce itself as distraction. It looks like confidence. It looks like you already know. It is the first and last time I signed off without both books open.

Learn SMS from life

To learn a safety management system, think about your life — not aviation examples. How you notice risk at home. How you decide what to fix first. How you check that a change actually stuck.

SMS does not need a huge budget or a shelf of consultants. It needs commitment, and enough love of the work that you would rather find the problem yourself than have it find you.

Where I am now

2024– · 34.91°N 33.63°E

Cyprus, an airline, and a company of my own.

TUS Airways, 2024–

Engineering Safety and Compliance Manager. Engineering compliance and safety across the CAMO, the Part-145 and the AOC.

I own the compliance-monitoring plan. I track (and issue) findings, corrective actions and evidence. I identify and manage engineering safety risk, liaise with the Cyprus DCA on engineering and continuing-airworthiness matters, and support accountable and safety management on engineering topics.

The work is quiet when it is working: a requirement is found, an audit follows it into practice, and an action closes with evidence. I keep the public account at that level. Internal findings, people and occurrences stay private.

AvioVerse

My own company, set up in the United Kingdom: consulting, human resources and training for aircraft operators.

Regulatory and management-system training, and CAR.66 and CAR.147 courses delivered to a CAR.147 organisation. Where an operator needs advice, or needs people rather than a document, AvioVerse is how I provide them.

Independent work

I have prepared organisations for the transition to Part-CAO and Part-CAMO, reviewed and amended maintenance programmes, and written a safety manual and a CAE for a declared training organisation. I am independent certifying staff for piston engine aeroplanes and R22 and R44 helicopters, and a mechanic on the R66. I deliver safety management training to maintenance organisations, CAMOs and operators.

For Avilaw I built the continuing-airworthiness section, pulling the regulation, the acceptable means, the guidance material and the EASA publications into one readable place — including one of the very few guides written for the Part-CAMO transition and Part-145 SMS implementation.

The Schoolhouse

37.20°N 18.40°E

I have never written software in my life. I have shipped it.

I am an aeronautical engineer. My hands know airworthiness and compliance, not compilers — so I learned to direct software instead of typing it. I say what the work needs, the machine proposes the code, and I refuse anything I do not understand. The product is mine; the typing is not.

Everything that taught me is here: forty cards in seven decks, from what a function actually is to how to keep an agent on a lead. Each one is short enough to read on the walk out to the aircraft, and each one stands on its own — so start with whichever deck sounds like your problem, not with the first card.

Teaching

2011– · 38.02°N 23.76°E

Years spent teaching the work — and finding out how much of it I actually understood.

If you cannot explain why a limit matters, you do not understand it yet. Teaching is where you find that out, in front of people who will ask.

Non-Commissioned Officers Academy

While I was still in uniform I taught at the Non-Commissioned Officers Academy: advanced mathematics, science of materials, and fluid dynamics, to first and second grade cadet officers.

Four years of standing in front of people who will ask the awkward question. That habit is older than any civil title on this map.

Aviatec, 2020–2024

Quality Manager of a Part-147 training organisation. A training organisation has two products: the course, and the evidence that the course is controlled.

I owned the MTOE and the audit programme, and wrote the procedures that kept us running through the COVID restrictions. I taught as well as managed — EASA regulations, safety management systems, human factors, maintenance programme development and planning — and the non-EASA courses too: human factors, EWIS, SMS, regulations.

I built the learning-management system (in Google Workspace) for developing instructors and examiners, and evaluated the new ones, permanent and ad-hoc. We expanded to satellite facilities in Doha and Morocco. I developed the Robinson R66 type rating course, along with Group 3 Level 1 training.

Training quality is easy to make bureaucratic. I tried to keep it usable: clear procedures, controlled questions, instructors who understood the rule, and a system that still worked when the classroom moved online.

Jetstream, 2021–2024

I qualified as an R22/R44 type rating instructor and delivered one course in 2021. What stayed with me was not the certificate — it was students contacting me long after the course ended, asking for guidance on modifications and troubleshooting. That is the test that matters: whether people still trust you when the classroom is over.

AvioVerse

Through AvioVerse I deliver regulatory and management-system training for aircraft operators — and, where needed, the people who make the system run rather than another document on the shelf. The same rule as every other deck on this map: if you cannot explain it, you do not understand it yet.

Control Tower

40.25°N 24.05°E

About Dennis Kefalas.

Aeronautical engineer. Four years at the Hellenic Air Force Academy and thirteen in uniform, then civil aviation from 2015 — the two ran alongside each other for nearly a decade. Safety, compliance, quality, airworthiness, training, operations and maintenance. EASA Part-66 licensed engineer, B1.2, B1.4 and C.

Now Engineering Safety and Compliance Manager at TUS Airways, in Cyprus.

In every organisation I have worked in I have tried to improve the system I found there — usually by automating whatever was in front of me, which for most of my career meant Excel. That habit is what eventually turned into building software.

I do not write software by typing it. I direct machines in plain language, review what they propose, and ship only what I understand. This site is one of those products.

What I trained in

A degree in aeronautical engineering from the Air Force Academy, and a master's in mathematical modelling from the National Technical University of Athens, taken while I was serving.

Accident investigation at Embry-Riddle. Flight safety, quality assurance, IOSA airline auditing with IATA, and the Inter-European Squadron Officer School. The certificates matter less than the habit they leave: find the requirement, follow it into practice, and check what actually happens rather than what is written down.

Family

Home is Cyprus, with the people I love. That is the part that matters, and the only part that needs a pin on this map. Names, ages and private detail stay off it — not because they are an afterthought, but because they are not for strangers. The people who matter already know where they sit in the story; visitors get the outline, and that is enough.

The Schoolhouse library

Short lessons about building software by directing machines in plain language.

The airframe

Every application has the same four parts

Learn the shape once and every new tool becomes their version of a part you know.

An aircraft has a cabin, a bay where the systems live, records that follow it for life, and defined joints between them. Software has the same gross anatomy, and it barely varies.

The frontend is what a person sees and touches. It runs on their device.

The backend runs on a computer you control. The rules live there, and so do the keys.

The API is the agreed joint between the two: which messages may pass, and in what shape.

The database keeps the permanent records. Everything else can crash and be restarted. This must not be lost.

Learn this once and new tools stop being alien. Whatever a framework calls itself, it is somebody's version of one of those four parts. It also tells you where an argument belongs: a colour change is a frontend question, a permission is a backend question, and what must survive a restart is a database question.

Do this: Take an application you use daily and name its four parts out loud.

Common mistake: Treating the screen as the whole system, and forgetting that rules and records have to live somewhere a user cannot reach.

The frontend is the cabin

Anything the passenger can reach, the passenger can tamper with.

The cabin is the part of the aircraft the passengers can get at. You design it knowing that. Nothing important is left loose in there.

The frontend is the cabin. It runs on somebody else's phone or browser — a machine you do not control and cannot inspect. Everything sent there can be read, and everything on it can be altered by anyone patient enough. So no keys, no prices, no permissions, no final decisions.

The hard problem a frontend solves is keeping the screen in step with the data. When one status changes, every place showing it has to change too. Doing that by hand is miserable, so frameworks do it for you: you declare that the screen is a picture of the data, you change the data, and the framework redraws whatever depended on it. That single idea is what React, Vue, Svelte and the rest are all selling.

Do this: Name one thing on a screen that must never be decided there.

Common mistake: Hiding a button and calling that a permission check.

The backend is the systems bay

Rules and keys belong behind a panel the passenger cannot open.

Behind the cabin lining is the bay where the systems actually live. Passengers do not go in there. It is not hidden to be mysterious; it is where things are enforced.

The backend is that bay. It is a program sitting on a computer you control, waiting for requests. When one arrives it checks who is asking, applies the rules, reads or writes the records, and sends an answer back.

Three things live here and nowhere else. The rules of the work, because this is the only place they cannot be bypassed. The keys and passwords, because this is the only place they are not handed to a stranger. And the decisions that count, because a decision made on a screen is a suggestion until the server agrees.

You do not need to read the code to ask which part owns a rule. If nobody can answer clearly, the design is mixing responsibilities that should stay apart.

Do this: For one feature, write down which rule the server must enforce even if the screen is bypassed entirely.

Common mistake: Letting the screen decide something and asking the server only to store the result.

The API is the interface control document

Two parts built by different people fit because a document says exactly how.

When two manufacturers build parts that have to bolt together, neither one guesses. A document specifies the joint: the exact dimensions, the fasteners, the loads it carries. Build to the document and the parts mate on the line.

An API is that document, for software. It says which addresses exist, what you may send to each, and what shape the answer comes back in. Send me this, and I will reply with that.

This is why parts written by strangers, in different languages, on different continents, fit together. The screen does not need to know how the server stores anything. The server does not need to know what the screen looks like. They only need to agree on the joint.

It is also why changing an API carelessly breaks things far away. Anyone built to the old document is now built to a joint that no longer exists.

Do this: For one feature, write the joint in a sentence: what goes in, what comes back.

Common mistake: Changing what an existing address returns, and being surprised when something else stops working.

The database is the technical record

Screens get repainted. The record outlives all of them.

An aircraft's records outlive its paint, its interior, its avionics and usually its owner. They are the thing you protect, because they are the only account of what happened that survives everything else.

The database is that record. It is structured, permanent storage, and it is the single source of truth. The screen only shows it. Closing a browser must not erase anything, and renaming a label on a page must not quietly rewrite history.

Two families exist. Relational databases keep data in tables with strict columns, linked by identifiers, and the database itself enforces the rules — no work order without a valid aircraft. That is the sensible default, because most real information is naturally relational. The other family trades those enforced rules for flexibility, which means the rules now live in somebody's head.

SQL is the small, stable language for asking these questions. It is fifty years old, it is not going anywhere, and it is genuinely worth learning.

Do this: List the real things one tool must remember, and treat that list as the design.

Common mistake: Designing the pages first and bending the records to fit them.

Two languages, and a pairing that already flies

You do not need a rating on every type. Two languages cover almost all of it.

Nobody flies every type. You get rated on what the job needs and you recognise the rest by name.

Two languages cover nearly all of this work. JavaScript, usually written as TypeScript, is the only language a browser runs — so anything on a screen is that, whatever else you choose. It runs servers too. Python is the language of automation, data and AI work. SQL asks the records questions. Everything else is a rating you take when a project demands it, which for most people is never.

TypeScript is JavaScript with a checking step. It catches a whole class of mistakes before the program runs — the difference between finding a defect at inspection and finding it in flight.

Then do not assemble from the full catalogue. Fly a proven pairing: a combination thousands of people already run, which is also the one AI assistants know best. Pick one and stay on it six months. Depth compounds.

Do this: Choose one combination for your first project and refuse to change it for six months.

Common mistake: A new set of tools every project, which makes you a permanent beginner in each.

How a request travels

A request goes out, a reply comes back

Press Enter and a strict little conversation happens, always in the same order.

A radio call has a format. You transmit in that format, and what comes back tells you how it went before it tells you anything else.

The web works the same way. You type an address and press Enter. The name is looked up and turned into a numeric address, the way a phone book turns a name into a number. An encrypted connection opens. Code on the server runs, usually asking the database something. A reply travels back and the browser draws it.

Every request carries a method saying what you want: read, create, update or delete. Every reply carries a status number saying what happened. Two hundred means fine. Four hundred and four means there is no such thing. Four hundred and one means you are not allowed. Five hundred means the server itself broke — that one is yours to fix, and the others usually are not.

The common style is to make addresses the things and methods the actions.

Do this: Next time a page fails, find the status number before theorising.

Common mistake: Treating every failure as one undifferentiated broken.

JSON is how systems speak to each other

One plain text format, readable by every language, is why strangers’ parts fit.

A shipping label works because everyone reads it the same way. Labelled lines, plain text, no interpretation required: part number here, quantity there, destination underneath.

JSON is the shipping label of software. It is text, laid out as labels and values, and it is what nearly every system sends when it has structured information to pass on. A report might travel as a number labelled id, a word labelled status, and a list labelled items. You can read it yourself without tools, which is the point.

Its value is that every language reads and writes it. A part written in Python and a part written in TypeScript exchange work through JSON without knowing anything about each other. That is what makes it possible to choose the right language for each part instead of one language for everything.

When somebody says an API returns JSON, they are saying the answers arrive as labelled text you can look at.

Do this: Find one API's example answer online and read it. You will understand more of it than you expect.

Common mistake: Assuming the format is technical and therefore unreadable.

Never trust the cabin

Anyone can send your server a hand-made request that never touched your screen.

This is the most common serious mistake beginners make, and it is worth its own card.

Your screen is a convenience for honest users. It is not a gate. Anyone can send requests to your server directly, without ever loading your page — the tools to do it are free and take a minute to learn. So every check that matters has to be made again on the server: is this person allowed, does this record belong to them, is this price the real price, is this quantity sane.

The screen may still do the check, because telling somebody early is kinder than telling them late. But the screen's check is a courtesy. The server's check is the control.

The same logic you already use elsewhere: a placard on the outside of a panel is information, not a lock.

Do this: For each rule in your feature, ask where it is enforced. If the answer is only the screen, it is not enforced.

Common mistake: Believing that because your page cannot send a bad value, nobody can.

The technical records

One table per real thing

Get the record shape right and everything is easier. Get it wrong and data piles on top of the mistake.

The data model is the one decision that gets ten times more expensive to change later, because by then real data has accumulated on top of it.

Four rules cover most of it. One table for each real kind of thing — aircraft, reports, people — and one row for each individual one. Every row gets its own identifier. Rows point at other rows by storing that identifier; where two things relate many to many, a small table in the middle holds one row per pairing. And do not store anything you can work out — totals, counts, ages. A stored copy drifts from the source, and then you have two versions of the truth.

The tell that you have it wrong: you are adding columns called item1, item2, item3. That is a second table trying to escape.

Sketch the tables before any code. Describe the work, ask for a proposed structure, then attack it with the awkward cases you know and the machine does not.

Do this: Sketch one tool's tables and test them against your two most awkward real cases.

Common mistake: Designing attractive pages first and forcing the records to fit them.

A migration is the modification record

You never modify an aircraft without recording the modification.

No modification goes onto an aircraft without paperwork that says what changed, when, and against what approval. The point is not bureaucracy. It is that the next person needs to know what state the aircraft is actually in.

As an application grows, the shape of its records has to change: a new column, a new table, a renamed field. A migration is a small numbered script that makes exactly that change and can be replayed, in order, on any copy of the database. Run them all against an empty database and you get today's structure. That is the modification record, and it is how a change made on your machine reaches the live one intact.

The rule that follows: never hand-edit a live database because it seems faster. The change then exists in one place only, recorded nowhere, and the next deployment has no idea it happened. Structural changes need a repeatable method and a backup you have actually restored.

Do this: Before any structural change, ask how it will be applied to the live records and by what script.

Common mistake: One quick manual fix in production that nobody can find six months later.

A backup you have never restored is a rumour

Backups are the fire bottle. You must have one, and you must have fired one.

A fire bottle you have never checked is not a fire bottle. It is a red cylinder you are hoping about.

Backups are the same. Almost everyone sets one up. Very few ever restore one, which means very few know whether the file being written every night is complete, readable, or in fact empty. The moment you find out is the worst possible moment to find out.

So restore one. Not in theory — take last night's backup, load it into a spare copy of the database, open it and look. Check the newest records are there, and that the awkward parts came across intact. Write down how long the whole thing took, because on the day it matters somebody will ask.

Do it again occasionally, because backups quietly break when the structure changes.

Do this: Restore one backup into a scratch copy this month and look at the newest record in it.

Common mistake: Treating the existence of a backup file as evidence that restoring it will work.

Why a page goes slow

Almost every slow page is one of three things, and none of them is a slow computer.

Slowness almost never comes from the language or the processor. It comes from waiting, and from repeating work.

The first cause is a missing index. An index is the index at the back of a manual: without one the database reads every row to find a match; with one it goes straight to the page. The difference is routinely a thousandfold. Index any column you regularly filter, sort or join by.

The second is fetching a list, then fetching each item's related record one at a time. Fifty rows becomes fifty-one trips to draw one page. It is invisible with five rows of test data and brutal with five thousand, which is why it survives to production. When you ask for a list view, say plainly: fetch the related data in one query.

The third is repeating expensive work — an outside lookup on every page view. Store the answer with an expiry and reuse it.

Then measure. Speeding up something that was never the bottleneck is the classic wasted week.

Do this: When a page feels slow, check for a missing index and repeated per-row queries before anything else.

Common mistake: Assuming a bigger machine will fix it.

Ground equipment

Version control is an audit trail that writes itself

It records every change, which is what makes bold experiments cheap.

Revision control on a manual is a discipline: every amendment recorded, every superseded page traceable. Git does that for code, and it does it automatically.

Git saves snapshots of the project. Each one carries who made it, when, and a message saying why — a full audit trail, without anyone maintaining it. You can branch, meaning work on a copy without touching the good version, merge it back when it is right, and return to any earlier point exactly as it was. GitHub is the hosted home for these projects, and adds a review step before changes are accepted.

The habit that matters is small and frequent snapshots, with messages that say why rather than what. And commit before every experiment. When an assistant mangles the code at midnight, one command throws away the mess and puts you back where you were. I have never once regretted a snapshot taken before trying something.

That is the whole trick: version control is what makes fearlessness cheap.

Do this: Commit before you let anything make a change you are unsure about.

Common mistake: One enormous snapshot at the end of the day, labelled “stuff”.

Packages are the parts catalogue

Modern software is mostly assembled from published parts.

Almost nothing is made from raw material any more. You order the part, you record which part number and which revision went on, and the record is what makes the next build identical.

A package manager is the catalogue and the stores counter together. It fetches published code your project depends on, and writes down the exact version of every single one in a lockfile, so that every machine builds the same aircraft. The JavaScript world uses npm; Python uses pip, or the newer and faster uv. The folder of fetched parts is large, regenerable, and never stored with your own code.

Version numbers read as three numbers: the last one is fixes, the middle one additions, the first one breaking changes. Treat a first-number change as an inspection task with a test flight, not a routine top-up.

And weigh each new dependency. Every one is somebody else's code you now depend on, update with, and inherit the faults of.

Do this: Before adding a part, ask what it does that twenty lines of your own would not.

Common mistake: Accepting a major version jump as if it were a patch.

Secrets live in a locked cabinet, never in the code

A key left in the code is found by automated scanners within minutes.

Keys are not left in the aircraft. They live in a cabinet, and taking one out is a recorded act. Nobody thinks this is excessive.

Software keys — database passwords, service credentials, API keys — follow the same rule and it is not optional. They never go in the code. On your own machine they live in a separate settings file that version control is explicitly told to ignore. In production they live in the hosting service's settings, handed to the program when it starts.

The reason for the strictness: public code is scanned continuously by people looking for exactly this. A key committed by accident is typically found and used within minutes, and it is usually your card that pays for the usage.

If a key does escape, the fix is not to delete the line. It is to revoke that key and issue a new one. The old one is in the history for ever.

Do this: Check that your settings file is on the ignore list before the first commit, not after.

Common mistake: Pasting a key in temporarily and meaning to move it later.

Deployment is getting it airborne

Three postures, and the choice is mostly about who does the maintenance.

Running on your own machine is a ground run. Deployment is putting the thing on a computer that is always on and reachable, which is a different set of concerns.

Three postures. A managed platform takes your code, builds it, serves it and handles the certificates — least work, least control, and free at small scale. Your own rented machine is cheap and will run anything, including jobs that never stop, but you are now the mechanic and the operations department: updates, firewall, certificates, monitoring, restarts. A managed backend service gives you database, logins and file storage ready-made, which is an enormous head start on their terms.

A container packages a program with its exact surroundings so it runs the same everywhere. It is the standard shipping container of software, and it is worth the extra learning only once deployment actually demands it.

Then automate the release: on every push, run the tests and deploy only if they pass. That is continuous inspection built into the process.

Do this: Choose your posture by who you want doing the maintenance at 2am.

Common mistake: Renting a machine to look after, for something a managed platform would have carried for nothing.

Airworthiness

Every input is unverified cargo

Nothing arriving from outside is trusted until it has been checked.

Cargo does not go in the hold because somebody says it is fine. It is checked, weighed and documented first. Anything arriving from outside your system deserves the same suspicion.

Check every incoming value on the server: the right type, a sensible range, present when it must be, allowed to be there at all. Libraries exist that do this for you from a description of the expected shape, so it costs a few lines rather than an afternoon.

The classic failure has a name. If you build a database question by gluing the user's text into it, a user can write their own question — and read or delete everything. The fix is free: never assemble a query by pasting text together. Pass the values separately, so data stays data and can never become an instruction. Any decent database library does this by default.

The same shape of attack now reaches AI systems, where hostile instructions hide inside content the model is asked to read.

Do this: For one form, write down every value it accepts and what would happen if each arrived as nonsense.

Common mistake: Checking the shape of data on the screen only.

Who you are is not what you may do

Checking the licence is not the same as checking the privilege.

A valid licence in somebody's pocket does not tell you they may sign for this particular work, on this particular type. Two questions, two checks, and mixing them up is how things get signed that should not have been.

Software has exactly the same pair. Authentication answers who are you — the login. Authorisation answers what may you do — the roles and the ownership. The classic hole is checking the first and forgetting the second: the user is definitely logged in, and is definitely reading somebody else's record because nobody asked whether it was theirs.

The check belongs on every request that touches a record, not once at the door.

And do not build the login itself. Password storage, resets, sessions, lockouts — this is solved, dangerous-to-improvise territory with a long history of expensive mistakes. Use an established service and spend your attention on the part only you can build.

Do this: For one record type, write the sentence: may this user do this to this specific record?

Common mistake: Assuming that being logged in is a permission.

Tests are the alarm, not the proof

Their job is not today. Their job is the day a distant change breaks something.

People assume tests exist to prove the thing works. That is the small part. A test is a permanently installed warning system: it rings when a change somewhere else breaks something over here, months after everyone forgot the two were connected.

Tests are code that checks your code, and they rerun in seconds, for ever. Some check one small piece in isolation. Some check that parts work together. Some drive the whole thing the way a user would.

This matters more with an AI assistant than without one. The assistant edits boldly and confidently across files; the tests are what catch the damage it did not know it was doing. Without them you are reviewing every line by eye for ever.

Two cheaper cousins do related work. Automatic style and defect checkers flag likely mistakes without running anything. A type-checking language catches a whole class of errors at build time — at inspection rather than in flight.

Do this: Ask for a test on the one path that must never break, before asking for more features.

Common mistake: Testing everything shallowly and the critical path not at all.

Logs are the flight recorder

When it misbehaves at two in the morning, logs are the only evidence you have.

After an event, you do not rely on anyone's recollection. You read the recorder. It was running before anybody knew there would be something to investigate, which is the entire point.

Logs are that. A program writes down what it did, in order, as it does it: this request arrived, this user, this record, this decision, this failure. When something goes wrong on a machine you cannot see, at an hour when nobody was watching, the log is the only account that exists.

So log what happened and which identifiers were involved — the record number, the user, the outcome. A line saying “error” with nothing attached is noise.

And never log secrets, or personal detail you do not need. Logs get copied, shipped to other services and read by people who were not thinking about privacy. Whatever is in them has effectively been published inside your organisation.

Do this: For one important action, write down the line you would want to read the next morning.

Common mistake: Adding logging after the incident that needed it.

Reading the code

Every value has a type, including nothing

A surprising share of all crashes is code expecting something that turned out to be nothing.

Syntax is costume. The ideas underneath are the same in every language you will ever meet, which is why learning them once makes every tutorial, error message and explanation readable.

All data is values, and every value has a kind. Numbers. Text, called a string, and always in quotes. True or false, called a boolean. And nothing at all, which different languages call null, undefined or None. A variable is simply a name bound to a value so the code can refer to it later.

That last kind causes more crashes than anything else. Code assumes something is there — a record, a field, an answer — and it is nothing. The message you will meet says it cannot read a property of undefined. You now know exactly what that means: something you expected to exist did not, and the line named in the error is where the assumption was made, not necessarily where the mistake was.

Do this: When something crashes, ask first which value turned out to be nothing.

Common mistake: Reading an error as gibberish when it is naming the exact file, line and assumption that failed.

Two shapes describe nearly all data

A list of things, and a thing with labelled parts. That is most of it.

There are two containers, and between them they describe almost every piece of information you will ever handle.

A list is an ordered many: three registrations, forty reports, the items on one work card. Use it when you have several of something.

A labelled set — languages call it a dictionary, an object or a map — is one thing with several named parts: registration, hours, status, owner. Use it when one thing has properties.

Now nest them. A fleet is a list of labelled sets. A work order is a labelled set containing a list of task sets. A month of reports is a list of those. That is genuinely the shape of most real data, and JSON is exactly these two shapes written down.

The moment you can look at real information — a form, a spreadsheet, a report — and see its list-and-labels shape, you can describe it to a machine precisely. That skill is worth more than any amount of syntax.

Do this: Take one document from your work and write it as lists and labels.

Common mistake: Describing data in prose when the shape would have said it exactly.

A function turns inputs into a result

A part you can bench-test beats a part you can only test installed.

A function takes inputs, does work, and returns a result. That is the whole idea, and code is mostly functions calling other functions.

Two habits make code dramatically easier to reason about, and you can ask for both without writing a line yourself.

Return early: handle the “no” cases at the top and get out, instead of nesting conditions five deep. Deeply nested code is where mistakes hide.

Prefer functions that touch nothing else. Give one the same inputs and it always gives the same output, changing nothing anywhere in the system. That kind of function can be bench-tested on its own, in seconds, in isolation. A function that quietly modifies things elsewhere can only be tested installed on the aircraft, with everything else running, and when it misbehaves you are searching the whole system for the cause.

Do this: Ask for the awkward part as a separate function that only takes inputs and returns a result.

Common mistake: Accepting one enormous function that does the fetching, the deciding and the displaying at once.

State is what changes while it runs

Most bugs are state bugs: something changed when you did not expect it to.

State is any data that changes while the program is running: what is currently selected, who is logged in, which items are loaded, whether the save is in progress.

Almost all difficult bugs are state bugs. Something changed when nothing should have. Something did not change when it should have. Two parts held different ideas of the same fact at the same moment. Nothing is broken in isolation — every piece works alone — and the fault only appears in a particular order of events. That is why “it only happens sometimes” is so common, and why writing down the exact sequence is half the fix.

Keeping a screen in step with changing data is the hardest version of this problem, which is precisely why frontend frameworks exist.

Scope is the smaller cousin: where a name is visible. A value created inside a function exists only there. When code cannot see something you are certain exists, scope is usually the answer.

Do this: When a bug is intermittent, write down the exact order of actions that triggers it.

Common mistake: Hunting for a broken part when the parts are fine and the sequence is not.

Standard equipment in standard places

Thirty unfamiliar files, and nearly all of them are standard fit.

Open an unfamiliar aircraft and you do not know that airframe — but you know where the manuals are, where the placards are, and what a fuse panel looks like. A project folder is the same. Almost everything in it is standard equipment in a standard location.

Start with the readme: what this is and how to run it. Always. Then the manifest — the file listing the project's parts and the commands it knows how to run, which is effectively its list of capabilities. Beside it sits the lockfile, which is machine-maintained and never hand-edited. An example settings file names which secrets the program expects without containing any. The ignore list says what must never be uploaded. The code itself lives in one folder, usually called src or app, and the tests in another.

Frameworks put files in fixed places on purpose: the location means something. When a folder structure seems rigid, the rigidity is usually load-bearing.

Do this: In any new project, read the readme and the list of commands before opening a single code file.

Common mistake: Browsing files alphabetically and mistaking volume for understanding.

Trace one action from end to end

One completed trace teaches more than an hour of browsing files.

Pick one button. Follow it the whole way: the piece of screen that holds it, the request it sends, the server route that receives it, the question that hits the database, the answer coming back, the screen changing. One trace, all the way through, and the shape of the whole project appears.

Doing it by reading takes a while. Asking is faster: give an assistant the project and say walk me through this — where is the screen, where are the rules, where are the records — then trace what happens when a user does this one thing. It is the fastest way into unfamiliar code, and you can check the answer by following it yourself.

A few words you will meet while tracing. A monolith is one program containing everything, and it is the right default. Layers means code separated by job: routes, then rules, then data. Middleware runs on every request before anything else — the guard at the single entrance. A background job is work done outside the request, so nobody waits for it.

Do this: Trace one action end to end before you change anything.

Common mistake: Reading everything and understanding nothing.

You are the pilot in command

Ask for the plan before the build

Fixing a plan costs one sentence. Fixing the wrong code costs an evening.

An AI assistant is a very fast, very well-read, occasionally overconfident first officer. The arrangement that makes that safe is the one cockpits already use: the pilot in command sets the destination, delegates the flying, and cross- checks continuously. Never sleeps in the seat.

So for anything beyond the trivial, say: propose a plan first, do not write code yet. Then read the plan. A wrong assumption in a plan is one sentence to correct. The same assumption discovered after four files were written costs an evening.

Then one task per request. Add the login page, not build the application. Small changes can actually be reviewed. Thousand-line changes get waved through, and waving things through is exactly how bad code gets in — not through malice, but through volume.

The pattern is a loop you run every time: plan, small step, read what changed, save it, check it actually works.

Do this: Say “propose a plan first, do not write code yet” and read what comes back.

Common mistake: Asking for the whole application in one go, and getting something too large to check.

Read what changed, not what you were told

You do not need to write code to smell trouble in it.

Every tool can show you exactly what changed: lines added, lines removed, file by file. That view is called a diff, and reading it is a skill you already have in another form — you have reviewed paperwork for things that should not be there.

You are not checking whether the code is elegant. You are looking for four smells, and all four are visible without knowing the language.

Code deleted that you did not ask about. A new outside dependency appearing from nowhere. A key, password or address typed directly into a file. And changes in files that have nothing to do with the task you asked for.

Any of those gets the same question: why did you change this? A good assistant answers it plainly. An unconvincing answer is itself information.

The discipline that makes this survivable is small changes. Nobody can read a thousand-line diff honestly, and pretending otherwise is how the smells get through.

Do this: Read the diff before saying yes, and ask about anything surprising.

Common mistake: Approving on the strength of a confident summary of what was done.

Done is a claim, not evidence

Run it, click it, look at it. Then it is done.

An assistant reporting that a feature is complete is making a claim about work it cannot fully observe. Sometimes the claim is true. Often it is true of the part it was thinking about and not of the whole.

So treat “done” the way you would treat any unwitnessed sign-off: it is the beginning of the check, not the end of the job. Run the thing. Click the button. Try the awkward case. Look at the record afterwards and confirm it says what it should.

Better still, invert the order. Ask for the test first, then the code that passes it. A test is a contract that cannot be quietly talked around, and it turns “it works” from an opinion into something that either passes or does not.

This is not distrust, it is verification, and it is the same reason a signature comes after an inspection rather than before it.

Do this: Before accepting any change, run the thing yourself and try one case nobody mentioned.

Common mistake: Accepting a summary of the work as evidence of the work.

Context is the fuel

It knows only what is in front of it. Nothing carries over by itself.

Nothing carries over between flights. You check the fuel before every one, because what was in the tanks last time is not there now. An assistant starts each session the same way: knowing nothing about your project except what it is handed, and everything you told it yesterday gone.

The fix is a standing orders file in the project: how to run it, the conventions you keep, what must never be touched, and why. It loads automatically into every session. Whenever you correct the same thing twice, the correction belongs in that file rather than in your repetition.

Keep it lean. A bloated file buries the rules that matter, and models over-obey shouted language: write “CRITICAL — you MUST” and you will get work contorted to satisfy the shouting. State each rule plainly and delete anything the assistant already gets right.

Within a session, name files exactly and paste errors verbatim. And start fresh when the topic changes: long conversations accumulate stale material and get worse.

Do this: Move your next repeated correction into the project's standing orders file.

Common mistake: Writing a long instruction file and wondering why the important rules get missed.

Where the first officer fails

Four failure modes, all four predictable, all four cheap to counter.

Assistants fail in patterns. Knowing the patterns is most of the defence.

Invented equipment. It produces a function or a setting that sounds exactly right for a library and does not exist, because it inferred it from the library's general feel. If code fails with “no such method”, suspect this first and tell it to check current documentation.

Sweeping changes. Large multi-file rewrites are where errors compound silently, because nobody can review them honestly. Break them into staged steps, saved between each.

Judgement calls. What to build, what good enough means, which trade-off suits your situation. It will pick one confidently and the pick is yours to make. Force the alternatives out: give me three approaches with trade-offs before choosing.

Agreeableness. It leans towards telling you your idea works. So do not ask whether it works. Ask what is wrong with this approach, and argue against it.

Do this: On your next design question, ask for three approaches with trade-offs before anything is written.

Common mistake: Reading confident agreement as a second opinion.

Writing the work card

Treat a prompt like a work card

If a colleague reading only your instruction would be confused, so is the model.

A work card does not say “look at the brakes”. It says who does it, to what, to which limits, and what is recorded at the end. That precision is what lets the work happen without the author standing there.

Write instructions to a model the same way. State the task, who the output is for, how long it should be, and what shape it should take. The test is exactly the one you would apply to a work card: if somebody competent read only this, would they know what to do?

Two more levers cost nothing. Give it a role — “you are a careful technical editor reviewing a procedure” — which measurably shifts rigour and vocabulary. And keep the material separate from the instruction: wrap anything you paste in simple tags, so pasted content cannot be mistaken for orders. When the material is long, put it first and your question at the very bottom. Models answer measurably better when they read the evidence before learning what to look for.

Do this: Rewrite one instruction to state task, audience, length and format explicitly.

Common mistake: A one-line request, then blaming the answer for guessing wrong.

Say what to do, not what not to do

“Write in flowing paragraphs” beats “do not use bullet points”, every time.

This is the smallest change with the largest effect, and it takes one rewrite to adopt.

A positive instruction gives a target to hit. A prohibition fences off one option and leaves everything else to guesswork. Tell a model not to be verbose and you have ruled out one style from a hundred. Tell it to reply in one sentence and there is nothing left to interpret.

So phrase your rules as the thing you want. Not “do not invent references” but “cite only passages that appear in the text I gave you”. Not “avoid technical language” but “explain it to somebody who has never opened a terminal”. Not “do not change other files” but “change only the file I named”.

The same instinct works on people, which is why briefings say what to do. It just matters more here, because a model has no shared context to fall back on when your instruction only rules one thing out.

Do this: Take your last prohibition and rewrite it as the behaviour you actually want.

Common mistake: A list of things not to do, and no description of the target.

Show it two worked examples

Two examples beat three paragraphs of description.

Describing a format is slow and lossy. Showing it is fast and exact. Two or three worked examples — this input, that output — are the single strongest lever you have for getting consistent results.

It works for the same reason a filled-in sample form works better than a page of notes on how to fill in the form. The pattern is visible, including all the small conventions you would never have thought to write down: how dates look, what to do with a missing value, how long an entry should be.

Choose the examples carefully, because they are your specification now. Include one straightforward case and one awkward one — the model will copy both the shape and the judgement. If your examples all show tidy inputs, you have taught it nothing about mess.

Where the output has to be machine-readable, examples and a defined shape work together: examples teach the judgement, the shape enforces the structure.

Do this: Add one clean example and one awkward one to an instruction that keeps producing the wrong shape.

Common mistake: Explaining the format at length instead of showing it once.

Let it think first, and let it say no

Thinking after the answer is worthless. Order matters.

Ask for the reasoning before the answer, in its own section. Accuracy climbs on anything that is not trivial, and you get something you can check rather than a conclusion you must take on faith.

The order is not a detail. A model writes one piece at a time and cannot go back and revise what it already said, so reasoning written after an answer is justification, not thinking. It will happily defend a conclusion it reached carelessly.

The second half of this card is just as valuable. Say explicitly that “I do not know, the text does not say” is an acceptable answer. Invented answers drop sharply once refusing is permitted, because you have removed the pressure to produce something.

For questions about a document, go further: require it to quote the relevant passage first and answer only from the quote. Then you can check the answer against its own evidence in seconds, without reading the document yourself.

Do this: Add “reason it through first, then answer” and “say if the text does not cover it” to your next question.

Common mistake: Asking for an answer with an explanation, which produces a guess with a defence attached.

An eval set is your test equipment

A prompt that worked three times is untested equipment.

You would not sign off a measurement taken with an instrument nobody had calibrated. A prompt that seemed to work on three tries is exactly that instrument.

Build a small set of representative inputs together with the answers you would accept. Twenty is enough to start. Then every time you change the instruction, run the whole set and grade the results — by exact comparison where the answer is a fact, by reading them where it is judgement, or by a second model checking against your criteria.

Two things happen. You stop guessing whether a change helped, because you can see it. And you find out that the change which fixed one case broke two others, which is otherwise invisible until a user finds it.

This is the difference between a demo and a product. Anyone can produce an impressive demonstration in an afternoon; the distance from works once to works reliably on real, messy input is covered by exactly this.

Do this: Write down ten real inputs and the answers you would accept, before tuning any instruction further.

Common mistake: Tuning by impression, one example at a time.

Operating the engine

It predicts the next piece of text

One fact explains every strength it has and every way it fails.

A large language model does one thing: it predicts the next small piece of text. Trained on an enormous amount of writing, it learned the shape of language, code and argument well enough that this prediction produces genuinely useful work.

The unit it predicts is called a token — roughly three quarters of a word. It produces one, appends it, and predicts the next from everything so far.

Everything else follows from that. Answers arrive a word at a time because they are being produced a piece at a time. It cannot go back and revise what it has already said, only continue — which is why reasoning before an answer helps and justifying afterwards does not. And because it is producing what is plausible, it is superb at pattern work and unreliable on exact recall.

Understanding this one mechanism is the difference between being impressed by the machine and being able to operate it.

Do this: Before any task, decide whether you are asking for a pattern or a fact.

Common mistake: Treating fluency as evidence of knowledge.

The context window is working memory

It is a clipboard, not a filing cabinet.

Everything a model can see in one conversation — your instructions, the documents you pasted, everything either of you has said — sits in one working memory called the context window. It is measured in tokens, and it is all the model knows about your session.

Three consequences matter.

Nothing persists. Between conversations, the clipboard is wiped. If it needs to know something today that you told it yesterday, you send it again. That is not forgetfulness in the human sense; there is simply nowhere for it to go.

It fills up. When it does, the oldest material effectively falls off the edge — including, quite often, the instruction you gave at the start.

And quality sags in the middle of a very full window. Relating everything to everything gets harder as the pile grows, so the material in the middle of an enormous context gets the least attention.

Do this: Start a fresh conversation when the topic changes, and re-state what matters.

Common mistake: A single enormous conversation that has quietly forgotten its own instructions.

There is no database inside

It has no store of facts to look things up in. That is why it invents.

This is the mental model that explains every failure you will meet.

There is no filing cabinet inside a model. Its knowledge is not stored anywhere retrievable — it is spread across billions of learned numbers as statistical tendencies. Nothing can be looked up, only reconstructed.

So it cannot tell you where it learned something, because there is no where. Its grasp of patterns is superb and its recall of specifics — part numbers, citations, figures, function names — is unreliable, exactly where accuracy matters most. Invention is intrinsic, not a bug somebody will fix. And its knowledge stops at the date its training stopped.

The working rule that follows: pattern work, trust it — drafting, classifying, explaining, restructuring. Specifics, make it look them up and quote what it found.

And no, training it on your own documents is almost never the answer. That is expensive, brittle and does not reliably add knowledge. Better instructions, examples, and handing it the right pages solve nearly every case.

Do this: Split your task into pattern work and exact facts, and treat the two differently.

Common mistake: Believing a confident, well-written specific because it is confident and well-written.

A model call is just a request

You pay by the word, in and out, on every single turn.

Using a model from code is not exotic. It is the ordinary request and reply of any other service: send the conversation so far, get the next message back. That loop is the entire foundation of every AI product you have used, which is worth knowing before anyone sells you a platform.

Cost follows the text. You pay for what you send and what comes back, every call — and a long conversation re-sends everything each turn, so cost grows with the square of a chat rather than in a line. Long documents in the instruction are the usual bill.

Two dials worth knowing. Model size trades intelligence for price: route the easy, high-volume work to the small cheap model and keep the expensive one for the hard part. And the same input can produce a different output each time, because probability is the mechanism — so for extraction and classification, ask for the most predictable setting, and never assume two runs will match.

Do this: Before automating anything, work out the cost of one call and multiply by the real volume.

Common mistake: Sending an entire document on every turn of a long conversation.

Structured output makes it a component

This is the step from chatbot to software part.

A chat window is a demonstration. A part your program can rely on is a product. The bridge between the two is telling the model the exact shape of the answer and requiring it to match.

Instead of asking for a summary and getting prose, you hand it a defined shape — these four labelled fields, this one from a fixed list of values — and the reply is constrained to that. Messy human text goes in; clean, labelled data comes out. A rambling report becomes a category, a severity, a component and an action, which your program can then store, count and route.

That is the moment a model stops being something a person talks to and becomes a function inside a system, sitting between two ordinary pieces of code that know nothing about it.

It also makes checking possible. A shape can be validated automatically before anything downstream trusts it, and anything that fails validation can be retried or set aside for a person.

Do this: For one task, write down the exact fields you want back before you write the instruction.

Common mistake: Parsing prose with clever text matching instead of asking for a defined shape.

Meaning as coordinates

An embedding is a position in meaning-space

Turn meaning into coordinates and “find things like this” becomes arithmetic.

Plot two numbers and you get a point on a chart. Plot several hundred and you get a point in a space nobody can picture, but the arithmetic works identically.

An embedding is a piece of text turned into a few hundred numbers — a position in that space — arranged so that texts meaning similar things land near each other. Brake wear limits sits close to minimum pad thickness before replacement, with not one word in common. Nobody named the axes; training discovered them.

Two consequences matter. Similarity becomes a distance you can measure, so finding things like this one is arithmetic rather than guesswork. And it is extraordinarily cheap: a whole document collection turns into coordinates on an ordinary laptop in minutes.

That is why the right shape is almost always coordinates find the material and the model reads it — never the reverse.

Do this: Think of one pile of documents where “find things like this one” would save real time.

Common mistake: Asking a model to read a whole archive when arithmetic could have found the ten relevant pages first.

Hand it the right pages

The model is not retrained. It is handed the relevant pages at question time.

The most useful pattern in this whole field, and the answer to “how do I make it know my documents”.

The question arrives. Your system searches your own documents, takes the best few passages, puts them in front of the model and asks it to answer from those, citing what it used. The model has learned nothing new — it has been handed the right pages, rather than sent on a course.

The consequence: answer quality is dominated by the search, not the model. When answers come back bad, debug the search first. Almost everybody blames the model, and almost everybody is wrong.

Two things decide the search. How documents are split — one position cannot represent a two-hundred-page manual, so split into coherent sections. And searching by words as well as by meaning: meaning blurs exact identifiers like part numbers and error codes, and keyword search catches precisely those. Run both, then re-score the top candidates with a second, more careful pass.

Do this: When answers are wrong, look at which passages were retrieved before touching the instruction.

Common mistake: Blaming the model for what the search failed to find.

Coordinates organise the haystack

Search is only the first thing coordinates are good for.

Once a pile of documents has coordinates, several things become nearly free that would otherwise be projects.

Grouping. Ten thousand reports fall into recurring themes nobody ever tagged, because similar ones sit together. Recurring problems are where useful ideas live.

Near-duplicates. The same issue reported twice in different words lands in almost the same place — deduplication without matching a single phrase.

Labelling without training. To classify a new item, find its nearest already- labelled neighbours and take their vote. Ten lines of code, no training, works from a handful of examples.

The odd one out. An item far from every group is unusual by definition. “Show me this month's entries least like anything we have seen” is a genuinely powerful review question.

Model calls are slow, expensive, and handle one item at a time. Geometry is instant and covers the whole pile. If an idea involves understanding a whole pile of documents, it is a coordinates problem first.

Do this: Ask what your largest pile of text would look like grouped by similarity.

Common mistake: Reaching for a model per item when the whole corpus could be organised at once.

Running the crew

An agent is a loop with tools

A model that can only talk becomes a model that can do.

On its own, a model produces text. Give it tools and a loop, and it acts.

The mechanism is simpler than the word suggests. You describe the tools available — search the documents, read this file, run this command. The model replies: call that one, with these arguments. Your code performs the action and feeds the result back. Round it goes until the job is done. The coding assistants people use daily are exactly this.

There is now a standard plug for the wiring: expose a tool once, in an agreed way, and any assistant can grip it, instead of a custom fitting for each one.

The caveats arrive with the capability. Errors compound across steps. Tokens burn fast, because every round re-sends the growing conversation. And something that can act needs limits on what it may do unsupervised.

Start with one plain request. Add the loop only when the task genuinely needs several steps.

Do this: For one task, write the list of tools it would need. If the list is one item, you do not need an agent.

Common mistake: Reaching for an agent when a single well-written request would have finished the job.

Stand on the lowest rung that works

Every rung up multiplies cost, delay and ways to fail.

There is a ladder. One request and one answer. A fixed chain of requests, each feeding the next. One agent with tools, deciding its own steps. A crew of agents coordinated by another. Each rung costs more, takes longer, and adds ways to go wrong.

The distinction worth learning: a workflow is a pipeline where your code decides the steps and the model fills them in — predictable, debuggable, cheap. An agent is a loop where the model decides what to do next. Workflows for known processes, agents for open-ended work. Most agent products that survive in production are, quietly, mostly workflow.

One piece of arithmetic explains most failures of ambitious systems. A step that succeeds ninety-five times in a hundred, run twenty times in sequence, produces a correct result about a third of the time. Reliability does not survive length. The answers are fewer steps and a check between stages. I have yet to meet a job where a crew beat a chain somebody had thought about properly.

Do this: Before building a crew, write down what one careful request cannot do.

Common mistake: Building the impressive version first, and discovering the arithmetic afterwards.

Scope the task, name the handback

A worker given “look into this” wanders. Given “return four fields” it delivers.

Think of each worker as a specialist hired for one task. It arrives with its instructions, does the work, hands back a result, and its memory is gone. Anything not in the handback is lost for ever. That single fact drives most of the design.

So write the work card properly. Scope small enough to finish in one sitting — one file, one chapter, one question. If a worker needs its own workers, the job was cut wrongly.

Then name the handback exactly. Between workers, prose is where information goes to die: a paragraph of summary loses precisely the details the next step needs. Say instead: return these fields, with these names, one entry per finding. The work card has always specified what gets recorded at the end; this is the same discipline.

Anything that must outlive a worker goes into a file or a database that others can read, never into a conversation, which evaporates.

Do this: Write the exact fields you want back before dispatching any task.

Common mistake: Asking a worker to “look into” something and receiving an essay you cannot use.

Never let the finder verify its own finding

The inspector is not the person who did the work. The rule transfers exactly.

Agents treat other agents' output as fact. That is the quiet failure mode of every multi-step system: one invented finding upstream becomes three paragraphs of confident analysis downstream, and nothing in between ever questions it. Garbage does not just propagate, it gets elaborated.

The counter is independent verification. Where a wrong finding is expensive — reviews, audits, safety claims, anything somebody will act on — send each finding to a separate sceptic whose only instruction is to refute it against the source text. Only findings that survive count.

The word separate is doing the work. A model asked to check its own output tends to agree with itself, in the same way and for the same reasons that people do. And a panel of identical checkers shares identical blind spots, so vary how they are asked.

This is the quality inspector, in software form, and it exists for the same reason: the person who did the work is the worst-placed person to find what they missed.

Do this: Add one independent checking step wherever a wrong answer would be acted on.

Common mistake: Asking the same model, in the same conversation, whether it is sure.

A human signs the release

The system can prepare the decision. A person releases it.

Decide in advance which actions require a person: anything irreversible, anything that leaves your organisation, anything expensive. Automation prepares the decision — gathers the evidence, drafts the assessment, lays it out — and a person releases it. That division is not caution. It is the arrangement that makes the automation usable at all.

Three limits belong with it.

Every loop gets a budget: a maximum number of steps, a maximum spend, and a defined path for giving up and reporting. Something unable to finish will happily burn money for ever. Nothing flies without a fuel gauge.

Anything reading untrusted text can be steered by instructions hidden inside it. Cap what it may do without confirmation, especially anything with write access.

And record every input, output and action. When a long run produces a wrong answer, “why did it do that?” is unanswerable unless it was all logged. Install the recorder before the incident, not after.

Do this: List the actions in your system that must never happen without a person, and make them impossible without one.

Common mistake: Discovering which actions needed approval by watching one happen.

Fault isolation

Reproduce it before you repair it

A fault you can trigger on demand is half fixed.

Debugging is fault isolation, and every habit that works in a hangar works here unchanged: reproduce, isolate, one change at a time, fix the cause and not the symptom. You are not learning a new discipline. You are applying one you have to a new system.

Start where you always start. Can you make it happen again, on demand? If yes, you can test whether your fix worked, which is the whole reason it matters. If it only happens sometimes, you have not found the real trigger yet — and “it is intermittent” is a description of your knowledge, not of the fault.

Write the exact steps down. Not roughly: exactly. Which record, which user, which order, which browser. Half of all intermittent faults become repeatable the moment somebody writes the steps down honestly, because the writing exposes the step everybody was doing without noticing. I have never regretted the ten minutes that took.

Do this: Write the reproduction steps before you touch anything.

Common mistake: Starting to fix something you cannot yet make happen, and never knowing whether you fixed it.

Read the error, bottom-up

It usually names the file, the line and the cause. Go and look.

When something impossible happens, the code raises an error. It travels back up through whatever called it until something handles it, or nothing does and the program stops, printing the trail of calls that led there.

That printout looks like noise and is not. It is a breadcrumb trail, and you read it from the bottom: the deepest line names the actual failure, and the lines above show how execution arrived there. It usually gives you the file and the line number. Go and look at that line before theorising about anything.

Reading a stack trace calmly is half of debugging, and it takes no programming ability at all — only the willingness to read something that looks unfriendly.

When you ask anybody for help, human or machine, give four things: the exact error text copied verbatim, the relevant code, what you expected against what happened, and what you already tried. Then say: diagnose the cause before proposing a fix. Vague descriptions get guesses; exact ones get diagnoses.

Do this: Copy the error verbatim and open the file and line it names.

Common mistake: Reporting that it is broken, and receiving a guess.

Suspect the last change first

It worked yesterday. What is different?

The most efficient question in troubleshooting is not what is wrong. It is what changed.

Version control answers it exactly: it will show you every line that differs from the last known-good state. Not what you remember changing — what actually changed, including the thing an assistant altered while you were looking elsewhere.

For something that broke further back, there is a tool that binary-searches your history: it checks out a point halfway back, you say whether the fault is present, and it halves again. A dozen answers finds the exact change among thousands. It is the same method you would use on a wiring run, applied to time instead of length.

This is also the practical argument for frequent small snapshots. When each one contains one small change, “what changed” has a short answer. When yesterday is one enormous snapshot, the tool can only tell you that something in it did it.

Do this: Before investigating, look at what changed since it last worked.

Common mistake: Theorising about the architecture when a one-line change three hours ago did it.

Find the fault by halving

Same logic as splitting a wiring run: check the middle, and you have halved the problem.

You know this method. You do not trace a circuit end to end; you check the midpoint, and whichever half is wrong, you halve that. Three or four checks isolate almost anything.

In software the equivalent is printing the suspect value at the midpoint of its journey. Correct there? The fault is downstream. Wrong there? Upstream. Repeat. Three checks find most faults, and each one is a single line you delete afterwards.

Four suspects come up again and again, and it is worth checking them before inventing anything more interesting. The data is not what you think — log the actual value, because it is empty, or missing, or text where you expected a number far more often than anybody believes. The outside part does not behave as you assumed — read its documentation for the two minutes it takes. Timing — something used a result before the slow step had finished. And drift between machines — a different version, a missing setting, a stale part.

Do this: Print the value at the midpoint before forming any theory.

Common mistake: Reading code for an hour instead of looking at one actual value.

Fix the cause, then trap it

One change at a time, and then a test so it can never come back unnoticed.

Two habits finish the job properly, and both come straight from the hangar.

One change at a time, retesting between changes. Shotgunning five fixes at once means that even success teaches you nothing: you do not know which one worked, whether two of them cancel out, or what the other four broke. It feels slower and it is not.

And fix the cause, not the symptom. Catching an error and hiding it makes the message go away and leaves the fault installed. Ask what allowed the bad value to exist in the first place.

Then trap it. Add the test that would have caught this fault. It costs a few minutes now, and that failure mode is permanently caught from here on — including by whoever, or whatever, edits this code in a year. This is how a system gets sturdier with age instead of more fragile: every fault that ever occurred leaves a permanent sentry behind it.

Do this: After every fix, add the test that would have caught it.

Common mistake: Five changes at once, and no idea afterwards which one mattered.

Learning and shipping

Build first, theory second

You learn this the way you learn an aircraft: on the floor, with the manual open.

Nobody learned an aircraft by reading the manual cover to cover. They learned it on the hangar floor with the manual open at the right page and a job in front of them. Software is the same, and the temptation to study first is the main thing that wastes people's first year.

So pick a real problem from your own work. Always your own — you can judge whether the result is right, which is the part no course can give you.

And use the machine as an instructor, not only as a mechanic. The highest-value request in the whole toolkit is: explain what you just wrote and why, line by line, to somebody who knows the basic concepts. Followed by: what are three ways to do this and their trade-offs? Every task becomes a lesson, at no extra cost.

Concepts over syntax. The machine writes the syntax now. Your job is the layer it cannot do: knowing what to build, how the parts should connect, and whether the output is right.

Do this: After the next thing gets built for you, ask for the line-by-line explanation.

Common mistake: Studying for months before building anything you care about.

The traps that claim beginners

Five traps, and every one of them feels like progress at the time.

Rewriting instead of debugging. When something misbehaves, the itch is to start over. Resist it. Debugging builds the skill you actually need; rewriting reliably reproduces the same fault in new words, three days later.

Tool tourism. A different set of technologies every project means being a permanent beginner in each. Pick a proven combination and stay on it at least six months. Depth compounds; novelty does not.

Tutorial paralysis. Watching courses feels productive and mostly is not. Enforce a ratio: one hour of learning to three hours of building.

Premature scaling. But will it handle a million users? You do not have ten. The simple version is correct at your size, and complexity bought for imaginary load is pure cost paid today.

Polishing before validating. Weeks perfecting features nobody asked for. Find out whether anybody wants it first.

I would add three operational sins to the list: accepting work unread, saving your progress rarely, and keys in the code.

Do this: Name which of these five you are doing right now.

Common mistake: Believing the trap you are in is the exception.

Mine the work you already know

The best problems are boring, specific, and already annoying somebody daily.

Innovation rarely comes from better technology. It comes from insight into a real problem multiplied by the ability to ship — and shipping just got cheap for everybody, which leaves the insight as the scarce half.

Insight is what you already have and a general-purpose developer does not. Nobody outside your trade knows which five-minute task everybody does eleven times a day. Hunt workflow pain, not app ideas: the spreadsheet everybody curses, the document nobody can find, the question asked five times a day. Boring pain, reliably felt, beats any brainstorm.

Three patterns are worth hunting in any industry. Turning documents into answers you can query, which is the largest open field right now. The queue behind one scarce expert, where a machine does the first pass and the expert signs. And glue work — anywhere a person retypes data from one system into another.

Your own trade is the first domain, not the ceiling. The instincts transfer: any regulated, document-heavy, safety-critical industry works the same way.

Do this: Write down three pieces of pain from your own week, not three app ideas.

Common mistake: Looking for a clever idea when a boring, daily irritation was the better business.

Five conversations before you build

Polite enthusiasm is worthless as evidence.

Talk to five people who actually have the problem before you build anything. Not to pitch — to find out whether the pain is real, frequent, and worth money. Those are three separate questions and an idea can pass one and fail the others.

Ask about the last time it happened, what they did instead, and what it cost them. Past behaviour is evidence; enthusiasm about a future product is not. Everybody is encouraging about an idea that costs them nothing.

Then build the crude version in a weekend and put it in front of one real person — even if that person is you. Watch what they actually do with it. Real usage steers better than any amount of planning, and a running ugly prototype beats a perfect plan every time.

Ask for money, or at least a concrete commitment, early. The moment you ask, the polite answers stop and the honest ones start.

Do this: Have five conversations about the pain before writing anything.

Common mistake: Treating “that sounds great” as validation.

Messages from visitors