Why viral AI demos impress by optimizing the unconstrained, and why real software starts when the constraints show up.

Picture the kind of clip that goes viral: someone types one sentence into a state-of-the-art large language model (LLM), “make me a neon space shooter with a boss fight,” and a minute later there's a playable game. It has particles, sound, a health bar, and a boss that changes attack patterns at half health. The replies write themselves: programmers are finished; anyone can build software now.

The demo is real. The model really did that. But the conclusion people draw depends on a quiet trick in how the task was set up. Once you see the trick, the demo tells you much less about software engineering than it appears to. It also points to what actually matters now: not the model but the prompt, which has quietly become the new programming language, and the person writing it.

Nothing to Get Wrong

“Make me a cool game” has no wrong answer. The person asking has no picture of the result beyond a vibe, so any competent output lands as a success. Controls, art style, enemy spawning, what happens when you die: the prompt leaves all of it open, and the model fills every gap with its own defaults. Those defaults are good because they’re distilled from thousands of similar projects.

That's not nothing, but it's a specific kind of success: optimizing the unconstrained. With no requirements, the model chooses the objective as well as the solution, and it naturally chooses objectives it's already good at. The viewer grades the result against nothing, and it passes.

“Make me a cool game” has no wrong answer.

Engineering is the opposite problem: optimizing the constrained. The file format is fixed, the performance budget is fixed, the old API isn't going anywhere, and the solution has to satisfy all of it at once. A demo that removes the constraints hasn't solved a harder version of the problem. It has quietly swapped in an easier one.

When the Target Closes

Now change the request. The game has to load levels from a file format your team already uses. Save files must stay compatible with last year's version. It has to hold 60 fps on a low-end Android phone, talk to a specific backend API, and support mods through a defined plugin interface.

Maintainability is one of these constraints too, not a universal goal. Some software is built to be maintained for years; some is written to be thrown away, and writing it that way on purpose is the right call. Nothing in “build me X” says which one you're making, so the model can't know unless someone tells it.

Nothing about the model changed. What changed is that the answer space collapsed from “anything plausible” to “the one thing that fits.” Every silent default is now a potential violation.

A user who states these constraints can often get a strong model to honor them. The problem is the user who doesn't know these are the things that need stating, who describes the game and assumes the rest will sort itself out. For them the model fills the gaps exactly as before. Now the guesses can be wrong, and nothing on screen says so.

What the Job Actually Is

This is the heart of it: software engineering was never mainly about writing correct code in a programming language. Syntax is the easiest part to look up, and it's the part LLMs have most thoroughly commoditized.

The actual job is decision-making under constraint. That means deciding what not to build, how the pieces relate and what owns what, and which trade-offs are acceptable: speed versus correctness, flexibility versus simplicity, shipping now versus maintaining later, and whether this code even needs a later.

None of these decisions appear in a casual prompt. When nobody makes them, the model makes them, invisibly and generically, tuned to the average request rather than your situation. The defaults become the architecture.

The defaults become the architecture.

Defaults That Don't Fit

Take something trivial: ask for Python code that reads user input and does something with it. You'll often get a wall of validation, exception handling, and edge-case coverage for a script maybe three people will ever run. An experienced engineer would write a fraction of that. It isn't laziness; they know which failures matter here and which never will. Knowing what to leave out is a subtraction skill, and it comes from having shipped and maintained real systems.

The reverse happens too, and it's more dangerous. The same model that overarmors a toy script can hand you a web app with no rate limiting, secrets in the source code, or an authorization check that trusts the client.

Understood.

Neither is a failure of capability. The model could have gotten it right. It just didn't know what “right” meant here, because nobody said.

The Robot in Your Head

Image generators make the pattern easy to see. Ask for “a cool robot” and you can get ten thousand striking designs. Try to get the specific robot you've already sketched: with its exact proportions, the way its shoulder plates overlap, that particular shade of rust. It becomes a fight, often harder than drawing it yourself. Generative models are superb at “surprise me well” and far weaker at “match what I already have in mind.”

Real software is almost always the second case, with one crucial difference. With the robot, anyone can see it's wrong. With code, the mismatch is usually invisible.

The app runs, and the tests the model wrote pass. The problem surfaces three months later as a corrupted save file, a security incident, or a feature that can't be added without rewriting everything. Evaluating the output takes the same expertise as specifying it.

Built for the Average Case

Take a handful of exceptional engineers from different corners of the industry. One builds engines for consoles: fixed hardware, a hard memory ceiling, every cycle counted. Another ships on PC, where the same code has to run across thousands of hardware configurations. A third has maintained an operating system kernel for decades, where breaking existing software is the cardinal sin. A fourth runs backend services where uptime and observability matter more than raw speed.

Their work looks very different, but not because they're different people. It's because they faced different constraints. Put any of them under the same constraints, with the same plan, and they'd converge on similar software. That's what engineering discipline is: the constraints drive the design, not the personality.

An LLM doesn't come with any of those constraints built in. What it has is its training data, and that data is dominated by the typical case. For every line written by an engineer squeezing a console to its limits, it has seen vastly more ordinary code written under no particular constraint at all. So when nothing tells it otherwise, it produces what satisfies the typical request: reasonable, common, middle-of-the-road code.

That code isn't bad. It's just aimed at nobody's actual situation, and the problem doesn't go away as models improve. A more capable model raises the ceiling of what it can produce when the constraints call for it, but the default still tracks the typical case, because the typical case is what an unconstrained request looks like. The constraints live in your head, not in the weights.

The Prompt Is the New Programming Language

If the model's defaults become the architecture whenever nobody intervenes, then the prompt is where the architecture gets decided. It's how you tell the model which constraints it's working under, and that makes prompting, done well, the most important engineering act in the whole process. In practice, the prompt has become the new programming language.

The comparison is useful precisely because of where it breaks. We've climbed this ladder before. Assembly gave way to C, and C to languages like Python, and each step let programmers say more with less: fewer lines, less detail, more of the busywork handled for them. Prompting looks like the next rung, where you state intent in plain language and the model fills in the code.

The difference is ambiguity. Different compilers can produce different machine code from the same C, but the language itself defines what the program means, so there's a firm contract underneath. Natural language has no such contract. The same sentence can be read a dozen ways, and the same model can pick a different reading on the next run, to say nothing of a different model. The less you write, the more is left to interpretation.

So how do you make a prompt deterministic? You can't, fully. What you can do is shrink the space of valid interpretations until almost every output the model could produce is one you'd accept. Every constraint you state rules out a whole family of wrong programs. That's what prompt engineering actually is: not magic phrases but the work of pushing an ambiguous language as close to an unambiguous one as you can get.. It's specification: stating interfaces, constraints, and invariants in English instead of code. The skill of software engineering didn't disappear. It moved up a layer, from writing the code to writing the specification the code is generated from.

The goal isn't to spell out every line, either. A prompt detailed enough to remove all ambiguity would just be the code written in English, which defeats the purpose, the same way nobody picks Python to write C-level detail. Good prompt engineering lives in between: precise about the constraints that matter, quiet about the details that don't. You keep the leverage of writing less without handing the important decisions to a guess.

It's also why oversight still matters. Compiler output isn't beyond question either, and people do inspect it when it counts, but the language gives everyone a shared definition of what the code should do. Model output has no such definition behind it, so someone who understands the problem needs to read it, test it, and judge it. One way to change that would be to step back and design an LLM-specific programming language, precise enough that the same program yields the same code whichever model runs it. If that ever happens, writing in it will probably look a lot like programming again.

This is the flip side of optimizing the unconstrained. A vague prompt leaves the space wide open, and the model chooses. A precise prompt closes it until the only thing left to choose is what you already meant.

Compare two requests for the same feature. One asks the model to add a save system and stops there. The other states which format saves must use, how far back old saves must still load, where save logic lives so storage can change later, what gets validated and what doesn't, and which dependencies are off-limits.

Both will produce working code. Only the second produces code that survives contact with real users, and every clause in it encodes a decision that came from experience: formats evolve, users do unexpected things, storage backends change, validation has a cost. The model didn't lack the ability to write any of this. It lacked the reasons.

An experienced engineer's prompt comes preloaded with fences like these, and the model's defaults land inside them. A newcomer's prompt has no fences. They aren't careless; they just don't yet know where fences go. Same model, same weights, opposite outcome. The differentiator was never the AI. It's the mental model the human brings to constrain it.

Two-Way Traffic

None of this means writing one perfect prompt and walking away. Constraints travel in both directions. The engineer states what matters, the model produces something, and the result talks back: a test fails, the compiler rejects a type, a profiler shows the frame budget blown, a linter flags a pattern the team banned. The engineer reads that feedback, decides what it means, and tightens or loosens the constraints for the next round.

Automated checks are a powerful part of that loop, but they don't replace the engineer in it. A test suite enforces only the constraints someone thought to encode. Tests the model writes for its own code mostly confirm its own assumptions. And a failing check says something is wrong, not which constraint was misunderstood. Deciding what the checks should check takes the same judgment the prompt needed in the first place.

The same goes for how a system grows. Good architecture is rarely built in one pass. An engineer usually builds the main pillars first (the core data model, the main loop, or the riskiest integration) tests them against reality, and only then fills in detail. Some constraints only show up along the way: a client changes a requirement, a measurement kills an approach, or an integration behaves differently than documented.

Models tend to attempt everything in a single step. With limited context and no view of constraints that haven't surfaced yet, that one big step often misses the mark, and a wrong foundation gets buried under a lot of plausible detail. An engineer paces the work, deciding what gets built now, what waits, and when complexity is allowed to grow. That pacing is a constraint too, and it's one the model can't set for itself.

This Isn't Gatekeeping

None of this argues that non-engineers shouldn't build things. Plenty of people are making genuinely useful tools for themselves, like a script that tidies a spreadsheet or a small app for a club, where “it works for me” really is the whole spec. That's a real and good change.

The critique is aimed at the pitch, not the people. “Anyone can build any software now” is a promise the demos don't support, because the demos are structured to avoid exactly the part that's hard.

When someone takes that promise at face value and hits a wall on a project with real requirements, they haven't failed. They were sold the wrong picture of where the difficulty lives, and nobody told them that guidance was the missing piece.

Closing the Gap

The missing ingredient is expert judgment in the loop. The practical question is how to get it there without an engineer typing every prompt, and there are a few concrete paths.

The first is encoding judgment where the model will always see it. Constraints like “never break save compatibility,” “no allocations inside the frame loop,” or “treat every modder-supplied file as untrusted” can live alongside the code as standing project instructions, and as tests and continuous integration (CI) checks that fail when they're violated. Written once by someone who knows why they matter, they shape every request after that, including requests from people who would never have thought to state them.

The second is tooling that asks instead of assuming. Picture a newcomer asking an agent for a save system. Instead of silently picking defaults, the agent replies with a short plan and the assumptions behind it: saves need to load only on the current version, nobody edits them by hand, the game runs only on desktop. Each assumption is a decision made visible. The newcomer may not know the right answer, but now they know a question exists, and they can take it to someone who does. Some coding agents already lean this way with planning modes and clarifying questions.

The third is structure. Engineers set the pillars and the order of work, and others build within that frame. Expert judgment stops being a bottleneck on every change and becomes the framework everyone else works inside.

The interesting question isn't whether AI can write software. It's how to get the judgment software requires into the process without requiring everyone to become an engineer first. This isn't about who's allowed to build software; it's about where the difficulty has always lived. One-shot demos are impressive because they optimize the unconstrained, routing around the hard part rather than solving it. Real software begins the moment the constraints show up, and deciding which ones matter, maintainability included, is still the engineer's job.

So the biggest lever isn't a better model. It's a better prompt, and a tighter loop around it. The prompt is the new programming language, and like every language before it, it's only as good as the person writing it and what they understand about the problem.

Understood.

What a Constrained Prompt States

Asking a model to “add a save system” and stopping there leaves every decision to the model's defaults. A constrained version of the same request spells out which format saves must use, how far back old saves must still load, where save logic lives so storage can change later, what gets validated and what doesn't, and which dependencies are off-limits.

Each clause encodes a decision that came from experience: formats evolve, users do unexpected things, storage backends change, and validation has a cost. Both prompts produce working code; only the second produces code that survives contact with real users.

Understood.