The Compiler That Never Says No
English is now a programming language whose compiler runs unclear instructions as readily as clear ones. Clear writing is the scarcest skill in technical work.

“Besides a mathematical inclination, an exceptionally good mastery of one’s native tongue is the most vital asset of a competent programmer.” - Edsger Dijkstra, 1975
Last week I read a New York Times guest essay by Joe Cruz, who chairs the philosophy department at Williams College, and it bothered me. Titled “I’m a College Professor. Writing Isn’t as Important as We Think.”, it went on to argue “What if we are mistaken in treating writing as the highest cultivation of thought?” “The evidence that supports this belief is thin at best.”
Cruz is asking how students learn to think, but he misses a much broader question. How do we most effectively translate thoughts into structured arguments, and ultimately use that to communicate? We work in very different fields, but I, too, spend my days wrestling with this problem. Since large language models went mainstream, people have debated about how much prompt engineering, which is mostly the art of communicating clearly, improves results. Today this question occupies a good deal of my wandering mind.
Those reflections began, as most of mine do, in a fit of mild exasperation. It had taken me five drafts to coax Claude into writing a succinct 150-word technical post.
On the surface, none of these five drafts was bad. Each was a fluent, confident answer to the request I had made. But with each new prompt, the request I gave changed, because I hadn’t decided precisely what the post was meant to communicate. As a result, the drafts did not improve in the linear fashion I had expected. At every step along the way, Claude encouraged me along, never pausing to complain that my brief was unclear. It wrote something plausible each time and waited for my next note. The drafts improved only as my own thought process iterated on the idea of the post, and could explain its goals more sharply.
Blaise Pascal described this same problem in 1656, apologizing to the reader of one of his Provincial Letters: “The present letter is a very long one, simply because I had no leisure to make it shorter.” Shortening takes time because it forces you to decide what matters. A model can give you unlimited drafts, but it can’t make that decision for you, and my five drafts were the time it took me to make it.
Since then, I’ve paid deliberately close attention to the writing that crosses my desk, and I have come to suspect that Professor Cruz has the matter almost exactly backwards. Working with AI has made writing part of programming, and the link between fast, productive engineering work and an ability to write well has never been stronger. To understand why, it helps to take a brief historical detour.
Every rung promised the end of the hard part
The history of programming is a ladder of abstraction. Machine code gave way to assembly, assembly to FORTRAN and the other early high-level languages, and those to the languages most software is written in today. Each rung let a programmer say more with less, and each arrived with a promise that the hard part was about to disappear. IBM’s 1954 report proposing FORTRAN said that “FORTRAN should virtually eliminate coding and debugging.” In 1981 a British magazine introduced a program called The Last One as “a program which could become the last one ever written by a human being.”
Neither came true. Each rung removed a kind of work, but the hard part simply moved further up the ladder with the programmer. FORTRAN freed people from managing the machine and left them the more challenging job of stating the problem precisely. John Backus, who led the FORTRAN team, looked back in 1978 and admitted they had been “hopelessly optimistic in 1954 about the problems of debugging FORTRAN programs.”
Grace Hopper wanted the top rung to be English. “What I was after in beginning English language,” she said in a 1980 oral history, “was to bring another whole group of people able to use the computer easily. The people that used English, not symbols.” Her FLOW-MATIC replaced symbols with English words and became one of the ancestors of COBOL. She met resistance for years. “I had a running compiler and nobody would touch it,” she recalled in 1986, “because, they carefully told me, computers could only do arithmetic, they could not write programs.”
Edsger Dijkstra, never one to keep an opinion to himself, thought the whole project a mistake. In a 1975 list he called “How do we tell truths that might hurt?”, he wrote that “the use of COBOL cripples the mind” and that “projects promoting programming in ‘natural language’ are intrinsically doomed to fail.” Yet tucked into the very same list is the line I stole for this essay’s epigraph.
Those claims seem to contradict each other until you see what Dijkstra was worried about. His target was the idea that English-looking syntax would make programming easy. He thought the hard part of programming was knowing exactly what you meant, and that only people with real command of a language could do that reliably. Formal languages, he argued three years later in “On the foolishness of ‘natural language programming’”, are “an amazingly effective tool for ruling out all sorts of nonsense that, when we use our native tongues, are almost impossible to avoid.”
In January 2023, Andrej Karpathy posted: “The hottest new programming language is English”. On the question of whether this rung could be built, Dijkstra, clearly, has lost. On the question of what it would demand of those who climb it, I suspect he will be proved right.
The first compiler that never says no
For most of its history the computer was a fussy correspondent: every compiler sent back what it could not understand, and every formal language carried a subtler discipline within it. A compiler couldn’t tell whether a program did what you meant. But a formal language couldn’t take a gesture in place of a decision. There was no way to tell C to handle the edge cases sensibly. Any case you wanted handled had to be written out, and writing it out meant deciding on what you wanted.
A large language model is a more obliging correspondent: it never refuses, and it takes the vague version without complaint. Hand it an ambiguous instruction and it returns the most plausible program, analysis or essay it can, written with the same confidence as if it were given a precise one. Ambiguity no longer fails loudly. It resolves silently into a default, and the default is whatever is most typical of the material the model learned from: roughly speaking, the average of everyone who has written something similar.
The ladder has also turned over. The top rung now composes and writes the lower ones: describe a system in English and a model will write it in Rust. The compiler underneath still refuses malformed code, but it can’t tell that the code isn’t what you meant.
The gap nobody can see
It will be objected that I overstate the case, since models grow more sophisticated at asking clarifying questions every month. That is true, so far as it goes. But a question requires someone to notice a gap in the first place.
To a model, missing context registers as a default. A model will ask about ambiguities it recognizes, like which quarter you meant, because those are ambiguous for everyone. It won’t ask about the ones specific to you: that your team defines churn its own way, that the customer already rejected the obvious answer, that the audience has read your last three memos. It has no way of knowing those facts exist, so it has no reason to ask. It fills the space with the likeliest answer and moves on, and nothing in the output tells you it did.
Part of the problem is shared context, which linguists call common ground. Herbert Clark and Susan Brennan observed that in conversation, people “cannot even begin to coordinate on content without assuming a vast amount of shared information or common ground.” Between colleagues, “do the usual cut” can carry a paragraph of meaning, because it’s an experience lived multiple times by both parties. Between you and a model, the only shared context is what you have written down. Writing for a model means decompressing: putting on the page everything you would normally leave to shared history.
Closing that gap with technology would require two things we don’t have. The first is a record of the context that mostly never gets written down. Scanning the world’s written materials was hard, but at least the books existed. Most of what a team knows lives in people’s heads. The philosopher Michael Polanyi argued in 1966 that much of it can’t even be fully put into words: “we can know more than we can tell.” The second challenge is the depth of context that can be held. The frontier models from Anthropic, OpenAI and Google top out around one million tokens, roughly 550,000 to 800,000 words. Researchers estimate that a child has heard up to 100 million words by age 13, on the order of a hundred times that, and Anthropic’s own documentation warns that “as token count grows, accuracy and recall degrade.”
Why files and rules don’t fix it
The obvious remedy is to write everything down once and be done with it, your context codified. I have tried. I keep a 200-line file that tells Claude who is on our team, what we’re building and how I like things written, and a separate file of about 90 lines of writing rules. They help, though not quite as much as I had hoped.
Context files go stale, and a model can’t tell. One of our engineers put it well in our engineering guide: “An AGENTS.md that describes code that no longer exists is worse than no file, because an agent will act on it confidently.” A context file describes the world on the day it was written, and the model treats it as true today.
Rules pile up like case law. Each of my writing rules fixed a specific past mistake. Together they collide: be concise but specific, cut the dashes but keep the rhythm. A model applies them literally and all at once, and some drafts now bend around my rules in ways no person would. The rules record what I disliked but say much less about what I want.
Armies learned this long ago. Detailed orders break on contact with events, so U.S. Army doctrine asks commanders for intent: a concise statement of purpose that helps subordinates “act to achieve the commander’s desired results without further orders, even when the operation does not unfold as planned.” A rulebook is a list of old orders. What a model needs for each piece of work is the intent: what this is for, who it’s for and what good looks like. Files can supply the background but the intent must be provided fresh for each new task, and it can only come from the human who writes it.
Why managers are good at this
If files cannot close the gap, it must instead be the writer. The last few weeks have convinced me that some of us were trained for this task long ago without knowing. Those people have spent years writing for readers who don’t share their full context, who had to be corrected regularly (and therefore benefited from improved instructions), and who would act confidently on tasks once they were given them (regardless of whether they should have been confident). For most of working history, two groups got that practice.
The first is managers. Management is a job built around handing intent to someone who doesn’t share your head, over and over, and living with the result. A vague request to a person doesn’t stay hidden for long: the wrong deliverable comes back a week later, and you learn what you left out. A good manager states the goal, the reason behind it, what good looks like and what to ignore.
Ethan Mollick, who teaches at Wharton, saw this in a class he taught this January. He gave his students four days to build a startup with AI. Most were executive MBAs working as doctors, managers and company leaders, and few had ever coded. He wrote that “in figuring out how to give these instructions to the AI, it turns out you are basically reinventing management,” and concluded: “My students figured this out in four days. Not because they were AI natives, but because they already knew how to manage.” Anthropic’s prompting guide makes the same point from the other side, telling you to “think of Claude as a brilliant but new employee who lacks context on your norms and workflows.”
The second group is engineers, who got the same training through a different loop. Formal languages made them write out every decision, and code reviewers kept asking what a function was for until the answer was obvious from the code. That is the practice Dijkstra had in mind when he ranked command of one’s native tongue above nearly everything else.
People who never had either loop are at a disadvantage they may never notice, because the model hides it from them. Their vague briefs come back as fluent drafts, and nothing tells them the draft is an average rather than an answer.
What a clear brief does
Four disciplines separate a brief that works from one that returns an average. None of them is new; a competent Victorian civil servant would have recognized (and indeed maybe even excelled at) all four.
It decides the point before any writing starts. In practice, you can state the argument or the decision in two sentences before you write anything for the model. This matters because a model will commit to some point whether or not you have chosen one, and it will choose the most common one. Skip it and the first draft becomes the thinking: you end up editing toward an idea you haven’t formed, which is how I spent five drafts on 150 words. None of this rules out using a model to stress-test the idea first. Just keep that conversation separate from the brief.
It gives an example before an abstraction. An adjective gives a model nothing to aim at. “Make it more sophisticated” was one of my earlier notes to Claude, and I got a kind of sophistication I didn’t want. One sample of the output you want carries context that a paragraph of description can’t, and if you can’t produce the example, you probably don’t understand the request yet.
It uses one name for each idea. A model, like any new reader, can take two words to mean two things. Call someone a customer in one paragraph and a client in the next, and a model may treat them as separate or invent a distinction you never meant. Engineers learn this as naming discipline in code, and it holds just as well for prose.
It says what the writer doesn’t know. A brief that marks its open questions lets the reader treat them as open. A brief that doesn’t gets its gaps filled with confident defaults. The same habit earns trust from human readers. Our engineers’ pull requests often say what a change doesn’t fix, and it makes the rest of the description easier to believe.
One test covers all four, and Anthropic’s guide calls it the golden rule: “Show your prompt to a colleague with minimal context on the task and ask them to follow it. If they’d be confused, Claude will be too.”
How to write now
The best company rule I’ve seen on writing with AI is Clay’s AI writing policy, which went viral in August. It allows AI for drafting, requires that “you must stand behind every idea and every sentence in your docs,” and holds that “more time should be spent authoring a document than consuming it.” It even, I was pleased to see, quotes Pascal. Its most revealing line is almost an aside: “If you are producing a longer piece of writing from a shorter prompt, consider instead just sharing the prompt itself.”
I’d take the prompt line further than Clay does. If the brief carries all the decisions, the brief is the document, and it deserves the care we used to give the draft. I’ve mostly stopped writing first drafts. I write briefs that follow the four disciplines above. Then I ask the model to attack the brief. It can only find the gaps anyone would notice, but answering its questions makes me reread the brief as a stranger would, and that is usually when I find the gaps specific to my work. I revise, let the model draft, and then edit hard. Pascal’s step still belongs to me: the cut is where I find out whether the brief really decided what matters.
The strongest objection is that the first draft is where the thinking happens. Ted Chiang put it memorably: “Using ChatGPT to complete assignments is like bringing a forklift into the weight room; you will never improve your cognitive fitness that way.” Mollick, as enthusiastic about AI as anyone, still writes every post as “a full draft entirely without any AI use at all.” Chiang is right that the lifting has to happen somewhere. I think it happens in the brief, for students as much as for anyone, because a brief has nowhere to hide an undecided idea, and a draft can cover one with a good paragraph.
That is where I part ways with how most universities have responded. Nearly every student now uses AI: a 2026 survey of UK undergraduates found that 94% use generative AI to help with assessed work. Schools have answered by moving assessment back into the room, in what Cruz calls an “arms race of countermeasures.” Cruz himself tried something better. Last semester he let students use AI on their final essays, then gave each of them a 45-minute oral exam, and found that “at least half of the class was more agile, more reflective and clearer during oral exams than they were in their papers.” He reads that as a sign that writing matters less. I am not convinced. To me, it is the result you’d get from students who had worked their thoughts out on the page first.
What comes next
In a twist of fate Cruz did not expect, we are about to see this battle between speaking and writing play out in the world. Voice tools like Wispr Flow let people speak their briefs instead of typing them, and speaking is fast: in one controlled study, people entered English text nearly three times faster by voice than on a keyboard. Spoken briefs will inevitably carry more context, since talking is how we are used to explaining things to people. However, I expect their content to be looser, because speech leans on a listener who can interrupt, and a model won’t. The first study I’ve found on spoken prompts, published this summer, reports that input modality “significantly influenced prompting behaviour but did not lead to measurable differences in subjective evaluations.” Which will prevail?
Professor Cruz enlists Socrates on his side, so it is worth recalling what Socrates actually complained of. Writing, he told Phaedrus, is like painting: its creations have “the attitude of life, and yet if you ask them a question they preserve a solemn silence.” The page cannot explain itself, so the writer has always had to do the explaining in advance. Language models make that burden heavier, because for the first time we are writing for a reader that acts on what we say and seldom stops to ask what we meant.
Professor Cruz thinks we have overrated writing. I think we are about to find it rather more important than we thought.
Sources
- Edsger W. Dijkstra, “How do we tell truths that might hurt?” (EWD498), 1975.
- Joe Cruz, “I’m a College Professor. Writing Isn’t as Important as We Think.”, The New York Times, September 29, 2026.
- Blaise Pascal, The Provincial Letters, Letter XVI, December 4, 1656.
- John Backus et al., “Preliminary Report: Specifications for the IBM Mathematical FORmula TRANslating System, FORTRAN”, IBM, 1954.
- David Tebbutt, “The Last One”, Personal Computer World, February 1981.
- John Backus, “The History of FORTRAN I, II, and III”, ACM History of Programming Languages conference, 1978.
- Grace Hopper, oral history interviewed by Angeline Pantages, Computer History Museum, 1980.
- Grace Hopper, interview with The New York Times, 1986, quoted in “Happy 109th birthday to Yale alumna Grace Hopper, a pioneer in computer science”, Yale News, December 9, 2015.
- Edsger W. Dijkstra, “On the foolishness of ‘natural language programming’” (EWD667), 1978.
- Andrej Karpathy, post on X, January 24, 2023.
- Herbert H. Clark and Susan E. Brennan, “Grounding in Communication”, in Perspectives on Socially Shared Cognition, APA, 1991.
- Michael Polanyi, The Tacit Dimension, 1966.
- Warstadt et al., “Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora”, CoNLL 2023.
- Anthropic, “Context windows”, Claude Platform Docs.
- U.S. Army, ADP 6-0, Mission Command: Command and Control of Army Forces, July 2019.
- Ethan Mollick, “Management as AI superpower”, One Useful Thing, January 27, 2026.
- Anthropic, “Prompting best practices”, Claude Platform Docs.
- Sophie Alpert, “Why Clay Has an AI Writing Policy”, Clay, August 19, 2026.
- Ted Chiang, “Why A.I. Isn’t Going to Make Art”, The New Yorker, August 31, 2024.
- Ethan Mollick, “Against ‘Brain Damage’”, One Useful Thing, July 7, 2025.
- Rose Stephenson and Charlotte Armstrong, “Student Generative Artificial Intelligence Survey 2026”, HEPI, March 12, 2026.
- Ruan et al., “Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones”, Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2018.
- Rathore, Billakanti, Ceha and Henderson, “A Comparison of Speech and Typing Input for Creative Generative AI Tasks”, CUI 2026, July 2026.
- Plato, Phaedrus, translated by Benjamin Jowett.