AI made me into a manager against my will
Epistemic status: Personal history anchored to dated records (chat exports, git log, daily notes). Public launch dates are from the vendors' own announcements. My own first-use dates are given to the month and are the earliest surviving evidence, so they can only be later than the truth, not earlier.
AI usage: The prose and the opinions are mine, written without AI. Claude Code reconstructed the timeline from my chat exports and git history, verified the public dates against the vendors' announcements, fact-checked the draft, wrote the footnotes and fixed typos and grammar.
I have recently noticed that the way I have been "coding" for the last year is much more similar to the way a manager at a software company would "code" than to the way a software engineer would. What's most striking to me is that I never explicitly decided to switch, it just gradually happened (like the Ship of Theseus). This is especially sad, because I've discovered that deep focus while coding actually felt very good, but I don't get it anymore when managing agents, especially when there are multiple of them. So I am curious to reconstruct the timeline of how this happened. With the way AI development is speeding up, it is very hard to remember accurately how things were at a given point in time in the past.
Before I proceed with the actual reconstruction, I want to reflect on "managing" AIs. I do feel like in my software job I was effectively my manager's Claude Code, but there were huge differences that made the experience very different for them. Even when I started and didn't know much:
- the code review was mandatory, so a more experienced engineer would read my code anyway
- there was an amazing testing system, so the majority of bugs were caught automatically
- I wouldn't do any catastrophic actions when I was unsure and would proactively reach out for guidance.
With AIs, (1) you basically cannot have, unless you ask a different agent to review the code. (2) you can build yourself using agents. (3) agents are getting better at this. So the way I manage agents now is much more frustrating than my manager managing me.
For context, my first usage of ChatGPT was in December 2022 (it must have been GPT-3.5).I think I played with the Playground via the API before then, and I vaguely remember that the first model I interacted with was GPT-3: you had to do prompt engineering, and overall it was extremely sensitive to the words you were using (e.g. making typos would decrease the quality of the output, i.e. it was more like autocomplete).
I think my manager metamorphosis started when I stopped using Google search for coding questions and used AI instead. According to my export of ChatGPT and Claude threads, the first example is from December 2022, when I asked a CSS question (based on how notoriously illogical CSS is, I am not surprised). It couldn't code that well yet, but you didn't need to go through Google search results manually. I was still worried about hallucinations back then, but in type-safe coding it wasn't a big deal. It is interesting that hallucinations just became not a thing anymore at some point.Funnily enough, I actually posted the entire code in question in the same session, but I don't count this as pasting the entire code yet, since this was CSS.
At some point I remember trying Copilot and it was terrible.The main issue was that it would autocomplete parameter names or function names in calls with something that looked extremely reasonable, but was still wrong. Eventually you would catch that, but it was still wasteful.
Then AI became better (or my image of it became more realistic) and you could actually paste entire pieces of code and it would sometimes find bugs. The first record of this is from March 2023. I didn't like this phase much, because integrating the AI's response was tedious, cumbersome and wasteful. At this point you still had to code; AI was mostly there for when you were stuck.
Soon I discovered Aider (fortunately my Aider commits are labeled as such, and they start in August 2024 and end in July 2025).In my mind Aider is an obvious ancestor of Claude Code, which was released almost two years later, but I have no clue whether the creators of Claude Code would agree with that.The main difference from Claude Code was that you had to use your own API key (but this meant you could use the best available model) and you had to actively choose the files to work with: files the AI was allowed to edit and files for it to read. This way you didn't have to copy the code into the chat and then integrate the response back into the code. I still did that sometimes, since API costs were high compared to a chat subscription (there were even sites like uithub.com that converted your entire public GitHub repo into one huge text file, so that you could feed it to the chat as context),but I would give the chat response to Aider and let it integrate it.This stage is actually fairly similar to how I am using harnesses today. I think the main difference is that the AI wouldn't discover context on its own as well, so it would make more mistakes. Also it wasn't as smart, so you couldn't delegate as much. And it couldn't do arbitrary tool calls: it had predefined commands like /run, /test, /lint, /git and /web, and apart from linting after each edit it didn't run them unless I asked. Shortly after I started, it learned to suggest shell commands, which I still had to confirm.
I was quite happy with Aider, but Claude Code was becoming bigger and bigger, so I decided to give it a go. According to my logs, I tried both Claude Code and Codex CLI in July 2025 (funnily enough, I asked ChatGPT how to install Claude Code). I remember that Codex CLI felt extremely weird: it kept asking for much more permissions compared to Claude Code (apparently the TypeScript CLI's default mode, "suggest", made you approve every write and every shell command), so I gradually switched to Claude Code.I also tried Gemini CLI at some point and it was even worse than Codex CLI.
It is crazy to think that I have been using Claude Code for only ~15 months. In my mind it felt like years.
Nowadays, I use both Claude Code and Codex. Codex is very good for /fast mode and back and forth. Claude Code is good for writing, taste and anything deep. I would copy an exception log from my website into Codex and it would fix it and deploy the fix. I even keep wondering whether I could just let exception logs go directly to Codex without me having to copy them.
My AI adoption overview
Tool | Publicly available | My earliest evidence | Delay |
|---|---|---|---|
GPT-3 Playground | API beta June 2020; no waitlist November 18, 2021 | undated memory, before ChatGPT | unknown |
ChatGPT | November 30, 2022 | December 2022 (first coding question the same month) | under a month |
GitHub Copilot | preview June 29, 2021; GA June 21, 2022 | undated memory | unknown |
Pasting code into chat | - | March 2023 (December 2022 if CSS counts) | - |
Aider | May 11, 2023 (Show HN) | August 2024 | about 15 months |
Claude Code | preview February 24, 2025; GA May 22, 2025 | July 2025 | about 4 months after the preview, under 2 after GA |
Codex CLI | April 16, 2025 | July 2025 | under 3 months |
Gemini CLI | June 25, 2025 | July 2025 | under a month |
As I discovered AI more and more, I kept noticing how much value it was providing, and my adoption lag was becoming shorter and shorter. Interestingly, after Claude Code and Codex I haven't tried anything substantially new, but I am not sure there was anything at all to try.
Against my will
You may object - no one forced you to use AIs, you are free to code manually or even in 0s and 1s if you like. And my honest answer is that I don't want to do it. It feels inefficient. I do care about results and my time. Before AIs I would code manually without knowing how inefficient it was, and as a side product I would have deep work flow and enjoy it a lot. Now this is wasteful, and I think I am already partially, gradually disempowered. I feel like I could pick up coding without AI again if needed, but it feels so cumbersome to do. I am also not complaining, just reflecting on my feelings. I suspect this is similar to people having to code in assembler or machine code at first and then stopping doing so (unless you really had to due to the nature of the task) once higher-level programming languages were available.Managing AI is the higher-level programming language. And I think once it stops making stupid mistakes, or I have a good enough test system to catch them early and easily, I will be happy. As for deep work, I've discovered the same feeling in writing, and here it makes sense not to use AI, so that's what I've been doing the last couple of days, and thus the increase in the volume of stuff I wrote.
- Yearly tasks like taxes actually work best for this: each year AI does more and more of my tax-related work, and even after the half-year wait between filing an extension and the actual submission, the latest AI is able to find additional things to do and mistakes the previous AI made.
- ChatGPT launched on November 30, 2022 as a free "research preview", "fine-tuned from a model in the GPT-3.5 series, which finished training in early 2022". https://openai.com/index/chatgpt/
- The GPT-3 API opened as a private beta on June 11, 2020 and dropped its waitlist on November 18, 2021; the Playground already existed by then. One caveat to my memory: instruction-tuned "InstructGPT" models were in beta on the API from early 2021 and became the default on January 27, 2022, so unless I picked the base
davincimodel on purpose, a 2021 or 2022 Playground session was probably already talking to an instruct model. On sensitivity to wording, Zhao et al. (2021) found that for GPT-3 "the choice of prompt format, training examples, and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art." https://openai.com/index/openai-api/ , https://openai.com/index/api-no-waitlist/ , https://openai.com/index/instruction-following/ , https://arxiv.org/abs/2102.09690 - Not in general: they stopped reaching me. Spracklen et al. measured package hallucination in generated code at "at least 5.2% for commercial models and 21.7% for open-source models" (USENIX Security 2025), and a May 2026 re-run on frontier models from late 2025 and early 2026 still finds "between 4.62% (Claude Haiku 4.5) and 6.10% (GPT-5.4-mini)". What changed is that the agent now runs the install, the compiler and the tests, so most hallucinations fail in front of the agent instead of in front of me. https://arxiv.org/abs/2406.10279 , https://arxiv.org/abs/2605.17062
- GitHub Copilot: technical preview June 29, 2021, paid release to all developers June 21, 2022. Both announcements describe only in-editor code suggestions; chat came later. I have no dated record of my own trial. https://github.blog/news-insights/product-news/introducing-github-copilot-ai-pair-programmer/ , https://github.blog/changelog/2022-06-21-github-copilot-is-now-available-to-individual-developers/
- Aider changed how it labels commits twice while I used it:
name (aider)as author or committer from v0.39.0 (June 2024),Co-authored-by: aidertrailers by default from v0.85.0 (June 27, 2025); I checked all formats. https://aider.chat/HISTORY.html - Aider's Show HN is from May 11, 2023: "aider will directly apply the changes to your source files. Each change is automatically committed to git with a sensible commit message." Its creator built it to escape "a somewhat klunky workflow where I had to cut and paste code into ChatGPT and then back into my source files", i.e. the previous phase of this post. Claude Code appeared as a "limited research preview" on February 24, 2025 and became generally available on May 22, 2025. Anthropic's own oral history of Claude Code names an internal CLI called clide and a two-day prototype from September 2024 as its origin; Aider is not mentioned anywhere in it. https://news.ycombinator.com/item?id=35901649 , https://www.anthropic.com/news/claude-3-7-sonnet , https://www.anthropic.com/news/claude-4 , https://www.anthropic.com/features/making-of-claude-code
- uithub.com (change the g in github.com to a u) appeared around October 2024, two months into my Aider period; gitingest.com followed in December 2024. In 2023 there were only command-line scripts (gpt-repository-loader, Show HN March 17, 2023) and a couple of small sites that are gone now. https://news.ycombinator.com/item?id=41797578 , https://news.ycombinator.com/item?id=42329071
- Aider later made this an official feature:
aider --copy-paste(v0.68.0, December 10, 2024) copies the code context to the clipboard for you and picks up the chat reply from it. The docs note that "Most LLM web chat TOS prohibit automating" the copy and paste steps themselves. https://aider.chat/docs/usage/copypaste.html - From Aider's changelog, with release dates: /run, /git, /test and /web all existed before I started (2023 to early 2024); automatic linting after every edit and optional automatic tests came in v0.36.0 (May 22, 2024); "Aider now offers to run shell commands" (install dependencies, run the program, run tests, after confirmation) came in v0.52.0 (August 23, 2024). Nothing through v0.86 (August 2025) lets the model run a tool without a confirmation step. https://aider.chat/HISTORY.html
- Codex CLI was open-sourced on April 16, 2025 as "a new experiment". Its README at the time: in the default "Suggest" mode it could read any file in the repo but needed approval for "All file writes/patches" and "Any arbitrary shell commands"; on macOS commands ran under Apple Seatbelt, and "Linux - there is no sandboxing by default. We recommend using Docker". The Rust rewrite, with a Linux sandbox and a more permissive default, was opt-in in June 2025 and the documented default by late August. Claude Code's own sandbox arrived on October 20, 2025, opt-in, and reduces prompts rather than adding them. https://openai.com/index/introducing-o3-and-o4-mini/ , https://raw.githubusercontent.com/openai/codex/ed5e848f3e5fffeaba7c0e4046e08928da50849c/README.md , https://www.anthropic.com/engineering/claude-code-sandboxing
- Gemini CLI launched on June 25, 2025, "Now in preview", Apache 2.0, with a free tier of 60 requests per minute and 1,000 per day. https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemini-cli-open-source-ai-agent/
- John Backus, who led FORTRAN, on the programmers of the 1950s: many "began to regard themselves as members of a priesthood guarding skills and mysteries far too complex for ordinary mortals" and "regarded with hostility and derision more ambitious plans to make programming accessible to a larger population." (Backus, "Programming in America in the 1950s", 1980, pp. 127-128.) https://softwarepreservation.computerhistory.org/FORTRAN/paper/Backus-ProgrammingInAmerica-1976.pdf