Skip to content
Analysis
Analysis: AI and Software Engineering — the 18-Month Reckoning
Mar 28, 2026 - Georg Zoeller

Analysis: AI and Software Engineering — the 18-Month Reckoning

In August 2024 we said agents were science fiction. In October 2024 we told companies to hold off on broad rollouts. It's March 2026. Time to revisit every prediction we made — and update the map.


In August 2024, we argued that AI coding assistance was real and measurable, that productivity gains would be uneven (favouring senior engineers), that the labour market impact would be serious, and that agents were science fiction — a distraction from the genuine, unrealised gains in developer experience and tooling.

In October 2024, we catalogued the tool landscape, warned about unsustainable pricing, and recommended that most companies hold back on broad rollouts until the dust settled.

Eighteen months is a long time in this field. Several of our predictions held up well. One failed spectacularly. The overall picture is more consequential — and more urgent — than either article suggested.


Part 1: Revisiting the Predictions

What We Got Right

Productivity gains are real and uneven. AI coding tool adoption has accelerated rapidly: 44% of developers used them in 2023, 62% in 2024, and by the 2025 Stack Overflow Developer Survey, 84% were using or planning to use AI tools — with 78% actively using them and 51% doing so daily. More telling: the productivity spread between senior and junior developers using AI has widened, not narrowed. Experienced engineers report acceleration on complex architectural tasks. Junior developers report acceleration on boilerplate and test generation — but also report increased difficulty understanding what their AI-generated code actually does. The “preventing learning” risk we flagged in October 2024 has matured into a documented pattern. Several organisations have reported onboarding cohorts where new engineers are proficient at AI-assisted delivery but struggle to debug or extend code without assistance.

The labour market signal is now unambiguous. In 2024, the data was suggestive. In 2026, it is explicit. Technology sector job postings for mid-level software engineers have fallen across North America and Europe. Headcount reductions at several large technology companies have been accompanied by explicit statements linking hiring freezes to AI-assisted productivity. More concerning: the junior developer pipeline has contracted. Fewer graduates are entering the field at the bottom rungs where they traditionally build foundational competence. The medium-term consequence — a shortage of senior engineers in the 2028–2030 window — is not yet visible in the data but is structurally predictable.

Open-source and pricing pressures did reshape the landscape. The consolidation we predicted in the October 2024 article arrived more or less on schedule. Several prominent AI coding startups from 2024 have been acquired. GitHub Copilot has iterated into a substantially more capable product. Cursor, despite predictions of Microsoft’s competitive response, has maintained a strong position by moving faster on agentic features than the platform incumbents.

Code quality risk from broad, unsupervised rollout is real. DORA’s 2025 State of AI-Assisted Software Development report found a 9% increase in bug rates in AI-heavy codebases, with AI-generated pull requests averaging 154% larger changesets and requiring 91% longer review times. AI acts as an amplifier — strengthening strong teams with established review practices, but accelerating technical debt creation in fragmented environments. Sixty-six percent of respondents cited “almost right” AI suggestions as their top frustration; 45% said debugging AI-generated code takes more time than writing it manually.


What We Got Wrong

Agents are no longer science fiction.

This was our most categorical prediction in August 2024, and it needs to be walked back clearly.

In August 2024, we wrote: “Agents are science fiction and unlikely to make an impact in the market in the next 1-2 years.”

It is now eighteen months later. That prediction was wrong. Not wrong in principle — the compounding error argument was sound, the reliability concerns were legitimate — but wrong on timeline and wrong on the rate of capability improvement.

The evidence:

  • SWE-bench Verified is the industry-standard benchmark for software engineering agents: real GitHub issues from real production repositories, with correctness judged by test suites. In early 2024, the best performing system solved approximately 19% of problems. By early 2026, leading systems claim over 80% on the Verified subset — a fourfold improvement in eighteen months. The caveat is important: researchers have flagged potential data contamination at these levels, and newer anti-contamination benchmarks (SWE-bench Pro, SWE-bench Live) show the same top models scoring only 19–23%. The capability is real; the magnitude of improvement is contested.

  • Claude Code (Anthropic, early 2025) demonstrated that a command-line coding agent, given access to a terminal, file system, and test runner, could autonomously complete multi-file refactoring tasks, debug failing test suites, and implement features from specification with minimal human intervention. It is the first tool we have used ourselves that meaningfully changed the productivity ceiling for a single engineer working alone.

  • GitHub Copilot Workspace extended Copilot from in-IDE suggestions to full task planning — from issue to pull request, including planning, implementation, and test generation across a full repository context.

  • Devin 2.0 (Cognition AI, April 2025), after an embarrassing 2024 that included credible allegations of demo manipulation, released a substantially more capable product with an agent-native IDE, parallel execution, and a price drop to $20/month. Its SWE-bench Verified score remains 13.86% — modest compared to frontier model scores — but in production use, Cognition claims it now generates a significant share of the company’s own internal pull requests. It still struggles with poorly-specified tasks and complex multi-system interactions — but that is a product and integration challenge, not a fundamental capability ceiling.

None of this means agents have arrived as general-purpose autonomous software engineers. The tasks that agents still fail on — whether that is 20% on the contamination-prone Verified subset or 80% on the cleaner benchmarks — are not random. They cluster around ambiguous specifications, cross-service state management, and tasks requiring genuine understanding of business context that is not encoded in the codebase. But the trajectory is unambiguous, and the practical utility is real enough that dismissing agents as fiction is no longer defensible.

We were wrong on timeline. We update our position accordingly.


Part 2: The Current Landscape

The Agent Spectrum

A framing that has proved useful: thinking about AI coding tools not as a binary (assistant vs. agent) but as a spectrum of autonomy, matched to task type and risk tolerance.

Autonomy LevelDescriptionBest fitMaturity
AutocompleteSuggestion-based, developer drivesAll developers, all tasksProduction-ready
In-context generationDeveloper prompts, AI generates, developer reviewsAll developers, well-specified tasksProduction-ready
Agentic IDEAI plans, executes, developer reviews diffsSenior developers, bounded tasksProduction-ready with governance
Terminal agentFull file system + test execution accessExperienced developers onlyEarly production
Autonomous PR agentIssue → PR with minimal human touchpointsSenior review requiredEmerging

The mistake in 2024 was treating “agents” as a single category and dismissing all of it. The mistake available today is treating it as a single category and adopting all of it. Most of the risk surface lives in the lower rows of the table above. Most of the current productiviy gain lives in the middle rows.

The Vibe Coding Phenomenon

2025 introduced a term that deserves attention: vibe coding. Coined by Andrej Karpathy in February 2025, it describes a mode of software development where the developer instructs AI in natural language, accepts outputs without rigorous review, and iterates based on whether the result feels correct rather than whether it is correct.

Karpathy’s framing was descriptive and arguably celebratory. The reality is more complicated.

For disposable prototypes, solo side-projects, and one-time data scripts, vibe coding is a legitimate and efficient mode. The risk-reward calculation is different: if the code is wrong, the cost is low.

For production systems, vibe coding is an accelerant toward the defect patterns DORA’s research documented. The speed is genuine. The technical debt accrual is also genuine. Several high-profile incidents in 2025 involved production code that was later traced to AI-generated implementations that had been accepted without adequate review — not because the engineers were incompetent, but because the output looked correct, was structured correctly, and passed surface-level review.

The Eloquence Trap — a concept we introduced in our Sovereign Command research to describe how polished AI output suppresses human verification — applies to code as much as to prose. A function that is cleanly formatted, well-commented, and passes the first test case creates a strong cognitive pull toward acceptance. The subtle boundary condition failure, the thread-safety issue, the SQL injection vector in the parameter handling — these do not announce themselves.

The Model Performance Plateau (That Isn’t)

In 2024, there was legitimate debate about whether model capability improvements would plateau. The “bitter lesson” optimists were right: the capability improvements did not plateau. Reasoning models applied to coding tasks — o3, Claude 3.7 Sonnet with extended thinking, and their successors — produce qualitatively different outputs on complex algorithmic and architectural tasks than their predecessors.

What has reached a point of diminishing returns, for most practical purposes, is raw code generation quality on well-specified tasks. The difference between the top models on a simple “write me a React component” prompt is now marginal. The differentiation has shifted to:

  • Multi-file coherence — maintaining consistent architecture across a large codebase during a refactoring operation.
  • Test generation quality — producing tests that actually catch the edge cases, not just the happy path.
  • Debugging under ambiguity — identifying the root cause when the error message is misleading.
  • Long-horizon task completion — executing a complex, multi-step change plan without losing context or contradicting earlier decisions.

These are the battlegrounds where Claude Code, Cursor with frontier models, and the emerging autonomous PR agents are competing. They are also the battlegrounds where the productivity ceiling for experienced engineers has risen most dramatically.


Part 3: The Labour Market Update

The Pipeline Problem Is Becoming Visible

The productivity gains from AI coding tools have been largely absorbed by reducing headcount growth, not by reducing existing headcount. Most companies have used AI productivity to avoid hiring rather than to execute large-scale layoffs — at least so far. The signals from 2025 suggest this is changing.

Anthropic’s Economic Index, published in early 2026, found software engineering among the occupations with the highest AI interaction rates — models are being asked to perform, augment, or assist with core software engineering tasks at a rate that suggests meaningful displacement of lower-complexity work. The index found that computer programming tasks are among the most exposed to AI substitution in the near term.

The structural concern is the pipeline. Junior roles have historically served two functions: delivering low-complexity work cheaply, and developing the senior engineers of the next decade. If AI substitution compresses the junior tier, the senior engineer shortage does not show up for five to seven years — long after the executives who made the hiring decisions have moved on.

This is not a prediction of catastrophe. It is a prediction of a skills transition that most organisations are not planning for.

The New Skill Stack

The skills that determine an engineer’s productivity ceiling have shifted measurably in eighteen months.

What has appreciated in value:

  • Systems architecture and specification writing. AI agents execute well when the task is well-specified. The skill of decomposing a complex requirement into unambiguous, testable sub-tasks — previously a senior engineer competence that was moderately valued — is now a primary productivity multiplier.
  • Adversarial code review. The ability to read AI-generated code with genuine critical attention, specifically looking for boundary conditions, security implications, and architectural inconsistencies that a plausible-but-wrong implementation will contain.
  • Test instrumentation design. If agents are going to make large, autonomous changes to codebases, test coverage is no longer a quality metric — it is an operational requirement. Designing test suites that actually catch the failure modes AI generates requires more sophistication than writing tests for known-correct code.
  • Prompt engineering for code contexts. Not the general “prompt engineering” that was declared obsolete in 2024, but the specific craft of providing AI coding tools with the context, constraints, and evaluation criteria they need to produce reliable outputs in a production codebase.

What has depreciated:

  • Boilerplate authorship — writing standard CRUD operations, standard UI components, standard test fixtures. This has not fully automated, but it is on a clear trajectory.
  • Single-function implementation from specification — the “here is the interface, implement it” task that historically occupied a significant portion of junior developer time.

Part 4: Updated Recommendations

Our October 2024 recommendations were conservative: build competence, establish governance, run small pilots, hold back on broad rollouts. That was the right call for 2024. In 2026, that posture is too conservative for most organisations and represents a meaningful competitive risk.

Updated recommendations, in order of priority:

1. Adopt AI-assisted coding tools across your developer base — now.

The risk of not adopting now outweighs the risk of adopting with imperfect governance. The productivity gap between AI-augmented engineering teams and unaugmented teams has widened to the point where it shows up in delivery cadence, time-to-market, and codebase maintainability. This is no longer a frontier technology requiring careful evaluation. It is operational infrastructure.

The caveat: adoption without governance creates the defect patterns DORA documented. The question is not whether to adopt but how to adopt responsibly.

2. Invest in senior engineer review capacity, not just developer seat count.

If AI tools are accelerating code production, the constraint shifts to review capacity. A senior engineer reviewing and validating AI-generated code is a higher-value activity than a senior engineer writing that code themselves. Organisational structures that treat review as a junior developer gate rather than a senior developer core competency are building in a quality risk.

3. Treat agents as a senior-supervised tool, not an autonomous system.

The autonomy gains are real. Deploy them with explicit governance: agent-generated pull requests require senior review, agent access to production systems requires explicit scope limiting, and agent-produced code changes must pass the same test instrumentation gates as human-written code. The SWE-bench numbers are impressive in a benchmark context. They do not transfer automatically to your specific codebase with its specific constraints and undocumented assumptions.

4. Rebuild junior developer pathways deliberately.

If AI tools are compressing the work traditionally used to develop junior engineers, organisations that want a senior engineer bench in 2030 need to deliberately reintroduce that development pathway. Pair junior engineers with AI tools on complex tasks under explicit senior mentorship, rather than assigning them autonomous AI-assisted work on simple tasks that teaches them neither the domain nor the fundamentals.

5. Audit your codebase for AI generation risk surface.

Not all codebases are equally exposed. High-traffic, security-sensitive, or regulatory-context code — authentication systems, payment processing, data handling, API surfaces — should be designated for higher review standards regardless of whether the implementation was human or AI generated. The distinction matters less than the risk tier.

6. GitHub Copilot remains the pragmatic starting point, but is no longer the ceiling.

For organisations that have not yet started: GitHub Copilot’s enterprise offering, combined with the governance features added in 2025 (audit logs, policy controls, IP indemnity on enterprise tiers), is a reasonable default starting point. It has the widest IDE coverage, the most enterprise-mature compliance posture, and the lowest change management friction.

For organisations with established AI tooling that have not upgraded since 2024: the gap between GitHub Copilot and the leading agentic tools (Claude Code for terminal-native workflows, Cursor for IDE-centric teams) is now large enough that a re-evaluation is overdue.

7. Claude Code deserves specific attention.

We do not recommend specific products lightly. Claude Code represents a qualitative shift in what a command-line coding agent can do. For engineering leaders who have not yet spent a day working with it on a real codebase, we recommend doing so before making any decisions about your organisation’s AI tooling strategy. The experience is genuinely clarifying about both the capability ceiling and the governance requirements.


Looking Ahead

In August 2024, the most aggressive prediction in our article was that “dramatic changes (with downstream effects on job profiles and the labour market) [lie] ahead of us.” That prediction looks conservative in retrospect.

The next phase of the capability curve — models with longer context windows that can hold entire large codebases in memory, tighter integration between agent systems and CI/CD pipelines, and autonomous testing frameworks that close the loop on AI-generated code — is already in early production at several organisations. The compounding error problem that made agents unreliable in 2024 has not been solved; it has been partially mitigated by better error recovery, better test instrumentation, and narrower task scoping.

The trajectory points toward a software engineering profession that looks less like a craft of writing code and more like an engineering discipline of designing systems, specifying requirements, governing automated execution, and reviewing outputs. That is a real profession. It requires real skills. It is not the same profession that most current software engineers trained for.

The organisations that will navigate this transition well are the ones building that competence now — not the ones waiting for the dust to settle, because in this field, the dust does not settle. It redistributes.


About the Author

Georg Zoeller is a Co-founder of Centre for AI Leadership and a former Facebook/Meta Business Engineering Director. He has decades of experience working on the frontlines of technology and was instrumental in building and evolving Facebook’s Partner Engineering and Solution Architecture roles.

This article is the third in the AI and Software Engineering series. Read Part 1 (August 2024) and Part 2 (October 2024).