Code Assistants and How They Changed How Software Gets Written
Software development was the first profession to absorb these tools at scale. The results are more nuanced than the productivity claims.
Programming was the natural first home for generative models. The output is text, correctness is testable, feedback is immediate, and the practitioners are already comfortable with automation.
Boilerplate, unfamiliar libraries, and cheaper questions about code
- Writing boilerplate, tests and configuration became dramatically faster.
- Unfamiliar languages and libraries became more approachable.
- Code review shifted toward checking generated code for subtle errors.
- Documentation and explanation improved, because asking about code became cheap.
The bottleneck was never typing
Overall delivery speed has improved less than the volume of code produced would suggest. The bottleneck in most software organisations is not typing; it is deciding what to build, agreeing on it, and verifying that it is correct. Generating code faster moves a fraction of the total effort.
There is also a measurable quality cost. Code that is plausible, compiles and passes a shallow test can still be wrong in ways that manifest months later. Several organisations have reported an increase in subtle defects and a decrease in the shared understanding that comes from writing code yourself.
Specifying, reviewing, and knowing when not to use it
The profession is converging on a new set of competencies: specifying precisely, reviewing critically, designing tests that catch generated errors, and knowing when not to use the tool. Those skills are harder to acquire than syntax, and they are now the differentiator.
Studies range from a gain to a slowdown
Productivity claims in software are notoriously difficult to verify. Studies find effects ranging from substantial improvement to a measurable slowdown, and the results correlate more with the experience of the developer and the complexity of the codebase than with the tool.
The most consistent finding is that these tools help most with unfamiliar languages and boilerplate, and help least with complex changes to a mature codebase where the bottleneck is understanding rather than typing.
Calls to functions that do not exist
Review practices are adapting. Reviewers now check for a class of errors specific to generated code: plausible-looking calls to functions that do not exist, subtly wrong parameter orders, and implementations that satisfy a test rather than the requirement.
Teams that adapted quickly formalised this into review checklists and expanded automated testing. Teams that did not found their defect rates rising in ways that were hard to attribute, because the code looked correct and demanded less scrutiny.
There is a countervailing effect that is easy to miss. Generating code is cheap; understanding it is not. Teams that increased their output without increasing their review capacity accumulated code that fewer people understood, which shows up later as slower changes and more incidents rather than as an immediate defect.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.