Overcut named in the 2026 Gartner® Innovation Insight: AI Software Factories reportRead more

How AI Agents Auto-Generate Technical Design Documents

Yuval Hazaz· Oct 8, 2026· 9 min read
Share
The design doc your team keeps skipping can now write its first draft: a race car parked in its grid box, seen from above

Key Takeaways

  • Most teams agree a technical design document is worth writing and skip it under deadline pressure, because the research behind a good one takes days.
  • An AI agent can draft a credible TDD by reading the ticket, its discussion, and the relevant repositories: goals and non-goals, current state, architecture diagrams, a phased plan, risks, and open questions.
  • Draft quality depends on context. An agent that sees only the ticket writes a template. One that reads the codebase writes a design that fits the system.
  • Generated designs fall short on business priorities, novel architecture, and tradeoffs that depend on numbers the agent cannot see. Those need an engineer.
  • In Overcut, every generated TDD is assigned to a person who owns it, nothing is approved by default, and only an approved design can drive AI code generation downstream.

Ask an engineering team whether design docs are worth writing and almost everyone says yes. Ask how many of last quarter's features had one and the answer drops. The features that got designs were the large, visible ones. The medium-sized changes, the ones that touched three services and a database migration, went straight to code because writing a design felt like a week the team did not have.

Those medium-sized changes are where design docs pay off most: big enough to go wrong expensively, small enough that nobody schedules time to think them through. We think AI agents change the economics here, because the expensive part of a technical design document was never typing the prose. It was the research behind it, and research is something agents are now good at.

Why Writing Technical Design Documents Slows Teams Down

A good software design document answers a short list of questions. What are we solving, and what are we deliberately not solving? How does the change fit the existing system? What alternatives did we reject, and why? What could go wrong? Malte Ubl, who spent more than a decade as an engineer at Google before becoming CTO of Vercel, wrote the widely cited account of how Google approaches this in Design Docs at Google. The structure he describes settles on roughly that shape: context and scope, goals and non-goals, the design, alternatives considered, and cross-cutting concerns like security and observability.

None of those questions is hard to write down once you know the answer. Getting to the answer is the cost. The author has to find every place the change touches, read how similar features were built, check which services call the API being modified, and track down the person who knows why the retry logic works the way it does. For a change spanning several repositories, that is days of reading before the first paragraph.

Coverage is also uneven from one design to the next. One engineer writes twelve pages with sequence diagrams, another writes four paragraphs and a task list. Reviewers spend the start of every review working out what the document skipped, and the gaps differ every time.

And design docs go stale. A system design document captures the system at the moment it was written, and the code keeps moving. Within a few months the doc describes a design that partly exists, engineers learn to trust the code over the doc, and they become less willing to write the next one.

What an AI-Generated Technical Design Document Includes

A generated TDD should read like something a senior engineer on the team would write, with the same sections every time.

Goal and scope. The problem, the intended outcome, and the explicit non-goals. Non-goals are the section reviewers argue about, so writing them down early surfaces disagreements before code exists.

Current state. How the relevant part of the system works today, with references to the actual files, modules, and services. This section separates a grounded draft from a generic one, because only something that read the code can write it.

Proposed architecture. The design itself, with diagrams generated as Mermaid so they render natively in GitHub, GitLab, and issue trackers and can be edited as text.

Phased implementation plan. Sequential phases with concrete tasks, written so each phase can land as its own reviewable change.

Risks and mitigations. Migration hazards, performance, backward compatibility, security exposure, and what the design does about each.

Open questions. The decisions the agent could not make from what it saw. We consider this the most important section, for reasons we come back to below.

A technical specification for a single API or component can skip some of these. A design that spans services should not.

How AI Agents Generate a Technical Design Document

The agent does what a careful engineer does before writing a design, in the same order, much faster.

It starts with the request: the ticket, its comments, linked issues, and any attached product requirements. Comments are easy to underrate. They are often where someone wrote "this also needs to work for the bulk import path," and a design that misses that line is wrong in a way the ticket alone would not reveal.

Next it works out where the change lands. In an organization with dozens of services this is the step most likely to go wrong, and a design aimed at the wrong repository can read perfectly and still be useless. In Overcut, the relevant repositories are identified from the organization's structure and the ticket's relationships, then cloned so the agent can read them directly.

Then it reads for prior art. How do existing endpoints handle pagination? How was the last schema migration staged? Is there already a caching layer this feature should reuse? Much of the value of a human-written design comes from the author knowing the codebase's conventions, and an agent can recover a large share of that by reading the code and its history, which is the same grounding every stage of the agentic software development lifecycle depends on.

Finally it writes. In our technical design proposal playbook, a senior developer agent reasons about the architecture and plan, and a technical writer agent shapes the result into a document a reviewer can read in one pass. Splitting the roles keeps design reasoning separate from presentation. A product manager agent then posts the draft back to the ticket, where the team already discusses the work, and opens a live session to answer the team's questions about it.

The workflow triggers from the ticket too, automatically when an issue is labeled needs-design or on demand with a /design comment. We have argued that the ticket is the real interface for AI development, and design is a natural place to prove it.

A ticket labeled needs-design with a draft technical design posted by the Overcut agent after reading payments-service, billing-api, and invoice-worker, listing goal and scope, current state, proposed architecture, phased implementation plan, and risks as complete, with three open questions remaining and the ticket marked design-needs-info

Where AI-Generated TDDs Fall Short

A generated design is only as good as what the agent could see, and some of what matters most is not written anywhere an agent can read.

Business context is the clearest gap. The agent does not know the team plans to deprecate this service next quarter, that a large customer is contractually promised a specific behavior, or that leadership has ruled out another vendor dependency. A design can be technically sound and organizationally wrong.

Novel architecture is a weak spot too. Agents are strongest when the codebase holds a precedent. A team's first event-sourced service has no prior art to read, so the agent falls back on general knowledge, and the draft lacks the judgment of someone who has run that architecture in production.

Then there are tradeoffs that depend on numbers. Whether to cache, shard, or precompute depends on traffic, latency budgets, and cost targets. If those live in a dashboard the agent cannot query, the design will guess or hedge.

That is why the open questions section carries so much weight. An agent that confidently fills every gap produces a document that reads well and hides its weakest assumptions. An agent that lists what it could not determine gives the reviewer a short, specific agenda. We would rather ship a draft with five real open questions than one that looks complete.

How to Use AI-Generated TDDs Without Removing Engineer Ownership

The risk with a generated document is that it gets approved because it looks finished. A design nobody really reviewed is worse than none, because it gives the team false confidence. The handoff from draft to approved document needs as much design as the generation.

It starts with ownership. Every generated design is assigned to the person who opened the ticket, and the ticket is labeled design-complete when the design is ready, or design-needs-info when open questions block progress. Nothing the agent writes counts as approved by default.

Review happens as a conversation in the ticket. The team asks follow-up questions, the agent answers them against the codebase, and the design is revised as decisions get made. The open questions get resolved here, usually by the people holding the business and operational context the agent lacked, and the design's owner edits the document directly where it is wrong.

Approval gates everything downstream, and this is the part we feel most strongly about. AI code generation is only as good as the plan it executes, and an unreviewed design fed straight to an implementation agent turns a wrong assumption into a pull request. In Overcut, the create PR from design workflow starts only when someone runs /pr on a ticket whose design has been approved. It drafts a phased implementation plan from the design, lands each phase as its own commit, adds tests, and opens the pull request for review.

We think this is the right division of labor. Agents take the reading, the cross-referencing, and the first draft, which is the work that made design docs too expensive for medium-sized changes. Engineers keep the judgment calls. When a draft costs minutes instead of days, the question stops being whether a change deserves a design and becomes which open questions need a human answer.

FAQs

Can AI write a TDD from a spec alone?

It can, but the result will be generic. A spec describes what to build, not how the existing system works, so an agent working only from the spec has to invent the current state and guess at conventions. The useful designs come from agents that also read the codebase, related tickets, and prior implementations. Without that context, expect a template with the right headings and few decisions you can act on.

Can AI-generated TDDs replace architect review?

No. Generated designs reduce the time architects spend on research and catch-up, so their review starts from a structured draft rather than a blank page. The judgment still belongs to them: weighing tradeoffs against business priorities, rejecting designs that fit the code but not the roadmap, and making calls on novel architecture. The better use is giving architects more designs to review, including for changes that previously had none.

How do AI-generated TDDs stay in sync with code changes?

Only if something checks them. Keep the design attached to the ticket the implementation references, then run a workflow on merge that compares the shipped change against the approved design and flags divergence. Small deviations can be recorded as amendments. Large ones should reopen the design for review. Without an automated check, generated TDDs go stale exactly like hand-written ones, and there will be more of them to go stale.

What is the difference between a TDD and a PRD?

A product requirements document defines what to build and why: the user problem, the desired outcome, success metrics, and scope from a product perspective. A technical design document defines how to build it: architecture, data model, interfaces, implementation phases, and risks. The PRD is usually written by product managers and comes first. The TDD is written or owned by engineers and turns those requirements into a buildable plan.

Your AI factory starts with one command

Book a 30-min demo call
Claude logoOpenAI logoCursor logoGitHub Copilot logo
~/acme-app
$ npx overcut init
✓ Signed in
✓ Current workspace: Acme
✓ Connected: Claude Code
> Set up Overcut for this repo.
Inspecting repo: GitHub remote, pnpm, vitest, GitHub Actions
Proposed workflows:
1. Code review
2. Fix CI
3. Auto docs update on merge
Draft run of "Code review" passed.
Publish and activate triggers? [y/N]