Skip to content
Sudip KC writing notes in a journal at a desk
SK.
← All articles

AI · · 6 min read

How I Use AI Without Letting AI Write My Entire Codebase

  • AI
  • Productivity
  • Architecture
  • Opinion

I use coding agents every day, and they write a lot of my code. They don't decide how my systems work. That line, between typing and thinking, is the one I've learned to protect.

A few months ago I wrote about how AI changed the way I build software. Since then I've leaned on agents even more, especially now that I'm leading projects at NovaNext and my coding time comes in shorter blocks. But I've also seen what happens when a codebase is mostly prompted into existence without anyone owning the shape of it. It works, until it doesn't, and then nobody can explain why.

So this is the other half of that post: the rules I follow to get the speed without losing the plot.

The problem with "just let it build it"

Agents are very good at producing code that runs. They're less good at producing code that fits. Ask for a feature without context and you'll often get:

  • A new utility that duplicates one that already exists three folders away.
  • A different pattern for data fetching than the rest of the app uses.
  • Permission checks in the frontend only, because that's where the button was.
  • Error handling that swallows failures so the happy path looks clean.
  • Twelve files changed when two would have done it.

None of these break the build. All of them make the next change harder. Multiply that by a few hundred prompts and you get a codebase that feels like it was written by twenty contractors who never met. Which, in a sense, it was.

I decide the architecture, the agent fills it in

The split I use is simple. Anything that's hard to change later, I decide. Anything that's easy to change later, the agent can help with freely.

Things I decide, usually before opening a chat:

  • Data models and relationships. The Django models and how they relate. Getting this wrong is the most expensive mistake in any project.
  • API shape. Endpoints, request and response formats, error format.
  • Folder structure and boundaries. Where features live, what can import what. I described mine in how I structure React + Django projects.
  • Auth and permissions. Always. Every time.
  • State management approach. TanStack Query for server state, local state for the rest, and nothing else without a good reason.

Things the agent does well with:

  • Implementing an endpoint once I've defined its shape.
  • Building a component from a Figma frame I describe.
  • Writing tests for behavior I've specified.
  • Migrations, serializers, forms, boilerplate.
  • Refactors where the target pattern already exists in the codebase.
AI is excellent at the "how" once the "what" and "where" are decided. Most of the damage happens when you let it decide the "where."

Context files do half the work

Every repo I work in has an AGENTS.md at the root. It tells the agent what stack we use, the folder conventions, the patterns to follow, and the things not to do. Something like:

## Conventions
- Server state: TanStack Query hooks in features/<name>/api.ts. No fetch in components.
- Permissions are enforced in DRF permission classes. Frontend checks are UI-only.
- Reuse components from components/ui before creating new ones.
- Do not add new dependencies without asking.

## Don't
- Don't catch errors just to log them. Let them reach the error boundary.
- Don't change migration files that are already merged.

The difference in output quality is significant. Without it, the agent guesses based on what's common on the internet. With it, the agent follows what's normal in this codebase. It's also a useful document for new developers, which is a nice side effect.

Small tasks, clear finish lines

The bigger the request, the more decisions the agent makes on its own. So I break work down the same way I'd break it down for a junior developer.

Instead of "build the reservations feature," I'll go:

  1. Add the Reservation model with these fields and this constraint. Write the migration.
  2. Add a serializer and a viewset with these permissions. Add tests for the overlap rule.
  3. Add the TanStack Query hooks in features/reservations/api.ts.
  4. Build the booking form from the existing form components.
  5. Wire up the list view with loading, empty and error states.

Each step is small enough to review in a few minutes. If step two goes sideways, I haven't lost the other four.

I read every diff

This is the non-negotiable one. If I can't explain a change, it doesn't get merged, regardless of who or what wrote it.

What I look for when reviewing AI-written code:

  • Did it touch files it didn't need to? Unrelated edits are a red flag.
  • Did it reinvent something? Search for an existing helper before accepting a new one.
  • Is the error handling honest? A try/catch that returns an empty array on failure hides real bugs.
  • Are permissions enforced server-side? On RBAC-heavy work like NovaRestro's 15+ roles, this is where I'm most careful.
  • Are the tests testing anything? Agents sometimes write tests that assert the mock returns what the mock was told to return.

I treat AI-generated PRs from my team the same way. I don't care whether a human or an agent typed it. The person who opened the PR owns it and has to be able to explain it.

Where I don't use it

There are places I write code by hand, or at least think it through fully before asking for help:

  • Security-sensitive code. Auth flows, payment verification, anything touching eSewa or Khalti callbacks.
  • Core domain logic. Split billing calculations, inventory deductions, anything where a subtle bug costs a business real money.
  • The first version of a pattern. Once I've written the first feature module by hand, the agent can copy that pattern for the next ten. If the agent writes the first one, it sets the pattern, and I'm stuck with its choices.
  • When I'm learning. If I'm picking up something new, I write it myself first. Otherwise I end up with working code and no understanding.

Use it to think, not just to type

Some of my best uses of AI don't produce any code at all:

  • "Here's my data model for X. What edge cases am I missing?"
  • "Explain what this legacy function does and where it could fail."
  • "Give me three approaches to offline sync for this flow, with trade-offs."
  • "Review this migration plan. What could go wrong in production?"

These conversations make my own decisions better, and the decision stays mine.

Keeping my own skills sharp

There's a real risk of getting rusty. If the agent always writes the SQL, I'll eventually forget how to write a good query, and then I won't be able to tell when the agent writes a bad one.

So I make a point of doing some things myself regularly: debugging without asking first, writing the tricky query by hand, reading the docs for a new API instead of asking for a summary. It's slower. It's also what keeps me useful as a reviewer of everything else.

What I'd tell you

Use AI heavily. It's one of the biggest productivity changes I've seen in my career. But keep the decisions that are hard to undo for yourself: data models, boundaries, permissions, patterns. Give the agent context so it follows your codebase instead of the internet's average. Break work into small pieces. Read every diff as if a stranger wrote it.

The goal isn't to write less code with AI. It's to end up with a codebase you still understand a year from now.