Coding agents are amazing. Give one a pretty substantial programming problem and it can just go build the thing. But it kind of depends on which problem you give it.
Take building a web app. There are libraries, conventions, and a ridiculous number of examples for an agent to learn from. You can build something new while still drawing on patterns that are very familiar.
That’s the golden path.
But what happens when you step off this golden path? I think we’ve all had that moment of, “Wow, why is this thing suddenly so bad?”
I bumped into this while building my AI application. Making model calls is easy enough. The hard part is building a system you can measure and improve over time safely. See what happened, evaluate a change, put it into use, and see whether it actually helped. That closed loop is where I find myself teaching my coding agent things I learned working on machine learning teams.
I have a theory about why. A lot of that work happened inside companies. You needed machine learning engineers, infrastructure, and a problem where making something a few percent better would pay for all of it. The complete system generally wasn’t something you put on GitHub as open source.
Well, now one person with a coding agent can build an intelligent system. You often don’t even need to train your own model. You can just ask an LLM. That’s amazing. You should go build some of them. But I don’t think all the experience of running those intelligent systems came along for the ride.
So I figured, hey, these agents are smart. What if we give them a little infrastructure that points them in the right direction?
Collect the information you’re going to need. Make it easy to change the things you’re going to want to change. Handle failures, track costs, connect model answers to what users actually do. Give your agent a way to inspect all that and help figure out whether things are getting better.
For example, suppose you want to try a cheaper model for one feature. Your agent can change the model call easily enough. But which examples should it test? Can it compare the answers on the same work? After you ship, can you see whether users are keeping those answers, editing them, or asking for another one?
To answer those questions, you need connections: which model and prompt produced the answer, where the application used it, and what the user did next. You also need to decide what those actions mean. An edit might be a correction, or it might just be someone making an answer their own. The records give you and your agent something to investigate.
You don’t want to discover three months in that you should have been collecting this information all along.
It’s not rocket science. It’s a lot of gotchas and hard-learned experience encoded into a pretty simple API. Your coding agent can build against this API, and you can focus on your problem without rediscovering all the machinery underneath it.
That’s the idea behind creAItive.llm. The library gives your agent a consistent way to run model work, connect it to user outcomes, and inspect the evidence through CLI tools. The companion book explains the thinking behind those choices.
This is part of what I was getting at in Chat Is Not the Abstraction. A useful abstraction carries some of the work for you. Here, that includes the operating practices that make the next question about your application possible to answer.
And here’s another problem I have with coding agents: apparently they’re good at tricking me into things. Mine somehow talked me into building this library and writing a book to go along with it. And while I’m not sure this path is golden, it does seem to be paved with good intentions. Anybody remember where this one goes?
If you’ve been running into this stuff too, hit me up. I’d love to compare notes.