This is an interesting idea. It's effectively what a good software engineer already does in their head, except it's doing it with a real compiler.
Richard Gabriel wrote something that has really stuck with me:
> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.
This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.
A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!
This is an application of the rule of three: you need at least three clients to write a good library, protocol, or other underlying abstraction. Except it proposes using a large language model to write the three clients and then throwing them away.
> A library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures.
Sigh.
Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
Unless I'm missing something, this doesn't help with preventing regressions. In the end, as the author already puts it, it's an integration test in the end, why not just write the integration tests directly?
The word "testing" is a bit misleading. It's not about test coverage, it's about experimenting to find architecture decisions that are suitable for further development. So, not generating automated tests, but instead building throw-away features on top of the new thing and seeing if they turn out okay. (Though the post remains a bit vague on how to judge "okay".) Maybe it should be called something like "ephemeral build-out" or "future usage trial" or so.
I've been using this pattern quite a bit recently for API design, and I really like it.
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
I had good results with a similar technique this summer.
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
Not unthinkable at all. Been doing this for several years as an opt-in on pr tags usually but now with agents writing code I have it on by default. I've got several skills files about accessing with a read only account and gitops done through the PR. It's an excellent dev environment for the agents.
We use NX in our monorepo, and it is great at determining which testing/linting tasks need to be run based on which libraries in the repo were "affected". We have a bunch of e2e tests, and I've been having a lot of success getting Claude with Opus to run only the relevant e2e tests when appropriate during development.
I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
That's a thing with AI-generated code. The lying machine happily reports that "boss, everything is clean, tests pass, code gate green", but once we need to build on top it's always "preexisting flaky tests, not related to this section" and trying to do commit --no-veriry before it's slapped on it's little robot hands twice.
There is a solution for green test suite - mutation tests. Testing your tests gives you solid ground to treat green tests as reliable. It takes time but also brings quality.
In terms of an agent reporting green tests never executed, a proper step in deployment workflow may be a good gate.
I'm just actually using the thing I'm building while building it, so I feel the sharp edges and the project quickly evolves based on actual need. Nothing new really.
Another trick I have found useful is to do integration testing with coverage enabled. One agent creates tasks for subagents that run the application with coverage enabled, it merges the reports, and based on that it comes up with new tasks for subagents, repeat until coverage no longer moves.
> In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
Seems like one more arrow in the toolkit. But the best testing looks at the sources and explores the cracks between the strata with edge cases, looks at limits, and where one method changes to another. (And hats off to Murphy, for waiting until after you ship...)
i'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
Richard Gabriel wrote something that has really stuck with me:
> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.
This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.
A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!
Sigh.
Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
Gets annoying pretty quickly
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.