Back to blog

How Our Engineering Team Uses AI, Part II: Meat Proxies

Eyal Bukchin · October 7, 2026 · 9 min read

In January we published How Our Engineering Team Uses AI. The tl;dr was that AI was useful around the edges of our work (getting oriented in unfamiliar code, exploring approaches, writing scripts) but not much use at the center of it, because mirrord is a low-level Rust tool with an unusual architecture, and general models couldn’t really deal with it. We wrapped up the post by calling AI “a powerful tool around the edges of software development.”

Eight months later, we asked the engineering team the same questions again. AI has since moved from the edges to the center: most of the engineers who replied barely write code by hand anymore. Someone still has to read all that code, and so the part of the job that grew the most is review, not only of pull requests, but also the inner loop of reading and correcting (or re-prompting) every chunk of code the agent writes.

I’m not really writing any code, I just review the slop and try to tame the beast.

We heard you like tools

In Part I, tools converged on ChatGPT, Gemini, Claude Code and Cursor. We still don’t mandate any specific tooling, and now there’s even less standardization among what the team uses. Claude Code and Codex are the most common agents, used through the CLI, the desktop apps, or an editor integration. Beyond that, it’s a long tail. Some examples:

  • Claude inside Emacs, through agent-shell, an Emacs package for running coding agents inside the editor.
  • Amp, a coding agent and development environment that can work in your local checkout from the terminal or on remote machines it spins up per task.
  • Orca, a desktop app for running several agents in parallel, each in its own git worktree.
  • pi, a minimal open-source agent harness that runs Codex, Claude and Copilot models from one place and is extended with plugins and your own markdown files.
  • Graphify, which turns a codebase into a graph agents can query, used to make cheaper Codex models more accurate on our code.
  • Gemma running locally through llama.cpp, for small chores like commits and rebases.
  • Zed keeps showing up as well, both for its agent integrations and as a fast, light editor for when a human needs to edit or clean up what the agent wrote.

Cost has entered the equation too. By one engineer’s estimate, Fable uses about ten times as much of a usage allowance as Opus, and for their work the quality gain was moderate at best, so they stay on Opus. Another engineer tried to save tokens by having a strong model write a detailed implementation guide and handing it to a cheaper model to write the code, since generating output costs more than reading input (alas, the cheaper models wrote code bad enough to make the savings pointless, so this experiment failed).

Where it helps

People now use agents for pretty much everything: planning, breaking tickets into subtasks, researching bugs, exploring unfamiliar parts of the codebase, writing code (both on the edges and in the core of the product), writing and running tests, pre-reviewing PRs before asking a human to look at it, diagnosing and fixing CI failures, and investigating production incidents. Some use cases in the general consensus:

  • Finding things by description rather than by name. When you’re working with e.g. an unfamiliar Kubernetes data model, you often don’t know what the thing you’re looking for is called, so you can’t grep for it, but you can describe it. Agents are very good at this kind of search.
  • Code outside the product. Build tooling, bash scripts, and throwaway tests for checking locally how our infrastructure behaves in an edge case (like how a queue handles some unusual sequence of events). The stakes are low, since a mistake breaks a build step or a local experiment rather than something a customer runs. Our xtask build scripts were described as something “no human hand has ever touched.”
  • Status updates. The age-old chore of keeping Linear issues and status posts current can finally be outsourced.
  • Questions about past decisions. Why a component was designed the way it was, or what was decided about a feature, is often faster to ask an agent than to track down the person who made the call.
  • Collecting and analyzing messy data. That’s how we found our flakiest tests and measured how our CI changes affected run durations.
  • Work a script should be doing. For repetitive tasks nobody has automated yet, having an agent do the work each time is slower and more expensive than a script would be, but we’re lazy.

Where it still struggles

Rust quality remains a recurring complaint. Agent-written Rust compiles, but it tends to be much longer than it needs to be, with unnecessary .to_owned(), .clone() and .collect() calls that someone then has to go through and remove. Agents are also eager to write tests, and love splitting code into tiny functions so they can add tiny tests for them. Even the engineers who have moved furthest toward AI still evaluate every line and the overall structure of a change themselves, and they often find it wanting.

Agents also tend to reinvent things. Ask one how to solve a problem and it will often come up with its own solution instead of finding the library or well-known approach that already solves it, so its research is a starting point rather than an answer. So most engineers prefer to do the initial research for a new area by hand. The point is for the engineer to understand the problem space deeply enough to make good calls on everything built on top of it.

Writing the PR description/commits makes me look at the problem from a higher level, which ensures I at the very least understand the solution and can judge it appropriately.

The strongest opinions were about writing style. Several engineers write every commit message, PR description and code comment themselves, because agents produce too much of it and too little of it makes sense. Those who let agents draft docs or RFCs usually end up rewriting most of it, because, as one put it, “AI generated text reads like word pulp.” But the best argument for doing it by hand is that writing the PR description forces you to look at the problem from a higher level, which can help ensure you understood the solution well enough.

AI generated text reads like word pulp.

We mostly just review now

Asked how their usage had changed since January, one engineer described much more parallelism and much more fatigue:

My week used to be 20-40% reviews, 60-80% code gen, now it’s 95% reviews, it’s so tiring.

Others described something close to burnout. A task now takes more rounds of reviewing and correcting before it’s actually finished, so finishing feels further away than it used to, and with several agents running, you’re in that state on several tasks at once. Some refuse the parallelism altogether, running a single agent thread with a narrow purpose and thinking carefully about what comes back.

i only use 1 thread at a time with a very specific purpose and think about things carefully. all those agent orchestrators and stuff are too much for me, it feels like the fastest way to get burnout

The most difficult part is keeping control over quality and maintaining an understanding of how the system works. Even the engineers who no longer write code by hand still go through several review passes after the code works, until they understand and have tested every line. Everyone works hard to not become meat proxies.

We’re trying to find solutions, though, and right now we’re hoping we can successfully automate the first pass. Greptile reviews our pull requests automatically (we’re currently evaluating it against Qodo, using AI to run the evaluation), and many of us have an agent review our own PR before a person sees it.

Building for the agent that will maintain it

There were also changes in how people design things. Some of us now build with the assumption that an agent will maintain the code, which in practice means thorough logging, so that an agent debugging it later has something to read (a courtesy I’m sure many humans would have also appreciated). Some write personal skills that encode their repeated rules and fixes.

That has a company-level version. The January post showed the CLAUDE.md in our mirrord repo, which back then included a long architecture walkthrough. It has since become an AGENTS.md, with CLAUDE.md as a link to it, and it got shorter: the walkthrough was deemed unnecessary and cut, leaving a one-line description of each core component plus the commands for building, testing and linting. It now sits alongside a company-wide one. Everyone at MetalBear imports one shared context repository from their personal ~/.claude/CLAUDE.md, which describes our teams, the tools we use and the known quirks of each, so every agent starts from the same background. The repository is private, but the pattern is simple enough to copy.

One engineer’s experiment with agents on a schedule

One of our engineers has been experimenting with agents that run on a schedule rather than on request. One agent triages CI every night. Another picks one small, open bug from Linear every morning and opens a PR to fix it. Tiny but non-critical is a category of work that people never get around to and agents don’t mind.

The same engineer also built a support bot that our solutions engineers use to draft answers, which a human then checks and sends. After it confidently described an unverified configuration as “confirmed working in a past similar case,” it got two standing rules: search past resolutions before reasoning from scratch, and end every answer with what it is based on.

So what changed?

In January, AI was a tool around the edges of our work. Now, thanks to models improving, our processes maturing, and simply the time it takes workflow changes to propagate, AI writes most of the code for many of the engineers who answered. The work that grew is making sure a person understood that code before it ships. It is certainly more tiring, and hopefully our next post on this will be about how we solved that.

What is mirrord?

mirrord is a Kubernetes development platform that lets developers and AI coding agents test code in a production-like environment before deploying it. Your service runs wherever you're working, locally, in CI, or in an agent's sandbox, while mirrord proxies its traffic, environment variables, and files to and from a shared staging cluster, so it behaves as if it were deployed without actually being deployed.

Engineering teams at companies like monday.com, National Australia Bank, and SurveyMonkey use mirrord to iterate and ship faster, while spending less on dev environment infrastructure.

Want to dig deeper?

With mirrord, cloud developers can run local code in the context of their Kubernetes cluster — streamlining coding, debugging, testing, and troubleshooting.