Ravin Elango

Writing

The QA org chart of 2030

July 9, 20263 min read

If generation is cheap — and it is getting cheaper — the QA org of 2030 will not look like a larger version of 2016. It will look like a smaller function with a different center of gravity.

I do not mean “testers go away.” I mean the job that was “turn this ticket into a script” goes away as a full-time identity. What remains is the work that was always the actual job, and that we under-titled because the scripts took so much time: deciding what must be true, deciding what evidence is enough, and signing.

Here is the sketch I keep drawing on whiteboards. It is a bet. It is dated. I will be wrong about the names.

The roles that remain

Quality lead. Still a manager of a system, not a manager of a pile of cases. Owns the release signature, the quality contract for agents, and the conversation with product when the contract is too expensive. Twenty years of this work does not vanish. It concentrates.

Eval engineer. A named role by 2028, I think, and ordinary by 2030. Designs goldens, failure taxonomies, graders. Treats the model and the judge as systems under test. Does not report to the team whose bonus depends on the eval passing. I would put this under quality. If you put it under the model org, you will get a leaderboard.

Domain oracle. The person who knows what “right” means in insurance, in a clinical workflow, in a published book. Not a tester who learned the domain. A domain person who learned how to sit inside a quality system. Agents make this person more valuable, because someone has to argue with the model when it is fluent and wrong.

Pipeline engineer. Keeps Playwright, Tosca, Python, CI, data, and permissions in a state that an agent can even see. This is the access problem from the other side. If staging is a rumor, you do not have an AI quality program. You have a demo.

Incident librarian. Someone has to keep the corpus of what actually broke. Agents are only as good as the failures you are willing to remember. Most orgs already have this person. They are called “the one who remembers 2019.” Give them a title and a budget.

The roles that dissolve

The layer that only produced script volume. The layer that only ran the same regression every Tuesday. The layer that existed because the tools were too clumsy for the domain oracle to use directly.

Those people are not surplus humans. They are candidates for eval work, oracle work, or pipeline work. The ones who wanted to own meaning will be fine. The ones who wanted to own keystrokes will need a different deal, and pretending otherwise is how you get a function that fights the tools.

I have managed both. The kind thing is to say this early.

Agents as staff

On the org chart I would draw the agents. Not as a logo in the corner. As named workers with managers.

regression-maintainer reports to the quality lead. It may open pull requests. It may not merge. It may not delete. Its weekly report is defects found, tests proposed, tests rejected by a human, and flake introduced.

grader reports to the eval engineer. Changing its prompt is a change request.

triage reports to whoever owns incoming failures. It may label. It may not close.

This sounds fussy. It is fussy on purpose. Unnamed automation becomes “the system,” and “the system” cannot be fired, coached, or deposed in an incident review. Named staff can.

Where this sits in the company

I do not think quality reports through engineering in 2030 in the companies that take AI risk seriously. Engineering wants to ship. Quality wants to be able to explain. Those are aligned until they are not, and models increase the number of days they are not.

In regulated work this is already true. In publishing it is becoming true, because the artifact is public. In consumer software it will be true the first time a generated suite signs off on a generated feature and the incident is bad enough to reach a board.

The org chart is a thesis about power. If quality does not own the contract the agents run under, the contract will be whatever was convenient for the last demo. I have watched that version. I am not interested in staffing it.