01Lead
Work
Outcomes over tasks. The through-line is rooms where a miss has a cost, and now a model in those rooms.
Progression
01
Early career →
02
Regulated / GxP
03
Markets & ops
04
2023 →
05
Now
Putting LLMs inside production QA — not a sandbox demo — with constrained roles, evals on the evaluator, and a human who still signs the release.
Outcomes
- ~70% reduction in manual effort on targeted ebook QA paths.
- ~$200K in cost avoidance by replacing a class of vendor-shaped work with a harness the team owns.
- A working prejudice: generation without an oracle is theater.
The model is staff. Staff need a job description, a manager, and a paper trail. That is the operating system I am now building twice: once inside a publisher, once as Nexacore.
Case studies
Problem, approach, result, meaning.
Digital publishing · production
AI-augmented ebook QA pipeline
- Problem
- Ebook quality is linguistic, structural, and high-volume. Click-scripts do not read a book. Adding people linearly does not survive the catalog. A miss is public.
- Approach
- Constrain a model (Claude in the loop) to specific checks with structured output. Keep deterministic oracles where they exist — metadata, links, navigation, format validity. Send the rest to a human with a diff, not a vibe. Sign the release as a person.
- Result
- ~70% less manual effort on the paths the harness owns. ~$200K in avoided cost against the vendor-shaped alternative. The surprising part was not generation. It was how much of the old suite was checking the wrong layer.
- What it means
- AI-in-QA works when the model has a job description. It fails when the job is “be a QA engineer.” Publishing made that obvious because the artifact cannot be patched after it ships.
Automation · shared harness
Python QA tooling for publishing workflows
- Problem
- Checks lived in people’s heads, vendor portals, and one-off scripts. Formats, devices, and OS versions multiplied. Nothing composed. Flakes were tribal knowledge.
- Approach
- A Python harness with Playwright where the UI actually mattered, shared fixtures, and oracles that fail for a reason a human can read. Tosca still in the estate where it earned its keep. The point was a system the team would run on a Tuesday, not a framework to present.
- Result
- [PLACEHOLDER] Critical-path regression pulled off the critical path of people’s weeks. Exact cycle-time figures on request. More useful: the suite became a place to put knowledge instead of a place knowledge went to die.
- What it means
- Automation that does not get run is documentation with extra steps. The quality of a framework is whether it is still alive after the person who wrote it has a different job.
AI evaluation methodology
Claude API QA capstone
- Problem
- If the model is the product — or a staff member in the product — string-matching the output is not a test. Teams were demoing “it wrote a test” and calling it quality.
- Approach
- An evaluation methodology: failure taxonomy, goldens that represent the incidents you actually fear, regression on behavior rather than wording, and a human-signed threshold for what “good” means this month. Python + the Claude API, treated like any other dependency.
- Result
- A repeatable way to ask whether the system got worse. Not a leaderboard. [PLACEHOLDER] Capstone figures and fixtures can be published once stripped of internal corpus.
- What it means
- QA for AI is still QA. Evidence, repeatability, ownership. The new part is that the oracle is harder — which is an argument for more discipline, not less.
A résumé exists for the rooms that still want one.
Download résuméLast updated May 2026