Harness engineering
How do AI agents work reliably instead of just impressively? I run measurement series across hundreds of agent runs and document the findings in a private research paper: test benches, sabotage probes, failure classes, and the rules that follow from them. That methodology carries every product I ship.
I'm happy to share the research document on request.