
Why Large Enterprise is Scared to Partner with Frontier Labs | Mercor CPO
The Twenty Minute VC
Hosted by Unknown
The 90% of enterprise workflows deemed AI-ready excludes every long-horizon agentic task that actually matters—and no enterprise has attempted a single one yet.
In Brief
The 90% of enterprise workflows deemed AI-ready excludes every long-horizon agentic task that actually matters—and no enterprise has attempted a single one yet.
Key Ideas
AI readiness misses long-horizon complex tasks
The 90% AI-ready workflow figure ignores every long-horizon agentic task that matters.
Token costs now exceed salary expenses
Mercor already spends more on tokens than salaries — 100% is the destination.
RL environments become preference data frontier
RL environments are the new preference ranking: the frontier data type nobody has solved yet.
AI services boom sustained by monopoly
The AI services boom is a decade-long SF knowledge monopoly, not a permanent business model.
Preserve human judgment; automate execution only
Never delegate judgment to a model — execution is fine, decision-making costs you the skill forever.
Why does it matter? Because the 90% AI-readiness figure everyone quotes is built on tasks nobody is actually attempting
Mercor's CPO Oswald Nitski spent this conversation dismantling the benchmarks enterprises use to feel confident about AI — and the frameworks investors use to size the market. What he reveals is that the most cited figure in enterprise AI is measuring the wrong thing, that security teams have their data placement exactly backwards, and that the next foundational data type is still largely unsolved.
- The "90% of workflows" readiness claim counts only tasks companies are already attempting — long-horizon agentic work is absent from every calculation, and Mercor's own Apex benchmarks put top models at just 50% there
- Enterprises are routing their most sensitive IP to self-hosted open-weight models while trusting Anthropic with HR and procurement — the precise opposite of logical security posture
- RL training environments — simulations of real apps that train agents under near-deployment conditions — are where preference ranking was in 2022: barely standardized, still hard to get right
- Mercor already spends more on tokens than on salaries; the AI services boom at Palantir and Microsoft has a decade-long expiration date baked in
The 90% AI-readiness figure measures tasks people already imagined — the real market is everything nobody has tried yet
The 90% figure is everywhere. Oswald doesn't buy it — not because models are weak, but because the calculation is built on the wrong inputs.
"These calculations might be based off of existing demand or things that come top of mind when current model users are thinking about what models could do," he says. "But there's a whole category of latent demand that people aren't even trying to do with models yet." His example: a procurement agent that automates an entire team for months, a human checking in once a week. Nobody's attempting that at enterprise scale. It doesn't appear in any benchmark.
Mercor's own Apex benchmarks put top models at around 50% on long-horizon workflows — and even that framing collapses for certain categories. Legal arguments, medical advice — these are uncapped-reward problems where the model can always get better. Treating them as pass/fail misses the structure entirely. "We need to be thinking more about continuous uncapped rewards," Oswald says.
The real implication: 90% is a snapshot of already-imagined tasks. The latent demand for long-horizon agentic automation hasn't been counted, which means the addressable market for frontier training data is far larger than current benchmarks suggest. Stop using the figure as a ceiling. It's a floor.
Enterprise security logic is backwards: the most sensitive IP goes to open-weight Chinese models, while HR data goes to Anthropic
Enterprises are putting their most competitively differentiated work — the legal memos, the proprietary workflows, the things that actually separate them from competitors — on self-hosted open-weight models they call "controlled." Generic HR and procurement, meanwhile, flows straight to closed frontier providers. The host catches the irony mid-conversation and can't let it go: "We put the sensitive sensitive data on open-source, most likely Chinese models — and we put the HR and procurement data on the closed model?"
Oswald's defense is nuanced: "The beauty about open-weight models is that the inference can happen in multiple places. You could make mistakes using them, but you have more control." Self-hosting does give infrastructure control. The reasoning isn't irrational. But the risk framework driving these decisions is built on vendor trust perception rather than actual data sensitivity analysis — companies worry about what Anthropic sees; they're less worried about where the open-weight model's weights came from, or what was in its training data.
It's a security posture optimized for the optics of control while potentially exposing the most valuable IP. Enterprise AI security teams need to interrogate which data goes where — not based on which vendor feels safer, but based on which data is actually most dangerous to lose.
RL training environments are the InstructGPT moment of 2025 — and almost nobody has cracked them yet
"The data type that's growing the fastest for us is environments."
The concept: high-fidelity simulations of real apps — Salesforce, enterprise file systems, business software stacks — that agents interact with during training. Pair that with a "rich start state" (a simulated world containing hundreds or thousands of files representing what you'd find on an actual machine) and training data finally looks like deployment conditions. If you want an agent to use Salesforce reliably, you need a mock that behaves exactly like Salesforce during eval and training.
The annotation complexity is severe. Agents must interact with the simulated world. The start state might run to thousands of files. Every edge case that surfaces in a real enterprise tool needs to be reproducible in the mock. "Just like years ago preference ranking was really hard to get set up — SFT was really hard to get set up when InstructGPT first came out — this is the frontier right now."
Labs are still figuring it out. Eventually, enterprises will run their own environment-based training projects. That window is years away. Whoever builds the infrastructure to deliver high-fidelity agent training environments at scale first will own the bottleneck for the next generation of capable AI — the same position InstructGPT's SFT data infrastructure held in 2022.
VC-funded founders are doing annotation themselves — and labs love the mispriced labor, right up until it can't scale
Mercor's sharpest competition right now isn't Surge. It's founders.
As the skill bar for annotation climbs — because models keep improving, meaning only higher-quality human feedback moves the needle — a cottage industry has emerged of startups where the founders themselves are making the training data. They raise venture capital, get lab access, and subsidize annotation with investor money. "Labs love this because it's just like totally mispriced," Oswald says. "They have loads of cash to blow... they're smart people, formerly great technical employees."
The scaling wall is obvious from the outside: one founder producing brilliant data points can't 10x throughput on demand. The operational machinery that handles hundreds of simultaneous heterogeneous projects doesn't exist at founder-annotation scale.
But the dynamic points toward where the field is actually heading: away from crowd labor, toward higher-skilled annotators, toward "the best people in the world" doing this work. The current VC subsidy is a price distortion, not a floor. If you're building in AI data, you're competing against founders burning investor capital below cost — design your moat around operational complexity and scale, not around matching their rates.
The AI services boom at Palantir and Microsoft is a San Francisco knowledge monopoly — and it has roughly a decade to live
Palantir skyrocketing on services. Microsoft standing up an AI deployment division. Oswald's read: transitional artifact, not structural feature.
"We have basically a concentration of a bunch of people in San Francisco who really know how to deploy agents, eval agents, be AI-first in engineering. And that knowledge just isn't out there yet." The services layer exists to bridge a geographic and temporal knowledge gap. Every enterprise that wants AI deployment expertise can't hire it because the talent barely exists outside one zip code. Matt from Factory's line gets quoted: services are an excuse for bad product. Oswald's reframe: they're an excuse for bad knowledge distribution, which is a solvable problem — just a slow one.
"Eventually it will be," he says of the knowledge spreading, "and maybe you won't need teams to go set up AI agents for every enterprise. It'll become more of a job function similar to software engineering." His timeline: roughly a decade.
Build products that assume the knowledge gap closes. Don't architect a business around services margins that exist only because San Francisco has a temporary monopoly on AI deployment know-how — because that monopoly is already dissipating.
When coding agents make engineering velocity unlimited, PMs — not engineers — become the binding constraint on growth
Skill issues have almost gone away. That's Oswald's read on what coding agents have done to product engineering at Mercor — a company that has 10x'd its headcount in under a year, with revenue rising commensurately.
When any competent engineer can execute at AI-assisted velocity, the constraint moves upstream. "We're bottlenecked by understanding business needs, user needs — and that's more of a PM job." The ratio of PMs to engineers is already shifting at Mercor; expect fewer engineers per PM as coding agent capability compounds.
What changes about what makes a great PM: tool mastery matters less (even Figma is being abandoned in favor of Claude's built-in design capabilities), business impact obsession matters more. "Now it's all about judgment and am I doing what is going to drive the most business value?"
The structural paradox: more velocity creates more surface area temptation. Engineers can ship faster, so teams do, and the product sprawls. The PM's hardest job in an AI-era org isn't accelerating output — it's ruthlessly preventing the product from ballooning into chaos. Mercor hires accordingly: biased toward senior candidates who can rapidly grok how the business makes money, because execution is no longer what's scarce.
Delegate execution to models, never decisions — the judgment you outsource is the judgment you permanently lose
Delegating decisions to a model doesn't just produce bad outputs occasionally — it atrophies the judgment muscle that made you valuable in the first place. That's Oswald's explicit warning to his team.
"I was very careful never to delegate judgment or decision-making to Miles" — his AI coding agent. "They make you think it's doing the right thing, but you have to be paranoid with them still." The fluency is the danger. Models produce confident, complete-feeling outputs. The atrophy happens quietly, beneath the surface of a working product.
"Don't delegate your decision-making, like your actual job, to a model — because you're going to lose that ability." At the organizational level, this shows up directly in hiring: Mercor's whiteboarding rounds exist to test whether candidates can reason through statistics, experimental design, and systems problems without offloading to Claude. The question beneath the question isn't whether you're AI-fluent. It's whether a human judgment capacity still exists underneath the tool use.
Draw an explicit line — AI handles execution, humans own judgment — and make it a hiring criterion. Interview for reasoning under ambiguity, not tool proficiency. Tool proficiency is now table stakes; it tells you almost nothing.
When execution gets cheap, judgment is the only thing left worth paying for
Every efficiency gain Oswald describes — faster engineering, cheaper annotation, broader AI deployment — drives execution cost toward zero. That convergence has one underappreciated consequence: judgment is what remains scarce. Not throughput, not headcount, not shipping velocity. The capacity to decide what to build, which workflows to automate, which data types will matter in two years — that's what Mercor charges a premium for, what its PMs are hired to supply, and what Oswald refuses to let a model touch. The organizations that thrive in that world won't be the ones that automated most aggressively. They'll be the ones that never let the model make the call.
Topics: AI data, human data annotation, enterprise AI, Mercor, frontier models, open source AI, RL environments, product management, AI services, token spend, agentic AI, data annotation market
Frequently Asked Questions
- What does the 90% AI-ready workflow metric miss about enterprise AI capability?
- The 90% AI-ready workflow figure ignores every long-horizon agentic task that matters, creating false confidence about enterprise AI readiness. No enterprise has yet successfully attempted even a single long-horizon agentic task in production. The fundamental gap exists between incremental workflow optimization and handling genuinely complex challenges requiring sustained reasoning, multi-step execution, and contextual judgment calls. These metrics measure something entirely different from actual frontier AI capabilities—they optimize existing processes rather than enabling transformative problem-solving. Understanding this distinction is critical for setting realistic enterprise AI expectations.
- How are AI token costs evolving compared to employment expenses for organizations like Mercor?
- Mercor already spends more on tokens than salaries—100% is the destination. This represents a fundamental restructuring of enterprise economics as AI capabilities expand and mature. Token costs are becoming the primary operational expense, gradually replacing traditional labor investments in organizational budgets. This trajectory demonstrates how AI-driven enterprises operate under entirely different economic models compared to legacy organizations. Companies must recognize this cost evolution when planning AI strategy and budget allocation, understanding that computational resources will increasingly dominate expenses over traditional headcount and human salaries.
- What is the critical frontier data type that AI development still needs to solve?
- RL environments are the new preference ranking: the frontier data type nobody has solved yet. Where preference rankings revolutionized AI training through human feedback, RL environments represent the next critical innovation frontier. These environments enable dynamic, interactive learning where AI agents continuously improve through iterative feedback loops and simulation-based scenarios. Solving RL environment data generation and utilization challenges is essential for advancing AI beyond static task automation toward genuine agentic reasoning, autonomous decision-making, and truly intelligent execution—capabilities enterprises need for meaningful transformation.
- Should enterprises delegate strategic decision-making to AI models?
- Never delegate judgment to a model—execution is fine, decision-making costs you the skill forever. This principle reflects a critical organizational risk: outsourcing judgment to AI erodes internal expertise and institutional capability developed over years. Enterprises can safely deploy AI for tactical execution, process automation, and information synthesis, yet must preserve human judgment for strategic decisions and complex contextual choices. Once decision-making authority transfers to models, rebuilding that organizational skill becomes extraordinarily difficult and expensive, leaving companies dependent on external systems for capabilities they once owned internally.
Read the full summary of Why Large Enterprise is Scared to Partner with Frontier Labs | Mercor CPO on InShort
