Own a Company? AI Labs Are Paying Up to $1M a Year for Your Workflow Data
Frontier labs ran out of public internet text to train on. Mercor and micro1 now pay ordinary mid-size companies six and seven figures a year to license the anonymized record of how their teams work.
Frontier AI labs have a problem money alone hasn't fixed. The open internet's stock of usable human-written text has already been scraped, re-scraped, and trained on past the point of real returns, a slowdown researchers have started calling the "data wall." The current scramble is for something the labs barely touched until recently: proof of how a real business operates day to day. A new category of intermediary has formed to package that proof and sell it.
The Data Wall
Public text taught early models what things are. A Wikipedia page can describe corporate finance in the abstract. It can't show the iterative path an accountant takes through an internal ERP system to clear a billing exception, or how a sales team moves a deal through a CRM under real pressure. The frontier labs' current focus, autonomous agents that complete multi-step work without a human in the loop, depends on exactly that kind of operational detail, and public data simply doesn't contain it.
Who's Buying It
Mercor and micro1 have built the brokerage layer connecting mid-size companies to that demand. Both pitch the arrangement as a recurring revenue line for the company, shaped more like a long-term licensing deal than a single data sale: anonymized contributions of Slack threads, CRM histories, code-review trails, and accounting reconciliation logs, paid out on an ongoing basis to the company that supplies them.
micro1 publishes tiered pricing publicly: standard workflow contributions starting around $100,000 a year, scaling past $500,000 for companies contributing across multiple workflow types, and reaching $1 million or more annually for highly specialized or proprietary operational data. Mercor runs a comparable model, connecting to a company's existing tools through read-only OAuth, extracting data structurally, anonymizing it, and paying out based on volume and depth, though it hasn't published flat tiers the way micro1 has.
What "Anonymized" Actually Means
Any company's security team will ask the obvious question: what exactly is leaving the building? Mercor describes its extraction pipeline as using deterministic pattern matching alongside generative-AI filters to detect and mask more than 60 categories of personally identifiable, health, and business-identifying information before anything leaves the company's environment, rendering the resulting tokens irreversible. Access is read-only and scoped to whatever tools and channels a company chooses to connect, rather than a full database export. These claims come from the vendors' own published security material. Independent third-party verification of these pipelines remains unavailable, leaving the masking guarantees resting on each platform's own assurances.
Provenance is a separate, and arguably harder, problem than anonymization: proving to a buyer, and eventually a regulator, that licensed training data was actually licensed rather than quietly scraped. That's the gap Story Protocol is now chasing. The company rebranded as the DATA Foundation on June 25, 2026, and launched Trace, an on-chain registry that timestamps a content hash, consent terms, license, and payment status for each data contribution. Its launch integration already covers more than a billion records from the human-data marketplace Kled, evidence that audit infrastructure for licensed data is being built ahead of a regulatory mandate, while the rules are still being written.
What Labs Are Actually Paying For
The premium tracks rarity and depth more than raw volume. Current high-value verticals include:
| Vertical | What's being licensed |
|---|---|
| Software engineering | How internal teams debug, review, and ship code across private repos |
| Finance & accounting | Reconciliation workflows, invoice-discrepancy resolution, regulatory reporting steps |
| Revenue / CRM | How a B2B deal moves from first outreach to signature |
| Operations & logistics | Exception handling, route changes, fulfillment decisions under real constraints |
The Bottom Line
For a mid-size company with genuinely documented workflows, this is a real and growing revenue line. It carries open questions, too: selling the record of how employees work raises consent and labor-relations concerns that few firms have policies for yet, and regulators have barely begun to address whether anonymized workflow data stays anonymous once a model has absorbed it. For everyone else, it's an early look at where AI training spend is heading next: toward the operational record of how work actually gets done, licensed and audited rather than scraped.
Do you own or run a company? You can browse current data partnership opportunities on aitrainer.work, or run a quick, free questionnaire first to see roughly where your company would land: check your eligibility.
Related reading
Mercor review β how Mercor pays individual contributors, for contrast with what it now pays companies for workflow data.
micro1 review β the platform publishing $1M/year enterprise data partnership tiers.
Mercor's supply-chain breach β the security incident that makes "anonymized" claims worth scrutinizing.
Scale AI's fallout β how Mercor and micro1 grew into the position to run this kind of brokerage.
Sources

Pietro R.
MSc Human-Computer Interaction | Founder & Product Owner
Pietro is the founder and technical lead of aitrainer.work. He builds and maintains the platform's data pipeline, certification infrastructure, and editorial standards.