Skip to content
aitrainer.work - AI Training Jobs Platform
Industry Analysis

The End of Crowd-Work: How Domain Experts Are Restructuring the Economics of AI Post-Training

We coded 1,610 active AI evaluation listings across 9 platforms. Domain Expert Track roles pay a median of $75/hr — 1.9x the Generalist Track median of $40/hr. Legal and Healthcare roles reach $100+ medians.

By Pietro R., aitrainer.work | Source: aitrainer.work proprietary data |
Beyond Crowd-Work: The Domain Expert Track Pay Premium — aitrainer.work

Abstract. Early AI data work was volume-driven: generalist contractors drew bounding boxes, tagged images, and verified simple classifications at scale. That model is dissolving. As frontier models approach ceiling performance on standard benchmarks, the bottleneck in AI development has shifted from data volume to evaluation quality -- specifically, to the kind of normative professional judgment that generalist annotators cannot provide. We examined 1,610 active AI evaluation listings across nine platforms to measure how this structural shift is registering in posted compensation. We find a measurable pay separation between what we term the Generalist Track (high-volume, low-credential platforms: Turing, Micro1, AfterQuery, Mindrift) and the Expert Track (credentialed-specialist platforms: Mercor, SME Careers, Ethos, Alignerr): a median hourly rate of $75 vs. $40, a 1.9x premium. Within Expert Track listings, pay varies sharply by domain: Legal roles post a median of $100/hr, Healthcare $85/hr, and Finance and Mathematics $80/hr each -- against Languages at $32.50/hr and Voice Acting at $35/hr at the other end. We report sample sizes with every finding, apply a $200/hr outlier screen to the platform-level figures, and note where domain-level samples fall below our n=10 reliability threshold.

Key takeaways

  • 1.9x pay premium: Expert Track platforms post a median hourly rate of $75 vs. $40 on Generalist Track platforms -- a gap that holds across every platform pair in the sample.
  • 32% of Expert Track listings clear $100/hr vs. 4% of Generalist Track listings. The high-pay tier is structurally concentrated on credential-gated platforms.
  • Domain determines pay more than platform. Legal listings ($100/hr median) earn more than Software Engineering listings ($55/hr median) on the same platform, and that domain gradient is consistent across SME Careers and Mercor independently.
  • Credential requirements are explicit and front-loaded on Expert platforms. SME Careers specifies a minimum degree on 93% of listings (43% Bachelor's, 37% Master's, 13% PhD). Mercor specifies a degree on 40% of listings.
  • The platform ceiling varies by 6x: Turing's maximum disclosed hourly rate is $65/hr. Mercor's is $237.50/hr (and its full-time network listings reach higher still). The practical implications for career strategy differ materially.

A note on terms. Throughout this piece, "domain expert" and SME -- subject matter expert -- refer to the same role: a credentialed professional (physician, attorney, engineer, PhD researcher) hired not to label data but to exercise professional judgment on AI model output within their field. This is distinct from the older crowd-work model, where a generalist rater without domain credentials could complete the task. The SME is the labor category this piece is about; the Expert Track / Generalist Track split below is our attempt to measure where and how much that category is being paid for.

Methodology

Population and sampling frame. The dataset comprises all listings marked active in our curated role intelligence system as of July 23, 2026. Coverage: Mercor (n=182), Micro1 (n=274), Turing (n=242), SME Careers (n=207), Ethos (n=67), Alignerr (n=34), AfterQuery (n=103), Terac (n=38), and smaller platforms (Vetto, Rex: n=15 combined); total N=1,610. All listings were denominated in USD at the time of collection; non-USD rates were normalized to USD using static exchange rates applied at scrape time.

Track classification. We classify platforms into two tracks based on their structural orientation: the Generalist Track (Turing, Micro1, AfterQuery, Mindrift) and the Expert Track (Mercor, SME Careers, Ethos, Alignerr). Expert Track platforms systematically require credentials as application prerequisites, gate listings through multi-stage vetting (AI interview, domain test, recruiter screen), and post rates denominated in hours rather than tasks. Generalist Track platforms accept applications without prior credential verification and primarily post task-based or flexible-hours roles accessible to any English speaker. We do not treat this as a quality claim; it is a structural descriptor.

Pay extraction and outlier handling. Pay figures are extracted from each listing's structured data fields (pay_amount, pay_min, pay_max, pay_rate_unit). We use pay_amount as the primary measure, which represents the midpoint of the stated range where one exists and the single stated figure where it does not. We exclude listings with pay_amount = 0 or where pay_rate_unit is not per_hour. For platform-level medians and averages, we apply a $200/hr outlier screen: listings above $200/hr are excluded from platform aggregates, flagged separately in Finding 4, and reported with their own sample sizes.

Missing-data handling and reporting threshold. We report domain-level medians only where n≥10 listings in that domain-platform cell disclose a non-zero hourly rate. Below that threshold we describe the cell as insufficient sample. All percentages in this piece are accompanied by their denominator (n) in the text itself, not only in a footnote.

Finding 1: The Expert Track Pay Premium

The most direct quantification of the structural shift in AI evaluation economics is the gap between what Expert Track platforms post and what Generalist Track platforms post. Across 1,235 listings with disclosed hourly rates, Expert Track listings (n=754) show a median of $75/hr and Generalist Track listings (n=726) a median of $40/hr -- a 1.9x premium.

Bar chart: median hourly rate by platform comparing Expert Track vs Generalist Track platforms.
Track Platforms N (with pay) Median/hr Maximum/hr
Expert TrackMercor, SME Careers, Ethos, Alignerr754$75$237.50
Generalist TrackTuring, Micro1, AfterQuery, Mindrift726$40$182

The high-pay threshold ($100/hr) reveals a sharper structural difference than the median alone. 32% of Expert Track listings (245/754) clear $100/hr versus 4% of Generalist Track listings (27/726). This 8x difference in the rate of high-pay listings is not an artifact of one outlier platform: every Expert Track platform in the sample has a higher share of $100+/hr listings than every Generalist Track platform in the sample.

Platform Track N Median/hr Avg/hr % above $100/hr Max/hr
MercorExpert180$70$7530%$237.50
EthosExpert67$80$8646%$150
AlignerrExpert34$80$9635%$190
SME CareersExpert207$55$5714%$150
Micro1Generalist229$47.50$516%$150
TuringGeneralist242$27$330%$65
AfterQueryGeneralist0*--------

*AfterQuery listings do not disclose hourly rates in a machine-readable field; the platform uses task-based and project-based compensation. Its 103 listings are included in the active listing count but excluded from hourly pay analysis.

SME Careers sits at a lower median ($55/hr) than the other three Expert Track platforms despite being credentialed-specialist-only by design. The reason is compositional: SME Careers' largest listing category by volume is quality assurance leads and language specialists (see Finding 2), which pull the platform median down even though its specialist medical, legal, and engineering roles are competitive with Mercor. Platform medians are a summary of whatever the platform happens to be hiring for at a given point in time, not a fixed property of the platform. Domain-level comparison is more informative.

Finding 2: The Domain Gradient

Pay varies more by domain than by platform. When we aggregate all 1,235 hourly-pay listings by domain category and apply the n≥10 threshold, a clear gradient emerges: Legal listings ($100/hr median, n=74) are the highest-compensated domain in the full dataset, followed by Healthcare ($85/hr, n=48), Accounting ($85/hr, n=20), Finance ($80/hr, n=61), and Mathematics ($80/hr, n=20). At the other end, Languages ($32.50/hr, n=163) and Voice Acting ($35/hr, n=83) post medians below the Generalist Track median.

Bar chart: median hourly rate by domain across 1,235 hourly pay listings in AI evaluation.
Domain N Median/hr Avg/hr Max/hr
Legal74$100$95$237.50
Healthcare48$85$92$155
Accounting20$85$75$175
Finance61$80$86$210
Mathematics20$80$74$175
Humanities9*$75$84$175
Business161$70$69$300
Design18$65$62$100
Writing29$60$75$350
Generalist138$60$65$300
STEM148$60$66$190
Software Engineering197$55$61$180
Data Science50$47.50$53$150
Voice Acting83$35$33$60
Languages163$32.50$37$115
Psychology11$21$46$115

*Humanities is reported with n=9, below our n≥10 threshold; included for reference but flagged as insufficient sample for inferential use. All other rows cleared n≥10.

The domain gradient holds when we isolate individual platforms. On SME Careers, Legal roles post a median of $115/hr (n=7) and Medical roles $100/hr (n=7) against the platform overall median of $55/hr -- a 2x within-platform gap driven by domain, not platform selection. On Mercor, Legal roles post $112.50/hr (n=7), Medical $115/hr (n=9), and Software Engineering $110/hr (n=16) at the top end, against the platform median of $70/hr. In both cases, the platform-level median is a weighted average of a heterogeneous mix: a listing for a Catalan Language Specialist at $25/hr exists on the same platform as a listing for a US Labor and Employment Lawyer at $237.50/hr.

The practical implication is that a professional choosing between platforms should be comparing at the domain level, not the platform level. A physician evaluating AI medical content is comparing Ethos, Mercor, and SME Careers' Healthcare cells -- all in the $85--$115/hr range -- not platform averages. A software engineer is comparing Software Engineering cells across the same platforms: $55--$110/hr, depending on seniority signals in the listing.

Finding 3: Credential Density on Expert Platforms

The pay premium on Expert Track platforms is structurally associated with explicit credential requirements. SME Careers is the most explicit about this in the dataset: 93% of its listings (193/207) specify a minimum degree requirement -- 43% (90) require a Bachelor's, 37% (77) a Master's, and 13% (26) a PhD. Only 7% (14) leave the credential field unspecified.

Platform Bachelor's Master's PhD Not specified Any degree required
SME Careers43% (90)37% (77)13% (26)7% (14)93%
Mercor20% (36)13% (24)7% (12)60% (110)40%
Ethos------100%0%*

*Ethos does not surface a structured degree field in its listing data; credential requirements are embedded in prose descriptions and were not parsed for this analysis. Ethos's vetting is known to include active professional licensure checks at the application stage, which is a functional equivalent not captured in this table.

Mercor's lower declared-credential rate (40%) reflects a different verification architecture: many Mercor listings foreground professional experience and peer-reviewed publication records (e.g., "ML Research PhD Experts (ICML / NeurIPS / ICLR Publications)") in prose without populating the structured degree field. The 60% "not specified" on Mercor does not indicate the absence of credential requirements; it indicates that Mercor routes verification through the interview and vetting process rather than through the listing form.

The degree requirement on SME Careers correlates with pay at the job-family level. Listings requiring a PhD have a median pay of $75/hr (n=26), those requiring a Master's $40/hr (n=77), and those requiring a Bachelor's $55/hr (n=90). The seemingly counterintuitive PhD-vs-Master's pattern reflects job mix rather than a negative return to a doctorate: the PhD pool is concentrated in STEM specialist roles (astronomer, biomedical engineer, data scientist team lead), while the Master's pool is heavily weighted toward language-specialist and QA-lead roles that carry lower base rates regardless of credential.

Finding 4: Where the $100/hr Ceiling Breaks

The $200/hr outlier screen applied to platform aggregates in Finding 1 excluded a small but real set of listings that represent the upper bound of what the AI evaluation market currently posts. We report them separately because they are substantively informative about where the structural ceiling is. Across the full 1,610-listing dataset, Mercor accounts for all listings above $200/hr.

Listing title Platform Hourly rate Domain
Regulatory Law ExpertMercor$1,950Legal
Cybersecurity ExpertMercor$1,950STEM
Legal Technology ExpertMercor$1,950Legal
Investment Banking ExpertMercor$1,950Finance
Private Detectives & InvestigatorsMercor$1,600Legal/Security
Health Insurance ExpertMercor$1,300Healthcare
Biotechnology Research ExpertMercor$1,300Healthcare
US & UK-Based Labor & Employment LawyersMercor$237.50Legal
Management Consulting ExpertMercor$185Business
Clinical Law Professor / Clinic DirectorMercor$180Legal
Physician Talent NetworkMercor$180Healthcare

Two observations. First, the listings in the $1,000+/hr range are structurally different from the hourly evaluation roles that make up the bulk of this dataset: they appear to represent project-rate or retainer engagements denominated as an hourly equivalent, not traditional gig-economy hourly work. Their inclusion in an hourly dataset warrants caution; we report them for completeness but do not treat them as representative of the Expert Track hourly market. Second, even setting aside the $1,000+ outliers, the range from $180--$237.50 at Mercor represents genuine hourly evaluation work performed by actively licensed professionals (physicians, attorneys, academic clinical directors). The TNW-cited aitrainer.work finding that Ethos posts rates in the $105--$225/hr range for credentialed expert evaluators is consistent with what we observe here, representing the realistic high end of the market for Expert Track work.

Discussion: What Is Actually Being Purchased

The pay gap between tracks is not simply a market quirk -- it reflects a structural shift in what AI labs are buying. Early RLHF data collection was designed around tasks where the correct answer was checkable by a naive rater: "does this image contain a cat," "rank these two responses by fluency." Human judgment served as a cheap, scalable approximation of ground truth on tasks with unambiguous answers.

The tasks that Expert Track platforms list have no such unambiguous ground truth. Whether a clinical diagnosis is sound, whether an employment contract contains an illegal restraint of trade clause, whether a biomedical engineering calculation meets regulatory standards -- these are normative judgments that require not just information but the professional infrastructure that converts information into defensible conclusions. This is consistent with how RLHF alignment itself is understood: model behaviors like safety, tone, and helpfulness are trained against normative, context-dependent conventions rather than fixed factual answers, which is why judging them correctly requires a rater who actually holds the relevant professional or cultural context. A physician evaluating an AI-generated differential diagnosis is not "labeling data." They are providing the ground truth reference against which the model's output is measured, in a domain where that reference is only available from someone who has completed the professional formation required to issue it.

This distinction between verification (can a generalist check this?) and authority (does this require professional formation to evaluate?) maps directly onto the Taxonomy Framework's three-tier model of RLHF annotation: Extension (volume), Evidence (factual accuracy), and Authority (professional standards). The domain gradient in Finding 2 is a market price signal for that authority: Legal and Healthcare posts clear $100/hr because courts, hospitals, and regulatory bodies recognize the cost of forming a professional capable of exercising that authority -- and AI labs are now paying to access it. Languages and Voice Acting clear $32--$35/hr because the authority required is different in kind and more broadly distributed.

The economic consequence for professionals is that the decision to enter the AI evaluation market should be made at the domain level, not the "AI training" level. A nurse deciding whether to supplement income through AI evaluation work is not choosing between "AI training" and "nursing" -- they are assessing whether the Healthcare cell of the Expert Track (median $85/hr, ceiling $155/hr based on current active listings) is an appropriate use of their credential. A software engineer evaluating the same question should compare the Software Engineering cell ($55/hr median on SME Careers, $110/hr median on Mercor for the same domain) against their opportunity cost.

Limitations

Selection and sampling frame sensitivity. While the macro trend (credentialed platforms paying a structural premium over generalist crowd platforms) is robust across every platform pair, specific numerical medians (e.g., $75 vs. $40/hr) are sensitive to listing aggregation and categorization. Platforms carry vastly different job mixes: for instance, SME Careers' lower overall median ($55/hr) relative to Mercor ($70/hr) is heavily influenced by a large volume of language-support and QA-lead roles in its active catalog rather than lower pay for equivalent specialist positions.

Nominal posted rates vs. realized hourly earnings. This dataset measures advertised hourly rates on active postings, not verified post-payout earnings. In practice, effective hourly rates frequently lag posted rates once unpaid administrative overhead (unpaid onboarding, qualification tests, task queue wait times, and rejected work) is factored in. Furthermore, posted ranges often reflect tier 1 (US/UK) ceiling rates, whereas geographic rate adjustments routinely lower effective pay for workers based in other regions.

Task-chunk distortions in niche domains. Domain-level anomalies (such as Psychology posting a $21/hr median across 11 listings) frequently reflect platform listing conventions rather than lower professional valuation. In specialized domains, platforms often post fractional task-chunk payouts, micro-screening assessments, or entry-level rater roles under broad academic category tags rather than full professional consulting contracts, artificially depressing the calculated hourly rate.

Cross-sectional data snapshot. This study reflects active public listings as of July 23, 2026. Listing inventories rotate rapidly based on AI lab demand cycles. The track classification (Expert vs. Generalist) describes platform structural characteristics and candidate entry gates, not worker quality or individual platform preference.

Conclusion

The structural separation between Expert Track and Generalist Track AI evaluation platforms is measurable in posted compensation at a scale that justifies the distinction. A 1.9x median pay premium ($75 vs. $40/hr), an 8x difference in the share of listings above $100/hr (32% vs. 4%), and a consistent domain gradient from $100/hr (Legal) to $32.50/hr (Languages) are all observable from a single snapshot of 1,610 active listings. The cross-platform consistency of the domain gradient -- Legal outperforming Software Engineering outperforming Languages on both SME Careers and Mercor independently -- suggests the gradient reflects genuine market pricing of domain authority rather than an artifact of how any single platform sets its rates.

What this means for the AI evaluation economy in practice: the transition from crowd-based annotation to domain-expert evaluation is not a gradual blending -- it is a structural fork. The platforms, the vetting mechanisms, the credential requirements, and the compensation levels are all distinct. Generalist Track platforms serve a labor supply function for tasks that do not require professional authority. Expert Track platforms serve a knowledge-acquisition function for tasks where professional formation is the prerequisite to providing usable ground truth. The two tracks compete for the same headline category -- "AI training jobs" -- but they are purchasing different things, and the price gap reflects that difference.

Sources

  1. aitrainer.work proprietary role intelligence extraction, active listings as of July 23, 2026: Mercor (182), Micro1 (274), Turing (242), SME Careers (207), Ethos (67), Alignerr (34), AfterQuery (103), Terac (38), other platforms (15); total N=1,610.
  2. Domain-level pay figures computed from listings that disclose a non-zero hourly rate (n=1,235 of 1,610 total). Per-domain n reported inline throughout. Cells below n=10 flagged as insufficient sample.
  3. Credential distribution data (min_degree field) extracted from SME Careers (n=207) and Mercor (n=182) formatted listing files. Ethos degree field not populated in structured data; omitted from credential table.
  4. Track-level comparison excludes AfterQuery from hourly pay figures (0 listings with disclosed hourly rate in the dataset) and applies a $200/hr outlier screen to platform medians and averages. Outlier listings reported separately in Finding 4.
  5. Taxonomy Framework: The Taxonomy Framework: Three Models of RLHF Annotation: Extension, Evidence, and Authority (arXiv, April 2026) -- referenced in the Discussion section as an external conceptual frame; this piece's empirical findings are independent of that framework.
  6. Toloka AI, Complete guide to RLHF for LLMs (Toloka AI, February 2026) -- referenced in the Discussion section on the normative, context-dependent nature of RLHF alignment judgments.
  7. TNW citation: Ana-Maria Stanciuc, "Ethos lands $22.75m Series A to fix what AI broke about hiring," The Next Web, May 6, 2026. The cited aitrainer.work per-hour rate finding ($105--$225 for Ethos) is consistent with the $80--$150/hr range observed for Ethos in this dataset (different listing cohort, different collection date).

Related reading

The AI Training Labor Market (2026): Cross-Platform Analysis of 1,615 Listings -- our earlier piece on interview requirements, weekly hours, and contract duration across the same platform set.

Best AI training platforms compared: full reviews of every platform in this analysis.

Mercor job listings -- see active Expert Track roles behind the Mercor figures above.

SME Careers job listings -- see active Expert Track roles behind the SME Careers figures above.

AI Training Pay Survey -- self-report your own rates to contribute to the next wave of this analysis.

Pietro R., founder of aitrainer.work

Pietro R.

MSc Human-Computer Interaction | Founder & Product Owner

Pietro is the founder and technical lead of aitrainer.work. He builds and maintains the platform's data pipeline, certification infrastructure, and editorial standards.

Comments

Loading comments…
💬

Share your thoughts on this article

Sign in to join the discussion.

Sign in to comment